Operation and maintenance robot, method, equipment, server, system and storage medium
By deploying operation and maintenance robots in the server cluster and using BMC for automatic monitoring and failure analysis, the problems of high cost and low efficiency of manual inspection in the existing technology are solved, and efficient and low-cost server operation and maintenance are achieved.
Patent Information
- Application Number
- CN202510527788.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-05-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, server operation and maintenance relies on manual inspection, which is costly and inefficient, making it difficult to meet the management needs of large-scale data centers.
By deploying a dedicated operation and maintenance robot in the server cluster, the server's baseboard management controller (BMC) is used for automatic monitoring and failure analysis, and accurate operation instructions are generated and executed to resolve the failure.
Real-time acquisition and intelligent analysis of multiple servers is realized, and operation instructions are automatically generated and executed, which improves operation and maintenance efficiency, shortens the average failure recovery time, reduces operation and maintenance costs, and reduces manpower dependence.
Smart Images

Figure CN120069851A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of server operation and maintenance technology, and in particular to operation and maintenance robots, methods, equipment, servers, systems and storage media. Background Art
[0002] Server operation and maintenance is a technical management system that ensures the stable and efficient operation of servers. With the acceleration of digital transformation, the scale of servers is growing exponentially. The manual operation and maintenance used in related technologies is costly and inefficient, and can no longer cope with complex challenges.
[0003] Therefore, how to realize automatic operation and maintenance of multiple servers at low cost and high efficiency is a technical problem that needs to be solved urgently. Summary of the invention
[0004] The present application provides an operation and maintenance robot, method, device, server, system and storage medium to at least solve the problem of high cost and low efficiency in related technologies.
[0005] In a first aspect, the present application provides an operation and maintenance robot, comprising: a first communication module and a management platform module;
[0006] A first communication module, connected to the management platform module, is used to receive operation data sent by multiple servers and send the operation data to the management platform module;
[0007] A management platform module is used to parse the operation data to obtain analysis results; the analysis results include fault information of multiple target servers among the multiple servers; generate a first operation instruction corresponding to a first server among the multiple target servers according to the analysis results, and send the first operation instruction to the first communication module;
[0008] The first communication module is further used to send a first operation instruction to the first server; the first operation instruction is used to instruct the first server to perform a corresponding first operation.
[0009] In a second aspect, the present application provides a server, including: a management controller and a second communication module;
[0010] A management controller connected to the second communication module, used to obtain the operation data of the server and send the operation data to the second communication module;
[0011] The second communication module is used to send the operation data to the operation and maintenance robot;
[0012] The second communication module is further used to receive the first operation instruction sent by the operation and maintenance robot;
[0013] The management controller is further used to execute the corresponding first operation according to the first operation instruction.
[0014] In a third aspect, the present application provides a server management system, including: an operation and maintenance robot as described in the first aspect above and multiple servers as described in the second aspect above.
[0015] In a fourth aspect, the present application provides an operation and maintenance method, including:
[0016] Obtaining operation data sent by multiple servers;
[0017] Parsing the operation data to obtain an analysis result; the analysis result includes fault information of multiple target servers among the multiple servers;
[0018] Generating a first operation instruction corresponding to a first server among the multiple target servers according to the analysis result; the first operation instruction is used to instruct the first server to perform a corresponding first operation.
[0019] In a fifth aspect, the present application further provides an operation and maintenance device, including:
[0020] An obtaining module, configured to obtain operation data sent by multiple servers;
[0021] A parsing module, configured to parse the operation data to obtain an analysis result; the analysis result includes fault information of multiple target servers among the multiple servers;
[0022] A generating module, configured to generate a first operation instruction corresponding to a first server among the multiple target servers according to the analysis result; the first operation instruction is used to instruct the first server to perform a corresponding first operation.
[0023] In a sixth aspect, the present application further provides an electronic device, including: a memory for storing a computer program; a processor for implementing the steps of any of the above operation and maintenance methods when executing the computer program.
[0024] In a seventh aspect, the present application further provides a computer-readable storage medium, in which a computer program is stored, and wherein the computer program, when executed by a processor, implements the steps of any of the above operation and maintenance methods.
[0025] In an eighth aspect, the present application further provides a computer program product, including a computer program, and the computer program, when executed by a processor, implements the steps of any of the above operation and maintenance methods.
[0026] Through this application, the operation data of multiple servers is collected in real time, intelligently analyzed, and accurate operation instructions are automatically generated and executed in a closed loop. First of all, it can improve efficiency, greatly compress the mean time to recovery, avoid service interruption. Secondly, it can reduce costs. A single operation and maintenance robot can manage thousands of servers, reducing the manpower dependence by 70%. It can also enhance reliability, avoid human errors, and provide a highly flexible and low-cost unmanned operation and maintenance solution for scenarios such as cloud data centers and edge computing. Brief Description of the Drawings
[0027] To more clearly illustrate the embodiments of this application, the following will briefly introduce the drawings required for the embodiments. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0028] Figure 1 Structural schematic of the server management system provided by the embodiment of this application Figure 1 ;
[0029] Figure 2 Structural schematic of the operation and maintenance robot provided by the embodiment of this application Figure 1 ;
[0030] Figure 3 Structural schematic of the operation and maintenance robot provided by the embodiment of this application Figure 2 ;
[0031] Figure 4 Structural schematic of the server management system provided by the embodiment of this application Figure 2 ;
[0032] Figure 5 Flow schematic of the operation and maintenance method provided by the embodiment of this application;
[0033] Figure 6 Structural schematic of the operation and maintenance equipment provided by the embodiment of this application;
[0034] Figure 7 Structural schematic of the electronic device provided by this application. Detailed Description of the Embodiments
[0035] The following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the drawings in the embodiments of this application. Obviously, the described embodiments are only some embodiments of this application, rather than all embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of this application.
[0036] It should be noted that the orientation or positional relationship indicated by terms such as "center", "longitudinal", "lateral", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", "axial", "radial", "circumferential", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation to the present application. The terms "mounted", "connected", and "coupled" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two elements. The terms "parallel", "perpendicular", and "equal" include the described situations and situations similar to the described situations, and the range of the similar situations is within an acceptable deviation range, where the acceptable deviation range is determined by those of ordinary skill in the art considering the measurements being discussed and the errors associated with the measurements of specific quantities (i.e., the limitations of the measurement system). For example, "parallel" includes absolute parallelism and approximate parallelism, where the acceptable deviation range of approximate parallelism may be, for example, within 5° deviation; "perpendicular" includes absolute perpendicularity and approximate perpendicularity, where the acceptable deviation range of approximate perpendicularity may also be, for example, within 5° deviation. "Equal" includes absolute equality and approximate equality, where the acceptable deviation range of approximate equality may be, for example, that the difference between the two equal ones is less than or equal to 5% of either one of them. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific circumstances.
[0037] Moreover, in the description of the present application, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0038] With the rapid increase in the scale of data centers and the number of servers, the importance of server operation and maintenance has leaped from technical support to a strategic cornerstone for the stable operation of the digital economy. The high-density deployment of a large number of devices and the requirements for business continuity pose many challenges to operation and maintenance.
[0039] In the related art, the management and maintenance of servers are usually carried out by manual inspection. However, this manual inspection method is not only inefficient, but also prone to human errors, and it is impossible to obtain the detailed status information of the servers in a timely and comprehensive manner, making it difficult to meet the management requirements of large-scale data centers. For example, in a data center with thousands of servers, a manual inspection may take several days, and some potential problems may be overlooked during the inspection process.
[0040] To solve the above technical problems, the inventors of the present application found that the function of the Baseboard Management Controller (BMC) of the server for monitoring the server hardware status can be fully utilized to achieve automatic monitoring. Considering that each server has its own BMC and the servers are relatively scattered, and also considering that although the status information of the servers can be obtained through a remote management platform for alarm in the related art, when the network is unstable or the platform fails, the timeliness and effectiveness of management will be affected. For example, when the network has a delay or is interrupted, the remote management platform may not be able to obtain the server status information in a timely manner. At the same time, data security cannot be guaranteed. Therefore, the inventors found that an exclusive operation and maintenance robot can be configured for a server cluster or data center deploying multiple servers. The robot establishes a communication link with the BMC of each server, obtains the operation data of each server, analyzes and obtains the fault conditions of each server, and then based on the fault conditions, returns an operation instruction to the BMC of the server that needs to be maintained, so as to execute the operation instruction through the BMC to quickly eliminate the fault, thereby realizing the in-depth control and management of the server cluster. Since the operation and maintenance robot can communicate with multiple servers at close range, there is no need to worry about problems such as network interruption, and data security can also be protected. Based on this, the embodiments of the present application provide an operation and maintenance robot, method, device, server, system and storage medium.
[0041] In order to enable those skilled in the art of the present technology to better understand the solution of the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0042] Combined with the specific application environment architecture or specific hardware architecture on which the execution of the operation and maintenance method depends, the specific application environment architecture or specific hardware architecture is described herein. Refer to Figure 1 , Figure 1 is the structural schematic diagram of the server management system provided by the embodiment of the present application Figure 1 . As Figure 1 shown, the system includes an operation and maintenance robot 101 and multiple servers 102. The operation and maintenance robot 101 is communicatively connected to the multiple servers 102. The communication between the operation and maintenance robot 101 and the servers 102 can be wired communication or wireless communication, and this embodiment does not make a limitation on this.
[0043] In the specific implementation process, multiple servers 102 are arranged in the computer room, and the operation and maintenance robot 101 moves within the scope of the computer room to manage and maintain the multiple servers 102 in the computer room. The operation and maintenance robot 101 obtains the operation data sent by the multiple servers 102, analyzes the operation data to obtain an analysis result, where the analysis result includes the fault information of multiple target servers 102 among the multiple servers 102, and generates a first operation instruction corresponding to the first server 102 among the multiple target servers 102 according to the analysis result. The first operation instruction is used to instruct the first server 102 to execute the corresponding first operation. The server management system provided in this embodiment can, by collecting and intelligently analyzing the operation data of multiple servers 102 in real time, automatically generate accurate operation instructions and execute them in a closed loop. First, it can improve efficiency, greatly compress the mean time to recovery, avoid service interruption, and second, it can reduce costs. A single operation and maintenance robot 101 can manage thousands of servers 102, reducing the manpower dependence by 70%. It can also enhance reliability and avoid human errors, providing a highly flexible and low-cost unmanned operation and maintenance solution for scenarios such as cloud data centers and edge computing.
[0044] Figure 2 The structural schematic of the operation and maintenance robot provided in the embodiment of the present application Figure 1 , as Figure 2 shown, the embodiment of the present application provides an operation and maintenance robot 101, and a detailed description of the operation and maintenance robot 101 is as follows:
[0045] The operation and maintenance robot 101 includes: a first communication module 201 and a management platform module 202.
[0046] The first communication module 201 is connected to the management platform module 202 and is used to receive the operation data sent by multiple servers and send the operation data to the management platform module 202.
[0047] The management platform module 202 is used to analyze the operation data to obtain an analysis result; the analysis result includes the fault information of multiple target servers among the multiple servers; generate a first operation instruction corresponding to the first server among the multiple target servers according to the analysis result, and send the first operation instruction to the first communication module 201.
[0048] The first communication module 201 is further used to send the first operation instruction to the first server; the first operation instruction is used to instruct the first server to execute the corresponding first operation.
[0049] In this embodiment, the first communication module 201 can be a wired communication module, such as an Ethernet interface, or a wireless communication module, such as a Wireless Fidelity (Wi-Fi) module.
[0050] In some embodiments, by setting up the first communication module 201, communication can be conducted with the BMCs of each server. The standard Intelligent Platform Management Interface (IPMI) protocol or Secure Shell (SSH) protocol can be adopted to collect relevant data. In practical applications, for the convenience of the operation and maintenance robot 101 to move, wireless communication can be preferentially adopted. If it is detected that the communication quality of the wireless communication is lower than a preset value, then wired communication can be switched. When the length of the network cable is insufficient, the operation and maintenance robot 101 can approach the server and insert the network cable to conduct wired communication to ensure the stability of data transmission.
[0051] Among them, the operation data of the server can include data such as hardware status information, component health status, and system operation logs. Among them, the hardware status information can include the real-time temperatures of key components such as the Central Processing Unit (CPU), memory, hard disk, and motherboard, the voltage values of each power supply module of the motherboard, fan speed, fault alarm, and speed regulation strategy, power input / output voltage, current, power, redundancy mode, etc. The component health status can include the error register status, the health degree, the number of bad sectors, and the temperature of the hard disk, and the power supply and link status of the expansion card. The system operation logs can include records of hardware failures (such as overheating, voltage overlimit), power-on / power-off events, firmware errors, etc.
[0052] The fault information can include information on the actual faults that occur in the corresponding server (such as fan anomalies, overheating, etc.), and can also include fault warning information (for example, it can be fault warning information obtained based on the analysis and prediction of historical operation data, or it can be fault warning information determined based on the deviation degree obtained by horizontally comparing the operation data of multiple servers).
[0053] In this embodiment, the first operation instruction can be a restart module instruction, a fan speed adjustment instruction, a standby device switching instruction, etc.
[0054] During the specific operation process, the operation and maintenance robot 101 establishes a communication connection with each server through the first communication module 201. The BMC of each server sends the collected operation data to the operation and maintenance robot 101 through the second communication module of the server. The operation and maintenance robot 101 receives the operation data of each server through the first communication module 201, and jointly analyzes the operation data of each server to obtain the fault information of each server. Furthermore, a corresponding first operation instruction is generated according to the fault information of each server, and the first operation instruction is sent to the BMC of the corresponding server through the first communication module 201. After receiving the first operation instruction, the BMC of the server executes the corresponding first operation according to the first operation instruction, such as adjusting the fan speed, etc.
[0055] As can be seen from the above description, the operation and maintenance robot provided in this embodiment can improve efficiency by collecting and intelligently analyzing the operation data of multiple servers in real time, automatically generating accurate operation instructions and executing them in a closed loop. Firstly, it can greatly compress the mean time to repair and avoid service interruption. Secondly, it can reduce costs. A single operation and maintenance robot can manage thousands of servers, reducing the manpower dependence by 70%. It can also enhance reliability and avoid human errors, providing a highly elastic and low-cost unmanned operation and maintenance solution for scenarios such as cloud data centers and edge computing.
[0056] In some embodiments, in order to facilitate moving to the positions of each server for close observation and detection of the server or performing contact operations such as plugging and unplugging hard disks, the operation and maintenance robot may further include a motion module and a control module. The management platform module, connected to the control module, is further configured to determine the processing priorities of multiple target servers according to the analysis results and send the processing priorities to the control module. The control module, connected to the motion module, is configured to perform path planning according to the processing priorities to obtain a first route. The motion module is configured to sequentially move to the target positions corresponding to multiple target servers according to the first route. The operation and maintenance robot provided in this embodiment can prioritize the processing of servers with more serious or urgent fault situations by determining the processing priorities and performing path planning based on the processing priorities, thereby improving the reliability of operation and maintenance and enhancing the operation and maintenance efficiency.
[0057] Among them, the control module can use a high-performance embedded processor to run a real-time operating system and is responsible for the overall control and task scheduling of the robot. Taking the control module based on the Advanced RISC Machine (ARM) architecture as an example, this control module can efficiently process sensor data, communication instructions, and the action control of the execution module to ensure the stable operation of each function of the robot.
[0058] In this embodiment, the motion module may include a mobile chassis. The mobile chassis can adopt a wheeled or caterpillar structure, and is equipped with a motor drive system and a navigation system, capable of realizing autonomous navigation and obstacle avoidance functions, enabling the robot to freely move to various server positions within the data center. For example, in a large data center where the ground is relatively flat, a wheeled mobile chassis can be selected and paired with a laser navigation system. The robot can avoid obstacles such as other devices and cables and move autonomously. Each server can be set with a Radio Frequency Identification (RFID) tag to uniquely identify each server and assist the operation and maintenance robot in locating the server.
[0059] To achieve precise path planning, after the computer room is built, the positions of each server can be recorded to generate a three-dimensional image of the computer room, and a high-precision map can be generated based on this three-dimensional image and imported into the operation and maintenance robot so that the operation and maintenance robot can plan routes based on this map.
[0060] In the specific implementation process, during the inspection process, the operation and maintenance robot can adopt a strategy of periodic inspection plus emergency inspection. In the periodic inspection, the operation and maintenance management of each server can be carried out according to a preset route. For example, taking a 12-hour cycle to collect the status of the indicator lights of each server at the position of each server. During the 12-hour interval, emergency inspections can be carried out. The operation and maintenance robot analyzes the real-time operation data sent by each server. If the analysis results show that multiple target servers have failed, the multiple target servers can be prioritized according to the severity and urgency of the failures of the multiple target servers, and the target servers with higher severity and urgency can be processed first to avoid irreparable losses. For example, if servers 1, 4, and 5 have failed, and the severity of the failures of these three servers is ranked from high to low as server 4, 1, and 5. Then the first route can be planned to move to the location of server 4 first, then move to the location of server 1, and then move to the location of server 5.
[0061] In some embodiments, in order for the operation and maintenance robot to perform contact operations such as hard disk plugging and unplugging on the server, the operation and maintenance robot may further include an execution module; a control module, connected to the execution module, and further configured to generate a second operation instruction according to the fault information corresponding to the second server after moving to the target position corresponding to the second server among multiple target servers, and send the second operation instruction to the execution module; the execution module is configured to perform a corresponding second operation on the second server according to the second operation instruction. The operation and maintenance robot of this embodiment realizes a closed-loop operation and maintenance process from fault identification to physical repair by integrating the execution module and the intelligent control logic, significantly improving the automation level. After the operation and maintenance robot accurately locates the target server, it can independently generate an adaptive operation instruction based on the fault information parsed in real time, and drive the execution module to complete precise contact operations such as hard disk plugging and unplugging, completely replacing traditional manual intervention. This mechanism not only avoids the risk of hardware damage caused by human misoperation, but also ensures the standardization and consistency of operations in complex environments, especially suitable for scenarios where it is difficult for humans to intervene, such as high-density cabinets or narrow spaces on the edge side, greatly enhancing the system's self-healing ability and operation and maintenance continuity.
[0062] Among them, the execution module may mainly include a robotic arm and a dexterous hand. The robotic arm and the dexterous hand adopt a multi-joint structure, with high flexibility and precision, and can complete complex physical operations such as plugging and unplugging hardware devices. For example, when the hard disk of the server fails, the robotic arm can accurately pull out the faulty hard disk and insert a new one.
[0063] In some embodiments, in order to facilitate the collection of data such as the environmental information of each server, the operation and maintenance robot may further include a sensor module; the sensor module is connected to the management platform module and is configured to send the collected sensing data to the management platform module; the sensing data includes the environmental information and physical states of multiple servers; the management platform module is configured to parse the sensing data and operation data to obtain an analysis result. The environmental information may include the temperature, humidity, and sound of the environment. The physical state may include the state of the indicator lights of the server. The operation and maintenance robot provided in this embodiment realizes a comprehensive perception of the server environment and physical state by integrating the sensor module, and combines the multi-dimensional parsing of operation data, significantly improving the operation and maintenance accuracy and proactive defense ability. The sensor collects the temperature, humidity, acoustic characteristics, and the state of the device indicator lights in real time, and collaborates with the management platform module for analysis, which can identify abnormal heat dissipation in the computer room, abnormal hardware noise, or component failures (such as abnormal flashing of the hard disk indicator light) in advance, and then trigger an adaptive adjustment strategy (such as starting a standby cooling system) or accurately locate the fault source. This solution breaks through the limitations of traditional single-dimensional operation and maintenance, especially suitable for data centers with high-density deployment or complex environments, effectively preventing chain failures caused by environmental factors, and enhancing the system's robustness.
[0064] In this embodiment, the sensor module can be installed at different positions of the robot. The sensor module can include a vision sensor, a temperature sensor, a humidity sensor, a sound sensor, etc., and can collect the environmental information and physical state of the server in all directions. For example, the vision sensor can be installed at the front end of the robot to identify the indicator light status, label information, etc. of the server; the temperature sensor can be installed near the server heat dissipation port to accurately monitor the temperature of the server; the sound sensor can detect abnormal fan speeds.
[0065] In some embodiments, in order to improve the parsing ability of the operation and maintenance robot, the operation and maintenance robot can further include an artificial intelligence algorithm module; a management platform module, connected to the artificial intelligence algorithm module, for sending operation data to the artificial intelligence algorithm module; an artificial intelligence algorithm module, for parsing the input operation data and outputting an analysis result, and sending the analysis result to the management platform module. The operation and maintenance robot provided in this embodiment, by introducing the artificial intelligence algorithm module, endows the operation and maintenance robot with deep learning and complex pattern recognition capabilities, enabling it to mine potential correlations from a large amount of heterogeneous operation data (such as performance metrics, log events, environmental sensing data), and accurately identify hidden faults that cannot be captured by traditional threshold methods (such as hardware degradation indicated by microsecond-level latency fluctuations). Through continuous training and optimization, the artificial intelligence model can adapt to the dynamic characteristics of different server clusters, improve the accuracy of fault root cause location to more than 95%, predict potential risks (such as the hard disk life is about to expire) and generate maintenance strategies in advance, upgrade the operation and maintenance mode from "post-event response" to "pre-event intervention", significantly reduce the probability of service interruption, and is especially suitable for intelligent autonomous operation and maintenance in ultra-large-scale data centers and hybrid cloud environments.
[0066] Among them, the artificial intelligence (AI) algorithm module can include a variety of algorithms. For example, it can include some basic image processing / audio processing algorithms. The image information collected by the vision sensor can be used to assist the robot in positioning. Furthermore, different entities can be distinguished through vision processing algorithms, and the specific server location can be located. The temperature data collected by the temperature sensor can be input into the AI algorithm to identify high-temperature locations, and the sound data collected by the sound sensor can be input into the AI algorithm to identify scenarios such as abnormal fan operation. In addition, it can also include AI warning algorithms, fault location algorithms, automated operation and maintenance task execution algorithms, heat dissipation control algorithms, etc.
[0067] Exemplarily, such as Figure 3As shown in the figure, the operation and maintenance robot includes a communication module, a management platform module, and an algorithm module. In the specific implementation process, the operation and maintenance robot obtains the server operation data from the BMC of the server through the communication module for raw data transmission. The communication module then performs local data transmission to send the operation data to the management platform module. The management platform module further performs local data transmission to send the operation data to the AI algorithm module. The AI algorithm module analyzes the operation data of each server to obtain the analysis result and returns the analysis result to the management platform module. The management platform module generates an operation instruction based on the analysis result and sends the operation instruction to the BMC of the server through the communication module. The BMC of the server executes the corresponding operation based on the operation instruction. The management platform module can also perform interface display in response to the user's query instruction, facilitating the user to understand the operation data and fault information of each server, and can receive the touch operation of the user on the display page to control the corresponding server, realizing the interaction between the robot and the user.
[0068] The embodiment of the present application also provides a server, including: a management controller and a second communication module; the management controller is connected to the second communication module and is used to obtain the operation data of the server and send the operation data to the second communication module; the second communication module is used to send the operation data to the operation and maintenance robot; the second communication module is also used to receive the first operation instruction sent by the operation and maintenance robot; the management controller is also used to execute the corresponding first operation according to the first operation instruction.
[0069] In this embodiment, the second communication module can be a wired communication module, such as an Ethernet interface, or a wireless communication module, such as a Wi-Fi module. The operation data of the server can include data such as hardware status information, component health status, and system operation logs. Among them, the hardware status information can include the real-time temperatures of key components such as the Central Processing Unit (CPU), memory, hard disk, and motherboard, the voltage values of each power supply module of the motherboard, the fan speed, fault alarm, and speed regulation strategy, the power input / output voltage, current, power, and redundancy mode. The component health status can include the error register status, the health degree, the number of bad sectors, and the temperature of the hard disk, and the power supply and link status of the expansion card. The system operation logs can include records of hardware failures (such as overheating, voltage overlimit), power-on / power-off events, firmware errors, etc. The first operation instruction can be a restart module instruction, a fan speed adjustment instruction, a standby device switching instruction, etc. In addition, each server can be set with a Radio Frequency Identification (RFID) tag to uniquely identify each server and assist the operation and maintenance robot in positioning the server.
[0070] As can be seen from the above description, the server provided in this embodiment can automatically generate accurate operation instructions and execute them in a closed loop by collecting and intelligently analyzing the operation data of multiple servers in real time. First, it can improve efficiency, greatly compress the mean time to recovery, and avoid service interruption. Second, it can reduce costs. A single operation and maintenance robot can manage thousands of servers, reducing the manpower dependence by 70%. It can also enhance reliability and avoid human errors, providing a highly elastic and low-cost unmanned operation and maintenance solution for scenarios such as cloud data centers and edge computing.
[0071] An embodiment of the present application also provides a server management system, including: an operation and maintenance robot as described in any of the above embodiments and multiple servers as described in the above embodiments.
[0072] Multiple servers send operation data to the operation and maintenance robot; the operation and maintenance robot receives the operation data sent by the multiple servers, analyzes the operation data, and obtains an analysis result; the analysis result includes the fault information of multiple target servers among the multiple servers; a first operation instruction corresponding to the first server among the multiple target servers is generated according to the analysis result, and the first operation instruction is sent to the first server; the first server receives the first operation instruction sent by the operation and maintenance robot and executes the corresponding first operation according to the first operation instruction. Among them, the first server can be one or more servers. For example, the operation and maintenance robot analyzes the operation data of multiple servers to determine the analysis result. The analysis result shows that three target servers, such as Server 1, Server 2, and Server 3, all have faults. Then, the three target servers can be used as the first server, and operation instructions corresponding to Server 1, Server 2, and Server 3 are generated respectively according to the analysis result. Since the fault types are different, some faults may not be solved by generating operation instructions. Then, operation instructions can also be generated only for some of the servers among the three target servers that can solve the faults by generating operation instructions.
[0073] Exemplarily, such as Figure 4As shown, the server cluster in the server management system includes one or more servers, and each server has a built-in BMC. The BMC can collect and monitor the running data of the server through various hardware communication methods such as the Inter-Integrated Circuit (I2C), for example, hardware status information such as CPU temperature and memory usage rate, and communicate with the operation and maintenance robot through the network interface to send the running data to the operation and maintenance robot. The operation and maintenance robot may include at least one of the following modules: motion module, control module, communication module, sensor module, execution module, management platform module. The operation and maintenance method provided in this embodiment has at least the following beneficial effects: improving server management efficiency: the operation and maintenance robot can automatically complete the server inspection task, and the management platform module of the operation and maintenance robot can process data and issue instructions in real time, greatly shortening the inspection and response time, improving the server management efficiency, and reducing the labor input cost. Enhancing the monitoring ability: through communication with the BMC, the management platform module can obtain the detailed hardware status information of the server in real time and discover potential fault hazards in time. For example, it can detect the read and write errors of the server hard disk in advance to avoid data loss; it can perform local real-time fault diagnosis or early warning based on historical data. Reducing network dependence: by setting up a management platform inside the robot, the problem of management interruption caused by unstable network or external remote management platform failure can be avoided, improving the reliability of the system. Even in the case of a short network interruption, the robot can still continue to complete the inspection task and upload the data after the network is restored. Implementing multiple controls: the management platform module can remotely control and physically operate the server according to the actual situation, reducing manual intervention and improving the response speed. When an emergency fault occurs in the server, measures can be taken quickly to handle it. Providing decision support: the management platform module analyzes and processes the collected server running data, provides a scientific decision-making basis for the management and maintenance of the server, and optimizes the running performance of the server. By analyzing the historical data of the server, the maintenance plan and hardware upgrade time of the server can be reasonably arranged.
[0074] In some embodiments, to ensure the management security of the server, a unified operation and maintenance account dedicated to intelligent inspection can be created in the BMC of each server, and a firewall can be configured in the BMC to only allow devices in the whitelist (i.e., specific operation and maintenance robots) to access and operate.
[0075] As can be seen from the above description, the server management system provided in this embodiment can collect and intelligently analyze the operation data of multiple servers in real time, automatically generate accurate operation instructions and execute them in a closed loop. First, it can improve efficiency, greatly compress the mean time to recovery, and avoid service interruption. Second, it can reduce costs. A single operation and maintenance robot can manage thousands of servers, reducing the manpower dependence by 70%. It can also enhance reliability and avoid human errors, providing a highly elastic and low-cost unmanned operation and maintenance solution for scenarios such as cloud data centers and edge computing.
[0076] Figure 5 It is a schematic flowchart of the operation and maintenance method provided by an embodiment of the present application. As Figure 5 shown, an embodiment of the present application provides an operation and maintenance method, and the method will be described in detail as follows:
[0077] 501. Obtain the operation data sent by multiple servers.
[0078] Specifically, each server can send the server operation data to the operation and maintenance robot in real time through its respective second communication module, and the operation and maintenance robot receives the operation data sent by each server in real time through the first communication module.
[0079] In this embodiment, the operation data of the server may include data such as hardware status information, component health status, and system operation logs. Among them, the hardware status information may include the real-time temperature of key components such as the Central Processing Unit (CPU), memory, hard disk, and motherboard, the voltage values of each power supply module of the motherboard, the fan speed, fault alarm, and speed regulation strategy, the power input / output voltage, current, power, redundancy mode, etc. The component health status may include the error register status, the health degree, the number of bad sectors, and the temperature of the hard disk, and the power supply and link status of the expansion card. The system operation logs may include records of hardware failures (such as overheating, voltage overlimit), power-on / power-off events, firmware errors, etc.
[0080] 502. Analyze the operation data to obtain an analysis result; the analysis result includes the fault information of multiple target servers among the multiple servers.
[0081] Specifically, after obtaining the operation data of each server, the operation data can be analyzed to obtain an analysis result. The analysis method can be to analyze and obtain the analysis result through statistical algorithms, or the operation data can be input into the AI algorithm model to output the analysis result.
[0082] In this embodiment, the fault information may include information about the actual faults that occur in the corresponding server (such as abnormal fan, over - high temperature, etc.), and may also include fault warning information (for example, it may be fault warning information obtained based on the analysis and prediction of historical operation data, or it may be fault warning information determined based on the deviation degree obtained by horizontally comparing the operation data of multiple servers).
[0083] In some embodiments, in order to improve the analysis accuracy, the operation data of multiple servers may be comprehensively considered to obtain an analysis result. Parsing the operation data to obtain the analysis result may include: parsing the operation data to determine the relative deviation degrees of multiple servers on multiple status items, and determining the analysis result according to the relative deviation degrees.
[0084] Specifically, obtain the real - time data of multiple status items of multiple servers and construct a population benchmark data set. Calculate the relative deviation degrees of the operation data of each server from the population benchmark through a statistical model (such as the percentile method or clustering analysis). Set a dynamic threshold, and mark the servers that continuously deviate from the dynamic threshold within a preset time period as abnormal to obtain the analysis result. Then, in the subsequent patrol inspection process, the processing priority of the servers with abnormalities can be increased, and they can be processed preferentially or a continuous attention period can be set, during which continuous key attention is paid. The operation and maintenance method provided in this embodiment can identify relative abnormalities missed by the traditional single - node threshold method (such as the temperature of a certain server is not higher than the safety value, but is 20% higher than that of similar devices for a long time) through the analysis of relative deviation degrees, and timely discover hidden faults. If the nodes at the same position of multiple servers all deviate from the population benchmark, it may indicate systematic risks such as uneven heat dissipation in the computer room or hardware batch defects, which can assist in locating the root cause. Exemplarily, taking the CPU temperature of the server as the status item, the CPU temperatures of each server may be less than the alarm threshold, but there is an individual server whose CPU temperature continuously exceeds the CPU temperatures of other servers for a preset time period. Then, this individual server can be set as abnormal and key - focused on in the subsequent operation and maintenance process, and its processing priority can be increased.
[0085] 503. Generate a first operation instruction corresponding to the first server among multiple target servers according to the analysis result; the first operation instruction is used to instruct the first server to perform the corresponding first operation.
[0086] Specifically, after obtaining the fault information of each server, corresponding first operation instructions can be generated according to the fault information of each server, and the first operation instructions are sent to the BMC of the corresponding server through the first communication module. After receiving the first operation instruction, the BMC of the server executes the corresponding first operation according to the first operation instruction, such as adjusting the fan speed, etc.
[0087] As can be seen from the above description, the operation and maintenance method provided in this embodiment can automatically generate accurate operation instructions and execute them in a closed loop by collecting and intelligently analyzing the operation data of multiple servers in real time. First, it can improve efficiency, greatly compress the mean time to recovery, and avoid service interruption. Second, it can reduce costs. A single operation and maintenance robot can manage thousands of servers, reducing the manpower dependence by 70%. It can also enhance reliability and avoid human errors, providing a highly elastic and low-cost unmanned operation and maintenance solution for scenarios such as cloud data centers and edge computing.
[0088] In some embodiments, to make reasonable plans, a processing order can be determined for each server. The method may further include: determining the processing priorities of multiple target servers according to the analysis results; the processing priorities are used to indicate the order of moving to the target positions corresponding to the multiple target servers in sequence; after moving to the target position corresponding to the second server among the multiple target servers, generating a second operation instruction according to the fault information corresponding to the second server; the second operation instruction is used to indicate performing a target operation on the second server. The operation and maintenance method provided in this example realizes efficient and accurate operation and maintenance through dynamic priority sorting and intelligent path planning, automatically assigns the processing order based on the fault severity and business impact, ensures that critical servers are repaired first, and minimizes the service interruption time to the greatest extent. After the operation and maintenance robot moves to the target position according to the optimized route, it autonomously generates an adapted operation instruction and executes the repair action, eliminating the manual decision-making delay and path redundancy, and significantly improving the fault handling efficiency. This mechanism is especially applicable to scenarios of high-density deployment or concurrent emergency faults, ensuring the optimization of resource scheduling, while reducing the safety risks of manual operations in complex environments, and providing stable and coherent automated operation and maintenance guarantee for large-scale server clusters.
[0089] In some embodiments, to closely monitor the environmental information of the server and improve the accuracy of fault analysis, parsing the operation data to obtain the analysis result may include: acquiring sensing data; the sensing data includes the environmental information and physical states of multiple servers; parsing the sensing data and the operation data to obtain the analysis result. The operation and maintenance method provided in this embodiment significantly improves the accuracy and forward-looking of fault diagnosis through the multi-dimensional analysis of fusing environmental sensing data and server operation data. Abnormal environmental temperature and humidity (such as local overheating or high humidity) can give early warnings of the failure of the cooling system or the risk of hardware corrosion. Combining server performance indicators (such as CPU load) and physical states (such as abnormal hard disk indicator lights), accurately locate the chain faults induced by environmental factors (such as high temperature causing a slowdown in disk input / output (I / O)), and avoid misjudgment caused by a single data dimension. At the same time, by establishing an environment-operation state association model, potential risks (such as accelerated aging of regional equipment caused by the imbalance of cold and hot channels in the computer room) can be predicted, guiding the dynamic adjustment of the cooling strategy or the migration of critical loads, and realizing the upgrade of the operation and maintenance mode from passive repair to active protection, which is especially suitable for reliability guarantee in high-density deployment or harsh edge environments.
[0090] In this embodiment, the sensing data may include image data containing indicator light states collected by a visual sensor, environmental temperature data collected by a temperature sensor, environmental humidity data collected by a humidity sensor, and server fan sound data collected by a sound sensor.
[0091] In some embodiments, to improve the parsing ability of the operation and maintenance robot, parsing the operation data to obtain the analysis result may include: inputting the operation data into a preset artificial intelligence model for parsing and outputting the analysis result. The operation and maintenance method provided in this embodiment deeply analyzes the operation data through the artificial intelligence model, endowing the operation and maintenance robot with complex pattern recognition capabilities beyond traditional rules, enabling it to capture implicit associations from a large amount of heterogeneous information (such as the coupling law between performance fluctuations and log events), dynamically learn the characteristics of different server clusters and adaptively optimize the analysis strategy, significantly improving the accuracy of fault tracing. At the same time, the model predicts the hardware degradation trend based on time-series data, triggers preventive maintenance instructions in advance, and promotes the transformation of the operation and maintenance mode from passive response to predictive intervention, effectively avoiding potential business interruptions. This technology is especially suitable for the intelligent management requirements of ultra-large-scale heterogeneous clusters, providing continuous evolving autonomous operation and maintenance capabilities for data centers in high-dynamic environments.
[0092] Among them, the artificial intelligence model can include various algorithm models. For example, it can include some basic algorithms for image processing / audio processing. The image information collected by the visual sensor can be used to assist the robot in positioning. Furthermore, different entities can be distinguished through visual processing algorithms, and the specific server location can be located. The temperature data collected by the temperature sensor can be input into the AI model to identify high-temperature locations, and the sound data collected by the sound sensor can be input into the AI model to identify scenarios such as abnormal fan operation. In addition, it can also include an AI warning model, a fault location model, an automated operation and maintenance task execution model, a heat dissipation regulation model, etc. In the specific implementation process, the operation data and sensing data can be input into the AI model, and each algorithm model in the AI model can analyze the operation data separately, and then obtain the results corresponding to each algorithm model respectively.
[0093] In some embodiments, to facilitate users to query data such as the operation data, analysis results, and processing priorities of each server, and to facilitate users to control the server through page operations, the method can further include: displaying a target page; the target page includes a data display panel and an operation execution panel; the data display panel includes the operation data and processing priorities of multiple servers; the operation execution panel includes operation options corresponding to multiple servers respectively; in response to a touch operation on the operation option for the third server among the multiple servers, sending a third operation instruction to the third server; the third operation instruction is used to instruct the third server to execute the corresponding third operation. The operation and maintenance method provided in this embodiment realizes the efficient collaboration between users and the server cluster through an integrated visual interaction interface. The data display panel centrally presents the real-time operation status and processing priorities of multiple servers, eliminating the problem of information fragmentation in cross-system queries. The operation execution panel provides fine-grained control options (such as service restart, firmware update). Users can trigger precise operation and maintenance actions with a single touch instruction, greatly reducing the operation threshold and response delay. This design takes into account both the overall situation awareness and the immediate intervention ability, and is especially suitable for scenarios such as multi-team collaboration or emergency fault handling, ensuring the transparency of operation and maintenance decisions and the traceability of operations. At the same time, it supports mobile device adaptation, enabling remote unmanned control, and providing an intuitive and secure interaction center for agile operation and maintenance in complex Information Technology (IT) environments.
[0094] Exemplarily, the CPU temperature of each server can be displayed in the data display panel, and can be displayed in the form of a heat map based on the location information of each server. The asset information of the server cluster, as well as the idle and busy status and load balancing situation of each server, can also be displayed. The operation options in the operation execution panel can include operations such as restarting the module, switching to the standby device, and firmware update.
[0095] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.
[0096] Figure 6 It is a schematic structural diagram of the operation and maintenance device provided by the embodiment of the present application. As Figure 4 shown, the embodiment of the present application also provides an operation and maintenance device, including:
[0097] An acquisition module 601, configured to acquire operation data sent by multiple servers;
[0098] An analysis module 602, configured to analyze the operation data to obtain an analysis result; the analysis result includes fault information of multiple target servers among multiple servers;
[0099] A generation module 603, configured to generate a first operation instruction corresponding to a first server among multiple target servers according to the analysis result; the first operation instruction is used to instruct the first server to execute a corresponding first operation.
[0100] The operation and maintenance device provided by this embodiment can improve efficiency by collecting and intelligently analyzing the operation data of multiple servers in real time, automatically generating accurate operation instructions and executing them in a closed loop, thereby greatly compressing the mean time to recovery, avoiding service interruption, reducing costs, enabling a single operation and maintenance robot to manage thousands of servers, reducing the manpower dependence by 70%, enhancing reliability, avoiding human errors, and providing a highly elastic and low-cost unmanned operation and maintenance solution for scenarios such as cloud data centers and edge computing.
[0101] In some embodiments, the analysis module 602 is specifically configured to: analyze the operation data to determine the relative deviation degrees of multiple servers on multiple status items, and determine the analysis result according to the relative deviation degrees; the generation module 603 is further configured to determine the processing priorities of multiple target servers according to the analysis result; the processing priorities are used to indicate the order of moving to the target positions corresponding to multiple target servers in sequence; after moving to the target position corresponding to a second server among multiple target servers, generate a second operation instruction according to the fault information corresponding to the second server; the second operation instruction is used to instruct to perform a target operation on the second server.
[0102] In some embodiments, the analysis module 602 is specifically configured to: acquire sensing data; the sensing data includes the environmental information and physical states of multiple servers; analyze the sensing data and the operation data to obtain an analysis result.
[0103] In some embodiments, the parsing module 602 is specifically configured to: input the operation data into a preset artificial intelligence model for parsing, and output an analysis result.
[0104] In some embodiments, the device 60 further includes a display module (not shown) configured to: display a target page; the target page includes a data display panel and an operation execution panel; the data display panel includes the operation data and processing priorities of multiple servers; the operation execution panel includes operation options corresponding to the multiple servers respectively; in response to a touch operation on the operation option corresponding to the third server among the multiple servers, send a third operation instruction to the third server; the third operation instruction is used to instruct the third server to execute the corresponding third operation.
[0105] For the descriptions of the features in the corresponding embodiments of the operation and maintenance device, reference can be made to the relevant descriptions of the corresponding embodiments of the operation and maintenance method, which will not be elaborated here one by one.
[0106] Figure 7 This is a schematic structural diagram of the electronic device provided by the present application. As Figure 7 shown, the electronic device 70 provided in this embodiment includes: at least one processor 701 and a memory 702. Optionally, the electronic device 70 further includes a communication component 703. Among them, the processor 701, the memory 702, and the communication component 703 are connected through a bus.
[0107] In a specific implementation process, at least one processor 701 executes the computer execution instructions stored in the memory 702, so that at least one processor 701 executes the above-mentioned operation and maintenance method embodiments.
[0108] For the specific implementation process of the processor 701, reference can be made to the above method embodiments. The implementation principle and technical effects are similar, and will not be elaborated here in this embodiment.
[0109] In the above embodiments, it should be understood that the processor may be a central processing unit (Central Processing Unit, abbreviated as: CPU), or other general-purpose processors, digital signal processors (Digital Signal Processor, abbreviated as: DSP), application specific integrated circuits (Application Specific Integrated Circuit, abbreviated as: ASIC), etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the application can be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor.
[0110] The memory may include a random access memory (RAM), and may also include non-volatile memory (NVM), such as at least one disk memory.
[0111] The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the buses in the drawings of this application are not limited to only one bus or one type of bus.
[0112] Embodiments of this application also provide a computer-readable storage medium storing a computer program, where the computer program is configured to execute the steps in any of the above-described embodiments of the operation and maintenance method when running.
[0113] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disc that can store a computer program.
[0114] Embodiments of this application also provide a computer program product. The above computer program product includes a computer program, and when the computer program is executed by a processor, the steps in any of the above-described embodiments of the operation and maintenance method are implemented.
[0115] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps in any of the above-described embodiments of the operation and maintenance method are implemented.
[0116] Those skilled in the art may further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0117] The above has introduced in detail an operation and maintenance robot, method, device, server, system, and storage medium provided by this application. Specific examples are used herein to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the protection scope of the claims of this application.
Claims
1. An operation and maintenance robot, characterized in that: include: A first communication module and a management platform module; The first communication module is connected to the management platform module and is used to receive operation data sent by multiple servers and send the operation data to the management platform module; The management platform module is used to parse the operation data to obtain analysis results; the analysis results include fault information of multiple target servers among the multiple servers; generate a first operation instruction corresponding to a first server among the multiple target servers according to the analysis results, and send the first operation instruction to the first communication module; The first communication module is further used to send the first operation instruction to the first server; the first operation instruction is used to instruct the first server to perform a corresponding first operation.
2. The operation and maintenance robot according to claim 1, characterized in that: Also includes motion module and control module; The management platform module is connected to the control module and is further used to determine the processing priorities of the plurality of target servers according to the analysis results and send the processing priorities to the control module; The control module is connected to the motion module and is used to perform path planning according to the processing priority to obtain a first route; The movement module is used to move to target locations corresponding to the multiple target servers in sequence according to the first route.
3. The operation and maintenance robot according to claim 2, characterized in that: Also includes an execution module; The control module is connected to the execution module and is further used to generate a second operation instruction according to the fault information corresponding to the second server after moving to the target position corresponding to the second server among the plurality of target servers, and send the second operation instruction to the execution module; The execution module is used to execute the corresponding second operation on the second server according to the second operation instruction.
4. The operation and maintenance robot according to any one of claims 1 to 3, characterized in that: Also included is a sensor module; The sensor module is connected to the management platform module and is used to send the collected sensor data to the management platform module; the sensor data includes the environmental information and physical status of the plurality of servers; The management platform module is used to analyze the sensor data and the operation data to obtain the analysis result.
5. The operation and maintenance robot according to any one of claims 1 to 3, characterized in that: It also includes artificial intelligence algorithm modules; The management platform module is connected to the artificial intelligence algorithm module and is used to send the operation data to the artificial intelligence algorithm module; The artificial intelligence algorithm module is used to parse the input operation data, output analysis results, and send the analysis results to the management platform module.
6. A server, characterized in that: include: A management controller and a second communication module; The management controller is connected to the second communication module and is used to obtain the operation data of the server and send the operation data to the second communication module; The second communication module is used to send the operation data to the operation and maintenance robot; The second communication module is further used to receive a first operation instruction sent by the operation and maintenance robot; The management controller is further configured to execute a corresponding first operation according to the first operation instruction.
7. A server management system, characterized in that: include: Operate robots and multiple servers; The plurality of servers send the operation data to the operation and maintenance robot; The operation and maintenance robot receives operation data sent by multiple servers, analyzes the operation data, and obtains analysis results; the analysis results include fault information of multiple target servers among the multiple servers; generates a first operation instruction corresponding to a first server among the multiple target servers according to the analysis results, and sends the first operation instruction to the first server; The first server receives a first operation instruction sent by the operation and maintenance robot, and performs a corresponding first operation according to the first operation instruction.
8. An operation and maintenance method, characterized in that: include: Get the running data sent by multiple servers; Analyzing the operation data to obtain analysis results; The analysis result includes fault information of multiple target servers among the multiple servers; generating a first operation instruction corresponding to a first server among the plurality of target servers according to the analysis result; The first operation instruction is used to instruct the first server to perform a corresponding first operation.
9. The operation and maintenance method according to claim 8, characterized in that: The method further comprises: Determine the processing priority of the plurality of target servers according to the analysis result; the processing priority is used to indicate the order of sequentially moving to the target positions respectively corresponding to the plurality of target servers; After moving to a target position corresponding to a second server among the multiple target servers, a second operation instruction is generated according to fault information corresponding to the second server; the second operation instruction is used to instruct to perform a target operation on the second server.
10. The operation and maintenance method according to claim 8, characterized in that: The step of parsing the operation data to obtain analysis results includes: Acquire sensor data; the sensor data includes environmental information and physical status of multiple servers; The sensing data and the operating data are analyzed to obtain the analysis result.
11. The operation and maintenance method according to claim 8, characterized in that: The step of parsing the operation data to obtain analysis results includes: The operating data is input into a preset artificial intelligence model for analysis, and the analysis result is output.
12. The operation and maintenance method according to any one of claims 8 to 11, characterized in that: The method further comprises: Displaying a target page; the target page includes a data display panel and an operation execution panel; the data display panel includes the operation data and processing priorities of the plurality of servers; the operation execution panel includes operation options corresponding to the plurality of servers respectively; In response to a touch operation on an operation option of a third server among the multiple servers, a third operation instruction is sent to the third server; the third operation instruction is used to instruct the third server to perform a corresponding third operation.
13. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, used to implement the steps of the operation and maintenance method as claimed in any one of claims 8 to 12 when executing the computer program.
14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the operation and maintenance method according to any one of claims 8 to 12.
15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the operation and maintenance method according to any one of claims 8 to 12 are implemented.
Citation Information
Patent Citations
Artificial intelligence machine room polling robot capable of supporting deep learning working principle
CN108490959A
Inspection tour maintaining robot control method and device, computer device and storage medium
CN109955242A
Machine room operation and maintenance system
CN114415674A
Fault positioning robot assisted regular inspection and patrol scheduling method based on electric power automatic operation and maintenance
CN116165484A
Data center intelligent operation and maintenance management system based on inspection robot
CN116600260A
Cited By
Passive indication system and method for fault component of server and related assembly
CN121434018A