A cabinet server monitoring operation and maintenance system, method and cabinet

By integrating a rack server monitoring and maintenance system into the switch's operating system, server nodes are automatically discovered and configured, and alarm information is monitored and processed in real time. This solves the problem of low efficiency in traditional monitoring methods and achieves efficient operation and maintenance management and stable operation.

CN116708161BActive Publication Date: 2026-05-29JINAN INSPUR DATA TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
JINAN INSPUR DATA TECH CO LTD
Filing Date
2023-06-30
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Traditional rack server monitoring and maintenance requires a lot of manual operation, which is inefficient, costly, and has a high rate of configuration errors.

Method used

By integrating a rack server monitoring and maintenance system into the switch operating system, the asset module monitors the switch port status, automatically discovers server nodes, the control module performs network configuration, the monitoring module performs real-time monitoring, and the alarm module processes the monitoring data to generate standardized alarm objects.

Benefits of technology

It enables automatic discovery, automatic configuration, and real-time monitoring of server nodes, improving operational efficiency, reducing failure rate and operational costs, and ensuring stable server operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116708161B_ABST
    Figure CN116708161B_ABST
Patent Text Reader

Abstract

The application relates to the field of servers, in particular to a cabinet server monitoring operation and maintenance system, a cabinet server monitoring operation and maintenance method and a cabinet. The system is connected to a switch operating system, is automatically started when the system is started and starts a DHCP service. The system comprises an asset module which is used for monitoring the port state of switches in a cabinet to determine whether a server node is connected, obtaining server configuration information when the server node is found and sending a node addition message; a control module which is used for performing network configuration on the connected server node after receiving the node addition message; a monitoring module which is used for monitoring the running state of the connected server node to generate a standardized alarm object and sending the standardized alarm object by using the configured network; and an alarm module which is used for receiving the standardized alarm object and performing alarm generation, upgrading, downgrading, merging and clearing operations on the standardized alarm object to generate alarm messages of each server node. The scheme of the application realizes automatic discovery and monitoring operation and maintenance of cabinet server nodes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of server technology, and in particular to a rack server monitoring and maintenance system, method, and rack. Background Technology

[0002] With the development of the Internet and big data, the demand for rack-mount products is gradually increasing. For traditional server monitoring and maintenance of such rack-mount products, manual configuration of the Baseboard Management Controller (BMC) and other configuration operations is usually required for each server before racking.

[0003] Currently, the traditional method of maintaining a server rack mainly involves manually adding server nodes to the monitoring platform. However, this method requires a lot of manual intervention, which is time-consuming, labor-intensive, inefficient, and has high maintenance costs, so it urgently needs to be improved. Summary of the Invention

[0004] In view of this, it is necessary to provide a server rack monitoring and maintenance system, method, and server rack to address the above technical issues.

[0005] According to a first aspect of the present invention, a rack server monitoring and maintenance system is provided, wherein the system is connected to the operating system of a switch, starts automatically upon power-on, and enables DHCP service; the system includes:

[0006] The asset module is used to monitor the port status of the switches in the cabinet to determine whether a server node is connected. When a server node is found to be connected, the module obtains the server configuration information and sends a node addition message.

[0007] The control module is used to perform network configuration on the access server node after receiving the node add message;

[0008] The monitoring module is used to monitor the operating status of the access server nodes to generate standardized alarm objects, and send the standardized alarm objects using the configured network;

[0009] The alarm module is used to receive the standardized alarm objects and perform alarm generation, escalation, downgrading, merging, and clearing operations on the standardized alarm objects to generate alarm messages for each server node.

[0010] In some embodiments, the asset module is further configured to:

[0011] Real-time monitoring of switch port status changes;

[0012] When the status of a port changes from DOWN to UP, the MAC address of the server on the other end of that port is obtained by querying the switch's MAC address table.

[0013] The server IP address is obtained by querying the switch's ARP table using the server's MAC address and finding the associated IP address.

[0014] The server node is added to the unified monitoring platform by obtaining the server manufacturer, model and serial number information using the factory default protocol configuration via the IPMI protocol.

[0015] After the addition is completed, a new message is sent to the control module to execute the corresponding automatic configuration. At the same time, the asset module's scheduled task starts to collect server asset data information periodically.

[0016] In some embodiments, the control module performs the following operations after receiving a new message:

[0017] Change the network mode of the BMC management IP from DHCP mode to static mode;

[0018] Configure the Trap alarm message, specifying the target address, port, and protocol information for sending Trap alarms, so that the monitoring module can listen for and receive alarm information sent by the server.

[0019] In some embodiments, the control module is also used to provide operations related to server firmware upgrades, operating system deployment, and network configuration.

[0020] In some embodiments, the monitoring module includes an active monitoring module and a passive monitoring module;

[0021] The active monitoring module is used to actively collect the server's operating status and encapsulate it into a standardized alarm object, which is then sent to the alarm module.

[0022] The passive monitoring module is used to receive the operating status reported by the server and encapsulate it into a standardized alarm object and send it to the alarm module.

[0023] In some embodiments, the active monitoring module has a built-in timed data acquisition task, which is responsible for periodically collecting sensor data from server nodes, analyzing the status values ​​of each sensor based on the collected data, and encapsulating abnormal information into standardized alarm objects and sending them to the alarm module for processing.

[0024] In some embodiments, the passive monitoring module is used to listen to the Trap messages sent by each server node in real time, and after parsing and processing the Trap messages, encapsulate them into standardized alarm objects and send them to the alarm module for processing.

[0025] In some embodiments, the alarm module performs the following operations when it receives a standardized alarm object sent by the active monitoring module and / or the passive monitoring module:

[0026] Determine if it is a newly generated alarm; if so, generate a new alarm message.

[0027] If the platform already has the corresponding alarm information, compare the alarm level of the alarm information with the level of the existing alarm information, and then perform the alarm upgrade or downgrade operation.

[0028] If the alarm levels are the same, alarm merging will be performed, and the alarm update time will be updated.

[0029] If it is an alarm recovery, the corresponding alarm information will be cleared, the alarm information will be inserted into the historical information table and the clearing time will be set.

[0030] According to a second aspect of the present invention, a rack server monitoring and maintenance method is provided, the method employing the rack server monitoring and maintenance system described above, the method comprising:

[0031] The asset module is used to monitor the port status of the switches in the cabinet to determine whether a server node is connected. When a server node is found to be connected, the server configuration information is obtained and a node addition message is sent.

[0032] After receiving the node add message, the control module performs network configuration on the access server node;

[0033] The monitoring module is used to monitor the operating status of the access server nodes to generate standardized alarm objects, and the standardized alarm objects are sent using the configured network.

[0034] The alarm module receives the standardized alarm objects and performs alarm generation, escalation, downgrading, merging, and clearing operations on the standardized alarm objects to generate alarm messages for each server node.

[0035] According to a third aspect of the present invention, a server rack is also provided, the server rack including the server rack server monitoring and maintenance system described in the foregoing embodiments.

[0036] The aforementioned rack server monitoring and maintenance system automatically discovers and adds server nodes by monitoring the status changes of switch ports. It also automatically configures server nodes through a control module, monitors the operational status of server nodes in real time through a monitoring module, and processes the data monitored by the monitoring module to generate alarm information. This system achieves automatic discovery and monitoring of rack server nodes, effectively improving the efficiency of server device discovery and maintenance, reducing maintenance costs, lowering the failure rate of server nodes, and ensuring the stable operation of the entire rack of server nodes.

[0037] In addition, the present invention also provides a rack server monitoring and maintenance method and a rack including the above rack server monitoring and maintenance system, which can also achieve the above technical effects, and will not be described in detail here. Attached Figure Description

[0038] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other embodiments can be obtained based on these drawings without creative effort.

[0039] Figure 1 This is a schematic diagram of a rack server monitoring and maintenance system according to an embodiment of the present invention;

[0040] Figure 2 This is a schematic diagram of a rack server connection provided in one embodiment of the present invention;

[0041] Figure 3 This is a schematic diagram of the rack server node discovery and addition process provided in one embodiment of the present invention;

[0042] Figure 4 This is a schematic diagram of a rack server monitoring and maintenance method according to an embodiment of the present invention.

[0043] [Explanation of Labels in the Attached Image]

[0044] 100: Rack server monitoring and maintenance system;

[0045] 110: Assets module;

[0046] 120: Control module;

[0047] 130: Monitoring module; 131: Active monitoring module; 132: Passive monitoring module;

[0048] 140: Alarm module;

[0049] 150: Server node. Detailed Implementation

[0050] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be further described in detail below with reference to specific examples and the accompanying drawings.

[0051] It should be noted that all uses of "first" and "second" in the embodiments of the present invention are for the purpose of distinguishing two entities or parameters with the same name but different names. It is clear that "first" and "second" are only for the convenience of expression and should not be construed as limiting the embodiments of the present invention. Subsequent embodiments will not explain this in detail.

[0052] In one embodiment, please refer to Figure 1 As shown, this invention provides a rack server monitoring and maintenance system 100. The system is integrated into the operating system of a switch, automatically starts upon power-on, and enables DHCP service. DHCP (Dynamic Host Configuration Protocol) is a local area network protocol that uses the UDP protocol. It has two main uses: automatically assigning IP addresses to internal networks or network service providers, and providing users or internal network administrators with a means of centrally managing all computers. The system includes:

[0053] The asset module 110 is used to monitor the port status of the switches in the cabinet to determine whether the server node 150 is connected. When the server node 150 is found to be connected, the module obtains the server configuration information and sends a node addition message.

[0054] Control module 120 is used to perform network configuration on access server node 150 after receiving the node add message;

[0055] The monitoring module 130 is used to monitor the operating status of the access server node 150 to generate standardized alarm objects, and send the standardized alarm objects using the configured network.

[0056] The alarm module 140 is used to receive the standardized alarm objects and perform alarm generation, escalation, downgrading, merging, and clearing operations on the standardized alarm objects to generate alarm messages for each server node 150.

[0057] The aforementioned rack server monitoring and maintenance system automatically discovers and adds server nodes by monitoring the status changes of switch ports. It also automatically configures server nodes through a control module, monitors the operational status of server nodes in real time through a monitoring module, and processes the data monitored by the monitoring module to generate alarm information. This system achieves automatic discovery and monitoring of rack server nodes, effectively improving the efficiency of server device discovery and maintenance, reducing maintenance costs, lowering the failure rate of server nodes, and ensuring the stable operation of the entire rack of server nodes.

[0058] In some embodiments, the asset module 110 is further configured to:

[0059] Real-time monitoring of switch port status changes;

[0060] When the status of a port changes from DOWN to UP, the MAC address of the server on the other end of that port is obtained by querying the switch's MAC address table.

[0061] The server IP address is obtained by querying the switch's ARP table using the server's MAC address and finding the associated IP address.

[0062] The server manufacturer, model and serial number information is obtained using the factory default protocol configuration via IPMI, and server node 150 is added to the unified monitoring platform.

[0063] After the addition is completed, a new message is sent to the control module 120 to execute the corresponding automatic configuration. At the same time, the asset module 110's scheduled task starts to collect server asset data information on a regular basis.

[0064] The rack server monitoring and maintenance system in this embodiment uses an asset module to overcome the shortcomings of traditional monitoring methods that require tedious manual configuration and addition. It can dynamically allocate, automatically discover, and automatically add BMC management IP addresses after the server nodes in the rack are powered on without manual operation, effectively improving maintenance efficiency and saving maintenance costs.

[0065] In some embodiments, the control module 120 performs the following operations after receiving a new message:

[0066] Change the network mode of the BMC management IP from DHCP mode to static mode;

[0067] Configure the Trap alarm message, specifying the target address, port, and protocol information for sending Trap alarms, so that the monitoring module 130 can listen for and receive alarm information sent by the server.

[0068] The rack server monitoring and maintenance system in this embodiment realizes automatic configuration of server node 150, avoiding manual configuration, reducing the workload, significantly reducing the configuration error rate, and having high accuracy.

[0069] In some embodiments, the control module 120 is also used to provide operations related to server firmware upgrades, operating system deployment, and network configuration.

[0070] The rack server monitoring and maintenance system in this embodiment also supports remote batch execution of server firmware upgrades, operating system deployments, and network configurations, making server maintenance management more comprehensive, greatly ensuring the stable operation of each server node in the rack, and reducing the probability of failure and the risk of losses caused by it.

[0071] In some embodiments, the monitoring module 130 includes an active monitoring module 131 and a passive monitoring module 132;

[0072] The active monitoring module 131 is used to actively collect the server's operating status and encapsulate it into a standardized alarm object, which is then sent to the alarm module 140.

[0073] The passive monitoring module 132 is used to receive the operating status reported by the server and encapsulate it into a standardized alarm object and send it to the alarm module 140.

[0074] The rack server monitoring and maintenance system in this embodiment achieves non-intrusive real-time monitoring of servers through two monitoring methods. The combination of active and passive monitoring modules greatly ensures the timeliness of obtaining server operating status, helps to avoid the omission of abnormal server information, and significantly improves the stability and security of rack servers.

[0075] In some embodiments, the active monitoring module 131 has a built-in timed acquisition task, which is responsible for periodically acquiring sensor data from the server node 150, analyzing the status values ​​of each sensor based on the acquired data, and encapsulating the abnormal information into standardized alarm objects and sending them to the alarm module 140 for processing.

[0076] In some embodiments, the passive monitoring module 132 is used to listen to the Trap messages sent by each server node 150 in real time, and after parsing and processing the Trap messages, encapsulate them into standardized alarm objects and send them to the alarm module 140 for processing.

[0077] In some embodiments, the alarm module 140 performs the following operations when it receives a standardized alarm object sent by the active monitoring module 131 and / or the passive monitoring module 132:

[0078] Determine if it is a newly generated alarm; if so, generate a new alarm message.

[0079] If the platform already has the corresponding alarm information, compare the alarm level of the alarm information with the level of the existing alarm information, and then perform the alarm upgrade or downgrade operation.

[0080] If the alarm levels are the same, alarm merging will be performed, and the alarm update time will be updated.

[0081] If it is an alarm recovery, the corresponding alarm information will be cleared, the alarm information will be inserted into the historical information table and the clearing time will be set.

[0082] The rack server monitoring and maintenance system in this embodiment analyzes and processes the alarm objects obtained by the monitoring module to identify valid alarm information and performs a level analysis on the alarm information. This allows maintenance personnel to prioritize the handling of higher-level alarm messages and avoid redundant alarm message reminders, thereby improving the efficiency and accuracy of subsequent fault elimination.

[0083] In yet another embodiment, for ease of understanding of the present invention, the following is used... Figure 2 Taking the illustrated application scenario as an example, assuming a rack contains n server nodes 150 (n is an integer greater than or equal to 2), the working principle and specific implementation of the rack server monitoring and maintenance system applied to this scenario will be explained in detail below. Its working principle involves embedding the system within the rack's access switch operating system, automatically starting with the switch upon power-on, and simultaneously enabling DHCP service for dynamic allocation of server BMC management IPs. The system includes an asset module 110, a control module 120, an active monitoring module 131, a passive monitoring module 132, and an alarm module 140. The asset module 110 is used for monitoring switch port status, discovering and adding server nodes 150, and collecting asset data; the control module 120 is used to receive server node 150 addition messages and complete network and Trap automatic configuration, as well as subsequent server operating system deployment and firmware upgrades; the active monitoring module 131 is used for periodic collection and analysis of sensor data from server node 150; the passive monitoring module 132 is used for real-time monitoring and processing of Trap messages pushed by each server node 150; and the alarm module 140 is used to complete alarm generation, escalation, degradation, merging, and clearing. The specific implementation method is as follows:

[0084] Step 1: Install this system in the cabinet and connect it to the switch operating system. It will start automatically when the switch is turned on. After startup, the asset module 110 will enable DHCP service. When the management port of server node 150 is connected to the switch port via network cable and started, it will dynamically assign a management IP address to server node 150BMC via DHCP for subsequent discovery and management.

[0085] Step two, the asset management module discovers and adds server node 150 by monitoring the switch port status, as follows: Figure 3 As shown, the system will start collecting server asset information periodically through the built-in timed collection task. The specific operation is as follows:

[0086] a. Monitor switch port status changes in real time;

[0087] b. When the status of a port changes from DOWN to UP, obtain the MAC address of the server on the other end of that port by querying the switch's MAC address table;

[0088] c. By obtaining the server's MAC address, look up the associated IP address in the switch's ARP table; this is the server's IP address.

[0089] d. Obtain server manufacturer, model, and serial number information using the factory default protocol configuration via IPMI;

[0090] e. Determine if the node exists by obtaining the serial number information. If it does not exist, add the device and save the device information.

[0091] f. After adding, send the server add message to control module 120 to execute the corresponding automatic configuration;

[0092] g. The built-in timed data collection task begins to collect server asset data information on a regular schedule.

[0093] Step 3: After receiving the new message from the server, the control module 120 performs network and Trap alarm configuration operations:

[0094] a. Change the network mode of the BMC management IP from DHCP to static mode to prevent the IP from changing after server node 150 restarts;

[0095] b. Configure the Trap alarm message, specifying the target address, port, and protocol information for sending Trap alarms, so that the passive monitoring module 132 can listen for and receive alarm information sent by the server.

[0096] In addition, the control module 120 supports subsequent user operations such as server firmware upgrades, operating system deployment, and network configuration via a web page.

[0097] Step four: The active monitoring module 131 is used for active monitoring of the server node 150. It collects sensor data of the server node 150 at regular intervals through the built-in timed monitoring and acquisition task, parses the status values ​​of each sensor, encapsulates the abnormal information into a standardized alarm body and sends it to the alarm module 140 for processing.

[0098] Step 5: The passive monitoring module 132 is used to listen in real time to the Trap alarm messages sent by each server node 150 to the rack server monitoring and maintenance system, and after parsing and processing the Trap messages, it encapsulates them into standardized alarm objects and sends them to the alarm module 140 for processing.

[0099] Step 6: The alarm module 140 receives the alarm objects sent by the active monitoring module 131 and the passive monitoring module 132, and performs corresponding processing after parsing out the alarm source, location, level and other information.

[0100] a. Determine if the alarm already exists on the platform; if not, generate a new alarm message.

[0101] b. If the platform already has corresponding alarm information, compare the alarm level of the alarm information with the level of the existing alarm information, and then perform alarm upgrade or downgrade operations; if the levels are the same, perform alarm merging and update the alarm update time.

[0102] c. If it is an alarm recovery, perform the corresponding alarm information clearing operation, delete the real-time alarm information and insert it into the historical information table and set the clearing time.

[0103] The rack server monitoring and maintenance system provided in this embodiment has the following beneficial technical effects: it overcomes the shortcomings of traditional monitoring methods that require tedious manual configuration and addition, and can achieve dynamic allocation, automatic discovery, automatic addition, automatic configuration and automatic monitoring of BMC management IP addresses of server nodes in the rack after power-on without manual operation. It achieves non-intrusive real-time monitoring of servers by combining active collection of sensor data and Trap alarm listening, ensuring the stable operation of each server node in the rack, reducing the probability of failure and the risk of loss caused by failure, and supporting remote batch execution of server firmware upgrades, operating system deployment and network configuration, effectively improving maintenance efficiency and saving maintenance costs.

[0104] In some embodiments, please refer to Figure 4 As shown, the present invention also provides a rack server monitoring and maintenance method 200, which uses the rack server monitoring and maintenance system described in the above embodiments, and the method includes:

[0105] Step 201: Use the asset module to monitor the port status of the switches in the cabinet to determine whether a server node is connected. When a server node is found to be connected, obtain the server configuration information and send a node addition message.

[0106] Step 202: After receiving the node add message, the control module performs network configuration on the access server node;

[0107] Step 203: Use the monitoring module to monitor the running status of the access server node to generate standardized alarm objects, and send the standardized alarm objects using the configured network;

[0108] Step 204: Receive the standardized alarm objects using the alarm module and perform alarm generation, escalation, downgrading, merging, and clearing operations on the standardized alarm objects to generate alarm messages for each server node.

[0109] The aforementioned rack server monitoring and maintenance method automatically discovers and adds server nodes by monitoring the status changes of switch ports, automatically configures server nodes through a control module, monitors the operating status of server nodes in real time through a monitoring module, and processes the data monitored by the monitoring module to generate alarm information. This method achieves automatic discovery and monitoring of rack server nodes, effectively improving the efficiency of server equipment discovery and maintenance, reducing maintenance costs, lowering the failure rate of server nodes, and ensuring the stable operation of the entire rack of server nodes.

[0110] In some embodiments, the asset module is further configured to:

[0111] Real-time monitoring of switch port status changes;

[0112] When the status of a port changes from DOWN to UP, the MAC address of the server on the other end of that port is obtained by querying the switch's MAC address table.

[0113] The server IP address is obtained by querying the switch's ARP table using the server's MAC address and finding the associated IP address.

[0114] The server node is added to the unified monitoring platform by obtaining the server manufacturer, model and serial number information using the factory default protocol configuration via the IPMI protocol.

[0115] After the addition is completed, a new message is sent to the control module to execute the corresponding automatic configuration. At the same time, the asset module's scheduled task starts to collect server asset data information periodically.

[0116] In some embodiments, the control module performs the following operations after receiving a new message:

[0117] Change the network mode of the BMC management IP from DHCP mode to static mode;

[0118] Configure the Trap alarm message, specifying the target address, port, and protocol information for sending Trap alarms, so that the monitoring module can listen for and receive alarm information sent by the server.

[0119] In some embodiments, the control module is also used to provide operations related to server firmware upgrades, operating system deployment, and network configuration.

[0120] In some embodiments, the monitoring module includes an active monitoring module and a passive monitoring module;

[0121] The active monitoring module is used to actively collect the server's operating status and encapsulate it into a standardized alarm object, which is then sent to the alarm module.

[0122] The passive monitoring module is used to receive the operating status reported by the server and encapsulate it into a standardized alarm object and send it to the alarm module.

[0123] In some embodiments, the active monitoring module has a built-in timed data acquisition task, which is responsible for periodically collecting sensor data from server nodes, analyzing the status values ​​of each sensor based on the collected data, and encapsulating abnormal information into standardized alarm objects and sending them to the alarm module for processing.

[0124] In some embodiments, the passive monitoring module is used to listen to the Trap messages sent by each server node in real time, and after parsing and processing the Trap messages, encapsulate them into standardized alarm objects and send them to the alarm module for processing.

[0125] In some embodiments, the alarm module performs the following operations when it receives a standardized alarm object sent by the active monitoring module and / or the passive monitoring module:

[0126] Determine if it is a newly generated alarm; if so, generate a new alarm message.

[0127] If the platform already has the corresponding alarm information, compare the alarm level of the alarm information with the level of the existing alarm information, and then perform the alarm upgrade or downgrade operation.

[0128] If the alarm levels are the same, alarm merging will be performed, and the alarm update time will be updated.

[0129] If it is an alarm recovery, the corresponding alarm information will be cleared, the alarm information will be inserted into the historical information table and the clearing time will be set.

[0130] It should be noted that the specific limitations on the rack server monitoring and maintenance methods can be found in the limitations of the rack server monitoring and maintenance system mentioned above, and will not be repeated here.

[0131] According to another aspect of the invention, please combine again Figure 2 As shown, this embodiment provides a server rack, which includes the server rack server monitoring and maintenance system described in the above embodiments. The system is connected to the operating system of the switch, starts automatically upon power-on, and enables DHCP service. The system includes:

[0132] The asset module is used to monitor the port status of the switches in the cabinet to determine whether a server node is connected. When a server node is found to be connected, the module obtains the server configuration information and sends a node addition message.

[0133] The control module is used to perform network configuration on the access server node after receiving the node add message;

[0134] The monitoring module is used to monitor the operating status of the access server nodes to generate standardized alarm objects, and send the standardized alarm objects using the configured network;

[0135] The alarm module is used to receive the standardized alarm objects and perform alarm generation, escalation, downgrading, merging, and clearing operations on the standardized alarm objects to generate alarm messages for each server node.

[0136] In some embodiments, the asset module is further configured to:

[0137] Real-time monitoring of switch port status changes;

[0138] When the status of a port changes from DOWN to UP, the MAC address of the server on the other end of that port is obtained by querying the switch's MAC address table.

[0139] The server IP address is obtained by querying the switch's ARP table using the server's MAC address and finding the associated IP address.

[0140] The server node is added to the unified monitoring platform by obtaining the server manufacturer, model and serial number information using the factory default protocol configuration via the IPMI protocol.

[0141] After the addition is completed, a new message is sent to the control module to execute the corresponding automatic configuration. At the same time, the asset module's scheduled task starts to collect server asset data information periodically.

[0142] In some embodiments, the control module performs the following operations after receiving a new message:

[0143] Change the network mode of the BMC management IP from DHCP mode to static mode;

[0144] Configure the Trap alarm message, specifying the target address, port, and protocol information for sending Trap alarms, so that the monitoring module can listen for and receive alarm information sent by the server.

[0145] In some embodiments, the control module is also used to provide operations related to server firmware upgrades, operating system deployment, and network configuration.

[0146] In some embodiments, the monitoring module includes an active monitoring module and a passive monitoring module;

[0147] The active monitoring module is used to actively collect the server's operating status and encapsulate it into a standardized alarm object, which is then sent to the alarm module.

[0148] The passive monitoring module is used to receive the operating status reported by the server and encapsulate it into a standardized alarm object and send it to the alarm module.

[0149] In some embodiments, the active monitoring module has a built-in timed data acquisition task, which is responsible for periodically collecting sensor data from server nodes, analyzing the status values ​​of each sensor based on the collected data, and encapsulating abnormal information into standardized alarm objects and sending them to the alarm module for processing.

[0150] In some embodiments, the passive monitoring module is used to listen to the Trap messages sent by each server node in real time, and after parsing and processing the Trap messages, encapsulate them into standardized alarm objects and send them to the alarm module for processing.

[0151] In some embodiments, the alarm module performs the following operations when it receives a standardized alarm object sent by the active monitoring module and / or the passive monitoring module:

[0152] Determine if it is a newly generated alarm; if so, generate a new alarm message.

[0153] If the platform already has the corresponding alarm information, compare the alarm level of the alarm information with the level of the existing alarm information, and then perform the alarm upgrade or downgrade operation.

[0154] If the alarm levels are the same, alarm merging will be performed, and the alarm update time will be updated.

[0155] If it is an alarm recovery, the corresponding alarm information will be cleared, the alarm information will be inserted into the historical information table and the clearing time will be set.

[0156] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0157] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A rack server monitoring and maintenance system, characterized in that, The system accesses the switch operating system, automatically starts upon power-on, and enables DHCP service. Once the server node's management port is connected to the switch port via a network cable and starts up, it dynamically assigns a management IP address to the server node's BMC via DHCP. The system includes: The asset module is used to monitor switch port status changes in real time. When a port's status changes from DOWN to UP, it retrieves the MAC address of the server on the other end of that port by querying the switch's MAC address table. It then looks up the associated IP address in the switch's ARP table using the obtained server MAC address; this is the server's IP address. The module uses the IPMI protocol and factory default configuration to obtain the server's manufacturer, model, and serial number information. It uses the obtained serial number to determine if the server node exists; if not, it adds the device and saves the device information. After adding the device, it sends a server addition message to the control module to execute the corresponding automatic configuration. The control module is used to change the network mode of the BMC management IP from DHCP mode to static mode after receiving the node add message; and to execute Trap alarm message configuration, configuring the target address, port and protocol information for sending Trap alarms so that the monitoring module can listen for and receive alarm information sent by the server. The monitoring module is used to monitor the operating status of the access server nodes to generate standardized alarm objects, and send the standardized alarm objects using the configured network; The alarm module is used to receive the standardized alarm objects and perform alarm generation, escalation, downgrading, merging, and clearing operations on the standardized alarm objects to generate alarm messages for each server node.

2. The rack server monitoring and maintenance system according to claim 1, characterized in that, The control module is also used to provide operations related to server firmware upgrades, operating system deployment, and network configuration.

3. The rack server monitoring and maintenance system according to claim 1, characterized in that, The monitoring module includes an active monitoring module and a passive monitoring module; The active monitoring module is used to actively collect the server's operating status and encapsulate it into a standardized alarm object, which is then sent to the alarm module. The passive monitoring module is used to receive the operating status reported by the server and encapsulate it into a standardized alarm object and send it to the alarm module.

4. The rack server monitoring and maintenance system according to claim 3, characterized in that, The active monitoring module has a built-in timed data acquisition task, which is responsible for periodically collecting sensor data from the server nodes, analyzing the status values ​​of each sensor based on the collected data, and encapsulating abnormal information into standardized alarm objects and sending them to the alarm module for processing.

5. The rack server monitoring and maintenance system according to claim 3, characterized in that, The passive monitoring module is used to listen to the Trap messages sent by each server node in real time, and after parsing and processing the Trap messages, it encapsulates them into standardized alarm objects and sends them to the alarm module for processing.

6. The rack server monitoring and maintenance system according to claim 3, characterized in that, When the alarm module receives a standardized alarm object sent by the active monitoring module and / or the passive monitoring module, it performs the following operations: Determine if it is a newly generated alarm; if so, generate a new alarm message. If the platform already has the corresponding alarm information, compare the alarm level of the alarm information with the level of the existing alarm information, and then perform the alarm upgrade or downgrade operation. If the alarm levels are the same, alarm merging will be performed, and the alarm update time will be updated. If it is an alarm recovery, the corresponding alarm information will be cleared, the alarm information will be inserted into the historical information table and the clearing time will be set.

7. A method for monitoring and maintaining rack-mounted servers, characterized in that, The method employs the rack server monitoring and maintenance system according to any one of claims 1-6, and the method includes: When the system is connected to the switch operating system, it starts automatically and enables DHCP service. After the management port of the server node is connected to the switch port through the network cable and starts up, it dynamically assigns a management IP address to the server node BMC through DHCP. The asset module monitors switch port status changes in real time. When a port's status changes from DOWN to UP, the MAC address of the server on the other end of that port is obtained by querying the switch's MAC address table. The associated IP address is then retrieved from the switch's ARP table using the obtained server MAC address; the server's manufacturer, model, and serial number are obtained using the factory default IPMI protocol configuration. The serial number is used to determine if the server node exists; if not, the device is added and its information is saved. After addition, a server addition message is sent to the control module to execute the corresponding automatic configuration. After receiving the node addition message, the control module changes the network mode of the BMC management IP from DHCP mode to static mode; it also executes Trap alarm message configuration, configuring the target address, port and protocol information for sending Trap alarms, so that the monitoring module can listen for and receive alarm information sent by the server. The monitoring module is used to monitor the operating status of the access server nodes to generate standardized alarm objects, and the standardized alarm objects are sent using the configured network. The alarm module receives the standardized alarm objects and performs alarm generation, escalation, downgrading, merging, and clearing operations on the standardized alarm objects to generate alarm messages for each server node.

8. A server rack, characterized in that, The rack includes the rack server monitoring and maintenance system as described in any one of claims 1-6.