Fault alarm method, device and storage medium

By building a global mapping table through switches and monitoring network connectivity, the problem of status silos and information isolation caused by BMC failures was solved. This enabled accurate location and rapid alarm of BMC, PDU and network cable failures, improving the efficiency of data center operation and maintenance.

CN122120101APending Publication Date: 2026-05-29SHENZHEN YIWANKE DATA EQUIP TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN YIWANKE DATA EQUIP TECH CO LTD
Filing Date
2026-03-20
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

In large data centers, BMC failures cannot be reported in a timely manner, resulting in the inability to monitor server status. Furthermore, the BMC, PDU, and network monitoring system are independent of each other, forming information silos and leading to low efficiency in troubleshooting.

Method used

The switch enables device identity detection for BMC and PDU, builds and dynamically updates a global mapping table, periodically monitors network connectivity, detects the physical connection status of network cables, and queries the power supply and energization status of PDUs and sockets based on the global mapping table to generate fault alarm information.

Benefits of technology

It enables automated construction and dynamic maintenance of server power link topology, accurately locates multi-dimensional faults, improves data center operation and maintenance efficiency, shortens fault repair time, and ensures stable operation of server clusters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122120101A_ABST
    Figure CN122120101A_ABST
Patent Text Reader

Abstract

The application discloses a fault alarm method and device and a storage medium, relates to the technical field of facility operation and maintenance, and comprises the following steps: based on a preset neighbor discovery protocol, completing equipment identity detection of a baseboard management controller (BMC) and a smart power distribution unit (PDU) by an exchange machine, constructing and dynamically updating a global mapping table corresponding to each device in a fault monitoring system; periodically monitoring network connectivity with the BMC through the exchange machine, detecting a network cable physical connection state when the connectivity fails, and based on the global mapping table, querying power supply and power-on states of the PDU and the jack associated with the BMC; performing fault root cause analysis on the network cable physical connection state and the power supply and power-on states, generating fault alarm information, and sending the fault alarm information to a preset management system. The application can realize automatic construction and dynamic maintenance of a server power supply link topology, break the information island of a traditional monitoring system, and realize accurate positioning and rapid alarm of multi-dimensional faults such as network cables, PDU jacks and BMCs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of facility operation and maintenance technology, and in particular to a fault alarm method, device and storage medium. Background Technology

[0002] In large data centers, the BMC (Baseboard Management Controller) acts as the server management system, responsible for monitoring hardware status and power management. However, when the BMC itself fails, server status cannot be reported in a timely manner, affecting system stability. Furthermore, the physical connection between the server and the PDU (Power Distribution Unit) sockets typically relies on manual recording, resulting in inefficient troubleshooting.

[0003] In existing technologies, BMC fault monitoring relies on its proactive reporting of heartbeat signals. Once the BMC crashes, the management system cannot detect the fault and requires additional hardware logic support, only able to detect local faults. Furthermore, the BMC, PDU, and network monitoring system are independent of each other, forming "information silos." Maintenance personnel must manually correlate information from multiple systems, lacking automated topology correlation, resulting in delayed and error-prone fault identification.

[0004] Therefore, there is an urgent need for a technical solution that can automatically detect equipment faults and dynamically maintain the server power link topology, while also enabling rapid detection and accurate location of BMC faults, network connectivity issues, and PDU socket status, in order to improve data center operation and maintenance efficiency and system reliability. Summary of the Invention

[0005] The main purpose of this application is to provide a fault alarm method, device and storage medium, which aims to automatically detect device faults and dynamically maintain the server power link topology.

[0006] To achieve the above objectives, this application proposes a fault alarm method applied to a fault monitoring system. The fault monitoring system uses a switch as its control core. The switch is connected via network cables to the management port of the Baseboard Management Controller (BMC) of at least one server and the network port of at least one Intelligent Power Distribution Unit (PDU). The server is equipped with a Power Supply Unit (PSU), which is connected to the PDU's socket for power supply via a power cord. The method includes: Based on the preset neighbor discovery protocol, the switch completes the device identity detection of BMC and PDU, and constructs and dynamically updates the global mapping table corresponding to each device in the fault monitoring system. The switch periodically monitors the network connectivity with the BMC. When connectivity fails, it detects the physical connection status of the network cable and queries the power supply and power status of the PDU and socket associated with the BMC based on the global mapping table. Perform root cause analysis on the physical connection status of the network cable and the power supply and energization status, generate fault alarm information, and send the fault alarm information to the preset management system.

[0007] In one possible implementation, the device identity detection between the BMC and PDU is performed through the switch based on a preset neighbor discovery protocol, and a global mapping table corresponding to each device in the fault monitoring system is constructed and dynamically updated, including: The switch sends router advertisement messages and receives router request messages from the BMC and PDU to complete the neighbor discovery protocol interaction. Generate a link-local address corresponding to each of the aforementioned devices, and bind the link-local address to the physical address of the BMC and the PDU respectively to maintain device identity consistency; Obtain the electrical parameters and PDU hardware logs of each socket in the PDU, obtain the PSU-related information and BMC hardware logs corresponding to the BMC, and obtain the preset physical layout information of the computer room. Based on the electrical parameters, the consistency of energy consumption data between the PSU and the PDU socket is verified. Based on the correlation of event timestamps between the PDU hardware log and the BMC hardware log, and combined with the physical layout information of the computer room, the association matching range of each device in the same physical area is constrained. The association matching result of each device is obtained through multi-dimensional consistency verification. The global mapping table is constructed based on the association matching results. The global mapping table records the association relationships between switch ports, BMC, PDU and each socket, and is updated at a preset period.

[0008] In one possible implementation, the process involves verifying the consistency of energy consumption data between the PSU and the PDU socket based on the electrical parameters, comparing the event timestamp correlation between the PDU hardware log and the BMC hardware log, constraining the association matching range of each device within the same physical area based on the data center physical layout information, and obtaining the association matching result of each device through multi-dimensional consistency verification, including: The electrical parameters of the PSU input reported by the BMC are compared with the electrical parameters to verify whether they are consistent within the preset engineering error tolerance. The PSU insertion / removal events in the BMC hardware log and the jack power-on / off events in the PDU hardware log are analyzed, and the event timestamps are compared to determine the matching pairs of events on the timeline. Based on the physical layout information of the data center, determine the physical area where the server is located, and prioritize filtering matching objects from PDUs within the physical area; By combining the consistency verification results of comprehensive energy consumption data, the correlation comparison results of event timestamps, and the physical area constraint screening results, the unique correspondence between the PSU and the PDU socket is determined, and the association matching results of each device are obtained.

[0009] In one possible implementation, the method further includes: After constructing the global mapping table, the association relationships in the global mapping table are verified, and BMCs that have not been associated with any PDUs or sockets are identified and marked as orphan nodes. Query the BMC hardware logs corresponding to the orphan node to check for any power supply failure events. If the power supply failure event occurs, it is determined that the server power supply corresponding to the orphan node is not connected to the PDU or the PDU socket group lacks a corresponding relationship, and the server power supply corresponding to the orphan node is marked as pending manual confirmation. If the power supply failure event does not exist, the re-association matching process is triggered, and the multi-dimensional consistency check is performed again to attempt to establish the association between the BMC and the PDU and socket. If the re-association and matching fails, an association error message is sent to the preset management system to prompt the maintenance personnel to check the physical connection status of the device.

[0010] In one possible implementation, when connectivity fails, detecting the physical connection status of the network cable and querying the power supply and energization status of the PDU and jack associated with the BMC based on the global mapping table includes: When a network connectivity failure with the BMC is detected, a preset number of retry probes are performed to eliminate false judgments caused by temporary network fluctuations. If connectivity is not restored after retrying the probe, a self-test is performed on the corresponding port of the switch to check the port configuration status and attempt to restore the port to normal operation. The physical layer register status is read from the switch to obtain link status parameters and negotiation status parameters, and the physical connection status of the network cable is detected based on the link status parameters and negotiation status parameters. Based on the global mapping table, query all PDUs and corresponding socket information associated with the BMC; The PDU's overall power supply status and the power supply and energization status of its corresponding sockets are queried in parallel using a preset protocol to collect complete power link status data.

[0011] In one possible implementation, detecting the physical connection status of the network cable based on the link status parameters and the negotiation status parameters includes: If the link status parameter shows a disconnected state, the physical connection of the network cable is determined to be abnormal. If the link status parameter shows a connected status and the negotiation status parameter shows a negotiation failure, then the physical connection of the network cable is determined to be normal, and the influence of the network cable itself is excluded. If both the link status parameter and the negotiation status parameter are displayed as normal, then the physical connection of the network cable is determined to be normal.

[0012] In one possible implementation, the step of performing root cause analysis on the physical connection status of the network cable and the power supply and energization status to generate fault alarm information includes: Determine whether there is a network cable fault based on the physical connection status of the network cable. If there is, determine that the root cause of the fault is an abnormal network cable connection. If the physical connection of the network cable is normal, the power supply and power status of the PDU and its corresponding socket are analyzed to determine whether there is a server power failure, partial power failure or PDU socket failure. If not, the switch is used to probe the designated management port of the BMC to determine whether the BMC service is in normal operation and to determine the root cause analysis results of the fault. Based on the root cause analysis results, the fault alarm information is generated, which includes the fault type, associated device identifier, fault occurrence time, and severity level.

[0013] In one possible implementation, if not, the switch is used to probe the designated management port of the BMC to determine whether the BMC service is in normal operation and to determine the root cause analysis results, including: When the physical connection of the network cable is normal and the power supply and power status of the PDU and the corresponding socket are normal, the switch initiates a connection request to the designated management port of the BMC. If the connection request is rejected or times out without a response, the BMC service is determined to be malfunctioning, and the root cause of the failure is identified as a BMC service failure. If the connection request is responded to normally, the system further obtains the BMC's operating status information and hardware logs through a preset protocol to comprehensively determine whether the BMC service is operating normally.

[0014] Furthermore, to achieve the above objectives, this application also proposes a fault alarm device, which includes: The construction unit is used to perform device identity detection between BMC and PDU through the switch based on the preset neighbor discovery protocol, and to construct and dynamically update the global mapping table corresponding to each device in the fault monitoring system. The detection unit is used to periodically monitor the network connectivity with the BMC through the switch. When the connectivity fails, it detects the physical connection status of the network cable and queries the power supply and power status of the PDU and socket associated with the BMC based on the global mapping table. The alarm unit is used to perform root cause analysis on the physical connection status of the network cable and the power supply and power-on status, generate fault alarm information, and send the fault alarm information to the preset management system.

[0015] In addition, to achieve the above objectives, this application also proposes a fault alarm device, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the fault alarm method as described above.

[0016] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the fault alarm method described above.

[0017] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the fault alarm method described above.

[0018] This application provides a fault alarm method, device, and storage medium. The fault alarm method, based on a preset neighbor discovery protocol, completes device identity detection of BMC and PDU through the switch, constructs and dynamically updates a global mapping table corresponding to each device in the fault monitoring system, and then periodically monitors network connectivity with the BMC through the switch. When connectivity fails, it detects the physical connection status of the network cable and queries the power supply and power status of the PDU and socket associated with the BMC based on the global mapping table. This allows for root cause analysis of the physical connection status of the network cable and the power supply and power status, generates fault alarm information, and sends the fault alarm information to a preset management system. This achieves automated construction and dynamic maintenance of the server power link topology, breaks down the information silos of traditional monitoring systems, and enables accurate location and rapid alarm of multi-dimensional faults such as network cables, PDU sockets, and BMC. This method does not require additional hardware logic, avoiding the tediousness and errors of manually recording the power link topology and solving the problem of state unawareness caused by BMC faults. This significantly improves data center operation and maintenance efficiency, shortens fault repair time, and ensures the stable operation of the server cluster. Attached Figure Description

[0019] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0020] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This is a flowchart illustrating an embodiment of the fault alarm method of this application. Figure 2 The device topology diagram provided for the fault alarm method of this application; Figure 3 The system architecture diagram provided for the fault alarm method of this application; Figure 4 A simplified flowchart illustrating the fault alarm method of this application; Figure 5 This is a schematic diagram of the device structure of the hardware operating environment involved in the fault alarm method in this application embodiment.

[0022] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0023] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0024] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0025] It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an electronic device, big data service platform, or fault alarm system capable of realizing the above functions. The following description uses a fault alarm system as an example to illustrate this embodiment and the subsequent embodiments.

[0026] Based on this, the embodiments of this application provide a fault alarm method, referring to... Figure 1 , Figure 1 This is a flowchart illustrating an embodiment of the fault alarm method of this application.

[0027] In this embodiment, the method is applied to a fault monitoring system. The fault monitoring system uses a switch as its control core. The switch is connected via network cables to the management port of the Baseboard Management Controller (BMC) of at least one server and the network port of at least one Intelligent Power Distribution Unit (PDU). The server is equipped with a Power Supply Unit (PSU), which is connected to the PDU's socket for power supply via a power cord. The fault alarm method includes steps S11 to S13: Step S11: Based on the preset neighbor discovery protocol, the device identity detection of BMC and PDU is completed through the switch, and the global mapping table corresponding to each device in the fault monitoring system is constructed and dynamically updated. It should be noted that the preset neighbor discovery protocol refers to the protocol used by nodes in a network to discover each other on the local link, obtain network configuration information, maintain neighbor status, and perform related network management functions. In this embodiment, it is specifically the neighbor discovery protocol under IPv6 network. A switch refers to the control core of the fault monitoring system, a network device that combines data forwarding, device discovery, data acquisition, topology calculation, and fault diagnosis functions. A BMC refers to the Baseboard Management Controller, an embedded management module of the server responsible for monitoring server hardware status, power management, logging, and other functions. A PDU refers to an Intelligent Power Distribution Unit, which has remote monitoring and management functions, provides multiple power sockets, and monitors the electrical parameters and switch status of each socket. Device identity detection refers to the process of confirming the unique identifiers and network addresses of the BMC and PDU in the network, ensuring that device identities are identifiable and do not conflict.

[0028] Furthermore, a fault monitoring system refers to a system centered on a switch, connecting the server BMC and PDUs to achieve power link topology construction, status monitoring, and fault alarms. A global mapping table is a structured data table recording the relationships between switch ports, BMCs, PDUs, and corresponding sockets; it is the core data support for fault location. Dynamic updates refer to refreshing the relationships in the global mapping table at preset intervals or when the device topology changes, ensuring that the table information is consistent with the actual physical connections.

[0029] The core objective of this step is to address the problems of traditional methods, such as reliance on manual recording of power link topologies, susceptibility to errors, and inability to dynamically adapt, through automated device identity detection and topology construction. In one possible implementation, a preset neighbor discovery protocol can be flexibly selected based on the network environment. In IPv4 networks, Address Resolution Protocol (ARP) can be used in conjunction with Dynamic Host Configuration Protocol (DHCP) to complete device discovery and identity binding, ensuring reliable device identity detection across different network environments.

[0030] It should also be noted that the core of device identity detection is to ensure the uniqueness and stability of the BMC and PDU in the network. By binding the link local address to the physical address, even if the device changes its physical network port or the network configuration changes slightly, the identity can still remain unchanged, thus avoiding the break of the topology relationship.

[0031] For details, please refer to steps S21-S45, which will not be repeated here.

[0032] Step S12: The switch periodically monitors the network connectivity with the BMC. When the connectivity fails, the physical connection status of the network cable is detected, and the power supply and power status of the PDU and socket associated with the BMC are queried based on the global mapping table. It should be noted that periodic monitoring refers to the switch continuously monitoring the network connection status with the BMC at preset time intervals to ensure timely detection of connectivity anomalies. Network connectivity refers to the smooth operation of the data transmission channel established between the switch and the BMC via the network cable, i.e., whether the switch can exchange data normally with the BMC. Connectivity failure refers to the state where the switch cannot establish a valid data transmission channel with the BMC and cannot obtain feedback information from the BMC. The physical connection status of the network cable refers to the actual connection condition of the network cable between the switch and the BMC, including whether the cable is plugged in tightly, whether it is broken, and whether the link is unobstructed.

[0033] Furthermore, the associated PDUs and sockets refer to the PDU devices and specific power supply sockets that are identified through a global mapping table and have a power supply connection with the PSU (Power Supply Unit) of the server to which the BMC currently has failed connectivity. Power supply status refers to the power switch status of the PDU socket, i.e., whether the socket is in an enabled power supply state. Power-on status refers to whether the PDU socket has actual current output, i.e., whether it has successfully supplied power to the connected PSU.

[0034] In one possible implementation, the time interval for periodic monitoring can be dynamically adjusted according to the data center's requirements for fault response speed. A shorter monitoring interval can be set for the BMC of core business servers, while the monitoring interval can be appropriately extended for the BMC of non-core business servers, thereby reducing system resource consumption while ensuring timely fault detection.

[0035] For details, please refer to steps S51-S63, which will not be repeated here.

[0036] Step S13: Perform root cause analysis on the physical connection status of the network cable and the power supply and power-on status, generate fault alarm information, and send the fault alarm information to the preset management system.

[0037] It should be noted that root cause analysis refers to the process of determining the specific cause of connectivity failure between the switch and the BMC based on collected data such as the physical connection status of network cables, the power supply and energization status of PDUs and sockets, and pre-defined fault judgment rules. Fault alarm information refers to structured data containing key fault-related information, used to inform maintenance personnel of the specific details of the fault for rapid troubleshooting and handling. The pre-defined management system refers to the data center platform used to centrally manage the status of various devices and receive fault alarms, enabling functions such as displaying, storing, and forwarding alarm information.

[0038] The core objective of this step is to transform the previously collected status data into clearly defined fault causes and promptly notify operations and maintenance personnel through alarm information, thus resolving the issues of vague causes and delayed responses in traditional fault diagnosis. The generation of alarm information follows standardized and structured principles, including key information such as fault type, associated device identifier, fault occurrence time, and severity level, enabling operations and maintenance personnel to quickly grasp the core fault situation and improve fault repair efficiency. In one possible implementation, fault judgment rules can be customized according to the data center's operational needs. For example, for core business servers, the severity level of a partial power outage can be set to high, while for non-core business servers, the level can be appropriately lowered. The system also supports adding new fault types and corresponding judgment logic, improving its adaptability.

[0039] Furthermore, during the root cause analysis of faults, each judgment step has a clear priority order. Network link faults have the highest priority because abnormal physical connections of the network cable will directly lead to data transmission interruption, and subsequent power link and BMC status detection cannot be performed normally. Power link faults have the next highest priority because abnormal server power supply will cause the BMC to malfunction, leading to connectivity failure. BMC service faults have the lowest priority and are only identified as such when both network and power links are normal. The severity levels of fault alarm information are divided into four levels: emergency, high, medium, and low. The emergency level corresponds to scenarios that seriously affect business operations, such as a complete server power outage or core PDU failure. The high level corresponds to scenarios such as partial power outages or BMC service failures. The medium level corresponds to scenarios such as single PDU socket failures. The low level corresponds to non-emergency scenarios such as associated anomalies. Alarm information is sent in various ways, including real-time push to the management system interface, SMS notifications, and email notifications. Operations personnel can set the receiving method according to the fault severity level to ensure that emergency faults are responded to in a timely manner.

[0040] Specifically, the process of performing root cause analysis on the physical connection status of the network cable and the power supply and energization status to generate fault alarm information can be referred to in steps S71-S83, and will not be repeated here.

[0041] Furthermore, through the SNMP (Simple Network Management Protocol) module built into the switch, SNMP Trap fault alarm information is sent to the preset management system, including key information such as the network port related to the switch fault, the IP address of the BMC, the link local address, the fault event, the PDU jack identifier, the timestamp, and the severity level (such as Critical, Warning).

[0042] This embodiment enables automated construction and dynamic maintenance of server power link topology, breaking down information silos in traditional monitoring systems. It achieves accurate location and rapid alarm for multi-dimensional faults such as network cables, PDU sockets, and BMC. This method requires no additional hardware logic; it integrates multiple industry standard protocols and hardware interface data through switch programming configuration. This avoids the tediousness and errors of manually recording power link topology and solves the problem of status unawareness caused by BMC faults. It significantly improves data center operation and maintenance efficiency, shortens fault repair time, and ensures the stable operation of server clusters.

[0043] In one feasible implementation, the device identity detection of the BMC and PDU based on the preset neighbor discovery protocol is completed through the switch, and the global mapping table corresponding to each device in the fault monitoring system is constructed and dynamically updated, including: Step S21: Send a router advertisement message through the switch, receive router request messages in response from the BMC and the PDU, and complete the neighbor discovery protocol interaction; It's important to note that a Router Advertisement (RA) message is a message actively sent by a switch to announce its existence and network configuration parameters to nodes within the local area network. It carries crucial information such as the routing prefix, link-layer address, and MTU (Maximum Transmission Unit). In essence, one or more servers' BMC management ports and one or more PDU ports are connected to different ports on the same switch via network cables, forming a local area network. The BMC's IP address is assigned by this switch. The servers may contain one or more PSU power supplies, connected to the PDU power jacks via power cables. (See reference...) Figure 2 as well as Figure 3 .

[0044] Furthermore, the Router Solicitation (RS) message refers to a request message sent by the BMC and PDU to trigger a quick response from the switch to the Router Advertisement (RS) message. Its core function is to obtain network configuration information to complete its own network access. The Neighbor Discovery Protocol (NDP) interaction refers to the process by which the switch, BMC, and PDU exchange bidirectional device identification and network configuration information through the sending and receiving of the aforementioned two types of messages.

[0045] The core objective of this step is to establish a basic communication link between the switch and the BMC and PDU, enabling the switch to identify target devices in the network, while allowing the BMC and PDU to obtain necessary network parameters for normal data transmission. In one possible implementation, the switch can dynamically adjust the sending period of router advertisement messages based on the number of devices in the network. In scenarios with dense devices, the period can be shortened to accelerate the access speed of new devices, while in scenarios with sparse devices, the period can be extended to reduce network bandwidth consumption.

[0046] Specifically, after the switch starts up, it immediately sends an initial router advertisement message to broadcast its own network configuration information to the local area network. After the BMC and PDU power on, if they have not obtained the network configuration, they will actively send a router request message. After receiving the request message, the switch will reply with a router advertisement message containing complete configuration parameters. The BMC and PDU parse the message and complete their own network configuration. At the same time, the switch records the link layer address and sending port information of the BMC and PDU. Thus, the neighbor discovery protocol interaction is completed, and a stable basic communication link is established.

[0047] For example, after a data center's fault monitoring system is started, the switch, as the core node of the local area network, immediately sends a router advertisement message. This message contains information such as the IPv6 routing prefix and an MTU value of 1500. At this time, the three newly powered-on server BMCs and two PDUs have not been configured with network parameters and each sends a router request message to the network. After receiving the messages, the switch replies to each device with a router advertisement message. The BMCs and PDUs extract the routing prefix from the message and configure their own IPv6 addresses. The switch then records the link layer addresses and corresponding ports of each device through message exchange, successfully completing the protocol interaction.

[0048] Step S22: Generate the link local address corresponding to each device, and bind the link local address to the physical address of the BMC and the PDU respectively to maintain device identity consistency; It should be noted that "equipment" refers to devices such as BMC and PDU participating in the fault monitoring system. Link-local address refers to a network address generated based on the EUI-64 standard, valid only within the local link range, and possessing local uniqueness. Physical address (MAC address) refers to the media access control address of the device's network interface card (NIC), a unique identifier fixed at the factory, used for link-layer device identification. Binding refers to establishing and storing a one-to-one correspondence between link-local addresses and physical addresses, ensuring the stability of their association. Device identity consistency means that the device's identity in the network remains unique and identifiable regardless of factors such as physical network port changes or topology adjustments.

[0049] The core objective of this step is to address the topology breakage caused by the volatile identity of devices. Through a stable address binding mechanism, the switch can accurately identify BMCs and PDUs over the long term, providing reliable identity support for subsequent association matching and dynamic topology maintenance. The implementation leverages the uniqueness of physical addresses, binding them to generated link-local addresses. This ensures both stability and communication functionality of the link-local address, guaranteeing that device identities remain valid even when the network environment changes. In one possible implementation, if an address conflict is detected (i.e., different devices generating the same link-local address), the switch will initiate an address migration protocol, assigning a temporary link-local address to the conflicting device and updating the binding relationship. Once the conflict is resolved, the original binding is restored, ensuring communication continuity.

[0050] Specifically, based on the EUI-64 standard, the switch expands and converts the physical addresses of the BMC and PDU to generate unique link-local addresses for each device. Then, an address mapping table is created to store the link-local address of each device in a one-to-one correspondence with its physical address. In subsequent communication and monitoring processes, the switch can use this mapping table to quickly locate the link-local address through the physical address, or identify the physical identity of the device through the link-local address, regardless of whether the device changes its physical network port, thus always maintaining the consistency of the device identity.

[0051] Additionally, if an address conflict is detected, such as a DAD (Duplicate Address Detection) failure, the switch initiates an address migration protocol (e.g., RFC 4429), allocates a temporary address, and updates the NDP cache to ensure communication continuity. Furthermore, it periodically (e.g., every 5 seconds) sends ICMPv6 Echo Requests to the BMC and PDU to monitor network reachability. Simultaneously, it maintains a switch port mapping table that records the switch port number, device type (BMC / PDU), device identifier (product serial number for BMC, device model and serial number for PDU), IPv6 link-local address, MAC address, and ICMP response status.

[0052] For example, a switch generates a link-local address (FE80::0200:25FF:FE96:1234) according to the EUI-64 standard for the physical address of a BMC; and generates a link-local address (FE80::0200:25FF:FEA8:4567) for the physical address of a PDU; and stores the mapping relationship between the two sets of addresses in a local mapping table. Subsequently, if the BMC changes its physical port connected to the switch for maintenance, the switch detects the port change corresponding to the physical address and quickly locates the original link-local address through the mapping table. Normal communication can resume without reconfiguration, ensuring device identity consistency.

[0053] Step S23: Obtain the electrical parameters of each socket in the PDU and the PDU hardware log, obtain the PSU related information and BMC hardware log corresponding to the BMC, and obtain the preset physical layout information of the computer room. It should be noted that each socket in the PDU refers to the power supply interface on the PDU used to connect the server PSU power plug. Each socket is independently controllable and monitorable. Electrical parameters refer to physical quantities reflecting the power supply status and energy consumption of the socket, including voltage, current, power consumption, and switch status. PDU hardware logs refer to the PDU's own operational event data, including socket power on / off, parameter anomalies, and equipment failures. The PSU-related information corresponding to the BMC refers to the server power supply unit data monitored and reported by the BMC, including the number of PSUs, their presence status, input electrical parameters, serial numbers, and plug / unplug events. PSU stands for Power Supply Unit, which is the server's power supply component responsible for converting the AC power provided by the PDU into DC power usable by the server. BMC hardware logs refer to the server hardware operational event data recorded by the BMC, including PSU status changes, hardware failures, and power management events. The preset data center physical layout information refers to the physical location distribution information of data center equipment pre-stored in the system, including server and PDU rack numbers, installation locations, and physical area divisions.

[0054] The core objective of this step is to collect multi-dimensional foundational data required to construct a global mapping table, providing data support for subsequent association matching. This is achieved by collecting operational data and logs from PDUs and BMCs using standardized protocols, retrieving preset physical location data, and forming a multi-dimensional data set of "electrical parameters + event logs + physical locations" to ensure the accuracy of association matching. In one possible implementation, electrical parameters are collected in real-time, hardware logs are collected using a combination of timed and event-triggered methods, and data center physical layout information is updated periodically, balancing data real-time performance and collection efficiency.

[0055] Specifically, the switch communicates with the PDU via Modbus TCP or SNMP protocols, polling and collecting real-time electrical parameters such as voltage, current, power consumption, and switch status of each socket, as well as socket identifiers. Simultaneously, it reads the PDU hardware logs and extracts key events. It also communicates with the BMC via IPMI protocols, subscribing to PSU-related status and events to obtain information such as PSU quantity, availability, input electrical parameters, and serial numbers. It reads the BMC hardware logs and extracts PSU-related events. Furthermore, it retrieves data from the system's pre-set database, including server and PDU rack numbers and installation locations. Finally, it stores all collected data categorized by device identifier, forming a structured data set.

[0056] Step S24: Based on the electrical parameters, verify the consistency of energy consumption data between the PSU and the PDU socket; based on the PDU hardware log and the BMC hardware log, compare the correlation of event timestamps; combine the physical layout information of the computer room to constrain the association matching range of each device in the same physical area; and obtain the association matching result of each device through multi-dimensional consistency verification. It should be noted that energy consumption data consistency refers to the consistency between the PSU input energy consumption data reported by the BMC and the corresponding socket output energy consumption data reported by the PDU within a preset engineering error tolerance, reflecting the continuity of energy transmission in the power supply link. Event timestamp correlation refers to the synchronous correspondence between socket events (power on / off, parameter anomalies) in the PDU hardware log and PSU events (plugging / unplugging, status change) in the BMC hardware log on the timeline. The same physical area refers to the physical space range (such as the same rack or rack group) including servers and corresponding power supply PDUs, defined according to the physical layout information of the data center. The association matching range refers to the range of potentially associated PDUs and sockets, narrowing the matching range through physical area constraints to improve efficiency. Multi-dimensional consistency verification refers to the process of determining the association between PSUs and PDU sockets by comprehensively verifying the results of energy consumption data consistency, event timestamp correlation, and physical area constraints. Equipment refers to the BMC, PSU, PDU, and PDU sockets. The association matching result refers to the unique correspondence between the determined PSU and PDU sockets, which is the core basis for constructing the global mapping table.

[0057] The core objective of this step is to automate and accurately match the association between PSU and PDU sockets, solving the problem that traditional methods cannot automatically establish this association. The approach involves first narrowing the matching range through physical area constraints, then verifying the energy transfer relationship through energy consumption data consistency, and finally verifying event synchronization through event timestamp correlation. These three methods mutually verify each other to ensure the accuracy and uniqueness of the matching result. In one possible implementation, the preset engineering error tolerance can be adjusted according to the equipment's precision, generally set at 3%-5%, and reduced to 1%-2% for high-precision equipment.

[0058] Specifically, firstly, based on the physical layout information of the data center, the physical area where each server is located is determined, and PDUs within that area are screened as potential matching targets, excluding PDUs that are physically far away; secondly, the input electrical parameters (focusing on power consumption) of each PSU reported by the BMC are compared with the output electrical parameters of the potential matching targets to verify whether they are consistent within the preset engineering error tolerance, and PDU sockets that meet the requirements are initially screened; then, the hardware logs of the PDUs and the BMC are analyzed, and the timestamps of the socket events and PSU events are compared to find the closest paired events on the timeline; finally, by combining the results of physical area constraints, energy consumption data consistency, and event timestamp correlation, a unique correspondence between each PSU and PDU socket is determined, forming an association matching result.

[0059] For example, based on the physical layout information of the data center, the physical area where Server-06 is located is determined to be rack 5, and the PDU in this rack is PDU-04. Only its 24 sockets are considered as potential matching targets for the two PSUs of Server-06. The input power consumption of PSU1 (800W) is compared with the output power consumption of each socket of PDU-04. It is found that the output power consumption of socket 12 is 785W, and the difference is within the 3% error tolerance. Log analysis shows that the power-on event of socket 12 of PDU-04 one hour ago (timestamp 10:05:23) and the PSU1 plug-in / unplug event recorded by BMC (timestamp 10:05:25) are 2 seconds apart, showing a strong correlation. Based on the above results, a unique correspondence is determined between PSU1 and socket 12 of PDU-04. Similarly, the correspondence between PSU2 and socket 13 of PDU-04 is determined, forming an associated matching result.

[0060] Step S25: Construct the global mapping table based on the association matching results. The global mapping table records the association relationships between switch ports, BMC, PDU and each socket, and updates the global mapping table according to a preset period.

[0061] It should be noted that the global mapping table refers to a structured data table that integrates all relationships between switch ports, BMC, PSU, PDU, and PDU sockets. It is the core data carrier for fault location and status monitoring. Relationships refer to the physical connections and logical correspondences between devices, including the network connection between switch ports and BMC, the affiliation between BMC and PSU, and the power supply connection between PSU and PDU sockets. The preset period refers to a fixed time interval pre-set by the system for updating the global mapping table, which can be adjusted according to the frequency of device changes. Updating refers to the process of adding, modifying, or deleting relationships in the global mapping table based on the matching results or device topology changes, ensuring that the table information is consistent with the actual physical connections.

[0062] The core objective of this step is to integrate scattered relationships into a unified, structured global mapping table, providing intuitive and accurate data support for subsequent fault detection and localization. The approach involves first constructing the mapping table based on the matching results and inherent device relationships, and then adapting to dynamic topology changes through a periodic update mechanism to ensure the mapping table remains effective in real time. In one possible implementation, the global mapping table is stored in a structured query language database table format, supporting fast querying and modification, while also providing a manual calibration interface so that maintenance personnel can manually adjust the relationships.

[0063] Specifically, the first step is to clarify the relationships between all devices: identify the BMC corresponding to each switch port, the server and PSU corresponding to each BMC, and the PDU and socket corresponding to each PSU; form a complete global mapping table containing switch port → server BMC → (PDU_A, socket X), (PDU_B, socket Y)... This table clearly records the correspondence between multiple PSUs and multiple PDU sockets of a single server, thus generating a mapping table of switch port → BMC → PDU + corresponding socket. The table fields include switch port number, BMC identifier, server number, PSU identifier, PDU identifier, PDU socket number, association status, etc.

[0064] Furthermore, a preset update cycle (e.g., 10 minutes) is set. Within each cycle, device identity detection and association matching are re-executed, and the mapping table association is refreshed. If events such as device hot-plugging or topology changes are detected, the update process is immediately triggered, and the mapping table is adjusted in real time. After the update, the mapping table is verified to ensure that there are no logical conflicts or missing data, and to ensure that it accurately reflects the actual physical connection status.

[0065] For example, a global mapping table is constructed based on the association matching results: Switch port 1 corresponds to BMC-01 (Server-01), BMC-01 corresponds to PSU1 and PSU2, PSU1 corresponds to PDU-01 jack 5, and PSU2 corresponds to PDU-02 jack 8; Switch port 2 corresponds to BMC-02 (Server-02), BMC-02 corresponds to PSU1, and PSU1 corresponds to PDU-01 jack 6. The system's preset update cycle is 10 minutes. At a certain moment, PSU2 on Server-01 is re-inserted into PDU-02 jack 9. Upon detecting the topology change, the system immediately triggers an update, re-matches and modifies the jack information corresponding to PSU2 in the mapping table to ensure that the mapping table is consistent with the actual connection.

[0066] This embodiment utilizes a neighbor discovery protocol to discover devices and establish basic communication. Address binding ensures stable device identities, and multi-dimensional basic data is collected. Device relationships are established through multi-dimensional consistency checks, and a global mapping table is finally built and dynamically updated. This achieves automated and precise construction and maintenance of server power link topology. The entire process requires no manual intervention, solving the problems of error-prone and difficult-to-maintain traditional manual topology recording. Furthermore, the dynamic update mechanism adapts to device topology changes, effectively breaking down information silos in device monitoring and significantly improving data center operation and maintenance efficiency and reliability.

[0067] In one feasible implementation, the process of verifying the consistency of energy consumption data between the PSU and the PDU socket based on the electrical parameters, comparing the event timestamp correlation between the PDU hardware log and the BMC hardware log, and constraining the association matching range of each device within the same physical area based on the physical layout information of the data center, and obtaining the association matching result of each device through multi-dimensional consistency verification, includes: Step S31: Compare the PSU input electrical parameters reported by the BMC with the electrical parameters to verify whether they are consistent within the preset engineering error tolerance. It should be noted that the PSU input electrical parameters reported by the BMC refer to the relevant physical quantity data monitored and uploaded by the server's baseboard management controller when the power supply unit (PSU) obtains power from the outside, including input voltage, input current, and input power consumption. The electrical parameters refer to the output electrical parameters of each socket of the intelligent power distribution unit (PDU), that is, the physical quantity data collected by the PDU through its own monitoring module when outputting power from the socket, also including output voltage, output current, and output power consumption. The preset engineering error tolerance refers to a reasonable range that is pre-set based on factors such as equipment hardware precision, circuit transmission loss, and environmental interference, allowing for differences between the two types of electrical parameters. Verifying consistency refers to judging whether the difference between the PSU input electrical parameters reported by the BMC and the corresponding socket output electrical parameters reported by the PDU is within the preset engineering error tolerance, thereby verifying the continuity of power transmission in the power supply link.

[0068] The core purpose of this step is to initially screen out PDU sockets that are associated with the PSU through the consistency verification of energy transmission data. The idea is to use the correlation of energy consumption data between the upstream and downstream of the power supply link to exclude PDU sockets that are not associated with the power supply, thereby improving the efficiency and accuracy of association matching.

[0069] Specifically, firstly, the BMC-reported input electrical parameters of the PSU to be matched are extracted from the stored structured data set, and the output electrical parameters of each socket of all PDUs within the potential matching range are also extracted. Then, key parameters such as voltage, current, and power consumption are compared and calculated to determine the ratio of the difference between the PSU input parameters and the PDU socket output parameters. Finally, it is determined whether the ratio of the difference of each key parameter is within the preset engineering error tolerance. If it is, the energy consumption data of the two are determined to be consistent, and the PDU socket is a candidate matching object for the PSU. If the ratio of the difference of any key parameter exceeds the tolerance, the matching qualification of the PDU socket is excluded.

[0070] Step S32: Analyze the PSU insertion / removal events in the BMC hardware log and the socket power-on / off events in the PDU hardware log, and compare the event timestamps to determine the matching pairs of events on the timeline; It should be noted that the PSU insertion / removal event refers to the event recorded in the log where the PSU is inserted into or removed from the server power interface. This event directly reflects the change in the physical connection status between the PSU and the server. The jack power-on / off event refers to the event recorded in the log where the PDU jack power switch is turned on (powered on) or off (powered off), reflecting the change in the jack's power supply status. The event timestamp refers to the specific time information of the event recorded in the log, accurate to the second or millisecond, and is the core identifier of the event occurrence sequence. The core purpose of this step is to further verify the actual correlation between the candidate matching object and the PSU through the correlation of the event occurrence sequence, and to eliminate false matches where the energy consumption data is consistent but there is no actual power supply connection. The implementation idea is to use the synchronicity of physical connection operation and power supply status change to lock the true related event pair by comparing timestamps. In one possible implementation, a time difference threshold (such as 5 seconds) can be set. When the timestamp difference between the PSU insertion / removal event and the jack power-on / off event is less than or equal to this threshold, it is determined to be a matched pair of events on the timeline.

[0071] Specifically, firstly, all PSU insertion / removal events are parsed from the BMC hardware logs, extracting the corresponding PSU identifier, event type (insertion / removal), and occurrence timestamp. Then, all socket power-on / off events are parsed from the PDU hardware logs, extracting the corresponding PDU identifier, socket number, event type (power-on / power-off), and occurrence timestamp. Next, the PSU insertion / removal events and socket power-on / off events within the candidate matching range are matched according to event type (insertion corresponds to power-on, removal corresponds to power-off). Finally, the timestamp difference between corresponding events is calculated one by one, and event pairs with timestamp differences within a preset threshold are selected as matched pairs on the timeline.

[0072] Step S33: Determine the physical area where the server is located based on the physical layout information of the computer room, and prioritize filtering matching objects from the PDUs within the physical area; It should be noted that the physical layout information of the data center refers to the structured information pre-stored in the fault monitoring system, reflecting the physical location distribution of various devices within the data center. This includes the rack numbers, installation locations within the racks, and physical area divisions for devices such as servers and PDUs. The physical area where the server is located refers to the specific spatial range where the server is actually installed, determined based on the data center's physical layout information. This range can be defined by racks, rack groups, or area division rules. Prioritizing matching objects within the physical area means that when matching PSUs and PDU sockets, PDUs deployed within that physical area are prioritized for matching. Only when no matching object meets the criteria within this range are PDUs outside the physical area considered, thus narrowing down the matching range.

[0073] The core objective of this step is to leverage the correlation between the physical locations of devices to constrain the matching range, reduce invalid matching calculations, improve the efficiency of correlation matching, and simultaneously reduce the probability of false matches caused by excessively distant physical locations. The underlying approach is based on the actual deployment scenario of data center equipment, where server power supply PDUs are typically deployed in the same physical area (such as the same server rack). This optimization of the matching process through physical location constraints is achieved. In one possible implementation, the data center physical layout information can be automatically updated. When a device is moved to a new physical location, the system can obtain location change information through relevant detection mechanisms and update the data center physical layout information in real time, ensuring the accuracy of physical area determination.

[0074] Specifically, the system first retrieves the physical layout information of the data center from the pre-set database. Based on the identifier (such as the server number) of the server to which the PSU to be matched belongs, it queries the physical information such as the rack number and installation location of the server. Then, based on the physical area division rules (such as the same rack being a physical area), it determines the physical area where the server is located. Next, it filters out the PDUs that are in the same physical area as the server from all the discovered PDUs, forming a priority matching PDU set. Finally, it performs subsequent energy consumption data consistency verification and event timestamp correlation comparison only in each socket of the priority matching PDU set to narrow down the matching range.

[0075] Step S34: Based on the energy consumption data consistency verification results, event timestamp correlation comparison results, and physical area constraint screening results, determine the unique correspondence between the PSU and the PDU socket, and obtain the association matching results of each device.

[0076] It should be noted that the energy consumption data consistency verification result refers to the judgment result obtained in step S31 regarding whether the electrical parameters of the PSU and PDU sockets are consistent within the preset engineering error tolerance (consistent / inconsistent). The event timestamp correlation comparison result refers to the judgment result obtained in step S32 regarding whether the PSU insertion / removal event and the PDU socket power-on / off event are paired events that match on the time axis (matched / not matched). The physical area constraint filtering result refers to the set of PDUs that are preferentially matched and located in the same physical area as the server, as determined in step S33. The unique correspondence between PSU and PDU sockets means that one PSU has an actual power supply connection with only one specific socket of one PDU. This relationship is unique, and there will be no situation where one PSU corresponds to multiple sockets or one socket corresponds to multiple PSUs.

[0077] The core purpose of this step is to accurately pinpoint the actual power supply connection between the PSU and PDU sockets through comprehensive analysis of multi-dimensional results, ensuring the accuracy and uniqueness of the correlation matching results. The approach is to use the results of three dimensions—physical location constraints, energy consumption data consistency, and event timing correlation—as the basis for judgment. These three dimensions are verified collaboratively, and only when the results of all three dimensions meet the matching conditions is a unique correspondence determined, thus avoiding mismatches caused by a single-dimensional judgment.

[0078] In one possible implementation, weighting coefficients can be set for the results of the three dimensions, and a comprehensive matching score can be obtained by weighted calculation. When the score is higher than a preset threshold, it is determined to be a unique correspondence, thereby further improving the reliability of the matching results.

[0079] Specifically, firstly, the results of each candidate matching combination (the combination of PSU and PDU sockets) are organized across three dimensions; then, each candidate combination is comprehensively evaluated, and is determined to be a valid matching combination only if it meets three conditions: "it is within the priority matching set of physical area constraints, the energy consumption data consistency verification is passed, and there are matching paired events on the time axis"; if there are multiple valid matching combinations, the difference in event timestamps is further compared, and the combination with the smallest difference is selected as the final matching result; if there is only one valid matching combination, it is directly determined as the unique correspondence between PSU and PDU sockets; finally, the unique correspondences of all PSUs are integrated to form the associated matching results of each device.

[0080] For example, the candidate matching combinations for PSU1 are (PSU1, PDU-04 jack 12) and (PSU1, PDU-04 jack 13). The combination (PSU1, PDU-04 jack 12) satisfies the following conditions: it is within the priority matching set of rack 5, its energy consumption data consistency verification passes (difference ratio 2.94% ≤ 3%), and there are paired events with a timestamp difference of 2 seconds. The combination (PSU1, PDU-04 jack 13), although within the priority matching set, has an energy consumption data difference ratio of 8% (exceeding the 3% tolerance) and no matching paired events. Therefore, it is determined that PSU1 and PDU-04 jack 12 have a unique correspondence. Similarly, the unique correspondences of other PSUs are determined, forming the associated matching results.

[0081] This embodiment narrows the matching range by using physical area constraints, verifies the power transmission relationship through energy consumption data consistency, confirms the timing of physical connection operations by using event timestamp correlation, and finally determines a unique association relationship by integrating multi-dimensional results. This achieves accurate and automated PSU and PDU socket association matching. This process requires no manual intervention, solving the problems of error-prone and inefficient traditional manual topology recording. Furthermore, multi-dimensional collaborative verification avoids the limitations of single-dimensional judgment, ensuring that the association matching results accurately reflect the actual power link connection relationship and effectively improving the intelligent management level of data center power links.

[0082] In one feasible implementation, the method further includes: Step S41: After constructing the global mapping table, the association relationships in the global mapping table are verified, and BMCs that have not been associated with any PDUs or sockets are identified and marked as orphan nodes. It should be noted that the association relationship refers to the physical connection and logical correspondence between devices, including the power supply association between the BMC and PDUs and sockets. A BMC that has not established an association with any PDU or socket refers to a BMC that only records its own basic information in the global mapping table, but does not correspond to any PDU identifier or socket number. An orphan node is a specific identifier for a BMC that has not established a valid association relationship with a PDU or socket, used to distinguish between normally associated BMCs and unassociated BMCs.

[0083] In one possible implementation, the association verification can be performed synchronously with the dynamic update of the global mapping table, automatically triggering the verification process after each update to identify orphan nodes in real time.

[0084] Specifically, firstly, the BMC information of all records in the global mapping table is traversed to extract the PDU and socket association fields corresponding to each BMC; then, the association fields of each BMC are checked one by one to see if there is valid data (i.e., whether it contains a clear PDU identifier and socket number); if the association field of a certain BMC is empty or there is no valid matching record, it is determined that the BMC has not established an association with any PDU or socket; finally, an orphan node identifier is added to the BMC and it is stored separately.

[0085] Step S42: Query the BMC hardware logs corresponding to the orphan node to check for power supply failure events; It should be noted that the BMC hardware logs corresponding to orphan nodes refer to the log data of server hardware operating status and events recorded by the BMC of the node marked as an orphan, including various event information such as power management and hardware failures. Power supply failure events refer to events recorded in the logs related to the server's inability to obtain normal power supply, including situations such as no power input to the PSU, abnormal power supply voltage, and power link interruption. In one possible implementation, keywords for power supply failure events (such as "Power Supply Failure" or "No Power Input") can be preset, and target events in the logs can be quickly filtered through keyword matching to improve detection efficiency.

[0086] Specifically, firstly, based on the orphan node identifier, the corresponding BMC is located, and a communication connection is established with the BMC. Then, the hardware log stored in the BMC is read via the IPMI protocol to obtain information such as event type, occurrence time, and event description. Next, according to preset event identification rules, events related to power supply in the log are filtered. Finally, it is determined whether any of the filtered events meet the definition of power supply failure, and the detection result is obtained. For example, for orphan node BMC-07, its hardware log is read via the IPMI protocol. The log records the event description "2024-05-20 14:30:00 Power SupplyFailure: PSU1 has no power input". According to preset rules, it is determined that there is a power supply failure event in this log.

[0087] Step S43: If the power supply failure event exists, it is determined that the server power supply corresponding to the orphan node is not connected to the PDU or the PDU socket group lacks a corresponding relationship, and the server power supply corresponding to the orphan node is marked as pending manual confirmation. It should be noted that "server power supply" refers to the PSU (Power Supply Unit) configured on the server corresponding to the orphan node, which is the core power supply component of the server. "Not connected to PDU" means that the server's PSU has not established a physical connection with any PDU jack via a power cable. "PDU jack group lacks corresponding relationship" means that although the server's PSU is connected to a PDU jack, the PDU or jack is not detected by the system, or no association record has been established in the global mapping table. "Pending manual confirmation" is an indicator set for server power supplies with abnormal power supply associations, used to prompt maintenance personnel to manually verify the actual physical connection status. In this state, the system does not automatically process the association relationship of the power supply.

[0088] The core purpose of this step is to identify the cause of the abnormality of orphan nodes based on power supply failure events, and guide maintenance personnel to conduct targeted investigations of physical connection problems by marking them as pending manual confirmation. The approach is to analyze the power supply abnormality events in the logs in conjunction with missing information to infer the possible problems in the actual physical connection.

[0089] Specifically, first, based on the specific description of the power supply failure event, determine the server power supply involved in the anomaly (such as PSU1, PSU2, etc.); then, comprehensively analyze the cause of the event to determine the anomaly type as the server power supply not being connected to a PDU or the PDU socket group lacking a corresponding relationship; next, find the record corresponding to the server power supply in the global mapping table and add a "pending manual confirmation" status label; finally, synchronize the status change information to the system database to ensure the consistency of status records.

[0090] Step S44: If there is no power supply failure event, trigger the re-association matching process, perform multi-dimensional consistency verification again, and attempt to establish the association between the BMC and the PDU and socket. It should be noted that triggering the re-association and matching process refers to the system automatically starting the preset association and matching program to re-execute the association and matching operation between the BMC, PDU, and socket. Multi-dimensional consistency verification refers to a verification process that integrates three dimensions: physical area constraints, energy consumption data consistency, and event timestamp correlation. It is the core method for establishing association relationships. Establishing association relationships refers to the process of determining the PDU and socket corresponding to the BMC through verification and writing the association information into the global mapping table.

[0091] The core objective of this step is to eliminate missing correlations caused by omissions in the initial matching or incomplete data collection. By re-executing the matching process, it attempts to automatically establish normal correlations, reducing manual intervention. The underlying idea is to infer that the missing correlation is a system-level matching issue, assuming the actual power supply is normal, and then repair the correlation by repeating the core verification process. In one possible implementation, the re-correlation matching process can optimize the data collection strategy, such as extending the data collection duration and increasing the frequency of electrical parameter collection, thereby improving the accuracy of the verification.

[0092] Specifically, the process begins by initiating a re-association and matching procedure to clear any invalid association records previously associated with the BMC. Then, relevant PSU information (such as electrical parameters, serial number, and hardware logs) corresponding to the BMC is re-collected, along with the electrical parameters, hardware logs, and physical layout information of all PDU sockets. Next, multi-dimensional consistency verification is performed again in the order of physical area constraints → energy consumption data consistency verification → event timestamp correlation comparison. Finally, based on the verification results, an attempt is made to determine the unique correspondence between the BMC and the PDU and sockets. If the verification passes, the global mapping table is updated.

[0093] Step S45: If the re-association and matching fails, an association error message is sent to the preset management system to prompt the maintenance personnel to check the physical connection status of the device.

[0094] It should be noted that "failed re-association matching" means that the multi-dimensional consistency check in step S44 failed to find a matching PDU and socket, making it impossible to establish a valid association between the BMC and the PDU. The preset management system refers to the platform in the data center used for centralized management of device status and receiving alarm information, enabling functions such as information display, storage, and notification. The association anomaly prompt information refers to structured information containing orphan node identifiers, server information, and speculative causes of the anomaly, used to clearly inform maintenance personnel of the specific circumstances of the association anomaly. Verifying the physical connection status of the equipment refers to maintenance personnel physically inspecting the physical connections of the server PSU and PDU sockets based on the prompt information (e.g., whether the power cord is firmly plugged in, whether it is plugged into the wrong socket, etc.). The core purpose of this step is to promptly notify maintenance personnel to intervene and ensure the integrity of the association relationship after the system's automatic repair fails.

[0095] In one possible implementation, the associated anomaly alert information can be sent in various ways (such as management system pop-ups, SMS, email), along with key information such as server location and BMC identifier, to facilitate maintenance personnel in quickly locating the device.

[0096] Specifically, the system first integrates information such as orphan node identifier, corresponding server number, rack location, and time of re-matching failure to generate standardized association anomaly prompts. Then, it activates the switch's built-in communication module and sends the prompts to the management system via SNMP or other preset communication protocols. After receiving the information, the management system displays it in real time on the operations and maintenance interface and sends reminders to operations and maintenance personnel according to preset notification rules. Finally, it waits for operations and maintenance personnel to verify the physical connection and updates the system status based on the feedback results.

[0097] This embodiment identifies orphan nodes in the global mapping table and analyzes the reasons for missing associations through log detection. For different reasons, it adopts different processing methods such as marking for manual confirmation, re-association and matching, and sending anomaly prompts. It can not only automatically repair association anomalies caused by system matching omissions, but also accurately locate physical connection problems that require manual intervention. It effectively ensures the integrity and accuracy of the global mapping table, avoids the failure of subsequent fault detection and location due to missing associations, and further improves the reliability and intelligence level of the fault monitoring system.

[0098] In one feasible implementation, when connectivity fails, detecting the physical connection status of the network cable and querying the power supply and energization status of the PDU and socket associated with the BMC based on the global mapping table includes: Step S51: When a network connectivity failure with the BMC is detected, a preset number of retry probes are performed to eliminate false judgments caused by temporary network fluctuations. It should be noted that network connectivity failure refers to the state where the switch, during periodic monitoring, does not receive a response to the BMC's probe packets and is unable to establish a valid data transmission channel with the BMC. The preset number of attempts refers to a fixed number of retry probes, pre-set based on network environment stability and device communication characteristics, intended to eliminate temporary connectivity anomalies caused by accidental factors. Retry probes refer to the operation where, after initially detecting a connectivity failure, the switch sends probe packets to the BMC again at preset time intervals to attempt to re-establish communication. Temporary network fluctuations refer to short-term network connectivity instability caused by non-substantial faults such as network signal interference and data transmission congestion.

[0099] The core objective of this step is to reduce the probability of false positives through multiple retries, ensuring that the detection results of connectivity failures are accurate and reliable. In one possible implementation, the preset number of retries can be dynamically adjusted according to the actual network environment. In scenarios with complex and fluctuating network environments, the number of retries can be increased appropriately, while in scenarios with stable networks, the number of retries can be reduced. At the same time, the time interval for retrying the probe can be set to gradually increase to balance detection efficiency and accuracy.

[0100] Specifically, when the switch detects a network connectivity failure with the BMC, it triggers a retry probe mechanism. Then, it continuously sends probe packets to the BMC at preset time intervals, accumulating the number of transmissions to a preset number. After each transmission, it waits for a preset duration to receive a response from the BMC. If a valid response is received from the BMC during the retry probe, it is determined to be a temporary network fluctuation, and the normal monitoring process is restored. If no response is received after completing the preset number of probes, the connectivity failure is confirmed, and subsequent processing steps are initiated.

[0101] Step S52: If connectivity is not restored after retrying the probe, perform a self-test on the port corresponding to the switch, check the port configuration status, and attempt to restore the port to normal operation. It's important to note that the port corresponding to the switch refers to the physical port of the switch directly connected to the BMC via a network cable, which is the key interface for data transmission. Self-test refers to the switch automatically checking the operating status and configuration parameters of its own ports to troubleshoot port faults. Port configuration status refers to key configuration parameters such as port activation status, speed configuration, and duplex mode, which directly affect the port's communication function. Attempting to restore normal port operation refers to the switch automatically performing operations such as restarting the port and restoring default configurations when an abnormal port configuration or operational failure is detected, attempting to repair the port's functionality. The core purpose of this step is to troubleshoot and eliminate connectivity failures caused by faults in the switch's own ports, ensuring that subsequent troubleshooting focuses on external links or devices. The approach is to first troubleshoot local port faults, then locate external problems, following an "inside-out" troubleshooting logic. In one possible implementation, port self-test can cover multiple dimensions, including port hardware status detection (e.g., whether the port is damaged), configuration parameter verification (e.g., whether the speed matches the BMC network port), and traffic statistics analysis (e.g., whether there are dropped packets), to comprehensively troubleshoot port anomalies.

[0102] Specifically, first, determine the switch port number connected to the malfunctioning BMC and identify the target for self-test; then check the port's activation status to confirm if it is active; next, verify if the port's speed and duplex mode are correctly configured and whether they match the BMC's network port configuration parameters; if a configuration anomaly is detected, automatically adjust the parameters to match; if the configuration is normal but the port's operating status is abnormal, attempt to restart the port or switch to restore it; after restarting, send an ICMPPing probe message to the BMC again; if connectivity is restored, the port fault has been repaired; if not, proceed to the next steps.

[0103] Step S53: Read the physical layer register status through the switch to obtain link status parameters and negotiation status parameters, and detect the physical connection status of the network cable based on the link status parameters and negotiation status parameters; It's important to note that the physical layer register refers to the register in the physical layer chip (such as the Ethernet PHY chip) corresponding to the switch's network port that stores port status information. It records key status data such as link connection and speed negotiation. Link status parameters reflect the physical connection of the network cable, used to determine if the cable is properly plugged in and if the link is established. Negotiation status parameters reflect the negotiation results between the switch port and the BMC port regarding speed and duplex mode, used to determine if the communication parameters match. The physical connection status of the network cable refers to the physical connection between the network cable and the switch port or BMC port, including states such as normal connection, disconnected cable, and poor contact.

[0104] The core objective of this step is to accurately determine whether there is a fault in the physical connection of the network cable by reading parameters at the hardware level. The approach is to utilize the characteristic that physical layer registers directly reflect the hardware connection status. By parsing link status and negotiation status parameters, network cable faults can be distinguished from other communication faults. In one possible implementation, the switch reads the physical layer register status through the MDIO (Management Data Input / Output) interface. This interface is an industry-standard interface that enables efficient management and status reading of the physical layer chip.

[0105] Specifically, the switch first establishes communication with the physical layer chip through the MDIO interface and sends a register read command; then it reads the link status parameters and negotiation status parameters from the specified register address and obtains the binary values ​​or status indicators of the parameters; next, according to the preset parameter parsing rules, it determines the physical connection status corresponding to the link status parameters (e.g., a parameter of "1" represents a connected link, and "0" represents a disconnected link); then it parses the negotiation status parameters to determine whether the speed and duplex mode negotiation is successful; finally, it combines the parsing results of the two parameters to determine the physical connection status of the network cable.

[0106] In one embodiment, the switch program checks the physical layer and obtains the link status. The switch's network port detects whether the network cable is inserted through physical layer signals (such as changes in the level of the Ethernet PHY chip). If the network cable is inserted and the signal is normal (even if the peer server is not powered on), the switch's PHY register group (Link Status, Auto-Negotiation Status) is read directly through the MDIO interface. Linkstatus represents the physical layer status of the network port. If Link status shows 1 (e.g., Ethernet0 / 1 Link Status is up), the switch and the peer server's BMC network port negotiate the speed and duplex mode through Fast Link Pulse (FLP). If the peer server is not powered on, is disabled, or damaged, the negotiation will fail (Auto-Negotiation Status is down), but the switch can still detect that the network cable is inserted (because the physical layer signal is present), indicating that the physical connection of the network cable is not a problem, ruling out network interface damage or disconnection causing the network port to be unreachable, and proceeding to step S54; if Link status shows 0, it indicates that the physical connection channel of the network cable is broken, directly marking it as a Network Cable Interface Failure event and generating a fault alarm.

[0107] Step S54: Based on the global mapping table, query all PDUs and corresponding socket information associated with the BMC; It should be noted that all PDUs and sockets associated with the BMC refer to all intelligent power distribution units and specific power sockets that provide power to the PSUs of the server to which the BMC belongs, as determined by the global mapping table. The corresponding information refers to key information such as the PDU's device identifier (e.g., model, serial number), socket number, and physical location. This information is used to accurately locate the power supply equipment and interfaces. In one possible implementation, the global mapping table adopts an indexed storage structure, using the BMC identifier as the index key, which allows for quick retrieval of the corresponding PDU and socket information, improving query efficiency.

[0108] Specifically, firstly, the unique identifier (such as product serial number or link local address) of the BMC with the current connectivity anomaly is extracted; then, using this identifier as the query condition, the global mapping table is retrieved; from the mapping table, all PDU records that are associated with the BMC are filtered out, and information such as the device identifier and location of each PDU is extracted; at the same time, information such as the socket number and socket identifier corresponding to the PSU server to which the BMC belongs is extracted from each PDU; and this information is organized into a structured query result set.

[0109] Step S55: The overall power supply status of the PDU and the power supply and power-on status of the corresponding sockets are queried in parallel through a preset protocol to collect complete power link status data.

[0110] It should be noted that the preset protocol refers to the standardized protocol pre-configured by the system for data communication with the PDU, including Modbus TCP protocol, SNMP protocol, etc., which support remote querying of PDU status. Parallel query refers to sending status query requests to all associated PDUs queried in step S54 simultaneously, rather than querying them one by one, in order to improve data acquisition efficiency. The overall power supply status of the PDU refers to whether the PDU is connected to the main power supply and whether it is in normal power supply mode, reflecting the power availability of the PDU itself. The power supply status of the corresponding socket refers to the power switch status (e.g., on or off) of the PDU socket associated with the server PSU, reflecting whether the socket allows power output. The power-on status of the corresponding socket refers to whether the PDU socket associated with the server PSU has actual current output, reflecting whether the socket has successfully powered the PSU. Complete power link status data refers to a comprehensive data set integrating the overall power supply status of the PDU, the power supply status of the associated socket, and the power-on status, which can completely reflect the operation of the power link from the PDU to the server PSU.

[0111] In one possible implementation, parallel queries can be achieved through multi-threading or asynchronous communication, while setting a query timeout. If a PDU query times out, it is marked as a query failure, and relevant information is recorded to ensure the integrity and timeliness of data collection.

[0112] Specifically, firstly, based on the query result set in step S54, the list of PDUs to be queried and the target sockets corresponding to each PDU are determined; then, multiple parallel query tasks are started, each task corresponding to one PDU, and a status query command is sent to the PDU through a preset protocol. The command explicitly specifies the overall power supply status to be queried and the power supply status and power-on status of the target socket; after receiving the command, the PDU returns the corresponding status data; the switch receives the return data from all PDUs, and if there is a query timeout or data abnormality, the corresponding PDU and socket status are marked as "unknown" and recorded; finally, all valid return data are integrated to form a complete power link status dataset.

[0113] This embodiment eliminates temporary network fluctuations through retry detection, then checks local faults by self-testing switch ports, followed by detecting the physical connection status of network cables, then locking the associated PDUs and sockets, and finally querying the power supply status in parallel. This ensures the accuracy of the detection results and improves data collection efficiency through parallel queries, comprehensively collecting multi-dimensional data such as network connectivity status, link physical status, and power link status. It effectively solves the problems of scattered data and low efficiency in traditional fault diagnosis, and improves the intelligence and accuracy of fault detection.

[0114] In one feasible implementation, detecting the physical connection status of the network cable based on the link status parameters and the negotiation status parameters includes: Step S61: If the link status parameter shows a disconnected state, it is determined that the physical connection of the network cable is abnormal. It should be noted that link status parameters, read from the switch's physical layer registers, reflect the connectivity status of the physical link between the network cable and the switch port / BMC port. These parameters are the core basis for determining whether the physical connection is functioning correctly. A disconnected state means that the binary value or status identifier corresponding to the link status parameter is a preset "disconnected" flag, indicating that the physical layer has not detected a valid link signal. An abnormal physical connection of the network cable refers to problems such as the cable not being properly plugged in, being broken, or having poor contact, which prevents the establishment of a physical transmission channel between the switch and the BMC. In one possible implementation, the "disconnected" flag of the link status parameter can be defined by a preset threshold or a specific binary code, and different models of physical layer chips can achieve standardized judgment through unified parameter parsing rules.

[0115] Specifically, the switch first reads the link status parameters in the physical layer register through the MDIO interface to obtain the raw data corresponding to the parameters; then, it decodes the raw data according to the preset parsing rules to determine whether it meets the identification characteristics of "disconnected state"; if the parsing result is "disconnected state", it directly determines that the physical connection of the network cable is abnormal, that is, no effective physical connection is established between the network cable and the switch port or BMC port, and there are faults such as loose plugging or breakage; finally, the determination result and the corresponding raw data of the link status parameters are recorded to provide a basis for subsequent fault alarms.

[0116] Step S62: If the link status parameter shows a connected state and the negotiation status parameter shows a negotiation failure, then the physical connection of the network cable is determined to be normal, and the influence of the network cable itself is excluded. It's important to note that "connection status" refers to the link status parameter being marked with the preset "connection" flag, indicating that the physical layer has detected a valid link signal and a physical connection has been established between the network cable and the ports of both ends. "Negotiation status parameters" are read from the physical layer registers and reflect the negotiation results of communication parameters such as speed and duplex mode between the switch port and the BMC port. "Negotiation failure" means the negotiation status parameter is marked with the preset "failure" flag, indicating that the two parties failed to reach an agreement on communication parameters. "Normal physical connection of the network cable" means that the network cable itself has no breaks, poor contact, or other problems and can transmit physical signals. "Rule the elimination of network cable faults" means confirming that the network cable is not the cause of network connectivity failure; the root cause should be checked from other aspects such as communication parameter configuration and BMC port status.

[0117] The core purpose of this step is to distinguish between physical network cable faults and communication parameter negotiation failures, avoiding misjudging negotiation failures as network cable faults and ensuring the accuracy of troubleshooting. The approach involves using link status parameters to determine the existence of a physical connection and negotiation status parameters to determine if the communication configuration matches. Combining these two methods enables precise differentiation of fault types. In one possible implementation, the determination of negotiation failure can be further refined into specific types such as rate negotiation failure and duplex mode negotiation failure, providing a more accurate reference for subsequent parameter adjustments.

[0118] Specifically, first, the link status parameters are read and parsed to confirm that they are displayed as "connected," indicating that the physical connection of the network cable is unobstructed. Then, the negotiation status parameters are read and parsed to confirm that they are displayed as "negotiation failed." Based on the parsing results of the above two parameters, it is determined that the network cable itself is not faulty and the physical connection is normal. At the same time, the specific type of negotiation failure (such as rate negotiation failure) is recorded to provide clues for subsequent troubleshooting of communication parameter configuration problems. Finally, the possibility of faulty network cable itself is ruled out, and the focus of troubleshooting is shifted to switch port configuration, BMC port status, etc.

[0119] Step S63: If both the link status parameter and the negotiation status parameter are displayed as normal, then the physical connection of the network cable is determined to be normal.

[0120] It should be noted that "both showing as normal" means that the link status parameter shows "connected" and the negotiation status parameter shows "negotiation successful," indicating that the physical connection and communication configuration meet the data transmission requirements. "Negotiation successful" means that the identifier corresponding to the negotiation status parameter is the preset "success" identifier, indicating that the switch port and the BMC port have reached an agreement on communication parameters such as speed and duplex mode, enabling normal data transmission. "Normal physical connection of the network cable" means that the network cable itself is in good condition and the physical connection to the ports of both ends is firm, enabling stable transmission of physical signals. The core purpose of this step is to confirm that there are no problems with the physical connection of the network cable and the configuration of communication parameters. This is achieved through dual parameter verification to ensure that there are no faults in the physical layer and data link layer related to the network cable, thus clarifying the direction for troubleshooting. In one possible implementation, after successful negotiation, the specific communication parameters (such as negotiated speed and duplex mode) can be further read and recorded to provide complete data for system maintenance and fault tracing.

[0121] Specifically, first, the link status parameters are read and parsed to confirm that they are displayed as "connected status," indicating that the physical connection of the network cable is valid. Then, the negotiation status parameters are read and parsed to confirm that they are displayed as "negotiation successful," indicating that the communication parameters match. Combining the normal status of the two parameters, it is determined that the physical connection of the network cable is normal and can support data transmission. Finally, the determination result and the negotiated communication parameters are recorded to rule out the possibility of network cable-related faults and proceed to the subsequent troubleshooting process for power link or BMC service status.

[0122] This embodiment achieves accurate judgment of the physical connection status of network cables by using different combinations of link status parameters and negotiation status parameters. It can quickly identify physical faults in the network cable itself and distinguish between communication parameter negotiation problems and network cable faults, avoiding misjudgments that could lead to incorrect troubleshooting directions. By clearly defining the physical connection status of the network cable, a clear scope for elimination is defined for subsequent root cause analysis. When the network cable is determined to be normal, the focus can be directly on troubleshooting other aspects such as the power link and BMC services; when the network cable is determined to be abnormal, the fault can be quickly located and addressed, significantly improving the efficiency and accuracy of troubleshooting.

[0123] In one feasible implementation, the step of performing root cause analysis on the physical connection status of the network cable and the power supply and energization status to generate fault alarm information includes: Step S71: Determine whether there is a network cable fault based on the physical connection status of the network cable. If there is, determine that the root cause of the fault is an abnormal network cable connection. It should be noted that the physical connection status of the network cable refers to the physical connection between the network cable and the device port, determined by reading the link status parameters and negotiation status parameters in the switch's physical layer registers. This includes states such as normal connection, disconnection, and poor contact. A network cable fault refers to a problem where the network cable itself is broken, damaged, or has poor contact with the switch port or BMC port, resulting in the inability to transmit physical signals normally. The root cause of the fault refers to the most fundamental reason for the network connectivity failure. An abnormal network cable connection means that the physical connection status of the network cable does not meet normal communication requirements, and is the direct cause of network connectivity failure. The core purpose of this step is to prioritize troubleshooting faults at the physical link level and quickly pinpoint network cable-related issues. The approach is to use the previously detected physical connection status results to directly correlate with the root cause of the fault, following the "physical first, logical second" fault diagnosis principle. In one possible implementation, the root cause can be further subdivided based on the specific abnormal type of the network cable physical connection status, such as "broken network cable," "network cable not properly plugged in," or "poor network cable contact," to improve the accuracy of fault location.

[0124] Specifically, first, the results of the physical connection status of the network cable determined in steps S61 to S63 are retrieved; if the result is "abnormal physical connection of network cable", it is directly determined that there is a network cable fault; then, combined with the physical layer register parameter parsing results, the abnormal scenario is further clarified (such as the network cable being broken or not plugged in properly when the link status parameter is in the disconnected state); finally, the root cause of the fault is determined to be the abnormal connection of the network cable, and key information such as the switch port and BMC identifier corresponding to the fault is recorded to provide a basis for generating alarm information.

[0125] Step S72: If the physical connection of the network cable is normal, analyze the power supply and power status of the PDU and the corresponding socket to determine whether there is a server power failure, partial power failure or PDU socket failure. It should be noted that a normal physical connection of the network cable means that the judgment results of steps S61 to S63 indicate that the network cable has no broken or poor contact issues, and the physical signal transmission is smooth. The power supply and power status of the PDU and its corresponding socket refer to the overall power supply status of the PDU associated with the target BMC, the power switch status (on / off) of the socket, and the actual power-on status (current output / no current output) obtained through a preset protocol. A complete server power failure means that all PDU sockets associated with all PSUs of the server are not effectively powered, and the server cannot obtain any power. A partial power failure means that when the server is configured with multiple PSUs, some PSUs are associated with PDU sockets that are not effectively powered, and the server relies only on some PSUs for power, losing redundancy. A PDU socket failure refers to an abnormal situation where the overall power supply of the PDU is normal, but a certain socket's power switch is on but there is no current output, or the switch status does not match the power-on status.

[0126] The core objective of this step is to investigate power link-level faults after ruling out network cable failures, accurately pinpointing the type of power supply-related anomalies. This is achieved by using a global mapping table to associate PDU and socket status, and employing multi-scenario logical judgment to differentiate between different types of power supply faults, providing precise information for fault handling. In one possible implementation, the judgment logic can be adjusted based on the number of PSUs configured on the server (single power supply / multiple power supply). Multi-power supply servers focus on detecting partial power outages, while single-power supply servers focus on detecting overall power outages.

[0127] Specifically, firstly, extract the power supply and power-on status data of all PDUs and corresponding sockets associated with the target BMC; if the power supply status of all associated sockets is ON, but there is no current output in the power-on state, it is marked as a ServerPower Loss event, indicating a complete server power failure; if the power supply status of some associated sockets is ON, but the power-on state is OFF (e.g., only one power supply of a dual-power server is out of power), it is marked as a Partial Power Loss event, with a severity level of Warning, indicating that the server may still be running but has lost redundancy; if the power supply status of some associated sockets is OFF, but the overall power supply status of the PDU is ON, it is marked as PDU Socket Failed, specifying which socket of which PDU is abnormal; if the power supply and power-on status of all associated sockets are normal, the power link failure is ruled out, and the subsequent troubleshooting steps are initiated.

[0128] Step S73: If not, then probe the designated management port of the BMC through the switch to determine whether the BMC service is in normal operation and determine the root cause analysis result of the fault. It should be noted that "no" refers to the judgment result of step S72 indicating that there is no overall server power failure, partial power failure, or PDU jack failure, and the power link status is normal. Switch detection refers to the operation where the switch actively sends a connection request to the designated management port of the BMC to check whether the port can respond normally. The designated management port of the BMC refers to a pre-defined specific port (such as TCP port 623) used for BMC service communication, and is the core port for BMC to provide management functions. BMC service refers to the software program running on the baseboard management controller, used for monitoring server hardware status, power management, fault reporting, etc. Normal operation status refers to the state in which the BMC service can normally receive and respond to external requests and complete various management functions. Fault root cause analysis result refers to the final cause of network connectivity failure determined after comprehensively considering all troubleshooting results.

[0129] The core objective of this step is to investigate BMC service failures after ruling out network cable and power link faults, thus completing a closed-loop troubleshooting process for the root cause of the fault. This is achieved by using port probing to verify the availability of the BMC service and directly linking the service response status to the root cause of the fault. In one possible implementation, a timeout period for port probing can be set; if no response is received within the timeout period, the service is considered abnormal. Multiple retry attempts can be made to eliminate the impact of temporary service fluctuations.

[0130] Specifically, the switch first sends a TCP connection request to the designated management port of the target BMC; if the connection request is rejected or a timeout occurs without a response, the BMC service is deemed to be malfunctioning, and the root cause of the fault is determined to be a BMC service failure; if the connection request is responded to normally, the BMC's operating status information and hardware logs are further obtained through the IPMI protocol to comprehensively determine whether there are any hidden faults in the BMC service; finally, based on the detection results and supplementary verification information, the final root cause analysis result is determined.

[0131] Step S74: Based on the root cause analysis results, generate the fault alarm information, which includes the fault type, associated device identifier, fault occurrence time, and severity level.

[0132] It should be noted that the root cause analysis result refers to the specific reasons for network connectivity failure identified in steps S71 to S73, including abnormal network cable connections, overall server power failure, partial power failure, PDU socket failure, BMC service failure, etc. Fault type refers to the specific category corresponding to the root cause, such as "abnormal network cable connection" or "PDU socket failure." Associated device identifier refers to the unique identifier of the device related to the fault, including switch port number, BMC identifier, PDU identifier, socket number, etc. Fault occurrence time refers to the specific time (accurate to the second) when the system detected the fault and determined its root cause. Severity level refers to the level of impact of the fault on server operation, such as "Critical" or "Warning," used to prioritize fault handling tasks.

[0133] In one possible implementation, the severity level can be dynamically mapped according to the fault type. For example, a complete server power outage or BMC service failure is mapped to the "emergency" level, a partial power outage is mapped to the "warning" level, and a PDU jack failure is mapped to the "medium" level.

[0134] Specifically, first, the fault type is determined based on the root cause analysis results; then, the identifiers of related devices are extracted (e.g., the switch port and BMC identifier for a network cable fault, and the PDU identifier and socket number for a PDU socket fault); the specific time of the fault occurrence is recorded; the fault severity level is determined according to the preset level mapping rules; finally, the above information is integrated in a standardized format to generate fault alarm information, ensuring that the information is complete, logically clear, and easy for maintenance personnel to quickly understand the fault situation.

[0135] This embodiment first checks for physical link (network cable) faults, then for power link (PDU and jack) faults, and finally for device service (BMC) faults, forming a complete troubleshooting logic "from the outside in, from the physical to the logical." This process, through precise status judgment and scenario classification, can accurately pinpoint the vague fault of "server disconnection" to specific fault types and associated devices, generating standardized alarm information containing key information. It not only solves the problems of vague causes and low efficiency in traditional fault diagnosis, but also optimizes fault handling priorities through severity level classification, significantly improving data center operation and maintenance efficiency, shortening the average fault repair time, and providing reliable assurance for the stable operation of server clusters.

[0136] In one feasible implementation, if not, the switch is used to probe the designated management port of the BMC to determine whether the BMC service is in normal operation and to determine the root cause analysis results, including: Step S81: When the physical connection of the network cable is normal and the power supply and power-on status of the PDU and the corresponding socket are normal, the switch initiates a connection request to the designated management port of the BMC. It should be noted that a normal physical network cable connection refers to the condition determined by reading the switch's physical layer register parameters, indicating no broken or poorly connected network cables, and a smooth physical link. Normal power supply and energization status of the PDU and its corresponding jacks means that the PDU associated with the target BMC is powered normally, all corresponding jack power switches are on and there is actual current output, and the server power supply unit (PSU) can obtain power normally. The designated management port refers to a pre-defined network port (such as TCP port 623) used by the BMC to provide management services externally; it is the communication entry point for BMC services. A connection request refers to a request message sent by the switch to the designated management port of the BMC based on the TCP protocol to establish a communication connection. The core purpose of this step is to verify the availability of BMC services through port connection probing after ruling out physical and power link failures, providing crucial evidence for root cause analysis. The implementation approach is to indirectly reflect the service operation status using the reachability of port communication, following the logic of "first eliminating external faults, then checking internal services."

[0137] In one possible implementation, the designated management port can be flexibly adjusted according to the BMC configuration. The system supports multiple preset candidate ports. When the first port probe fails, other candidate ports are automatically tried to improve the probe success rate.

[0138] Specifically, first, confirm that the physical connection status of the network cable is normal, and that the power supply and power-on status of the PDU and its corresponding socket meet the normal power supply requirements; then, the switch constructs a TCP connection request message based on the preset BMC designated management port information, the message containing key information such as the switch identifier and the probe purpose; next, the switch sends a connection request to the designated management port through the network link established with the BMC; finally, wait for the BMC's response, and at the same time start a timeout timer to reserve response time for subsequent judgment.

[0139] Step S82: If the connection request is rejected or does not receive a response after a timeout, it is determined that the BMC service is malfunctioning and the root cause of the fault is a BMC service failure. It should be noted that a connection request being rejected means that the designated management port of the BMC is open, but it explicitly returns a response message that refuses to establish a connection (such as a TCP RST message), indicating that the BMC actively refuses to establish a connection with the switch. No response received within a timeout means that after the switch sends a connection request, it does not receive any response message from the BMC within a preset timeout period, indicating that the BMC has not processed the connection request. BMC service malfunction means that the BMC's management service has not started normally, the service process is stuck, or the service function is malfunctioning, making it unable to provide normal communication services. The root cause of the fault refers to the most fundamental reason for the network connectivity failure. A BMC service failure means that the BMC's own management service is malfunctioning, which is the direct cause of the network connectivity failure.

[0140] The core purpose of this step is to quickly determine the BMC service status through port connection results, pinpoint the root cause of service-level faults, and provide a clear direction for fault alarms and handling. The approach is to directly correlate port connection rejections or timeout responses with the service's running status, leveraging the communication characteristics of the TCP protocol to achieve rapid service availability detection. In one possible implementation, the specific fault scenarios corresponding to "connection rejected" and "timeout without response" can be distinguished. For example, a connection rejection might be due to BMC service configuration limitations, while a timeout without response might indicate that the BMC service is frozen or the process has crashed, providing operations and maintenance personnel with more accurate troubleshooting clues.

[0141] Specifically, firstly, monitor the response of the switch after sending a connection request; if a connection rejection message is received from the BMC, it is determined that the connection request has been rejected; if no response is received after the timeout timer is triggered, it is determined that no response has been received after the timeout; based on either of the above two situations, it is marked as a BMCFault event, and it is determined that the BMC service is malfunctioning; finally, based on the overall investigation results of network connectivity failure, determine that the root cause of the fault is a BMC service failure, and record key information such as the BMC identifier, the designated management port, and the response status.

[0142] Step S83: If the connection request is responded to normally, the BMC's operating status information and hardware logs are further obtained through a preset protocol to comprehensively determine whether the BMC service is operating normally.

[0143] It should be noted that a normal connection request response means that after the designated management port of the BMC receives the connection request, it returns a response message agreeing to establish a connection (such as a TCP SYN+ACK message), and the switch and the BMC successfully establish a TCP connection. Pre-configured protocols refer to pre-configured standardized protocols used to communicate with the BMC and obtain its internal status information, such as the IPMI protocol (Intelligent Platform Management Interface Protocol). BMC operational status information refers to the BMC's own service operation parameters, including service process status, CPU utilization, memory usage, network configuration status, etc. Hardware logs refer to the server hardware operation event data recorded by the BMC, including PSU status changes, sensor data anomalies, service start / stop events, etc. Comprehensive judgment refers to combining the connection response results, operational status information, and hardware logs to verify the actual operation of the BMC service from multiple dimensions, avoiding the limitations of judging by a single indicator. Normal BMC service operation means that the BMC's management service functions are complete, the processes are stable, and it can normally monitor the server hardware status and respond to external management requests.

[0144] The core objective of this step is to further verify the functional integrity of the BMC service, based on a normal port connection, to eliminate hidden faults such as "port accessible but service malfunctioning," and to ensure the accuracy of fault root cause identification. The approach involves deeply understanding the internal state of the BMC through protocol interaction and combining this with log information for multi-dimensional analysis. In one possible implementation, a preset protocol can support batch acquisition of runtime status parameters and log data, while setting parameter anomaly thresholds. When a parameter exceeds the threshold, it is automatically marked as a suspicious anomaly to aid in comprehensive judgment.

[0145] Specifically, after the switch successfully establishes a TCP connection with the BMC, it first sends a data query command to the BMC through a preset protocol. The command explicitly requests to obtain the operating status information and hardware logs within a specified time period. After receiving the command, the BMC returns the corresponding status data and log content. The switch parses the received data to check whether the service process is active and whether the critical resource usage is within a reasonable range. At the same time, it filters events related to the BMC service in the hardware logs to determine whether there are records of abnormal service restarts, functional errors, etc. Finally, by combining the rationality of the operating status parameters and the completeness of the hardware logs, the switch obtains the final judgment result on whether the BMC service is operating normally.

[0146] This embodiment first initiates a port connection request after ruling out external faults, then makes a preliminary judgment on the service status based on the connection response, and finally conducts in-depth analysis of internal data through protocol interaction, forming a complete BMC service fault diagnosis process of "port reachability detection → preliminary status judgment → in-depth data verification". This process not only utilizes port probing for rapid preliminary judgment, but also avoids missing hidden faults through in-depth data collection, ensuring the accuracy of BMC service fault diagnosis. Simultaneously, this process forms a closed loop with the previous network cable and power link fault diagnosis process, comprehensively covering the main fault scenarios of server disconnection, providing crucial support for accurate root cause location, and effectively improving the efficiency and reliability of data center fault diagnosis.

[0147] For example, to help understand the implementation process of the fault alarm method, please refer to... Figure 4 , Figure 4 A simplified flowchart illustrating the fault alarm method of this application.

[0148] Specifically, after the server powers on and the BMC is running normally, the fault monitoring system centered on the switch starts working. The switch first sends IPv6 router advertisement messages, and the BMC and PDU respond to the router request messages to complete the neighbor discovery protocol interaction. It also strongly binds the IPv6 link local address generated by EUI-64 with the device MAC address to maintain identity consistency. It periodically sends ICMPv6 probe messages to monitor network reachability and maintain the switch port mapping table. Then, it collects the socket electrical parameters by polling the PDU and subscribes to the PSU related information of the BMC. After multi-dimensional verification of electrical parameter consistency, hardware event sequence association and physical area constraints, it builds and dynamically updates a global mapping table that records the relationship between switch port → BMC → PDU and corresponding socket. At the same time, it detects orphan nodes and marks them as pending manual confirmation.

[0149] Furthermore, when the switch periodically detects network connectivity failures with the BMC, it first performs a preset number of retry probes to eliminate temporary network fluctuations. After consecutive timeouts, it performs a self-test on the corresponding port and attempts to recover. Then, it checks the physical connection status of the network cable by reading the link status and negotiation status parameters of the physical layer register. If the link status parameters show a disconnection, the network cable is marked as faulty. If the connection is normal, it queries the associated PDU and socket based on the global mapping table and obtains their power supply and power status in parallel through Modbus TCP to determine whether there is a server power failure, partial power failure, or PDU socket failure. If the power and link status are normal, it probes the designated management port of the BMC and determines whether the BMC service is abnormal based on the connection response. Finally, the switch activates the SNMP module and sends an alarm message containing the fault type, associated device identifier, timestamp, and severity level to the preset management system.

[0150] This solution uses a switch as the core control hub to locate faults in network cables, PDU sockets, or BMC services in a hierarchical manner, and sends precise alarms through SNMP Trap. In this way, the switch integrates multiple protocols to realize the automated construction of power link topology and precise fault location in a hierarchical manner, breaking the traditional information silos and manual dependence.

[0151] It should be noted that the examples in the figure are only for understanding this application and do not constitute a limitation on the fault alarm method of this application. Any simple modifications based on this technical concept are within the protection scope of this application.

[0152] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0153] This application also provides a fault alarm device, the fault alarm device comprising: The construction unit is used to perform device identity detection between BMC and PDU through the switch based on the preset neighbor discovery protocol, and to construct and dynamically update the global mapping table corresponding to each device in the fault monitoring system. The detection unit is used to periodically monitor the network connectivity with the BMC through the switch. When the connectivity fails, it detects the physical connection status of the network cable and queries the power supply and power status of the PDU and socket associated with the BMC based on the global mapping table. The alarm unit is used to perform root cause analysis on the physical connection status of the network cable and the power supply and power-on status, generate fault alarm information, and send the fault alarm information to the preset management system.

[0154] The fault alarm device provided in this application, employing the fault alarm method in the above embodiments, can solve the technical problems in the background art. Compared with the prior art, the beneficial effects of the fault alarm device provided in this application are the same as the beneficial effects of the fault alarm method provided in the above embodiments, and other technical features in the fault alarm device are the same as the features disclosed in the methods of the above embodiments, and will not be repeated here.

[0155] This application provides a fault alarm device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the fault alarm method in Embodiment 1 above.

[0156] The following is for reference. Figure 5 The diagram illustrates a structural schematic suitable for implementing a fault alarm device according to embodiments of this application. The fault alarm device in these embodiments may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 5 The fault alarm device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0157] like Figure 5As shown, the fault alarm device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in the read-only memory 1002 or a program loaded from the storage device 1003 into the random access memory 1004. The random access memory 1004 also stores various programs and data required for the operation of the fault alarm device. The processing unit 1001, the read-only memory 1002, and the random access memory 1004 are interconnected via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. Communication device 1009 allows the fault alarm device to communicate wirelessly or wiredly with other devices to exchange data. Although the figures show fault alarm devices with various systems, it should be understood that implementation or possession of all the systems shown is not required. More or fewer systems may be implemented alternatively.

[0158] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from read-only memory 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0159] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0160] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the fault alarm methods provided by the above methods.

[0161] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0162] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0163] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.

Claims

1. A fault alarm method, characterized in that, The method is applied to a fault monitoring system, which uses a switch as its control core. The switch is connected via network cables to the management port of the Baseboard Management Controller (BMC) of at least one server and the network port of at least one Intelligent Power Distribution Unit (PDU). The server is equipped with a Power Supply Unit (PSU), which is connected to the PDU's socket for power supply via a power cord. The method includes: Based on the preset neighbor discovery protocol, the switch completes the device identity detection of BMC and PDU, and constructs and dynamically updates the global mapping table corresponding to each device in the fault monitoring system. The switch periodically monitors the network connectivity with the BMC. When connectivity fails, it detects the physical connection status of the network cable and queries the power supply and power status of the PDU and socket associated with the BMC based on the global mapping table. Perform root cause analysis on the physical connection status of the network cable and the power supply and energization status, generate fault alarm information, and send the fault alarm information to the preset management system.

2. The fault alarm method as described in claim 1, characterized in that, The preset neighbor discovery protocol, through the switch, completes device identity detection between the BMC and PDU, and constructs and dynamically updates the global mapping table corresponding to each device in the fault monitoring system, including: The switch sends router advertisement messages and receives router request messages from the BMC and PDU to complete the neighbor discovery protocol interaction. Generate a link-local address corresponding to each of the aforementioned devices, and bind the link-local address to the physical address of the BMC and the PDU respectively to maintain device identity consistency; Obtain the electrical parameters and PDU hardware logs of each socket in the PDU, obtain the PSU-related information and BMC hardware logs corresponding to the BMC, and obtain the preset physical layout information of the computer room. Based on the electrical parameters, the consistency of energy consumption data between the PSU and the PDU socket is verified. Based on the correlation of event timestamps between the PDU hardware log and the BMC hardware log, and combined with the physical layout information of the computer room, the association matching range of each device in the same physical area is constrained. The association matching result of each device is obtained through multi-dimensional consistency verification. The global mapping table is constructed based on the association matching results. The global mapping table records the association relationships between switch ports, BMC, PDU and each socket, and is updated at a preset period.

3. The fault alarm method as described in claim 2, characterized in that, The process involves verifying the consistency of energy consumption data between the PSU and the PDU socket based on the electrical parameters, comparing the event timestamp correlation between the PDU hardware log and the BMC hardware log, and constraining the association matching range of each device within the same physical area based on the data center physical layout information. The association matching results for each device are obtained through multi-dimensional consistency verification, including: The electrical parameters of the PSU input reported by the BMC are compared with the electrical parameters to verify whether they are consistent within the preset engineering error tolerance. The PSU insertion / removal events in the BMC hardware log and the jack power-on / off events in the PDU hardware log are analyzed, and the event timestamps are compared to determine the matching pairs of events on the timeline. Based on the physical layout information of the data center, determine the physical area where the server is located, and prioritize filtering matching objects from PDUs within the physical area; By combining the consistency verification results of comprehensive energy consumption data, the correlation comparison results of event timestamps, and the physical area constraint screening results, the unique correspondence between the PSU and the PDU socket is determined, and the association matching results of each device are obtained.

4. The fault alarm method as described in claim 2, characterized in that, The method further includes: After constructing the global mapping table, the association relationships in the global mapping table are verified, and BMCs that have not been associated with any PDUs or sockets are identified and marked as orphan nodes. Query the BMC hardware logs corresponding to the orphan node to check for any power supply failure events. If the power supply failure event occurs, it is determined that the server power supply corresponding to the orphan node is not connected to the PDU or the PDU socket group lacks a corresponding relationship, and the server power supply corresponding to the orphan node is marked as pending manual confirmation. If no power supply failure event occurs, the re-association matching process is triggered, and the multi-dimensional consistency check is performed again to attempt to establish the association between the BMC and the PDU and socket. If the re-association and matching fails, an association error message is sent to the preset management system to prompt the maintenance personnel to check the physical connection status of the device.

5. The fault alarm method as described in claim 1, characterized in that, When connectivity fails, the physical connection status of the network cable is detected, and based on the global mapping table, the power supply and energization status of the PDU and socket associated with the BMC are queried, including: When a network connectivity failure with the BMC is detected, a preset number of retry probes are performed to eliminate false judgments caused by temporary network fluctuations. If connectivity is not restored after retrying the probe, a self-test is performed on the corresponding port of the switch to check the port configuration status and attempt to restore the port to normal operation. The physical layer register status is read from the switch to obtain link status parameters and negotiation status parameters, and the physical connection status of the network cable is detected based on the link status parameters and negotiation status parameters. Based on the global mapping table, query all PDUs and corresponding socket information associated with the BMC; The PDU's overall power supply status and the power supply and energization status of the corresponding sockets are queried in parallel using a preset protocol to collect complete power link status data.

6. The fault alarm method as described in claim 5, characterized in that, The step of detecting the physical connection status of the network cable based on the link status parameters and the negotiation status parameters includes: If the link status parameter shows a disconnected state, the physical connection of the network cable is determined to be abnormal. If the link status parameter shows a connected status and the negotiation status parameter shows a negotiation failure, then the physical connection of the network cable is determined to be normal, and the influence of the network cable itself is excluded. If both the link status parameter and the negotiation status parameter are displayed as normal, then the physical connection of the network cable is determined to be normal.

7. The fault alarm method as described in claim 1, characterized in that, The step of performing root cause analysis on the physical connection status of the network cable and the power supply and energization status to generate fault alarm information includes: Determine whether there is a network cable fault based on the physical connection status of the network cable. If there is, determine that the root cause of the fault is an abnormal network cable connection. If the physical connection of the network cable is normal, the power supply and power status of the PDU and its corresponding socket are analyzed to determine whether there is a server power failure, partial power failure or PDU socket failure. If not, the switch is used to probe the designated management port of the BMC to determine whether the BMC service is in normal operation and to determine the root cause analysis results of the fault. Based on the root cause analysis results, the fault alarm information is generated, which includes the fault type, associated device identifier, fault occurrence time, and severity level.

8. The fault alarm method as described in claim 7, characterized in that, If not, then the designated management port of the BMC is probed through the switch to determine whether the BMC service is in normal operating condition, and the root cause analysis results are determined, including: When the physical connection of the network cable is normal and the power supply and power status of the PDU and the corresponding socket are normal, the switch initiates a connection request to the designated management port of the BMC. If the connection request is rejected or times out without a response, the BMC service is determined to be malfunctioning, and the root cause of the failure is identified as a BMC service failure. If the connection request is responded to normally, the system further obtains the BMC's operating status information and hardware logs through a preset protocol to comprehensively determine whether the BMC service is operating normally.

9. A fault alarm device, characterized in that, The fault alarm device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the fault alarm method as described in any one of claims 1 to 8.

10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the fault alarm method as described in any one of claims 1 to 8.