Node fault locating method and supernode server

CN122547589APending Publication Date: 2026-08-11GUANGDONG PURUI YUNCHUANG TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-19
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0004]然而,对于大规模部署的服务器集群中,由于节点数量可能达到数十甚至上百个,因此,在对故障节点定位时,存在定位效率不高的问题

Benefits of technology

[0033] The node fault location method and supernode server provided in this application embodiment, by establishing a communication connection between the first baseboard controller and the second baseboard controller, can read the second information representing the health status of the second server node from the second baseboard controller of the second server node, and determine the fault status of the second server node by analyzing the second information, thereby achieving rapid location of node faults and improving the efficiency of fault node location.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122547589A_ABST
    Figure CN122547589A_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a node fault positioning method and a supernode server. The method comprises: establishing a communication connection with a second baseboard controller, and reading second information from the second baseboard controller; the second baseboard controller is a baseboard controller of a second server node; the first server node and the second server node are server nodes in a supernode server, and the first server node and the second server node are on the same board card; the second information represents state information of the second server node; according to the second information, a fault state of the second server node is determined and reported. The method is used to improve the positioning efficiency of the fault node.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a node fault location method and a supernode server. Background Technology

[0002] With the rapid development of technologies such as data centers, cloud computing, and artificial intelligence, super node servers are gradually becoming the core hardware architecture for high-performance computing and large-scale data processing. Super node servers typically consist of multiple boards (nodes) with identical functions. These nodes communicate at high speed through a backplane or interconnect bus to jointly complete complex computing tasks.

[0003] Currently, when a node fails during the use of a supernode server, the status of each node needs to be checked manually.

[0004] However, in large-scale server clusters, where the number of nodes can reach dozens or even hundreds, there is a problem of low efficiency in locating faulty nodes. Summary of the Invention

[0005] The node fault location method and supernode server provided in this application embodiment are used to improve the efficiency of locating faulty nodes.

[0006] In a first aspect, embodiments of this application provide a node fault location method, applied to a first baseboard controller in a first server node, the method comprising:

[0007] Establish a communication connection with the second baseboard controller and read the second information from the second baseboard controller; the second baseboard controller is the baseboard controller of the second server node, the first server node and the second server node are server nodes in the super node server, and the first server node and the second server node are on the same board, and the second information represents the status information of the second server node.

[0008] Based on the second information, determine and report the fault status of the second server node.

[0009] Optionally, a communication connection is established with the second baseboard controller, and second information is read from the second baseboard controller, including:

[0010] Identify and connect the target extended communication bus for connection with the second baseboard controller;

[0011] Second information is read from the second baseboard controller via the target extended communication bus.

[0012] Optionally, after establishing a communication connection with the second substrate controller and reading the second information from the second substrate controller, the method further includes:

[0013] Write first information to the second baseboard controller so that the second baseboard controller can determine and report the fault status of the first server node based on the first information; the first information is characterized as the status information of the first server node.

[0014] Optionally, when the first server node is the master node, the method further includes:

[0015] A communication connection with the third baseboard controller is established by extending the communication bus. The third baseboard controller is the baseboard controller of the third server node, and the third server node is the server node of the secondary node in the super node server.

[0016] The first information of the first server node is sequentially written into the third baseboard controller, and the third information is read from the third baseboard controller. The third information represents the status information of the third server node.

[0017] Optionally, the method further includes:

[0018] Read the ID information from the CPLD register; the ID information is the information sent to the CPLD register by the identity setting interface.

[0019] Based on the ID information, the first server node is determined to be the master node.

[0020] Optionally, when the second server node is a secondary node, the method further includes:

[0021] A communication connection with the third baseboard controller is established by extending the communication bus;

[0022] The second information of the second server node is written sequentially into the third baseboard controller, and the third information is read from the third baseboard controller.

[0023] Optionally, after establishing a communication connection with the second substrate controller and reading the second information from the second substrate controller, the method further includes:

[0024] The second information is sent to the fourth baseboard controller, which is the baseboard controller of the fourth server node. The fourth server node is the other server node in the supernode server besides the second server node.

[0025] Optionally, the second information is information read by the second baseboard controller from the CPLD register in the second server node;

[0026] And / or,

[0027] The second information includes at least one of temperature parameter information, voltage parameter information, and operating status parameter information.

[0028] Optionally, a communication connection is established with the second baseboard controller, and second information is read from the second baseboard controller, including:

[0029] Based on a preset polling frequency, a communication connection is established with the second baseboard controller, and second information is read from the second baseboard controller; wherein, the polling frequency is adjusted according to the load status of the server nodes in the supernode server.

[0030] Secondly, embodiments of this application provide a supernode server, including a first baseboard controller and a second baseboard controller. The first baseboard controller is installed on a first server node, and the second baseboard controller is installed on a second server node. The first server node and the second server node are server nodes on the same board in the supernode server; wherein...

[0031] The second baseboard controller is used to acquire second information characterizing the health status of the second server node;

[0032] The first baseboard controller is used to establish a communication connection with the second baseboard controller to read second information from the second baseboard controller and determine and report the fault status of the second server node based on the second information.

[0033] The node fault location method and supernode server provided in this application embodiment, by establishing a communication connection between the first baseboard controller and the second baseboard controller, can read the second information representing the health status of the second server node from the second baseboard controller of the second server node, and determine the fault status of the second server node by analyzing the second information, thereby achieving rapid location of node faults and improving the efficiency of fault node location. Attached Figure Description

[0034] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0035] Figure 1 A flowchart illustrating the node fault location method provided in this application;

[0036] Figure 2 Module illustration of the supernode server provided in this application Figure 1 ;

[0037] Figure 3 Module illustration of the supernode server provided in this application Figure 2 ;

[0038] Figure 4 This is a schematic diagram of the node structure provided in an embodiment of this application;

[0039] Figure 5 This is a schematic diagram illustrating the working principle of communication between nodes in an embodiment of this application.

[0040] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation

[0041] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0042] First, let me explain the terms used in this application:

[0043] A hypernode server can refer to a high-density computing architecture that tightly couples multiple independent server nodes within the same physical chassis through a high-speed interconnect backplane or switching network. These nodes share power, heat dissipation, and management network, presenting themselves as a logically unified and powerful computing cluster. It is often used in scenarios that require horizontal scaling and high reliability, such as high-performance computing and cloud computing hyperconverged infrastructure.

[0044] A server node can refer to an independent computing unit in a supernode server, which may include a CPU, memory, storage, baseboard controller and operating system, and can run tasks independently or work in collaboration with other nodes.

[0045] The baseboard controller refers to the baseboard management controller (BMC). The baseboard management controller can be a dedicated management microcontroller on the server node, which runs independently of the main CPU. It is responsible for collecting health information such as node temperature, voltage, and fan speed, providing remote management interfaces (such as IPMI and Redfish), and can exchange data with the BMCs of other nodes through an extended communication bus, thereby realizing cross-node status monitoring and fault reporting.

[0046] A board can refer to a printed circuit board (PCB) within a supernode server that carries server nodes and their related circuitry. It can be a large backplane or carrier board that integrates multiple nodes. In other words, multiple server nodes are soldered or plugged into the same physical PCB and directly connected to the management bus via internal wiring, without the need for external cables.

[0047] Channel selection chip can refer to a hardware multiplexer (such as...) Multiplexer / demultiplexer (e.g., PCA9548) can be connected to a single extended communication bus interface of the main BMC and branch out multiple downstream channels. By writing control words to this chip, the main BMC can selectively connect the bus to any downstream channel, thereby establishing communication with multiple target board controllers in sequence, avoiding bus address conflicts and expanding the fan-out capability of the management network.

[0048] An extended communication bus can refer to multiple communication lines extended through a channel selection chip. Bus links, these buses originate from the main board controller, pass through the channel selection chip, and then branch off to the respective target board controllers. Each extension The bus has independent electrical paths, allowing the master controller to communicate with multiple slave devices (such as the BMCs of other nodes) at different time slices, thereby constructing a star or tree-type board-level management network.

[0049] CPLD registers refer to the storage units within a Complex Programmable Logic Device (CPLD) integrated on a server node. They are used to store node hardware configuration information, such as slot numbers, primary / secondary role flags, hardware version numbers, and production / test bits. The baseboard controller communicates via a low-speed bus (e.g., ...). (or LPC) reads these read-only or read-write registers to obtain immutable underlying identity and status information, which serves as the basis for node startup, fault determination, and bus arbitration.

[0050] In existing technologies, communication between boards (nodes) is a common scenario in servers or heterogeneous computer systems. With the increasing demand for full-rack systems, communication between identical boards will become the norm. However, online problem localization remains a significant challenge. For example, in a supernode server, multiple identical boards are typically inserted, each equipped with an external communication connector bearing a corresponding node ID. This ID determines whether the current board is the master node. The problem is: when one node fails, how to quickly locate the faulty node online? This remains an unsolved issue.

[0051] The node fault location method provided in this application embodiment can utilize the BMC communication mechanism between the primary node and secondary nodes, combined with... Bus expansion and register read technologies enable active polling and synchronization of status information between nodes, thereby quickly locating faulty nodes.

[0052] The specific application scenarios of this application can be applied to supernode server architectures, with typical scenarios including AI training servers, distributed storage servers, and high-performance computing clusters. In a supernode server, multiple nodes with identical functions (such as GPU cards and storage boards) are connected via a backplane or interconnect bus to form a unified computing or storage unit. For example, in an AI training scenario, eight nodes connect via high-speed... The system uses a bus interconnect, with the master node responsible for coordinating task allocation and status monitoring. Within the storage server, redundancy is used between nodes to improve data reliability. This technical solution utilizes the master node's BMC's active polling mechanism to achieve real-time monitoring of the health status of secondary nodes, thereby quickly locating faulty nodes and triggering alarms or repair processes.

[0053] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.

[0054] Figure 1 This is a flowchart illustrating the node fault location method provided in this application, as shown below. Figure 1 As shown, the method is applied to a first baseboard controller in a first server node, and the method includes:

[0055] S101. Establish a communication connection with the second baseboard controller and read the second information from the second baseboard controller; the second baseboard controller is the baseboard controller of the second server node, the first server node and the second server node are server nodes in the supernode server, and the first server node and the second server node are on the same board, and the second information represents the status information of the second server node.

[0056] The first server node and the second server node can refer to two independent and physically separated server nodes in the supernode server. Each of them has its own independent processor, memory, storage and operating system, and can be tightly coupled to the same supernode server through a high-speed backplane or internal interconnection network.

[0057] Having the first server node and the second server node on the same board means that the two server nodes are integrated on the same physical printed circuit board. This can reduce communication latency between nodes, reduce connector failure points, and facilitate board-level management.

[0058] The communication connection between the first baseboard controller and the second baseboard controller can use the low-speed, high-reliability management bus inside the board, for example, Inter-Integrated Circuit (IBC), PMBus (Power Management Bus), LPC (LowPin Count), or dedicated sideband management channels (such as NC-SI, RBT (Remote Block Transfer)). In some embodiments, the first baseboard controller, acting as the master node, can send read requests to the second baseboard controller via the bus. The second baseboard controller, acting as the slave node, responds and returns the health status data it has collected. This connection method is independent of the node's main CPU or operating system. Even if the CPU of the second server node completely fails, as long as the second baseboard controller is still powered by a backup power supply, the first baseboard controller can still read the second information through this connection.

[0059] The second information can be a set of monitoring data characterizing the health status of the second server node, which can be continuously collected by the second baseboard controller and stored in its own registers or shared memory. The second information includes, but is not limited to: the temperature, voltage, power consumption, fan speed, memory error count, PCIe link status, operating system heartbeat signal, BIOS self-test code, and various sensor threshold alarms of the second node CPU.

[0060] In some embodiments, the first baseboard controller can be accessed via a pre-established management bus (e.g., The first baseboard controller initiates a read operation according to the agreed communication protocol. For example, the first baseboard controller first addresses the second baseboard controller via the bus address, and then sends a command to read the second information-specific register or data block (such as the "Get Sensor Reading" command of the IPMI (Intelligent Platform Management Interface) protocol or a custom private command). After receiving the command, the second baseboard controller retrieves the latest health status data from its internal memory and sends it back to the first baseboard controller via the bus.

[0061] S102. Based on the second information, determine and report the fault status of the second server node.

[0062] After reading the second information, the first baseboard controller can compare the real-time data in the second information with preset health thresholds or normal operating status models item by item. Once it finds that a certain indicator exceeds the tolerance range (such as temperature exceeding 85℃, voltage dropping below 90% of the rated value, or three consecutive heartbeat timeouts), it can determine that the second server node is in a specific fault state (such as overheating degradation, power supply abnormality, system hang). Subsequently, the first baseboard controller actively reports the fault state to the chassis management module or the upper-level network management system through its own management network interface (such as the BMC's web interface, SNMP Trap, or Redfish events), and records the fault type, occurrence time, and key on-site data in the local log for maintenance personnel to locate and repair later.

[0063] Therefore, the node fault location method provided in this application embodiment can realize the problem location and fault diagnosis of each node through mutual communication between the baseboard controllers.

[0064] Optionally, a communication connection is established with the second baseboard controller, and second information is read from the second baseboard controller, including:

[0065] Identify and connect the target extended communication bus for connection with the second baseboard controller;

[0066] Second information is read from the second baseboard controller via the target extended communication bus.

[0067] The target extended communication bus may refer to the extended communication bus in the first baseboard controller used to connect with the second baseboard controller.

[0068] In some embodiments, the first substrate controller can be based on a channel selection chip (such as...) The multiplexer or PCIe switching chip dynamically selects the target bus that is actually connected to the second baseboard controller from multiple candidate extended communication buses, and electrically connects the two through the switching logic inside the chip.

[0069] Therefore, the first baseboard controller, acting as the master device, sends a request to read the second information to the slave device address of the second baseboard controller according to the bus protocol. After receiving the command, the second baseboard controller transmits the node health data (such as temperature, voltage, and fault codes) collected and stored internally back to the first baseboard controller through the same bus, thereby completing the operation of reading the second information.

[0070] Optionally, after establishing a communication connection with the second substrate controller and reading the second information from the second substrate controller, the method further includes:

[0071] Write first information to the second baseboard controller so that the second baseboard controller can determine and report the fault status of the first server node based on the first information; the first information is characterized as the status information of the first server node.

[0072] In addition to obtaining the second information of the second server node from the second baseboard controller to determine the fault status of the second server node, the first baseboard controller can also send the first information of the first server node to the second baseboard controller so that the second baseboard controller can judge the health status of the first server node based on the first information.

[0073] Therefore, when the first baseboard controller is unable to report its node status to the external management platform due to firmware abnormalities, network interface failures, or power fluctuations, the second baseboard controller can use the received first information to complete a health assessment of the first server node and trigger alarms or perform node reset operations on its behalf. This effectively avoids fault omissions caused by misjudgments or communication blind spots of a single baseboard controller, ensuring that any abnormality of any node can be promptly detected and reported by the peer node. This provides management redundancy for the supernode server without single points of failure and significantly reduces the risk of downtime.

[0074] Optionally, when the first server node is the master node, the method further includes:

[0075] A communication connection with the third baseboard controller is established by extending the communication bus. The third baseboard controller is the baseboard controller of the third server node, and the third server node is the server node of the secondary node in the super node server.

[0076] The first information of the first server node is sequentially written into the third baseboard controller, and the third information is read from the third baseboard controller. The third information represents the status information of the third server node.

[0077] The third baseboard controller can be any baseboard controller other than the first baseboard controller in the supernode server. In some embodiments, the second baseboard controller and the third baseboard controller can be the same baseboard controller or different baseboard controllers.

[0078] When the primary server node is the master node, it can write primary information to and read secondary information from each of the third-level baseboard controllers. This allows the master node to centrally acquire health data (third information) from all secondary nodes, enabling a unified assessment of the cluster status and reporting to higher layers. Simultaneously, each secondary node is empowered to independently determine the master node's health status using the received primary information. For example, by analyzing the master node's temperature, heartbeat, or voltage, a secondary node can proactively trigger an alarm or initiate a master-slave switchover request if it detects an anomaly in the master node (such as a heartbeat timeout or temperature exceeding limits). This leverages the advantages of centralized management and aggregation by the master node while preventing it from becoming a single point of failure. It ensures that any anomaly in any node within the supernode can be promptly detected by other nodes, thereby improving the overall robustness and high availability of the supernode server's fault detection.

[0079] In some embodiments, when there are multiple third baseboard controllers, that is, multiple corresponding third server nodes, the first baseboard controller can establish communication connections with the third baseboard controllers respectively through an extended communication bus, and then write the first information of the first server node into the third baseboard controller in sequence, and read the third information from the third baseboard controller.

[0080] In this embodiment of the application, the method further includes:

[0081] Read the ID information from the CPLD register; the ID information is the information sent to the CPLD register by the identity setting interface.

[0082] Based on the ID information, the first server node is determined to be the master node.

[0083] In determining whether the first server node is a master node, the first baseboard controller accesses the internal register address space of the CPLD (Complex Programmable Logic Device) via the low-speed management bus on the board to read pre-programmed node identification information (such as slot number, role type, or master / slave flag). This ID information can be hardened into the CPLD's read-only register by hardware jumpers, DIP switches, or firmware during system startup. The controller obtains the value of this register by sending a read command, and then compares it with preset judgment rules (such as ID=0x01 indicating a master node, ID=0x02 indicating a slave node) to determine whether the node should assume the master or slave role, thereby determining the subsequent read / write permissions and monitoring behavior on the management bus. In this embodiment, when the second server node is a slave node, the method further includes:

[0084] A communication connection with the third baseboard controller is established by extending the communication bus;

[0085] The second information of the second server node is written sequentially into the third baseboard controller, and the third information is read from the third baseboard controller.

[0086] When the second server node acts as a secondary node, it can also proactively initiate communication connections with other third baseboard controllers (i.e., other secondary nodes) within the cluster via the extended communication bus. After establishing a connection, the second node sequentially writes its second information into the designated storage area of ​​each third baseboard controller, allowing other secondary nodes to monitor its health status in real time. Simultaneously, the second node also reads the corresponding third information from each of these third baseboard controllers. Through this bidirectional information exchange between secondary nodes, even if the primary node temporarily loses connection or malfunctions, the secondary nodes can still maintain local cross-monitoring and fault awareness, providing a data foundation for subsequent autonomous alarms or temporary arbitration, and enhancing the robustness of the entire supernode in the event of partial management link failures.

[0087] Optionally, after establishing a communication connection with the second substrate controller and reading the second information from the second substrate controller, the method further includes:

[0088] The second information is sent to the fourth baseboard controller, which is the baseboard controller of the fourth server node. The fourth server node is the other server node in the supernode server besides the second server node.

[0089] The fourth baseboard controller can be the baseboard controller of the fourth server node, and the fourth baseboard controller can be a different baseboard controller from the second baseboard controller.

[0090] After successfully reading the second information representing the health status of the second server node from the second baseboard controller, the information is proactively sent to the fourth baseboard controller. This enables distributed sharing and redundant backup of health status across the cluster. In this way, even if the first baseboard controller (master node) subsequently fails or the management link is interrupted, the fourth baseboard controller can still independently determine the operating status of the second server node based on the received second information, and trigger alarms or perform isolation operations on its behalf. This avoids information silos caused by single points of failure and enhances the overall supernode's rapid response capability and reliability to abnormal events.

[0091] Optionally, the second information is information read by the second baseboard controller from the CPLD register in the second server node;

[0092] And / or,

[0093] The second information includes at least one of temperature parameter information, voltage parameter information, and operating status parameter information.

[0094] In the second server node, the CPLD register serves as the central hub for board-level management. Its internal registers contain static configuration information determined by hardware jumpers, DIP switches, or firmware during node startup, such as node slot number, hardware version, vendor ID, and primary / secondary role flags. The second baseboard controller can read this low-level identity and status data from the CPLD register via the low-speed management bus to obtain the hardware identity fingerprint and basic operating mode of the second server node.

[0095] Temperature parameters reflect the operating temperatures of critical heat sources such as the CPU, memory, and motherboard in the second server node, used to determine if there is an overheating risk or heat dissipation failure. Voltage parameters include the deviations between the current values ​​of each voltage rail, such as the core power supply and auxiliary power supply, and preset thresholds, used to detect power supply anomalies or voltage drops. Operating status parameters cover dynamic operating indicators such as CPU load, memory error count, operating system heartbeat, watchdog timeout flag, and PCIe link status, used to assess whether the node is in a normal, degraded, or suspended state. These three types of information comprehensively characterize the node's health level from three dimensions: thermal, electrical, and logical. An abnormality in any of these parameters may trigger a fault alarm, thereby achieving comprehensive and multi-layered real-time monitoring of the second server node.

[0096] In some embodiments, when determining the fault state based on three types of information—temperature, voltage, and operating status—the first baseboard controller compares the real-time collected values ​​with preset threshold ranges or normal behavior models: if any temperature sensor reading continuously exceeds the high-temperature alarm line (e.g., CPU exceeding 95°C), it is determined to be an overheating fault; if any voltage rail deviates from the rated value by more than the allowable deviation (e.g., ±5%), it is determined to be a power supply abnormality; if the operating status parameters show signs such as operating system heartbeat stoppage, watchdog timeout, uncorrectable memory error, or PCIe link disconnection, it is determined to be a system hang or hardware failure. When multiple dimensions are abnormal simultaneously (e.g., excessively high temperature accompanied by voltage drop), the fault chain can be further inferred (e.g., heat dissipation failure leading to overheating and triggering voltage reduction protection), thereby comprehensively outputting a clear fault type.

[0097] Optionally, a communication connection is established with the second baseboard controller, and second information is read from the second baseboard controller, including:

[0098] Based on a preset polling frequency, a communication connection is established with the second baseboard controller, and second information is read from the second baseboard controller; wherein, the polling frequency is adjusted according to the load status of the server nodes in the supernode server.

[0099] During initialization, the super server node can be configured with a default polling frequency (e.g., once every 2 seconds). The first baseboard controller actively initiates a bus connection establishment process with the second baseboard controller at this fixed time interval, and reads the second information after completing the handshake. In some embodiments, the preset polling frequency can be arbitrarily set according to user needs, such as once every 2 seconds, once every 1 second, or twice every 1 second.

[0100] When the computational load of server nodes within a supernode increases (e.g., CPU utilization exceeds 80% or memory bandwidth is tight), the polling frequency can be dynamically reduced (e.g., extended from 2 seconds to 10 seconds) to reduce potential interference from management bus access to the main business data channel and reduce the CPU interrupt usage of the baseboard controller. Conversely, when a node is idle or lightly loaded, or when a node's health indicator is detected to be close to the threshold edge, the polling frequency can be increased (e.g., shortened to 500 milliseconds) to achieve a rapid response to potential faults.

[0101] In some embodiments, in addition to detection according to a preset polling frequency, the first baseboard controller can also perform on-demand actions when it detects its own anomaly or receives a trigger signal. For example, when the first baseboard controller detects a sharp rise in the CPU temperature or a drastic fluctuation in the core voltage of its node, it will immediately initiate a non-periodic read operation on the second baseboard controller to confirm whether the second node has also encountered a similar power supply or heat dissipation failure, thereby determining whether the fault is local or global. Alternatively, when the first baseboard controller receives an immediate diagnostic command (such as an OEM command via IPMI) from the chassis management module or the upper-layer network management system, and detects a hardware interrupt signal output by the second server node (such as the ALERT# pin being pulled low), it will also temporarily break the preset polling interval and read the second information in real time. This can shorten the fault confirmation delay in critical anomaly scenarios and avoid missing the best handling opportunity due to waiting for the next polling cycle.

[0102] Figure 2 Module illustration of the supernode server provided in this application Figure 1 ,like Figure 2 As shown, the supernode server includes a first baseboard controller and a second baseboard controller. The first baseboard controller is installed on a first server node, and the second baseboard controller is installed on a second server node. The first server node and the second server node are server nodes on the same board in the supernode server; wherein,

[0103] The second baseboard controller is used to acquire second information characterizing the health status of the second server node;

[0104] The first baseboard controller is used to establish a communication connection with the second baseboard controller to read second information from the second baseboard controller and determine and report the fault status of the second server node based on the second information.

[0105] On the same board of the supernode server, a first server node and a second server node are installed, each equipped with its own baseboard controller. The second baseboard controller is responsible for collecting and storing the status information (i.e., secondary information, such as temperature, voltage, and operating status) of the second server node, acting as its local management agent. The first baseboard controller, as the active monitoring party, can establish a communication connection with the second baseboard controller according to a preset strategy (such as polling or event triggering) to read the collected secondary information from the second baseboard controller. Then, the first baseboard controller analyzes the secondary information (e.g., comparing it with thresholds, determining if the heartbeat has timed out) to determine if the second server node has malfunctioned and the specific type of malfunction. Finally, it reports the malfunction status to the upper-level network management system or chassis management module through its own management network interface, completing a closed-loop health monitoring system across nodes.

[0106] In one possible implementation, the first baseboard controller can also be used to acquire first information characterizing the health status of the first server node;

[0107] The second baseboard controller can also be used to establish a communication connection with the first baseboard controller to read first information from the first baseboard controller and, based on the first information, determine and report the fault status of the first server node.

[0108] The supernode server provided in this embodiment can execute the methods provided in the above method embodiments. Its implementation principle and technical effect are similar, and will not be described in detail here.

[0109] Figure 3 Module illustration of the supernode server provided in this application Figure 2 ,like Figure 3 As shown, the supernode server includes 8 nodes (i.e., server nodes) and corresponding input / output (IO) boards for each node. Each node includes a BMC (Baseboard Controller), a CPLD (CPLD Register), a Node ID connector, and a connector. Each IO board includes a connector and expansion slots. Channel selection chip (i.e., channel selection chip). Among them...

[0110] The CPLD can obtain the node ID through the Node ID connector to determine whether it is a master or slave node. It can also obtain temperature, voltage, and operating status parameters. This allows the BMC to obtain common node information (status information). During fault location, the BMCs between nodes can be extended... The channel selection chip and connector are connected to enable the BMC to determine the fault status of other nodes.

[0111] In some embodiments, the process may be as follows:

[0112] 1. BMC reads CPLD to determine whether it is currently a primary node or a secondary node;

[0113] 2. If the BMC confirms that it is currently the master node, the BMC will open the corresponding extension through the connector. The channel selection chip is used, and the master node BMC polls the other nodes. At the same time, the master node BMC reads and writes public information to the other node BMCs.

[0114] 3. Following the procedure in step 2, the BMC then opens the corresponding extension via the connector. The channel selects the chip, polls other secondary nodes, and simultaneously reads and writes public information to other node BMCs.

[0115] 4. If node 2 fails, the BMC of node 1 can obtain information about node 2, and node 1 can report the failure of node 2; conversely, node 2 can also check whether node 1 has failed.

[0116] Figure 4 This is a schematic diagram of the node structure provided in the embodiments of this application, such as... Figure 4 As shown in the figure, this illustrates a board communication architecture based on the collaboration between the IO board and the Base board. The IO board receives 8 channels of signals via an external connector. After selection by an extended communication channel chip, it connects to the Base board's SLIMSAS connector via a cable. The Base Board's BMC communicates with the IO board via the IPMB protocol. Simultaneously, the BMC communicates with the Base board via… Connect the CPLD, which reads the node ID information from the Node ID connector to determine whether the board is the master node.

[0117] In some embodiments, the extended communication channel chip can also be replaced with Channel switching chip.

[0118] Its working principle is as follows: After the BMC of the master node board determines its own role based on the Node ID, it polls and opens the extended communication channel chip to access the other 7 slave nodes to realize the reading and writing of public information; the slave nodes open the extended communication channel chip to access the status information of other nodes, thereby realizing efficient communication and management between nodes in the multi-board scenario of supernode server through node ID identification and channel control.

[0119] Figure 5 The diagram illustrates the working principle of communication between nodes in the embodiments of this application, as follows: Figure 5 As shown, the cable originates from the connector on the first I / O board and, following the path indicated by the arrow, sequentially connects the connectors of multiple subsequent I / O boards, forming a chain connection. This cascading method enables signal transmission and communication between multiple I / O boards, ensuring data interaction and control command transmission between nodes.

[0120] The supernode server provided in this embodiment utilizes the BMC on each node. The communication method enables the location of problems and troubleshooting of faults at each node.

[0121] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.

[0122] Finally, it should be noted that other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This invention is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein, and is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.

Claims

1. A method for locating node faults, characterized in that, The method, applied to a first baseboard controller in a first server node, includes: Establish a communication connection with the second baseboard controller and read second information from the second baseboard controller; the second baseboard controller is the baseboard controller of the second server node, the first server node and the second server node are server nodes in the supernode server, and the first server node and the second server node are on the same board, and the second information represents the status information of the second server node; Based on the second information, determine and report the fault status of the second server node.

2. The method of claim 1, wherein, The establishment of a communication connection with the second baseboard controller and the reading of second information from the second baseboard controller include: Identify and connect the target extended communication bus for connection with the second baseboard controller; The second information is read from the second baseboard controller via the target extended communication bus.

3. The method of claim 1, wherein, After establishing a communication connection with the second substrate controller and reading second information from the second substrate controller, the method further includes: Write first information to the second baseboard controller so that the second baseboard controller can determine and report the fault status of the first server node based on the first information; the first information is characterized as the status information of the first server node.

4. The method of claim 1, wherein, When the first server node is the master node, the method further includes: A communication connection with the third baseboard controller is established by extending the communication bus. The third baseboard controller is the baseboard controller of the third server node, and the third server node is the server node of the secondary node in the super node server. The first information of the first server node is sequentially written into the third baseboard controller, and the third information is read from the third baseboard controller, wherein the third information represents the status information of the third server node.

5. The method of claim 4, wherein, The method further includes: Read the ID information from the CPLD register; the ID information is the information sent to the CPLD register by the identity setting interface. Based on the ID information, the first server node is determined to be the master node.

6. The method of claim 4, wherein, When the second server node is a secondary node, the method further includes: A communication connection with the third baseboard controller is established by extending the communication bus; The second information of the second server node is sequentially written into the third baseboard controller, and the third information is read from the third baseboard controller.

7. The method of claim 1, wherein, After establishing a communication connection with the second substrate controller and reading second information from the second substrate controller, the method further includes: The second information is sent to the fourth baseboard controller, which is the baseboard controller of the fourth server node, and the fourth server node is another server node in the supernode server besides the second server node.

8. The method according to any one of claims 1 to 7, characterized in that, The second information is the information read by the second baseboard controller from the CPLD register in the second server node; And / or, The second information includes at least one of temperature parameter information, voltage parameter information, and operating status parameter information.

9. The method according to any one of claims 1 to 7, characterized in that, The establishment of a communication connection with the second baseboard controller and the reading of second information from the second baseboard controller include: Based on a preset polling frequency, a communication connection is established with the second baseboard controller, and second information is read from the second baseboard controller; wherein, the polling frequency is adjusted according to the load status of the server nodes in the supernode server.

10. A supernode server, characterized by It includes a first baseboard controller and a second baseboard controller. The first baseboard controller is installed on a first server node, and the second baseboard controller is installed on a second server node. The first server node and the second server node are server nodes on the same board in the supernode server. The second baseboard controller is used to acquire second information characterizing the health status of the second server node; The first baseboard controller is used to establish a communication connection with the second baseboard controller to read the second information from the second baseboard controller and determine and report the fault status of the second server node based on the second information.