Rack server, fault diagnosis method and storage medium for rack server
By introducing pre-installed switching components and control components into the entire cabinet server, automatic diagnosis and information summary of link status data is achieved, and the problem of difficult to locate cable link failures is solved, improving the convenience and accuracy of fault diagnosis.
Patent Information
- Application Number
- CN202510396371.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2045-03-31
AI Technical Summary
It is difficult to directly locate the cable link faults in the entire cabinet server, resulting in difficulty in troubleshooting.
The pre-installed switching components are introduced into the entire cabinet server, and the link status data is obtained through the control components of the functional nodes for troubleshooting, and the fault diagnosis information of other functional nodes is collected through the designated functional nodes to realize the summary of fault information.
It improves the convenience and accuracy of fault diagnosis, reduces the difficulty of fault location, and improves the maintainability and operation efficiency of the server.
Smart Images

Figure CN119902922B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a whole-cabinet server, a fault diagnosis method for the whole-cabinet server, and a storage medium. Background Art
[0002] In the prior art, rack-mounted servers can be equipped with compute nodes and switch nodes, with each compute node interconnected to the switch using cable assemblies. Given that the cables are located inside the rack, the connections are extremely complex, and the links are bidirectional via the switch, making it difficult to directly locate a cable link failure. This demonstrates the difficulty of locating faults in rack-mounted servers. Summary of the Invention
[0003] The present application provides a whole cabinet server, a fault diagnosis method for the whole cabinet server, and a storage medium, so as to at least solve the problem of difficulty in locating faults in the whole cabinet server in the related art.
[0004] The present application provides a whole-cabinet server, comprising: a group of functional nodes, the group of functional nodes comprising computing nodes and switching nodes, the computing nodes and the switching nodes being interconnected via a communication link on a cable assembly, the cable assembly comprising a pre-installed switching component, and each functional node in the group of functional nodes having a communication link with the pre-installed switching component; wherein the control component of each functional node is configured to obtain link status data of each functional node and perform fault diagnosis on the link status data of each functional node to obtain fault diagnosis information of each functional node, wherein the link status data of each functional node is used to represent the link status of the communication link to which each functional node is connected; and the control component of a specified functional node in the group of functional nodes is configured to collect fault diagnosis information of the other functional nodes from the control components of the other functional nodes via the pre-installed switching component, wherein the other functional nodes are functional nodes in the group of functional nodes other than the specified functional node.
[0005] The present application also provides a fault diagnosis method for a whole-cabinet server, wherein a group of functional nodes on the whole-cabinet server includes computing nodes and switching nodes, the computing nodes and the switching nodes are interconnected through a communication link on a cable assembly, the cable assembly includes a preset switching component, and each functional node in the group of functional nodes has a communication link with the preset switching component; the method includes: obtaining link status data of each functional node through the control component of each functional node, and performing fault diagnosis on the link status data of each functional node through the control component of each functional node to obtain fault diagnosis information of each functional node, wherein the link status data of each functional node is used to represent the link status of the communication link to which each functional node is connected; and collecting fault diagnosis information of the other functional nodes from the control components of the designated functional node in the group of functional nodes via the preset switching component, wherein the other functional nodes are functional nodes in the group of functional nodes other than the designated functional node.
[0006] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned methods for diagnosing the fault of a whole-cabinet server are implemented.
[0007] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned methods for diagnosing faults of whole-cabinet servers when executed by a processor.
[0008] According to the present application, for a whole cabinet server, a group of functional nodes thereon include computing nodes and switching nodes, the computing nodes and switching nodes are interconnected through a communication link on a cable assembly, the cable assembly includes a preset switching component, and each functional node in a group of functional nodes has a communication link with the preset switching component; wherein, the control component of each functional node is used to obtain the link status data of each functional node, and perform fault diagnosis on the link status data of each functional node to obtain fault diagnosis information of each functional node, wherein the link status data of each functional node is used to represent the link status of the communication link to which each functional node is connected; the control component of a specified functional node in a group of functional nodes is used to obtain the link status data of each functional node through the preset switching component. A component collects fault diagnosis information of other functional nodes from the control components of other functional nodes, wherein the other functional nodes are functional nodes other than the specified functional node in a group of functional nodes. Since the control component in each functional node can read the link status data in the corresponding node, the faulty communication link can be identified based on the read link status data, and automatic diagnosis of the faulty link can be achieved; and the setting of the preset switching component enables the specified functional node to collect the fault diagnosis information of other functional nodes, and completes the aggregation and integration of the fault diagnosis information. Therefore, the problem of difficulty in fault locating of the whole cabinet server in the related technology can be solved, and the technical effect of improving the convenience of fault diagnosis of the whole cabinet server can be achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0010] Figure 1 A schematic structural diagram of an optional whole-cabinet server provided in an embodiment of the present application.
[0011] Figure 2 A schematic structural diagram of another optional whole-cabinet server provided in an embodiment of the present application.
[0012] Figure 3 A schematic structural diagram of another optional whole-cabinet server provided in an embodiment of the present application.
[0013] Figure 4 A schematic structural diagram of another optional whole-cabinet server provided in an embodiment of the present application.
[0014] Figure 5 A schematic structural diagram of another optional whole-cabinet server provided in an embodiment of the present application.
[0015] Figure 6 A schematic structural diagram of another optional whole-cabinet server provided in an embodiment of the present application.
[0016] Figure 7 A schematic structural diagram of another optional whole-cabinet server provided in an embodiment of the present application.
[0017] Figure 8 A schematic structural diagram of another optional whole-cabinet server provided in an embodiment of the present application.
[0018] Figure 9 A flowchart of an optional method for diagnosing faults of a whole-cabinet server provided in an embodiment of the present application.
[0019] Figure 10 A structural block diagram of an optional electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0020] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0021] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.
[0022] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0023] According to one aspect of an embodiment of the present application, a whole-rack server is provided. As used below, the term "module" may refer to a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.
[0024] Figure 1 FIG. 1 is a schematic diagram of an optional whole-cabinet server according to an embodiment of the present application. Figure 1As shown in FIG, a whole-rack server 101 includes: a group of functional nodes 102, wherein the group of functional nodes 102 includes a computing node 1021 and a switching node 1022. The computing node 1021 and the switching node 1022 are interconnected via a communication link on a cable assembly 103. The cable assembly 103 includes a pre-installed switching component 1031. Each functional node in the group of functional nodes 102 has a communication link with the pre-installed switching component 1031. The control component of each functional node is configured to obtain link status data of each functional node and perform fault diagnosis on the link status data of each functional node to obtain fault diagnosis information for each functional node. The link status data of each functional node is used to represent the link status of the communication link to which each functional node is connected. The control component of a specified functional node in the group of functional nodes 102 is configured to collect fault diagnosis information of other functional nodes from the control components of other functional nodes via the pre-installed switching component 1031. The other functional nodes are functional nodes in the group of functional nodes 102 other than the specified functional node.
[0025] With the continuous development of artificial intelligence (AI) technology, AI (Artificial Intelligence) model training and inference are gaining increasing attention. As the critical infrastructure supporting AI model training and inference, the next-generation computing cluster architecture faces severe performance challenges. The marginal benefits of individual server processor performance are diminishing, approaching physical limits, making it difficult to meet the continuously growing computing power demands of large AI models. Furthermore, the strategy of horizontally scaling computing clusters by simply increasing the number of servers also encounters efficiency and scalability bottlenecks, facing challenges including cost, data synchronization, and energy consumption, hindering the efficient execution of large-scale parallel computing.
[0026] Against this backdrop, the AI super-node whole-cabinet system emerged. Through an innovative topology architecture, it systematically integrates accelerator modules and switching nodes, comprehensively optimizing aspects such as multi-computing power integration, high-speed advanced interconnection, heat dissipation, high-power density power supply, and whole-cabinet management. It provides a scalable high-bandwidth domain (HBD) super-node system with 32 cards or more to meet the cutting-edge training and inference needs of large models.
[0027] The whole cabinet system may include at least one whole cabinet. In a whole cabinet, multiple functional nodes may be set up. Here, a functional node refers to a physical device capable of performing specific business functions, including but not limited to computing, storage, network communication, and management control, etc. Among them, a computing node may be a physical device with computing functions, which may be a physical device equipped with a high-performance processor, used to perform complex computing tasks, such as AI model training, reasoning, data processing, etc. Each computing node may contain one or more computing units; a switching node may be a physical device with data transmission and network communication functions, which may be a physical device equipped with high-speed network switching components, responsible for establishing and managing data transmission links between computing nodes and between computing nodes and external networks. It can be used to process large amounts of data streams, forward data packets, perform flow control, error detection and repair, etc.
[0028] Optionally, multiple accelerator modules and one or more switching nodes can be installed in a single cabinet. The nodes containing the accelerator modules can be connected to connectors (e.g., high-speed connectors) on the switching nodes via connectors. The accelerator modules can be OAMs (Open Accelerator Modules), which can be GPUs (Graphics Processing Units) computing units. The switching nodes can also be called switching units, switches, etc.
[0029] The node where the accelerator module resides can be a compute node (computing server). Accordingly, a whole cabinet can include one or more compute nodes, and each compute node can include multiple accelerator modules. The connector on a compute node can be called a compute connector, and the connector on a switch node can be called a switch connector. The connection between a compute connector and a switch connector can be one-to-one, many-to-one, or many-to-many. This embodiment does not limit the connection between compute nodes and switch nodes.
[0030] In related technologies, multiple computing nodes (i.e., computing server nodes) equipped with OAM (Open Accelerator Module) cards can be connected through a switch node as a data link relay, using a specific physical connection form to achieve full data interconnection between the OAM cards. The nodes can also be integrated into the same cabinet, providing centralized power supply and centralized heat dissipation, thereby obtaining a whole-cabinet server, such as Figure 2 As shown, it includes 8 computing nodes at the top and bottom, and 12 switching nodes in the middle. High-speed connectors are provided in the computing nodes and the switching nodes. The computing nodes and the switching nodes adopt a certain connection topology and are connected by cables.
[0031] For example, an AI cabinet design is as follows: Figure 3 As shown, the entire cabinet includes N compute nodes and M switching nodes arranged in between, where N and M are positive integers greater than or equal to 1. High-speed connectors are installed in both the compute nodes and the switching nodes, and the compute nodes and switching nodes are connected by cables using a specific connection topology. A single compute node has four OAM2.0 GPUs, and the entire cabinet has 4*N GPUs.
[0032] For a whole cabinet, a distribution of the cabinet front can be as follows Figure 4 As shown, the computing nodes are placed on the upper and lower sides of the cabinet, and the switching node is placed in the middle. Therefore, some computing nodes can be installed in the top area of the entire cabinet, and the other part can be installed in the bottom area of the entire cabinet. This layout is conducive to dispersing the heat of the server and avoiding all computing devices being concentrated in one area, which leads to excessive heat dissipation pressure. At the same time, arranging the switching node in the middle of the cabinet can ensure that the communication distance between all computing nodes and the switching node is close, reduce network delay, make the wiring of power cables more balanced, avoid extra energy consumption and complex layout caused by excessive cables on one side, so that the power supply and communication lines can be more balanced. Figure 4 This represents an optional configuration. The number of switching nodes and compute nodes can be adjusted based on actual needs. For example, if more computing power is required, the number of compute nodes can be increased; if greater network bandwidth is required, the number of switching nodes can be increased. The height of each U-slot (i.e., the location where the node device is placed) is also adjustable. The square brackets that secure each node in the cabinet can be set to one or more optional sizes, such as OU (48mm height), RU (44.45mm height), and SU (46.5mm height).
[0033] In addition to the above configuration, such as Figure 4As shown, the whole cabinet server in this embodiment may also include: a management switch, which can serve as the network management center of the whole cabinet server, for managing all computing nodes, switching nodes and other related equipment inside the cabinet, collecting the operating status and health information of each node in the cabinet, and reporting this information to a remote management platform through a specific management protocol; a cable management tray, for organizing and managing the cables inside the cabinet, ensuring that the cables are laid out reasonably and neatly, preventing the cables from being entangled and causing signal interference or poor heat dissipation, and also facilitating maintenance personnel to access, remove and troubleshoot cables. It can be designed with multiple compartments, ring binders or wire clamps to accommodate PowerShell (power distribution unit) for providing power and monitoring and managing power; RackStiffener (cabinet reinforcement) for enhancing the structural strength and stability of the cabinet; CDU (Cooling Distribution Unit) for distributing coolant.
[0034] One distribution at the back of the cabinet can be as follows Figure 5 As shown, the back of the cabinet includes three components: the busbar, cable tray (a structure used to support and manage cables), and manifold. The busbar is located in the middle or on one side, and power is supplied to the entire cabinet by connecting powershelf nodes ("power supply racks" or "power module racks"). The cable trays are located on the left and right sides, connecting high-speed signals from compute nodes and switching nodes according to the designed interconnection topology. The manifolds are located on the left and right sides of the rear of the cabinet, providing water inlet and outlet, creating a liquid path for liquid cooling throughout the cabinet.
[0035] An interconnection relationship between computing nodes and switching nodes can be as follows Figure 6 As shown, Figure 6The upper area of the cabinet contains N compute nodes (#1 to #N), and the lower area contains M switch nodes. Each compute node has a central processing unit (CPU), which can connect to four OAM 2.0 processors (GPUs) via PCIe (Peripheral Component Interconnect Express) switches. If there are N compute nodes in the cabinet, the total number of GPUs is 4*N. For example, with 16 compute nodes, the total number of GPUs is 64 (16*4). The compute nodes also have data processing unit (DPU) network cards, storage devices, and baseboard management controllers (BMCs). Storage devices can be, but are not limited to, solid-state drives (SSDs). Each switch node has a media access control (MAC) component. High-speed signals between the switch nodes and compute nodes are interconnected via cable trays.
[0036] In AI whole-rack servers, computing nodes and switching nodes can all be connected to RackCable Tray (cable bridges in or near the rack) through high-density connectors, enabling Serdes (Serializer / Deserializer) high-speed interconnection between all GPUs in the AI whole-rack server. Here, Serdes (Serializer / Deserializer) technology is used to convert parallel data into serial data for easy transmission through cables, and then restore it to parallel data at the receiving end, thereby significantly improving the speed and efficiency of data transmission without affecting data integrity. Figure 3 As shown in the figure, the high-speed signal resources of each OAM (GPU) in each computing node can be evenly distributed to each page switching node, realizing a scale-up topology.
[0037] For whole-cabinet servers, a single server requires thousands of cables. Once a fault occurs, it will affect the entire cabinet. Therefore, it is necessary to diagnose the cable backplane fault. However, because the computing nodes and switches are interconnected using a cable tray, the connectors and cables are all located inside the cabinet, and the connection relationship is extremely complex. In addition, because the link transmits bidirectionally through the switch, if a cable link fails, it is difficult to directly locate it. It can only be removed completely and then tested and located one by one. Fault diagnosis in whole-cabinet servers is still not perfect.
[0038] It can be seen that the whole cabinet server in the related art has the problem of difficulty in fault location, and a whole cabinet server with convenient fault diagnosis, high maintainability, strong scalability, high fault location accuracy and fast fault location speed is needed.
[0039] In order to at least partially solve the above technical problems, this embodiment provides a whole cabinet server, including a group of functional nodes, which include computing nodes and switching nodes. The computing nodes and switching nodes are interconnected through a communication link on a cable assembly. The cable assembly includes a pre-set switching component, and each functional node has a communication link with the pre-set switching component; on each functional node, the control component of each functional node performs fault diagnosis on the link status data of each functional node, so as to determine whether the communication link to which each functional node is connected has a fault, thereby realizing automatic fault location; and a designated functional node obtains the fault diagnosis information of other functional nodes from other functional nodes via the pre-set switching component, thereby realizing the aggregation of fault diagnosis information, making it convenient for relevant personnel to obtain fault diagnosis information, while reducing the difficulty of fault location and improving the convenience of fault diagnosis, and improving the convenience of information acquisition.
[0040] Here, a communication link can be a cable on a cable assembly. The link status data for each functional node is used to indicate the link status of the communication link to which each functional node is connected. The communication links to which each functional node is connected can include communication links between the functional node and other functional nodes other than itself, as well as communication links between the functional node and pre-configured switching components. The link status of a communication link refers to the current operating status of the communication link, which may include the integrity and efficiency of received signal transmission, such as the signal bit error rate (BER), latency, jitter, signal strength, signal integrity, and whether the link is active or disconnected. This data can be collected by the corresponding functional node to assess the health of the link.
[0041] Optionally, the link status data of each functional node can be used to indicate the link status of all or part of the communication links to which each functional node is connected, that is, it can respectively include the link status of at least part of the communication links to which the corresponding functional node is connected, for example, including the bit error rate of at least part of the communication links. Here, although the link status data of each functional node can record the status data of a failed communication link, it can indicate that the communication link that is not recorded has not failed. Optionally, each functional node can include a control component, which can be a BMC, that is, a baseboard management controller, for monitoring and managing the health status of the server, including but not limited to temperature, power status, fan speed and link status.
[0042] Optionally, the control component may be used to read data stored in the corresponding node and report the data to the remote management software or the local control center via a communication protocol for fault analysis.
[0043] Optionally, the control component may obtain the link status data of the corresponding node in real time or periodically to ensure the real-time nature and accuracy of the link status data.
[0044] In this embodiment, the control component of each functional node, such as BMC_RT (the control component of a computing node) and BMC_SW (the control component of a switching node), can collect link status data of the communication link to which it is directly connected. Based on the collected link status data, fault diagnosis can be performed to identify possible link problems, such as signal distortion and abnormal bit error rate, and generate fault diagnosis information. For example, the link status information of the corresponding node can be obtained, and based on the bit error rate in the link status information, it can be determined whether the communication link has a fault and which communication link has the fault, so as to obtain fault diagnosis information for the node.
[0045] In this embodiment, the cable assembly includes a preset switching component, and each functional node in a group of functional nodes has a communication link with the preset switching component. The preset switching component can support multi-channel communication and can reasonably distribute and redirect data streams from different functional nodes to ensure efficient data exchange between communication links.
[0046] Due to the characteristics of distributed management, the BMCs of each node do not have direct mutual communication management. One functional node can be designated as a designated functional node among multiple functional nodes, which can be configured as a resident management node.
[0047] Optionally, the designated function node may be a computing node or a switching node.
[0048] Optionally, the other functional nodes are functional nodes other than the designated functional node in a group of functional nodes, and the other functional nodes may be computing nodes or switching nodes.
[0049] In this embodiment, the control component of the designated functional node can be used to collect fault diagnosis information of other functional nodes from the control components of other functional nodes regularly or based on event-driven by pre-set switching components, so as to realize the aggregation and integration of fault diagnosis information.
[0050] For example, Figure 7As shown in the figure, the BMC of the designated functional node (switch node 0) is equipped with an I2C (Inter-Integrated Circuit) switch component board through a cable assembly to collect information from other nodes. The BMC of switch node 0 can read fault diagnosis information from the BMCs of other nodes.
[0051] Optionally, after the fault of the cable assembly is located and diagnosed, the above-mentioned whole cabinet server can be used to generate a corresponding fault diagnosis report and provide processing suggestions. For example, it can diagnose the specific communication link location of the serious signal interruption and recommend triggering the backup path switching or performing physical replacement and repair.
[0052] Through the embodiment provided by the present application, the whole cabinet server includes: a group of functional nodes, a group of functional nodes includes computing nodes and switching nodes, the computing nodes and switching nodes are interconnected through a communication link on a cable assembly, the cable assembly includes a preset switching component, and each functional node in a group of functional nodes has a communication link with the preset switching component; wherein, the control component of each functional node is used to obtain link status data of each functional node, and perform fault diagnosis on the link status data of each functional node to obtain fault diagnosis information of each functional node, wherein the link status data of each functional node is used to represent the link status of the communication link to which each functional node is connected; the control component of a specified functional node in a group of functional nodes is used to collect fault diagnosis information of other functional nodes from the control components of other functional nodes through the preset switching component, wherein the other functional nodes are functional nodes other than the specified functional nodes in a group of functional nodes, which can solve the problem of difficulty in fault locating of the whole cabinet server in the related art and improve the convenience of server fault locating.
[0053] In an exemplary embodiment, a communication link between a computing node and a switching node is connected to a signal channel of the computing node and a signal channel of the switching node; the control component of each functional node is also used to read the link status data stored by a designated component of each functional node to obtain the link status data of each functional node, wherein the link status data of each functional node is used to represent the link status of the communication link to which each functional node is connected in the direction toward each functional node.
[0054] Here, the signal channel refers to a small channel inside the computing node and the switching node for data transmission. It can be a physical pin on an integrated circuit or a circuit path connected through a signal enhancement device such as a retimer. It is determined by the configuration of the computing node or the switching node and is not limited to this in this embodiment.
[0055] Optionally, each functional node (including computing nodes and switching nodes) is equipped with a control component (such as a BMC) for reading link status data stored in a designated component, such as signal strength, bit error rate, link bandwidth usage, signal delay, etc.
[0056] Optionally, the control component of each node may determine whether the link status deviates from a normal range based on a preset threshold and a fault algorithm, so as to obtain fault diagnosis information of each functional node.
[0057] Optionally, the above-mentioned fault diagnosis information may be used to indicate whether a corresponding node has a fault, and may also be used to indicate the specific content of the fault of the corresponding node, for example, whether the node is disconnected.
[0058] Optionally, the link status data of each functional node is used to represent the link status of the communication link to which each functional node is connected in the direction toward each functional node. For example, the link status data of each functional node is used to evaluate the status information of the signal received by the functional node as the receiving end of the signal.
[0059] Optionally, the above-mentioned fault diagnosis information can also be used to indicate the presence of a faulty signal channel on the corresponding node. For example, if the link status data indicates the presence of a fault or anomaly, the control component can determine the type of fault and the signal channel where the fault occurs. For example, if the signal bit error rate received by the 58th signal channel of GPU2 on the computing node is too high, the control component can diagnose that a fault has occurred on this specific link, determine the fault type and the signal channel where the fault occurs, and generate corresponding fault diagnosis information.
[0060] Similar to the above embodiment, the control component may report the generated fault diagnosis information to the control component of the designated functional node through the preset switching component.
[0061] Through this embodiment, fault diagnosis is performed based on the link status data of each functional node, and fault diagnosis information of each functional node is generated, which can improve the real-time and accuracy of fault diagnosis and enhance the maintainability and operating efficiency of the entire cabinet server.
[0062] In an exemplary embodiment, a group of functional nodes includes a first functional node that stores topological data of a cable assembly, where the topological data of the cable assembly is used to describe the topological relationship of the cable assembly; a control component of the first functional node is further used to perform fault diagnosis on the link status data of the first functional node based on the topological relationship of the cable assembly to obtain fault diagnosis information of the first functional node, wherein the fault diagnosis information of the first functional node is used to indicate the signal channels connected to both sides of the first communication link to which the first functional node is connected and where a fault occurs.
[0063] Optionally, the first functional node may be a designated functional node or other functional nodes. Furthermore, it may be a computing node or a switching node, which is not limited in this embodiment.
[0064] In this embodiment, the topological data of the cable assembly is used to describe the topological relationship of the cable assembly, which may include the physical location of each node (such as the U position), the connector type, the cable length, the interconnection relationship of the signal channels, etc., and may be pre-stored in the functional node, for example, it may be stored in the first functional node. Based on the topological data of the cable assembly, the first functional node can obtain the connection relationship of all communication links in the entire cabinet server, including the signal channels corresponding to all communication links. That is, based on the topological data of the cable assembly, the first functional node can construct a virtual connection map of the signal channels in the entire cabinet server, and based on this map, analyze the connectivity and data transmission status of each communication link.
[0065] Optionally, the control component of the first functional node can be used to perform fault diagnosis on the link status data of the first functional node based on the topological relationship of the cable assembly to obtain fault diagnosis information of the first functional node, wherein the fault diagnosis information of the first functional node is used to indicate the signal channels connected to both sides of the first communication link to which the first functional node is connected and which has failed. For example, based on the topological relationship of the cable assembly, the control component of the first functional node can compare the link status data of a specific link with the expected ideal state. If the signal delay of a communication link is found to be abnormal, the communication link fault can be determined, and the signal channel positions on both sides of the faulty link can be accurately located in combination with the topological data of the cable assembly, including the specific faulty computing node position, the faulty signal channel position on the faulty computing node, the faulty switching node position, and the signal channel position on the faulty switching node.
[0066] Through this embodiment, fault diagnosis is performed based on the topological structure of the cable assembly, which can improve the accuracy of cable assembly fault diagnosis and enhance the maintainability of the entire cabinet server.
[0067] In an exemplary embodiment, a first storage component is provided on the first connector of the first functional node; the first storage component stores the topology data of the cable assembly; and the control component of the first functional node is further used to obtain the topology data of the cable assembly stored in the first storage component from the first connector.
[0068] Optionally, the topology data of the cable assembly can be the topology relationship of the Cable Tray pre-stored in the first storage component, which can be burned into the designated storage component during the factory delivery process of the entire cabinet server to provide accurate cable topology information based on the actual hardware configuration without the need to manually record complex cable topology information.
[0069] Optionally, the first storage component can be set on the first connector, and then the control component of the first functional node can directly access the first storage component through the control interface of the connector (such as I2C), and actively read the cable assembly topology data in the first storage component regularly or as needed.
[0070] Optionally, the topology of the cable assembly may be a scale-up topology, a scale-out topology, or other topologies, which are not limited in this embodiment.
[0071] Taking the scale-up topology as an example, scale-up refers to improving the processing power of the entire system by increasing the resources (such as CPU, memory, storage, etc.) of a single node in the computer system architecture. In the design of a whole-cabinet server, scale-up means that the high-speed signal resources of each accelerator module are evenly distributed to each switching node to ensure balanced data exchange and the overall high performance of the whole-cabinet server. In other words, no matter how large the amount of data is, it can be transmitted quickly, and the data transmission speed of the entire cabinet server will not be slowed down by the large amount of data transmission of a certain node.
[0072] In the above-mentioned whole-cabinet server, the control component of the first functional node can directly access the corresponding stored topology data, thereby improving the efficiency and speed of fault diagnosis.
[0073] Optionally, obtaining the topological data of the cable assembly stored in the first storage component can be performed before reading the link status data of the corresponding node, or it can be performed after reading the link status data of the corresponding node. That is, the topological data of the cable assembly can be obtained in advance, and when abnormal data in the link status data is read, the fault source can be quickly located based on the pre-acquired topological data of the cable assembly. Alternatively, the topological data of the cable assembly can be obtained after the abnormal link status data is read for fault diagnosis. The timing of obtaining the topological data of the cable assembly stored in the first storage component can be determined according to the application scenario.
[0074] Through this embodiment, by pre-storing the topology data of the cable assembly, the efficiency and accuracy of the fault diagnosis of the cable assembly can be improved.
[0075] In an exemplary embodiment, the first storage component is a field replaceable unit (FRU). The control component of the first functional node is further configured to read the topology data of the cable assembly stored in the FRU via a communication bus between the control component of the first functional node and the first connector.
[0076] Optionally, based on the topological data of the cable components, a global view including high-speed interconnection mapping relationships can be formed, which can be maintained by the upper-level management software. However, due to the configuration differences of the AI cabinets, this global view may become complex and prone to errors in a cluster environment. That is, the high-speed interconnection mapping relationships are maintained by the upper-level management software, and the high-speed interconnection topology is maintained in the upper-level management software. If there are different types of AI cabinets in the cluster and the interconnections within the cabinets are different, mapping errors are likely to occur.
[0077] In this embodiment, the topology data of the cable assembly may be stored in the first field replaceable unit. This design allows for dynamic adjustment of the stored data according to the actual configuration, thereby reducing mapping errors caused by configuration differences.
[0078] Here, a field-replaceable unit (FRU) refers to a hardware component that can be replaced independently without damaging or affecting other components. In this embodiment, due to the complex and diverse cable connections within the entire cabinet, by storing the topology data of the cable assemblies in the FRUs of the compute nodes, this information can be accessed to adapt to changes in the hardware configuration within the cabinet.
[0079] In order to enable the control component to read data in the field replaceable unit, a standard communication bus can be used for data transmission between the control component on the first functional node and the first connector, such as an I2C bus, an SPI (Serial Peripheral Interface) bus, etc. These bus protocols have the advantages of point-to-point communication, high data transmission rate, and simple control, and are suitable for reading and updating data stored in the first storage component.
[0080] Here, the Inter-Integrated Circuit Bus (I2C bus) is a bidirectional two-wire serial bus for short-distance communication, mainly used for brief digital communication between devices.
[0081] Optionally, the first control component may directly read the topology data of the cable assembly stored in the first field replaceable unit, or read a routing table containing connection information of all communication links between computing nodes and switching nodes generated based on the topology data of the cable assembly.
[0082] For example, the first control component can send a read command to the FRU on the connector of the computing node through the I2C bus. The FRU can respond to the read command and return the stored topology data of the cable assembly to the BMC through the I2C bus. The use of this communication mechanism ensures the reliability and efficiency of reading data, and can also easily read information without disassembly or significantly interfering with system operation.
[0083] According to this embodiment, by storing the topology data of the cable assembly in the first field replaceable unit, the flexibility and durability of data storage can be improved.
[0084] In an exemplary embodiment, a whole cabinet server includes multiple node placement positions, a functional node in a group of functional nodes is placed in a node placement position in the multiple node placement positions; an accelerator module of a computing node has multiple module signal channels, a media access control component of a switching node has multiple component signal channels, a communication link between a computing node and a switching node is connected to a signal channel of the computing node as a module signal channel of the accelerator module, and is connected to a signal channel of the switching node as a component signal channel of the media access control component; in the topology data of the cable assembly, the fields recording each functional node include a node type field and The placement position identification field, the field for recording the computing node also includes the module number field and the module channel number field, and the field for recording the switching node also includes the switching node number field and the component channel number field, wherein the node type field is used to record the node type of the corresponding functional node, the placement position identification field is used to record the placement position identification of the node placement position where the corresponding functional node is placed, the module number field is used to record the module number of the corresponding accelerator module, the module channel number field is used to record the channel number of the corresponding module signal channel, the switching node number field is used to record the node number of the corresponding switching node, and the component channel number field is used to record the channel number of the corresponding component signal channel.
[0085] Here, the media access control component can be used to process the reception and transmission of data packets, perform link layer protocol encapsulation and decapsulation, and ensure the correct transmission of data in the network. In the whole cabinet server, the MAC component can be used to support high-speed data communication between computing nodes; the component signal channel is used for communication between the MAC component and other components. Each channel corresponds to a physical port or a group of logical ports, which can support high-bandwidth data transmission. For example, a 512-channel MAC component can process 512 independent links at the same time, providing low-latency, high-throughput network support for large-scale AI computing clusters.
[0086] In this embodiment, multiple rows of node placement positions with uniform heights can be designed inside a whole cabinet server. The height of each node placement position is a certain number of OUs (i.e., cabinet units, for example, the height of each OU unit is 48 mm). A functional node in a group of functional nodes can be placed in a node placement position among multiple node placement positions. It can be a computing node or a switching node. For example, the 40 OU and 30 OU positions can be used to place computing nodes and switching nodes, respectively.
[0087] Here, a communication link refers to the physical channel connecting the computing node and the switching node. Each communication link is composed of a module signal channel of the accelerator module of the computing node and a component signal channel of the media access control component of the switching node, realizing point-to-point signal transmission.
[0088] Optionally, the accelerator module of the computing node can include multiple module signal channels, and the media access control component of the switching node can have multiple component signal channels. Through a predefined topological relationship, each communication link can accurately connect the module signal channel of the computing node and the corresponding component signal channel of the switching node, ensuring that the signal channels in the accelerator module of the computing node correspond one-to-one with the signal channels of the media access control component of the switching node. For example, GPU2 (having module signal channel #45) on computing node #3 is connected to the MAC component (having component signal channel #312) of switching node #7 via a communication link, realizing directional data transmission.
[0089] Optionally, the topology data of the cable assembly can record the connection relationship between the computing node and the switching node through a series of key fields. These fields may include node type, node position number, module number, module channel number, switching node number, and component channel number. Among them, the node type field is used to record the node type of the corresponding functional node, for example, it can record whether it is a computing node or a switching node; the placement position identification field is used to record the placement position identification of the node placement position where the corresponding functional node is placed; the module number field is used to record the module number of the corresponding accelerator module; the module channel number field: used to record the channel number of the corresponding module signal channel; the switching node number field is used to record the node number of the corresponding switching node; the component channel number field is used to record the channel number of the corresponding component signal channel. Among them, in the whole cabinet server, the aforementioned placement position is called the U position, so the placement position identification field can be called the U position identification field or the U position field.
[0090] Optionally, the topological data of the cable assembly can be stored in the form of a table or database, where each row of data corresponds to a communication link and includes key information such as the node type field, placement position identification field, module number field, module channel number field, switching node number field, and component channel number field.
[0091] Optionally, the value range and encoding scheme of the fields in the topology data of the cable assembly can be pre-set. For whole-cabinet servers with different configurations, different field value ranges can be set. For example, for a whole-cabinet server with 16 computing nodes, each computing node has 4 GPUs, each GPU has 96 signal paths, these signals are connected to 12 switching nodes, each switching node has a MAC component, and each MAC component has 512 signal channels. When setting the value range of its module channel number field, it needs to be able to cover 96 signal paths, and when setting the value range of its component channel number field, it needs to be able to cover 512 signal channels.
[0092] Optionally, the topology data of the cable assembly may further include a check code field for verifying whether the read data is complete.
[0093] Optionally, the value range and coding scheme of the fields in the topology data of the cable assembly constructed based on the above ideas can be as shown in Table 1.
[0094] Table 1
[0095]
[0096] Through this embodiment, link information of the communication link can be quickly acquired based on the fields of the read topology data of the cable assembly, thereby improving the efficiency, accuracy and reliability of fault diagnosis of the cable assembly.
[0097] In an exemplary embodiment, the control component of the other functional node is further used to receive a data acquisition request sent by the specified functional node via the preset switching component; in response to the data acquisition request, a fault message is generated based on the fault diagnosis information of the other functional node, and the fault message is sent to the control component of the specified functional node via the preset switching component; wherein the fault message carries the following information: the node identifier of the other functional node; the channel identifier of the signal channel connected to both sides of the second communication link to which the other functional node is connected and to which the fault occurs; and the fault type of the second communication link.
[0098] Similar to the above embodiments, the control component of a designated functional node may collect data from the control components of other functional nodes through a preset switching component periodically or based on an event trigger (eg, when a signal anomaly or link failure is detected).
[0099] Optionally, the control component of the designated functional node may send a data acquisition request to other functional nodes via the preset switching component. The data acquisition request may include an identifier of the requested functional node and the type of requested information (eg, fault information).
[0100] Optionally, in response to the data acquisition request, the control components of other functional nodes may be configured to directly send the fault diagnosis information to the control component of the designated functional node via the preset switching component.
[0101] Optionally, in response to the data acquisition request, the control components of other functional nodes may also be configured to generate a fault message based on the fault diagnosis information of the other functional nodes, and send the fault message to the control component of the designated functional node via the preset switching component.
[0102] In this embodiment, the fault message may carry the following information: node identifiers of other functional nodes, which are used to clarify the source node of the fault information, facilitating fault location and subsequent processing; channel identifiers of the signal channels on both sides of the second communication link where the fault occurs, including the module signal channel numbers of the accelerator modules on both sides of the faulty communication link and the component signal channel numbers of the media access control components, which provide accurate information on the specific location of the fault; the fault type of the second communication link, which is used to indicate the status of the second communication link, helping to quickly determine the nature of the fault and take targeted repair measures.
[0103] The fault type of the second communication link may be determined based on a preset trigger condition, which may include but is not limited to: speed reduction, disconnection, etc.
[0104] Optionally, when the second communication link does not fail, the control components of other functional nodes can still generate fault messages in response to data acquisition requests, and send the fault messages to the control components of designated functional nodes via preset switching components, wherein the fault type of the second communication link can indicate that the corresponding communication link is normal.
[0105] For example, the link status can be judged based on the bit error rate (BER). The bit error rate (BER) is an important parameter for measuring data transmission quality. It represents the ratio of erroneous bits to the total number of transmitted bits. A low BER indicates good signal transmission quality, while a high BER may mean that the link is faulty or about to fail, requiring timely intervention and repair.
[0106] Optionally, the fault message may further include handling suggestions.
[0107] Optionally, fault diagnosis codes can be preset for generating fault messages for different fault types. For example, as shown in Table 2, the code corresponding to the normal state is 0x00, indicating that the signal integrity meets the standard, the link is in the normal working state, and no intervention is required; the code corresponding to the speed reduction state is 0xFE, indicating that the bit error rate (BER) is higher than normal but has not reached the level of complete interruption, generally occurring when 1e-12 < BER < 1e-9. This means that the signal quality has deteriorated, the data transmission rate has decreased, and performance may be affected. At this time, the whole cabinet server needs to start link retraining to try to restore the normal transmission speed of the link; the code corresponding to the disconnection state is 0xFF, indicating that the bit error rate is extremely high or the physical layer of the link is unlocked, the link is completely interrupted, and data cannot be transmitted. In this case, the whole cabinet server needs to trigger the switching of the standby path and perform physical replacement and repair because the link cannot work properly and immediate measures need to be taken to avoid more serious impacts.
[0108] Table 2
[0109]
[0110] Optionally, the fault diagnosis codes in the above-mentioned whole cabinet server can be modified as needed. For example, new fault types and corresponding fault codes can be added, or new trigger conditions can be added for a certain fault type.
[0111] Optionally, the node identifiers of other functional nodes and the channel identifiers of the signal channels on both sides of the second communication link where a fault occurs in the above-mentioned fault message can be generated based on the coding scheme of the topology data of the cable assembly in the foregoing embodiments. For example, a fault message can be generated: 0x01 0x12 0x2 0x3A 0x7 0x01FE 0xFE; correspondingly, this fault message can indicate that the computing node (0x01) is at the 18U position (0x12), and the 58th path (0x3A) of GPU2 (0x2) to the 510th channel (0x01FE) of the switching node 7 (0x7) has a speed reduction (0xFE).
[0112] Through this embodiment, fault messages can be generated based on fault diagnosis information and aggregated to a specified functional node, which can improve the accuracy and reliability of fault diagnosis of the cable assembly.
[0113] In an exemplary embodiment, the whole cabinet server includes multiple node placement positions, and one functional node in a group of functional nodes is placed on one of the multiple node placement positions; the node identifiers of other functional nodes include the placement position identifiers of the node placement positions where the other functional nodes are placed.
[0114] In order to achieve efficient and orderly internal network communication and maintenance work, multiple standardized node placement positions can be designed inside the entire cabinet server. The height of each placement position is marked in U (cabinet unit, taking OU as an example, 1OU=48mm).
[0115] Optionally, a functional node in a group of functional nodes is placed in a node placement position among a plurality of node placement positions, and the functional node may be physically positioned based on the node placement position identifier.
[0116] Optionally, the node identifier of other functional nodes includes not only their node type (such as computing node, switching node, etc.), but also the placement position identifier of the node placement position where the other functional nodes are placed, so as to achieve accurate physical positioning of the nodes. For example, the node identifier of a functional node may include "38OU", indicating that this is a functional node placed in a node placement position at a height of 38 OU.
[0117] Furthermore, based on the placement position identifier in the functional node's node identifier, the specific node placement position can be located, identifying the functional nodes at both ends of the communication link where the fault occurred. The addition of the placement position identifier allows fault location to be directly pinpointed to the specific location in physical space, not just the logical node type and number, improving the efficiency and accuracy of fault location.
[0118] Through this embodiment, accurate physical positioning of the functional node is achieved by locating the node placement position, thereby improving the efficiency and accuracy of fault positioning.
[0119] In one exemplary embodiment, the accelerator module of a compute node is connected to the connector of the compute node via the compute node's retimer, forming multiple module signal channels of the accelerator module. A communication link between the compute node and the switch node connected to a signal channel of the compute node constitutes a module signal channel of the accelerator module. Correspondingly, the control component of the compute node is further configured to read link status data stored in a first register of the retimer to obtain the link status data of the compute node.
[0120] To ensure the transmission quality of high-speed signals in connectors and cable assemblies, the accelerator module can be connected to the connector through a retimer to form multiple module signal channels of the accelerator module. Here, the retimer can receive the high-speed signal sent by the accelerator module. During the high-speed signal transmission process, the signal quality will degrade due to reflection, crosstalk, etc., resulting in signal errors at the receiving end. The retimer can detect the signal changes and enhance the signal quality through adaptive signal shaping technology to ensure the integrity of the signal after long-distance transmission. In addition, it can also adjust the signal time to ensure signal synchronization at the receiving end, reduce signal attenuation and crosstalk, and ensure the integrity of the signal during long-distance transmission.
[0121] In the above-mentioned whole-cabinet server, to achieve high-density and high-performance computing, the multiple module signal channels on each accelerator module can establish high-speed interconnection with the switching node through a cable assembly. In this connection mode, each module signal channel of the accelerator module directly corresponds to a specific communication link through the connector of the computing node, ensuring point-to-point transmission of the signal.
[0122] Optionally, the retimer can include a corresponding monitoring mechanism to monitor the status of the module signal channel connected to it in real time and obtain the corresponding cable link status data, including but not limited to: signal strength, bit error rate, signal integrity, and link activity status. This status information can be stored in its internal registers and can be easily queried by the control component of the computing node at any time.
[0123] Through this embodiment, the link status data stored in the retimer of the computing node is read by the control component of the computing node, which can improve the real-time performance and accuracy of fault diagnosis.
[0124] In an exemplary embodiment, the accelerator module of the computing node is connected to the connector of the computing node to form multiple module signal channels of the accelerator module. A communication link between the computing node and the switching node is connected to a signal channel of the computing node, which is a module signal channel of the accelerator module. Correspondingly, the control component of the computing node is also used to read the link status data stored in the accelerator module to obtain the link status data of the computing node.
[0125] In this embodiment, the control component of the computing node can also directly read the link status data stored in the accelerator module, thereby obtaining real-time signal link status feedback, reducing the level and time of data transmission, and simplifying the fault diagnosis process of the cable assembly.
[0126] Optionally, the control component of the computing node may include a preset management interface protocol for directly communicating with the accelerator module.
[0127] Optionally, a register may also be provided in the accelerator module for storing link status data, which can be read by the control component of the computing node through a preset management interface protocol.
[0128] For example, in a computing node, the accelerator module can store link status data in its internal link status register, and the control component (BMC) of the computing node can read the data in the link status register of the accelerator module directly through the communication link periodically or under event-driven conditions.
[0129] Through this embodiment, the link status data stored in the accelerator module is directly read by the control component of the computing node, which can improve the efficiency and accuracy of the fault diagnosis of the cable assembly.
[0130] In an exemplary embodiment, the switching node includes an Ethernet switching component. The control component of the switching node is further configured to read the link status data stored in the second register on the Ethernet switching component to obtain the link status data of the switching node.
[0131] An Ethernet switch is an integrated circuit that implements Ethernet switching functionality, processing and forwarding high-speed network data packets. In a rack-mount server, it can store real-time status information about connected links through registers, including but not limited to link activation status, data transmission rate, and bit error rate.
[0132] Optionally, the control component of the switching node can be used to read the link status data stored in the second register on the Ethernet switching component, so as to monitor the status of all connected links in real time and identify failed links or links that are about to fail.
[0133] Through this embodiment, the link status data stored in the Ethernet switching component is read by the control component of the switching node, thereby improving the reliability of the entire cabinet server.
[0134] In an exemplary embodiment, the control component of each functional node is further configured to read the link status data stored in the designated component of each functional node via the management data input / output interface of each functional node to obtain the link status data of each functional node.
[0135] The Management-Data-Input / Output (MDIO) interface is a standard management interface used to connect management controllers (such as BMCs) and network devices (such as retimers). It is widely used in the configuration, monitoring, and maintenance of network devices. In the aforementioned rack-mount server, each accelerator module (such as a GPU) can be connected to the cable tray through a retimer, and the retimer supports the MDIO interface for link status monitoring.
[0136] Optionally, the control component of each functional node may directly access a designated component of each functional node through the MDIO interface to read the link status data of the corresponding node.
[0137] For example, a BMC of a computing node may read link status data stored in a Retimer of the computing node through an MDIO interface, or a BMC of a switching node may read link status data stored in a second register of an Ethernet switching component of the switching node through an MDIO interface.
[0138] According to this embodiment, the reliability of fault diagnosis of the cable assembly can be improved by reading the first link status data stored in the first designated component through the management input / output interface.
[0139] In an exemplary embodiment, the accelerator module of the computing node has multiple module signal channels, the media access control component of the switching node has multiple component signal channels, and a communication link between the computing node and the switching node is connected to a signal channel of the computing node, which is a module signal channel of the accelerator module, and is connected to a signal channel of the switching node, which is a component signal channel of the media access control component.
[0140] Similar to the aforementioned embodiment, each computing node may include an accelerator module, which may have multiple high-speed signal transmission channels, i.e., module signal channels; each switching node may include a media access control component, and the MAC component may have multiple component signal channels.
[0141] In the communication architecture of the whole cabinet server, each communication link between the computing node and the switching node can be connected to a computing node and a switching node. In some embodiments, it can be connected to a module signal channel of the accelerator module of the computing node and a component signal channel of a media access control component of the switching node. This one-to-one mapping relationship can ensure the accurate transmission of data between the computing node and the switching node, and can also accurately locate the fault point when performing cable assembly fault diagnosis.
[0142] Through this embodiment, through a one-to-one connection between a module signal channel of a computing node and a component signal channel of a MAC component of a switching node, the signal transmission efficiency of the entire cabinet server can be improved, and the efficiency and accuracy of fault diagnosis of the cable assembly can be improved.
[0143] In an exemplary embodiment, a communication link between a computing node and a switching node is connected to a signal channel of the computing node and a signal channel of the switching node; the control component of the designated functional node is also used to locate the signal channel connected to the other side of the second communication link based on the topological relationship of the cable assembly to update the fault diagnosis information of the second functional node when the fault diagnosis information of the second functional node among other functional nodes is only used to indicate that the second communication link connected to the second functional node has a fault and is connected to the signal channel on one side of the second functional node.
[0144] Similar to the aforementioned embodiment, the control component of the second functional node among the other functional nodes can perform fault diagnosis based on the link status information in the second functional node to obtain fault diagnosis information. In this case, the fault diagnosis information may only include signal channel information of the faulty communication link on the side of the functional node, but lack signal channel positioning of the other side of the faulty link, i.e., the remote node; in this case, the fault diagnosis information sent by the control component of the second functional node to the control component of the designated functional node via the preset switching component may only be used to indicate the signal channel to which the second communication link to which the second functional node is connected and which has failed is connected on the side of the second functional node.
[0145] Optionally, the control component of the second functional node may proactively send the fault diagnosis information to the control component of the designated functional node via the preset switching component, or may send the fault diagnosis information to the control component of the designated functional node via the preset switching component in response to a data acquisition request.
[0146] In this embodiment, when the fault diagnosis information of the second functional node among other functional nodes is only used to indicate the signal channel to which the faulty second communication link connected to the second functional node is connected on one side of the second functional node, the control component of the designated functional node can locate the signal channel to which the other side of the second communication link is connected based on the topological relationship of the cable assembly.
[0147] Optionally, the topological relationship of the cable assembly acquired by the control component of the designated functional node may be stored in the designated functional node, or stored in other functional nodes, or sent by other terminal devices.
[0148] Optionally, the control component of the designated functional node can feed back the updated fault diagnosis information to the control component of the second functional node. The updated fault diagnosis information not only includes the signal channel information of the fault link on the side of the second functional node, but also adds the signal channel information of the fault link on the other side of the fault link, thereby realizing two-way confirmation of fault location and improving the accuracy and efficiency of fault diagnosis.
[0149] Through this embodiment, by specifying a functional node to supplement unidirectional fault diagnosis information based on the topological relationship of the cable assembly, the reliability of fault diagnosis of the entire cabinet server can be improved.
[0150] In an exemplary embodiment, a second storage component is provided on the second connector of the designated functional node; the second storage component stores the topology data of the cable assembly; and the control component of the designated functional node is further used to obtain the topology data of the cable assembly stored in the second storage component from the second connector.
[0151] Optionally, the second connector of the designated functional node may be provided with a storage component for storing topology data of the cable assembly, which may include layout information of all communication links within the entire cabinet server, including but not limited to signal channel numbers at both ends of the link.
[0152] Through the topology data of the cable assembly stored in the second storage component, the control component of the specified functional node can access the complete cable assembly topology information without obtaining it through network communication, reducing the dependence on the top-level management software, so that the control component can quickly respond to and handle potential link failures.
[0153] Through this embodiment, by directly acquiring the topology data of the cable assembly through the control component, the dependence of fault diagnosis on network communication can be reduced, and the efficiency and accuracy of fault diagnosis of the cable assembly can be improved.
[0154] In an exemplary embodiment, the whole cabinet server also includes: a rack top switch; wherein the control component of each functional node is also used to report the fault diagnosis information of each functional node to the rack top switch, wherein the fault diagnosis information of each functional node is only used to indicate the signal channel to which the faulty communication link connected to each functional node is connected on one side of each functional node; the rack top switch is used to locate the signal channel to which the other side of the faulty communication link connected to each functional node is connected based on the topological relationship of the cable assembly represented by the preset routing table, so as to update the fault diagnosis information of each functional node.
[0155] Here, the Top-of-Rack Switch (TOR switch) is usually located at the top of the rack and is used to connect servers and other network devices in the rack. It can aggregate the traffic in the rack and connect it to the upper-level network.
[0156] Optionally, a preset routing table can be stored in the TOR switch, which can be used to represent the topological relationship of the cable assembly. The preset routing table can be a routing table of all communication links from the accelerator module of the computing node to the Ethernet switching component of the switching node. It can record the signal channel number of each communication link on both sides of the communication node, as well as the topological relationship and status information of these links.
[0157] Optionally, the control component of each functional node in the above-mentioned whole cabinet server can be used to report the fault diagnosis information of each functional node to the rack top switch. In this case, the fault diagnosis information of each functional node is only responsible for collecting and reporting the fault diagnosis status of one side of the node, that is, it only includes the signal channel number of the fault link on the side of the node.
[0158] The existence of the preset routing table provides comprehensive link layout information for fault diagnosis of the entire cabinet server, enabling it to quickly identify the connection status of the other end of the fault link when receiving single-side fault diagnosis information, achieving two-way confirmation of fault location.
[0159] In order to accurately determine the status of both ends of the faulty link, the top-of-rack switch can parse the fault diagnosis information reported by each functional node based on the preset routing table after receiving the information, and supplement the signal channel number of the faulty link on the remote node side. That is, by comparing the link layout data in the routing table with the link status data reported by the functional node, the top-of-rack switch can accurately identify the signal channel to which the faulty link is connected on the other side, thereby identifying both ends of the faulty link and realizing two-way supplementation and confirmation of fault information.
[0160] Optionally, the rack top switch can update the fault diagnosis information of each functional node based on the signal channel to which the other side of the faulty communication link to which each functional node is connected is connected, and feed back the updated fault diagnosis information. For example, the updated fault diagnosis information of each functional node can be sent to the corresponding functional node.
[0161] Through this embodiment, the top-of-rack switch parses the link status data based on the routing table, which can achieve bidirectional fault location and improve the accuracy of fault diagnosis.
[0162] In an exemplary embodiment, the top-of-rack switch is further configured to aggregate the fault diagnosis information of each functional node by calling a designated management interface to form a fault diagnosis log of the cable assembly.
[0163] Optionally, the top-of-rack switch is also used to summarize the fault diagnosis information of each functional node by calling a specified management interface, and extract the fault diagnosis results based on the acquired fault diagnosis information to generate a fault diagnosis log for the cable assembly. The log may include but is not limited to specific information of all faulty links, such as the signal path number, link type, fault type (speed reduction, disconnection), and the physical location of the faulty link in the entire cabinet.
[0164] Optionally, the designated management interface may be a Redfish interface. Here, Redfish is a management interface based on a RESTful API for managing and monitoring servers. Through this standardized calling interface, the top-of-rack switch can access and control various resources in the server in a unified manner. For example, in this embodiment, the top-of-rack switch can use the Redfish interface to call the control components of each functional node and obtain their fault diagnosis information.
[0165] For example, Figure 8 As shown in the figure, the BMC of each node can report fault diagnosis information to the TOR switch. The BMC of the computing node (switch node) reads the link status register of the retimer (Ethernet switching component) through MDIO to implement interconnect signal status monitoring in the B (A) direction (if there is no retimer, it directly reads the GPU data). The TOR switch can read and summarize the abnormal information of each node through the management interface, form an error log, and report it to the WEB (World Wide Web).
[0166] Through this embodiment, the fault diagnosis information is aggregated through the designated management interface, which simplifies the information collection process, improves the information processing efficiency, and ensures the accuracy of the fault diagnosis information.
[0167] In an exemplary embodiment, the control component of the designated functional node is further used to generate a fault diagnosis log of the cable assembly based on the fault diagnosis information of other functional nodes and the fault diagnosis information of the designated functional node; and transmit the fault diagnosis log of the cable assembly to a terminal device connected to the designated functional node.
[0168] In the whole cabinet server, the designated functional node can also be responsible for the aggregation and processing of fault diagnosis information. Optionally, the control component of the designated functional node can obtain fault diagnosis information from other functional nodes and combine this information with the fault diagnosis information collected by this node to form a comprehensive fault diagnosis log of the cable assembly.
[0169] Optionally, after collecting fault diagnosis information from other functional nodes, the control component of the designated functional node can generate a fault diagnosis log for the cable assembly according to preset rules and algorithms. The log may include details of all identified communication link faults in the entire cabinet server, including but not limited to the signal channel numbers at both ends of the fault link, fault type, fault occurrence time, impact range and possible causes of the fault, etc.
[0170] The generated fault diagnosis log of the cable assembly can be transmitted by the control component of the designated functional node to the terminal device connected to the designated functional node. The terminal device can be a remote management tool or a server operation and maintenance platform to read, parse and display the fault diagnosis log, analyze the cause of the fault, and formulate a repair plan.
[0171] Through this embodiment, by specifying a functional node to generate a fault diagnosis log and transmitting it to the connected terminal device, the management process of fault information can be simplified and the efficiency and accuracy of fault diagnosis can be improved.
[0172] According to another aspect of the embodiment of the present application, a fault diagnosis method for a whole-cabinet server is also provided. The fault diagnosis method for a whole-cabinet server is applied to the above-mentioned product embodiment and optional implementation mode, and will not be repeated hereafter.
[0173] The fault diagnosis method of the whole cabinet server in the embodiment of the present application can be applied to the whole cabinet server in the aforementioned embodiment. The whole cabinet server includes a group of functional nodes, and the group of functional nodes includes computing nodes and switching nodes. The computing nodes and switching nodes are interconnected through communication links on the cable assembly. The cable assembly includes pre-installed switching components, and each functional node in the group of functional nodes has a communication link with the pre-installed switching components.
[0174] Figure 9 FIG. 1 is a flow chart of an optional method for diagnosing a fault of a whole-cabinet server according to an embodiment of the present application. Figure 9 As shown, the process of the fault diagnosis method for the above-mentioned whole cabinet server may include the following steps: step S902, obtaining the link status data of each functional node through the control component of each functional node, and performing fault diagnosis on the link status data of each functional node through the control component of each functional node to obtain fault diagnosis information of each functional node, wherein the link status data of each functional node is used to represent the link status of the communication link connected to each functional node; step S904, collecting the fault diagnosis information of other functional nodes from the control components of other functional nodes through the control component of a specified functional node in a group of functional nodes via a preset switching component, wherein the other functional nodes are functional nodes other than the specified functional node in a group of functional nodes.
[0175] In this embodiment, step S902 can be executed by the control component of each functional node in the aforementioned embodiment, and step S904 can be executed by the control component of a specified functional node in a group of functional nodes in the aforementioned embodiment. The manner in which the control component of each functional node executes step S902 and the manner in which the control component of a specified functional node in a group of functional nodes executes step S904 are similar to those in the aforementioned embodiment, and have been explained and will not be repeated here.
[0176] Through the embodiments of the present application, the link status data of each functional node is obtained through the control component of each functional node, and the link status data of each functional node is diagnosed through the control component of each functional node to obtain fault diagnosis information of each functional node, wherein the link status data of each functional node is used to represent the link status of the communication link to which each functional node is connected; the fault diagnosis information of other functional nodes is collected from the control components of other functional nodes via the preset switching component through the control component of a designated functional node in a group of functional nodes, wherein the other functional nodes are functional nodes other than the designated functional node in a group of functional nodes, thereby solving the problem of difficulty in fault locating of the whole cabinet server in the related art and improving the convenience of fault locating of the whole cabinet server.
[0177] In an exemplary embodiment, a communication link between a computing node and a switching node is connected to a signal channel of the computing node and a signal channel of the switching node. Acquiring link status data of each functional node through a control component of each functional node includes: reading, through the control component of each functional node, link status data stored by a designated component of each functional node to obtain the link status data of each functional node, wherein the link status data of each functional node represents a link status of a communication link connected to each functional node in a direction toward each functional node.
[0178] In an exemplary embodiment, a group of functional nodes includes a first functional node storing topology data of a cable assembly, where the topology data of the cable assembly is used to describe a topological relationship of the cable assembly. Performing fault diagnosis on the link status data of each functional node by a control component of each functional node to obtain fault diagnosis information for each functional node includes: performing fault diagnosis on the link status data of the first functional node by the control component of the first functional node based on the topological relationship of the cable assembly to obtain fault diagnosis information for the first functional node, where the fault diagnosis information of the first functional node indicates signal channels connected to both sides of a first communication link to which the first functional node is connected and to which a fault has occurred.
[0179] In an exemplary embodiment, a first storage component is provided on a first connector of a first functional node; topology data of a cable assembly is stored in the first storage component. The method further includes: obtaining, via a control component of the first functional node, the topology data of the cable assembly stored in the first storage component from the first connector.
[0180] In an exemplary embodiment, the first storage component is a field replaceable unit (FRU). Obtaining, by the control component of the first functional node, the topology data of the cable assembly stored in the first storage component from the first connector includes: reading, by the control component of the first functional node, the topology data of the cable assembly stored in the FRU via a communication bus between the control component of the first functional node and the first connector.
[0181] In an exemplary embodiment, a whole cabinet server includes multiple node placement positions, a functional node in a group of functional nodes is placed in a node placement position in the multiple node placement positions; an accelerator module of a computing node has multiple module signal channels, a media access control component of a switching node has multiple component signal channels, a communication link between a computing node and a switching node is connected to a signal channel of the computing node as a module signal channel of the accelerator module, and is connected to a signal channel of the switching node as a component signal channel of the media access control component; in the topology data of the cable assembly, the fields recording each functional node include a node type field and The placement position identification field, the field for recording the computing node also includes the module number field and the module channel number field, and the field for recording the switching node also includes the switching node number field and the component channel number field, wherein the node type field is used to record the node type of the corresponding functional node, the placement position identification field is used to record the placement position identification of the node placement position where the corresponding functional node is placed, the module number field is used to record the module number of the corresponding accelerator module, the module channel number field is used to record the channel number of the corresponding module signal channel, the switching node number field is used to record the node number of the corresponding switching node, and the component channel number field is used to record the channel number of the corresponding component signal channel.
[0182] In an exemplary embodiment, the above method also includes: receiving, through the control components of other functional nodes, a data acquisition request sent by a designated functional node via a preset switching component; in response to the data acquisition request, generating, through the control components of other functional nodes, a fault message based on the fault diagnosis information of the other functional nodes, and sending the fault message to the control component of the designated functional node via the preset switching component; wherein the fault message carries the following information: the node identifier of the other functional node; the channel identifier of the signal channel connected to both sides of the second communication link to which the other functional nodes are connected and to which the fault occurs; and the fault type of the second communication link.
[0183] In an exemplary embodiment, the whole cabinet server includes multiple node placement positions, and a functional node in a group of functional nodes is placed on a node placement position in the multiple node placement positions; the node identifiers of other functional nodes include the placement position identifiers of the node placement positions where the other functional nodes are placed.
[0184] In an exemplary embodiment, an accelerator module of a computing node is connected to a connector of the computing node via a retimer of the computing node, forming multiple module signal channels of the accelerator module. A communication link between the computing node and the switch node is connected to a signal channel of the computing node as a module signal channel of the accelerator module. The link status data of each functional node is obtained by reading link status data stored in a designated component of each functional node through a control component of each functional node, including: reading link status data stored in a first register of the retimer through the control component of the computing node to obtain the link status data of the computing node.
[0185] In an exemplary embodiment, an accelerator module of a computing node is connected to a connector of the computing node to form multiple module signal channels of the accelerator module. A communication link between the computing node and the switch node is connected to a signal channel of the computing node as a module signal channel of the accelerator module. The link status data of each functional node is obtained by reading link status data stored in a designated component of each functional node through a control component of each functional node, including: reading the link status data stored in the accelerator module through the control component of the computing node to obtain the link status data of the computing node.
[0186] In an exemplary embodiment, the switching node includes an Ethernet switching component. Reading, by the control component of each functional node, link state data stored in a designated component of each functional node to obtain the link state data of each functional node includes: reading, by the control component of the switching node, the link state data stored in a second register on the Ethernet switching component to obtain the link state data of the switching node.
[0187] In an exemplary embodiment, the link status data stored in the designated component of each functional node is read by the control component of each functional node to obtain the link status data of each functional node, including: reading the link status data stored in the designated component of each functional node via the management data input / output interface of each functional node by the control component of each functional node to obtain the link status data of each functional node.
[0188] In an exemplary embodiment, the accelerator module of the computing node has multiple module signal channels, the media access control component of the switching node has multiple component signal channels, and a communication link between the computing node and the switching node is connected to a signal channel of the computing node, which is a module signal channel of the accelerator module, and is connected to a signal channel of the switching node, which is a component signal channel of the media access control component.
[0189] In an exemplary embodiment, a communication link between a computing node and a switching node is connected to a signal channel of the computing node and a signal channel of the switching node. The method further includes: when the fault diagnosis information of a second functional node among other functional nodes is used only to indicate the signal channel to which the second functional node is connected that has failed, and the control component of the designated functional node locates the signal channel to which the second communication link is connected on the other side based on the topological relationship of the cable assembly, thereby updating the fault diagnosis information of the second functional node.
[0190] In an exemplary embodiment, a second storage component is provided on a second connector of the designated functional node; the second storage component stores topology data of the cable assembly. The method further includes: obtaining, via a control component of the designated functional node, the topology data of the cable assembly stored in the second storage component from the second connector.
[0191] In an exemplary embodiment, the above method also includes: reporting the fault diagnosis information of each functional node to the rack-top switch of the entire cabinet server through the control component of each functional node, wherein the fault diagnosis information of each functional node is only used to indicate the signal channel to which the faulty communication link connected to each functional node is connected on one side of each functional node; locating the signal channel to which the other side of the faulty communication link connected to each functional node is connected based on the topological relationship of the cable components represented by the preset routing table through the rack-top switch to update the fault diagnosis information of each functional node.
[0192] In an exemplary embodiment, the method further includes: calling a designated management interface through a top-of-rack switch to aggregate fault diagnosis information of each functional node to form a fault diagnosis log of the cable assembly.
[0193] In an exemplary embodiment, after collecting fault diagnosis information of other functional nodes from the control components of other functional nodes, the above method also includes: generating a fault diagnosis log of the cable assembly based on the fault diagnosis information of other functional nodes and the fault diagnosis information of the designated functional node through the control component of the designated functional node; and transmitting the fault diagnosis log of the cable assembly to a terminal device connected to the designated functional node.
[0194] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.
[0195] According to another aspect of the embodiments of the present application, a computer-readable storage medium is provided. The computer-readable storage medium includes a stored program, wherein the program executes the steps of any of the above method embodiments when it is run.
[0196] In an exemplary embodiment, the computer-readable storage medium is a non-volatile storage medium, which may include but is not limited to: a USB flash drive, ROM, RAM, a mobile hard disk, a magnetic disk, or an optical disk, etc., various non-temporary storage media that can store computer programs.
[0197] According to another aspect of the embodiments of the present application, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor is configured to execute the steps of any of the above-described method embodiments through the computer program. In an exemplary embodiment, the electronic device may further comprise a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.
[0198] For specific examples in this embodiment, reference may be made to the examples described in the above embodiments and exemplary implementation modes, and this embodiment will not be described in detail here.
[0199] According to another aspect of an embodiment of the present application, a computer program product is also provided, comprising a computer program / instruction containing program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication portion 1009, and / or installed from a removable medium 1011. When the computer program is executed by the central processing unit 1001, the various functions provided by the embodiments of the present application are performed. The serial numbers of the embodiments of the present application are for descriptive purposes only and do not represent the merits of the embodiments.
[0200] Figure 10 The following schematically shows a block diagram of a computer system structure of an electronic device for implementing an embodiment of the present application. Figure 10As shown, computer system 1000 includes a CPU (Central Processing Unit) 1001, which can perform various appropriate actions and processes according to programs stored in ROM 1002 or programs loaded from storage 1008 into RAM 1003. Random access memory 1003 also stores various programs and data required for system operation. CPU 1001, read-only memory 1002, and random access memory 1003 are connected to each other via bus 1004. An I / O (Input / Output) interface 1005 is also connected to bus 1004.
[0201] The following components are connected to the I / O interface 1005: an input section 1006 including a keyboard, mouse, and the like; an output section 1007 including devices such as a CRT (Cathode Ray Tube), an LCD (Liquid Crystal Display), and speakers; a storage section 1008 including a hard disk; and a communication section 1009 including a network interface card such as a local area network card or a modem. The communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to the I / O interface 1005 as needed. Removable media 1011, such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, is installed in the drive 1010 as needed, so that computer programs read from the media can be installed in the storage section 1008 as needed.
[0202] In particular, according to an embodiment of the present application, the processes described in the various method flow charts can be implemented as computer software programs. For example, an embodiment of the present application includes a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for executing the methods shown in the flow charts. In such an embodiment, the computer program can be downloaded and installed from a network via the communication portion 1009 and / or installed from a removable medium 1011. When the computer program is executed by the central processing unit 1001, the various functions defined in the system of the present application are performed.
[0203] It should be noted that Figure 10 The computer system 1000 of the electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0204] Obviously, those skilled in the art should understand that the modules or steps of the present application described above can be implemented using a general-purpose computing device, they can be concentrated on a single computing device, or distributed across a network composed of multiple computing devices, they can be implemented using program code executable by the computing device, and thus, they can be stored in a storage device and executed by the computing device, and in some cases, the steps shown or described can be performed in a different order than herein, or they can be fabricated into separate integrated circuit modules, or multiple modules or steps can be fabricated into a single integrated circuit module for implementation. Thus, the present application is not limited to any specific combination of hardware and software.
[0205] The above are only preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various modifications and variations of the present application are possible. Any modifications, equivalent substitutions, improvements, etc. made within the principles of the present application shall be included in the scope of protection of the present application.
[0206] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0207] The above is a detailed introduction to a whole cabinet server provided by this application. This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core ideas of this application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the scope of protection of the claims of this application.
Claims
1. A whole cabinet server, characterized in that: include: A group of functional nodes, the group of functional nodes including computing nodes and switching nodes, the computing nodes and the switching nodes being interconnected via a communication link on a cable assembly, the cable assembly including a pre-installed switching component, each functional node in the group of functional nodes having a communication link with the pre-installed switching component; wherein, The control component of each functional node is configured to obtain link status data of each functional node and perform fault diagnosis on the link status data of each functional node to obtain fault diagnosis information of each functional node, wherein the link status data of each functional node is used to indicate the link status of a communication link to which each functional node is connected; The control component of the designated functional node in the group of functional nodes is used to collect fault diagnosis information of the other functional nodes from the control components of the other functional nodes through the preset switching component, wherein the other functional nodes are functional nodes in the group of functional nodes other than the designated functional node.
2. The whole cabinet server according to claim 1, characterized in that: A communication link between the computing node and the switching node is connected to a signal channel of the computing node and a signal channel of the switching node; The control component of each functional node is further used to read the link status data stored in the designated component of each functional node to obtain the link status data of each functional node, wherein the link status data of each functional node is used to represent the link status of the communication link connected to each functional node in the direction toward each functional node.
3. The whole cabinet server according to claim 2, characterized in that: The group of functional nodes includes a first functional node storing topological data of the cable assembly, where the topological data of the cable assembly is used to describe a topological relationship of the cable assembly; The control component of the first functional node is further used to perform fault diagnosis on the link status data of the first functional node based on the topological relationship of the cable assembly to obtain fault diagnosis information of the first functional node, wherein the fault diagnosis information of the first functional node is used to indicate the signal channels connected to both sides of the first communication link to which the first functional node is connected and to which the fault occurs.
4. The whole cabinet server according to claim 3, characterized in that: A first storage component is provided on the first connector of the first functional node; the topology data of the cable assembly is stored in the first storage component; The control component of the first functional node is further configured to obtain, from the first connector, the topology data of the cable assembly stored in the first storage component.
5. The whole cabinet server according to claim 4, characterized in that: The first storage component is a field replaceable unit; The control component of the first functional node is further configured to read the topology data of the cable assembly stored in the field replaceable unit via a communication bus between the control component of the first functional node and the first connector.
6. The whole cabinet server according to claim 3, characterized in that: The whole cabinet server includes a plurality of node placement positions, and a functional node in the group of functional nodes is placed in a node placement position in the plurality of node placement positions; The accelerator module of the computing node has multiple module signal channels, and the media access control component of the switching node has multiple component signal channels. A signal channel of a communication link between the computing node and the switching node connected to the computing node is a module signal channel of the accelerator module, and a signal channel connected to the switching node is a component signal channel of the media access control component. In the topology data of the cable assembly, the fields for recording each functional node include a node type field and a placement position identification field, the fields for recording the computing nodes also include a module number field and a module channel number field, and the fields for recording the switching nodes also include a switching node number field and a component channel number field, wherein the node type field is used to record the node type of the corresponding functional node, the placement position identification field is used to record the placement position identification of the node placement position where the corresponding functional node is placed, the module number field is used to record the module number of the corresponding accelerator module, the module channel number field is used to record the channel number of the corresponding module signal channel, the switching node number field is used to record the node number of the corresponding switching node, and the component channel number field is used to record the channel number of the corresponding component signal channel.
7. The whole cabinet server according to claim 2, characterized in that: The control component of the other functional node is further configured to receive a data acquisition request sent by the designated functional node via the preset switching component; in response to the data acquisition request, generate a fault message based on the fault diagnosis information of the other functional node, and send the fault message to the control component of the designated functional node via the preset switching component; The fault message carries the following information: the node identifier of the other functional node; the channel identifier of the signal channel on both sides of the second communication link to which the other functional node is connected and to which the fault occurs; and the fault type of the second communication link.
8. The whole cabinet server according to claim 7, characterized in that: The whole cabinet server includes multiple node placement positions, and a functional node in the group of functional nodes is placed on a node placement position among the multiple node placement positions; the node identifiers of the other functional nodes include the placement position identifiers of the node placement positions where the other functional nodes are placed.
9. The whole cabinet server according to claim 2, characterized in that: The accelerator module of the computing node is connected to the connector of the computing node via the retimer of the computing node to form multiple module signal channels of the accelerator module, and a communication link between the computing node and the switching node connected to a signal channel of the computing node is a module signal channel of the accelerator module; The control component of the computing node is further configured to read the link status data stored in the first register of the retimer to obtain the link status data of the computing node.
10. The whole cabinet server according to claim 2, characterized in that: The accelerator module of the computing node is connected to the connector of the computing node to form multiple module signal channels of the accelerator module, and a communication link between the computing node and the switching node is connected to a signal channel of the computing node as a module signal channel of the accelerator module; The control component of the computing node is further configured to read the link status data stored in the accelerator module to obtain the link status data of the computing node.
11. The whole cabinet server according to claim 2, characterized in that: The switching node includes an Ethernet switching component; The control component of the switching node is further configured to read the link status data stored in the second register on the Ethernet switching component to obtain the link status data of the switching node.
12. The whole cabinet server according to claim 2, characterized in that: The control component of each functional node is further configured to read the link status data stored in the designated component of each functional node via the management data input / output interface of each functional node to obtain the link status data of each functional node.
13. The whole cabinet server according to claim 2, characterized in that: The accelerator module of the computing node has multiple module signal channels, and the media access control component of the switching node has multiple component signal channels. A communication link between the computing node and the switching node is connected to a signal channel of the computing node, which is a module signal channel of the accelerator module, and a signal channel connected to the switching node is a component signal channel of the media access control component.
14. The whole cabinet server according to claim 1, characterized in that: A communication link between the computing node and the switching node is connected to a signal channel of the computing node and a signal channel of the switching node; The control component of the designated functional node is further configured to locate the signal channel to which the other side of the second communication link is connected based on the topological relationship of the cable assembly, when the fault diagnosis information of the second functional node among the other functional nodes is only used to indicate the signal channel to which the second communication link connected to the second functional node is connected on one side of the second functional node, so as to update the fault diagnosis information of the second functional node.
15. The whole cabinet server according to claim 14, characterized in that: A second storage component is provided on the second connector of the designated functional node; the second storage component stores the topology data of the cable assembly; The control component of the designated functional node is further configured to obtain, from the second connector, the topology data of the cable assembly stored in the second storage component.
16. The whole cabinet server according to claim 1, characterized in that: The whole cabinet server also includes: a rack top switch; wherein, The control component of each functional node is further configured to report fault diagnosis information of each functional node to the top-of-rack switch, wherein the fault diagnosis information of each functional node is only used to indicate the signal path to which the failed communication link connected to each functional node is connected on one side of each functional node; The rack top switch is used to locate the signal channel connected to the other side of the failed communication link connected to each functional node based on the topological relationship of the cable assembly represented by the preset routing table, so as to update the fault diagnosis information of each functional node.
17. The whole cabinet server according to claim 16, characterized in that: The top-of-rack switch is further configured to aggregate the fault diagnosis information of each functional node by calling a designated management interface to form a fault diagnosis log of the cable assembly.
18. The whole cabinet server according to any one of claims 1 to 17, characterized in that: The control component of the designated functional node is also used to generate a fault diagnosis log of the cable assembly based on the fault diagnosis information of the other functional nodes and the fault diagnosis information of the designated functional node; and transmit the fault diagnosis log of the cable assembly to the terminal device connected to the designated functional node.
19. A method for diagnosing faults in a whole cabinet server, characterized in that: A group of functional nodes on the whole cabinet server includes computing nodes and switching nodes, the computing nodes and the switching nodes are interconnected via a communication link on a cable assembly, the cable assembly includes a pre-installed switching component, and each functional node in the group of functional nodes has a communication link with the pre-installed switching component; the method includes: Acquiring link status data of each functional node through the control component of each functional node, and performing fault diagnosis on the link status data of each functional node through the control component of each functional node to obtain fault diagnosis information of each functional node, wherein the link status data of each functional node is used to indicate the link status of a communication link to which each functional node is connected; The control component of the designated functional node in the group of functional nodes collects fault diagnosis information of the other functional nodes from the control components of other functional nodes via the preset switching component, wherein the other functional nodes are functional nodes in the group of functional nodes other than the designated functional node.
20. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein the computer program implements the steps of the method according to claim 19 when executed by a processor.
Citation Information
Patent Citations
Switch dynamic link backup and expansion method and device, equipment and storage medium
CN113992605A
Server, computer system and multi-node management method
CN119376995A