A server, its device monitoring system and method

By introducing interconnected resource modules and computing resource modules into the server, and using the firmware of the first controller and the exchange component to obtain the identification information of the accelerator card, the problem that unallocated accelerator cards cannot be identified and monitored by the system is solved, and dynamic identification and monitoring of all accelerators is realized, ensuring the safe and stable operation of the server.

CN119906687BActive Publication Date: 2025-06-20LANGCHAO ELECTRONIC INFORMATION IND CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510387250.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-06-20
Estimated Expiration
2045-03-31

AI Technical Summary

Technical Problem

In the converged architecture, unassigned acceleration cards cannot be identified by the system, nor can they establish monitoring links, resulting in blind spots in device status monitoring and thermal regulation, affecting the safe operation of the server.

Method used

By introducing an interconnected resource module and a computing resource module into the server, the firmware interaction of the first controller with the exchange component is used to obtain identification information of the acceleration card, and the second controller determines the monitoring mode of the acceleration card based on these identification information, so as to realize dynamic identification and monitoring of all acceleration cards.

Benefits of technology

Dynamic identification and monitoring of unassigned accelerator cards is realized, eliminating blind spots in equipment operation status monitoring and thermal regulation, and ensuring the safe and stable operation of the server.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119906687B_ABST
    Figure CN119906687B_ABST
Patent Text Reader

Abstract

The present invention discloses a server, its device monitoring system and method, which relate to the technical field of servers. It includes a first controller and a second controller. Through the firmware interaction between the first controller in the interconnection resource module and the switching component connected with an acceleration card in the interconnection resource module, the identification information of the acceleration cards connected under multiple switching ports of the switching component is obtained. The second controller determines the monitoring modes of all local acceleration cards within the first computing resource module where it is located based on the identification information of all acceleration cards obtained by the first controller, and monitors each local acceleration card according to the corresponding monitoring mode, solving the technical problem that unallocated acceleration cards cannot be recognized by the system and cannot establish a monitoring link, and achieving the technical effect of dynamically identifying the types of acceleration cards under different heterogeneous acceleration card resource pool configurations and implementing out-of-band monitoring by adopting corresponding monitoring logics.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of servers, and in particular, to a server, an equipment monitoring system and method thereof. Background Art

[0002] Data centers are transforming from a computing center architecture to a data-centric converged architecture, enabling dynamic configuration and rapid upgrade of CPUs (Central Processing Units) and accelerator cards through physical decoupling. Traditional servers adopt a single chassis integrated design, where the Host BIOS (Host Basic Input / Output System) can obtain all accelerator card information and establish a BMC (Baseboard Management Controller) monitoring link. In the converged architecture, unallocated accelerator cards are physically isolated from the host, resulting in the Host BIOS only being able to recognize allocated devices, causing unallocated accelerator cards to fall outside the BMC monitoring scope, creating blind spots in device status monitoring and heat dissipation regulation, and affecting the safe operation of the server.

[0003] Therefore, how to provide a solution to the above technical problems is an issue that those skilled in the art need to solve currently. Summary of the Invention

[0004] The present invention provides a server, an equipment monitoring system and method thereof, so as to at least solve the problem in related technologies that unallocated accelerator cards can neither be recognized by the system nor establish a monitoring link.

[0005] The present invention provides an equipment monitoring system for a server. The server includes an interconnected resource module and at least one first computing resource module. The equipment monitoring system of the server includes:

[0006] A first controller disposed in the interconnected resource module, connected to at least one switching component in the interconnected resource module. The first controller is configured to interact with the firmware of at least one of the switching components to obtain the identification information of the accelerator cards connected under the switching ports of the switching components.

[0007] A second controller disposed in the first computing resource module, connected to a local accelerator card. The second controller is configured to determine the monitoring mode of the local accelerator card according to the identification information of the local accelerator card obtained by the first controller, and perform monitoring operations on the local accelerator card according to the monitoring mode of the local accelerator card. The local accelerator card is the accelerator card disposed in the same first computing resource module as the second controller.

[0008] The present invention also provides a server, which includes an interconnection resource module, a plurality of first computing resource modules, a plurality of second computing resource modules, and the device monitoring system of the server as described above.

[0009] The interconnection resource module includes a plurality of switching components and a plurality of first communication ports. The first computing resource module includes a plurality of acceleration cards and a plurality of second communication ports. The second computing resource module includes a processor and a third communication port. The interconnection resource module and the first computing resource module are connected through the first communication port and the second communication port. The interconnection resource module and the second computing resource module are connected through the third communication port and the first communication port.

[0010] The present invention also provides a device monitoring method for a server, which is applied to a first controller. The first controller is disposed in the interconnection resource module of the server. The first controller is connected to at least one switching component in the interconnection resource module. The server further includes a first computing resource module. The device monitoring method for the server includes:

[0011] Interact with the firmware of at least one of the switching components;

[0012] Obtain the identification information of the acceleration cards connected under the switching ports of the switching components, so that a second controller disposed in the first computing resource module determines the monitoring mode of the local acceleration cards according to the identification information of the acceleration cards obtained by the first controller, and performs monitoring operations on the local acceleration cards according to the monitoring mode of the local acceleration cards; the local acceleration cards are the acceleration cards disposed in the same first computing resource module as the second controller.

[0013] The present invention also provides a device monitoring method for a server, which is applied to a second controller. The second controller is disposed in the first computing resource module of the server. The second controller is connected to local acceleration cards. The server further includes an interconnection resource module. The device monitoring method for the server includes:

[0014] Obtain the identification information of the acceleration cards connected under the switching ports of the switching components obtained by the interaction between the first controller in the interconnection resource module and at least one switching component;

[0015] Determine the monitoring mode of the local acceleration cards according to the identification information of the acceleration cards obtained by the first controller;

[0016] Perform monitoring operations on the local acceleration cards according to the monitoring mode of the local acceleration cards; the local acceleration cards are the acceleration cards disposed in the same first computing resource module as the second controller.

[0017] Through the present invention, a first controller in an interconnected resource module interacts with the firmware of at least one switching component connected with an acceleration card in the interconnected resource module, accesses the configuration space of the acceleration cards connected under multiple switching ports of the switching component, dynamically obtains the identification information of the acceleration cards, and summarizes the identification information of all the acceleration cards in the server. A second controller disposed in the first computing resource module determines the monitoring modes of all local acceleration cards in the first computing resource module where it is located based on the identification information of all the acceleration cards obtained by the first controller, and monitors each acceleration card according to the corresponding monitoring mode, realizing the dynamic identification of the identification information of all the acceleration cards, solving the technical problem that unallocated acceleration cards can neither be recognized by the system nor establish a monitoring link, realizing the dynamic identification of the types of acceleration cards under different heterogeneous acceleration card resource pool configurations and adopting corresponding monitoring logics to realize out-of-band monitoring, and ensuring the technical effect of the safe and stable operation of the server. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the embodiments of the present invention, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0019] Figure 1 The structural schematic diagram of the device monitoring system of the first server provided by the embodiment of the present invention;

[0020] Figure 2 The structural schematic diagram of the device monitoring system of the second server provided by the embodiment of the present invention;

[0021] Figure 3 An acceleration card monitoring flowchart provided by the embodiment of the present invention;

[0022] Figure 4 The structural schematic diagram of the device monitoring system of the third server provided by the embodiment of the present invention;

[0023] Figure 5 The structural schematic diagram of the device monitoring system of the third server provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0024] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present invention.

[0025] It should be noted that in the description of the present invention, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present invention are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0026] In order to enable those skilled in the art of the present technology to better understand the solution of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0027] In the fusion architecture, the server consists of an IO Box (Input / Output Box, input / output resource box) and multiple device chassis including a Host Box (computing resource box) and a heterogeneous acceleration card Box (resource box), and realizes the dynamic configuration and rapid upgrade of the CPU (Central Processing Unit) and acceleration cards through physical decoupling. The traditional general server adopts a single chassis integrated architecture, and all hardware devices are integrated in the same chassis. The Host BIOS (Host Basic Input / Output System) can obtain the identification information of all acceleration cards in the heterogeneous acceleration card resource pool, construct complete device asset information and transmit it to the Host BMC (Host Baseboard Management Controller), and the Host BMC parses and establishes a monitoring link for each acceleration card. However, in the fusion architecture, due to the physical isolation between the unallocated acceleration cards in the heterogeneous acceleration card resource pool and the Host (host), the Host BIOS can only identify the allocated acceleration cards, resulting in that the unallocated acceleration cards cannot be recognized by the system and the BMC monitoring link cannot be established either.

[0028] To solve the above problems, please refer to Figure 1 , the present invention provides a device monitoring system for a server. The server includes an interconnected resource module 1 and at least one first computing resource module 2. The device monitoring system of the server includes:

[0029] A first controller 11 disposed in the interconnected resource module 1, connected to at least one switching component 12 in the interconnected resource module 1. The first controller 11 is configured to interact with the firmware of at least one switching component 12 to obtain the identification information of the acceleration cards connected under the switching ports of the switching components;

[0030] The second controller 21 disposed in the first computing resource module 2 is connected to the local acceleration card. The second controller 21 is configured to determine the monitoring mode of the local acceleration card according to the identification information of the local acceleration card obtained by the first controller 11, and perform monitoring operations on the local acceleration card respectively according to the monitoring mode of the local acceleration card; the local acceleration card is an acceleration card disposed in the same first computing resource module 2 as the second controller 21.

[0031] The server in this embodiment may specifically be a server with a converged architecture. The design feature of the converged architecture is resource decoupling, that is, the entire system is decoupled into different independent pooled modules, and the independent pooled modules exist in the form of separate Boxes. The interconnection resource module 1 and the first computing resource module 2 in this embodiment are both an independent pooled module.

[0032] The interconnection resource module 1 is specifically an IO Box, which includes a first controller 11, a plurality of switching components 12, and a plurality of first communication ports. The PCIe (Peripheral Component Interconnect Express) resources of multiple CPUs in the server are expanded, networked, and allocated through the switching components 12 in the IO Box. In this embodiment, the first controller 11 is connected to the corresponding communication ports on each switching component 12. To reduce the occupation of the communication ports on the first controller 11, as an optional embodiment, an expansion component may also be provided in the interconnection resource module 1. The first end of the expansion component is connected to the first controller 11, and the plurality of second ends of the expansion component are respectively and correspondingly connected to the corresponding communication ports on the plurality of switching components 12. Specifically, the first controller 11 and each switching component 12 can communicate through the I2C (Inter-Integrated Circuit) protocol, and the switching component 12 can select a PCIe Switch chip. The firmware of the switching component 12 provides basic commands for reading and writing registers. The first controller 11 sends corresponding commands to the firmware of the switching component 12 to implement the function of reading and writing the registers inside the switching component 12, realize the unified management and configuration of all switching components 12, ensure the normal operation of the switching component 12 and its data exchange ability. The switching component 12 also includes a plurality of switching ports, and the plurality of switching ports are respectively and correspondingly connected to the plurality of first communication ports on the interconnection resource module 1.

[0033] After the entire server rack is powered on in coordination, the first controller 11 is configured to interact with the firmware of at least one switching component 12, access the configuration space of the acceleration card connected under the switching port of the switching component 12, so as to obtain the identification information of the corresponding acceleration card. Specifically, the first controller 11 establishes a communication connection with the firmware of the switching component 12, and through the command of the basic read-write register provided by the switching component 12, interacts with the firmware of multiple switching components 12 in sequence. The first controller 11 accesses the configuration space of the acceleration card connected under each switching port in sequence. Specifically, the first controller 11 sends a command through the firmware of the switching component 12 and uses the previously established communication mechanism to access the configuration space of the acceleration card corresponding to each switching port. After accessing the configuration space of the acceleration card, the first controller 11 reads the unique identification information of the acceleration card therefrom, such as Vendor ID (Vender Identifier, supplier identifier) and Device ID (Device Identifier, device identifier). These information are stored in specific registers of the configuration space, and the first controller 11 obtains the identification information by reading these registers. It can be understood that the configuration space of the acceleration card is accessed through the PCIe bus, and each acceleration card has a unique memory mapping address for its configuration space. The first controller 11 calculates these addresses to access the corresponding registers. Since all the acceleration cards in the server are mounted on the switching component 12, therefore, the identification information of the acceleration cards obtained by the first controller 11 through accessing the firmware of multiple switching components 12 is the identification information of all the acceleration cards in the server, ensuring that the first controller 11 can comprehensively identify the types and identities of all the acceleration cards in the server.

[0034] Considering that in the actual use process, according to different customer requirements, there will be various different heterogeneous accelerator card resource pool configurations in the fusion architecture server. For example, Configuration 1 is the heterogeneous accelerator card of Manufacturer A, and Configuration 2 is the heterogeneous accelerator card of Manufacturer B. There are also scenarios where different heterogeneous accelerator cards are mixed and inserted in some working conditions. Among them, a heterogeneous accelerator card refers to a hardware acceleration device used in a server to accelerate specific computing tasks. These accelerator cards are based on different architectures and technologies and have different functions and performance characteristics, including but not limited to GPU (Graphics Processing Unit), FPGA (Field-Programmable Gate Array), ASIC (Application-Specific Integrated Circuit), etc. Correspondingly, a heterogeneous accelerator card resource pool refers to integrating multiple different types of heterogeneous accelerator cards into a resource pool, and through a unified management and service platform, realizing dynamic allocation and efficient utilization of resources. The accelerator cards in the resource pool can come from different suppliers, have different architectures and functions, but through the management of the resource pool, resource sharing and optimized scheduling can be achieved.

[0035] The second controller 21 in the first computing resource module 2 can perform out-of-band monitoring on the heterogeneous accelerator card through the I2C channel. However, the I2C slave addresses and monitoring methods of different types of heterogeneous accelerator cards will also vary. For example, the accelerator card of Manufacturer A may use the I2C slave address 0x50, while the accelerator card of Manufacturer B may use the I2C slave address 0x51. Then, based on the determined I2C slave address, the corresponding monitoring method needs to be adopted to monitor the accelerator card. For example, for the accelerator card of Manufacturer A, it may be necessary to read specific registers to obtain temperature and power consumption information, while for the accelerator card of Manufacturer B, different registers may need to be read to obtain this information. Therefore, the purpose of obtaining the identification information of the accelerator card in this embodiment is to determine the type of the accelerator card, so as to determine the I2C slave address and monitoring method of the accelerator card, so that an out-of-band monitoring channel can be successfully established for different types of accelerator cards. Refer to Figure 1 As shown, an expansion chip can also be arranged in the first computing resource module 2, and the second controller 21 is connected to each local accelerator card through the expansion chip.

[0036] Specifically, the second controller 21 determines the identification information of the local acceleration card from the identification information of multiple acceleration cards obtained from the first controller 11. The local acceleration card is an acceleration card within the same first computing resource module as the second controller 21, and the number of local acceleration cards can be one or more. For each local acceleration card, the second controller 21 determines the monitoring mode of the local acceleration card according to the identification information of the local acceleration card. Among them, the second controller 21 can passively receive the identification information of the acceleration card obtained by the first controller 11, or actively obtain the identification information of multiple acceleration cards obtained by the first controller 11. After the second controller 21 obtains the identification information of the local acceleration card, it can determine the type of the local acceleration card according to the identification information, and further determine the monitoring mode of the local acceleration card. The monitoring mode is determined at least according to the I2C slave address corresponding to the acceleration card type and the monitoring method. The monitoring method includes, but is not limited to, the specific register address to be accessed, the parameters to be monitored, etc. After determining the monitoring mode, perform the monitoring operation on the corresponding local acceleration card according to the monitoring mode. The monitoring operation includes, but is not limited to, establishing an out-of-band monitoring channel for the acceleration card according to the I2C slave address and performing monitoring according to the corresponding monitoring method. It can be understood that for the working conditions of the same type of acceleration card in a heterogeneous acceleration card resource pool, since the acceleration card models are the same, the monitoring methods are the same, but the I2C slave address can be adjusted according to the actual working conditions. For example, the I2C address of the same type of acceleration card is fixed during hardware design, and the addresses of the same type of acceleration cards are the same. At this time, the monitoring modes corresponding to the same type of acceleration cards are the same, or the addresses of the same type of acceleration cards are dynamically set in different slots, and the monitoring modes of each acceleration card are also different. For the working conditions of different types of acceleration cards in a heterogeneous acceleration card resource pool, the monitoring methods corresponding to any two different types of acceleration cards may be the same or different. The I2C slave addresses corresponding to any two different types of acceleration cards usually vary. At this time, the monitoring modes corresponding to different types of acceleration cards are different. In some special working conditions, such as following the same standard or reusing components, there may be a situation where the I2C slave addresses of different types of acceleration cards are the same. When the monitoring methods are also the same, the monitoring modes corresponding to different types of acceleration cards are the same. Set the monitoring modes of each acceleration card according to the actual engineering needs. This embodiment does not make specific limitations here. The solution of this embodiment is not restricted by whether the acceleration card is allocated, and can realize the dynamic identification of the heterogeneous acceleration card device types under different heterogeneous acceleration card resource pool configurations, and can adopt corresponding monitoring logics for different device types to achieve out-of-band monitoring.

[0037] It can be seen that in this embodiment, the first controller 11 in the interconnection resource module 1 interacts with the firmware of multiple switching components 12 connected with acceleration cards in the interconnection resource module 1, accesses the configuration space of the acceleration cards connected under multiple switching ports of the switching components 12, dynamically obtains the identification information of the acceleration cards, and summarizes the identification information of all acceleration cards in the server. The second controller 21 in the first computing resource module 2 determines the monitoring modes of all acceleration cards in its own first computing resource module 2 based on the identification information of all acceleration cards obtained through the first controller 11, and monitors each acceleration card according to the corresponding monitoring modes, realizing the dynamic identification of the identification information of all acceleration cards, so as to establish a monitoring link that meets the monitoring requirements of heterogeneous acceleration cards, eliminate the blind areas of key functions such as the monitoring of the device operation status and heat dissipation regulation of the server, and ensure the safe and stable operation of the server.

[0038] Based on the above embodiment:

[0039] In an exemplary embodiment, the interconnection resource module 1 includes multiple first communication ports, the first computing resource module 2 includes multiple second communication ports, the multiple first communication ports and the multiple second communication ports are correspondingly connected, and the multiple first communication ports are correspondingly connected with the multiple switching ports;

[0040] The first controller 11 is further configured to establish and send a first mapping relationship according to the corresponding relationship between the first communication ports and the switching ports and the identification information of the acceleration cards connected under the switching ports;

[0041] The second controller 21 is further configured to obtain the identification information of the local acceleration cards according to the corresponding relationship between the second communication ports and the first communication ports and the first mapping relationship.

[0042] In this embodiment, the first communication ports on the interconnection resource module 1 are connected to one switching port on the switching component 12 in a one-to-one correspondence internally, and the first communication ports on the interconnection resource module 1 are connected to one second communication port on the first computing resource module 2 in a one-to-one correspondence through cables, which are used to transmit hardware signals such as I2C and PCIe. The second communication ports on the first computing resource module 2 are connected to one acceleration card internally. It can be understood that one first communication port is connected to one switching port internally and one second communication port externally. The first controller 11 also obtains the identification information of the acceleration cards connected under each switching port. Based on this, a first mapping relationship between the first communication ports, the switching ports, and the identification information of the acceleration cards connected under the switching ports can be established. The first controller 11 can send the first mapping relationship to the second controller 21. The second controller 21 matches the identification information of the acceleration card corresponding to the first communication port from the first mapping relationship according to the information of the first communication port corresponding to its second communication port, so as to determine the type of the acceleration card connected to the second communication port.

[0043] By establishing the first mapping relationship, the first controller 11 can dynamically identify the identification information of the acceleration cards connected under each switching port and transmit this information to the second controller 21. This enables the system to automatically adapt to different types of acceleration cards without manual configuration, and can identify heterogeneous acceleration cards from different manufacturers and models, ensuring the compatibility and universality of the system with various acceleration cards. When a new acceleration card needs to be added or an existing acceleration card needs to be replaced, the system can automatically identify and configure the new device without a complex reconfiguration process. Based on the first mapping relationship and the correspondence of the second communication ports, the second controller 21 can accurately determine the type of each acceleration card and the corresponding monitoring method, thereby achieving precise monitoring.

[0044] In an exemplary embodiment, multiple first communication ports are each configured with an independent first identifier, and multiple switching ports are each configured with an independent second identifier;

[0045] The first controller 11 is further configured to pre-store the correspondence between the first communication ports and the switching ports established through the first identifier and the second identifier.

[0046] In this embodiment, each switching component 12 has 5 external switching ports HP (High-Performance Port) and DP (DisplayPort), and a total of 5 CDFP ports (high-speed signal ports, that is, the first communication ports in this embodiment) are led out. Assuming that the IO Box has a total of four switching boards, and each switching board is provided with two switching components 12, the entire IO Box can finally give out 40 (8×5) CDFP ports externally. Independent first identifiers (Switch_CDFPID) are assigned to the 40 CDFP ports, and can be defined as follows:

[0047] The first identifiers of the first communication ports corresponding to the first layer are respectively: 0xA0, 0xA1, 0xA2, 0xA3, 0xA4; 0xA5, 0xA6, 0xA7, 0xA8, 0xA9; the first identifiers of the first communication ports corresponding to the second layer are respectively: 0xB0, 0xB1, 0xB2, 0xB3, 0xB4; 0xB5, 0xB6, 0xB7, 0xB8, 0xB9; the first identifiers of the first communication ports corresponding to the third layer are respectively: 0xC0, 0xC1, 0xC2, 0xC3, 0xC4; 0xC5, 0xC6, 0xC7, 0xC8, 0xC9; the first identifiers of the first communication ports corresponding to the fourth layer are respectively: 0xD0, 0xD1, 0xD2, 0xD3, 0xD4; 0xD5, 0xD6, 0xD7, 0xD8, 0xD9.

[0048] The five external ports of each switching component 12 are numbered software-wise, i.e., a second identifier (Switch_PortID) is assigned to the switching ports of each switching component 12, such as xP64, xP80, xP96, xP112, xP128, where x represents the number of the switching component 12 (x = 0, 1, 2, 3, 4, 5, 6, 7). After the above hardware design is determined, the corresponding relationship between the second identifier (Switch_PortID) of the switching ports of the switching component 12 and the first identifier (Switch_CDFPID) of the first communication ports on the IO Box can be determined.

[0049] After the corresponding relationship is determined, it can be stored in the firmware of the first controller 11 in the form of a configuration file as follows:

[0050] {

[0051] "Config":

[0052] {

[0053] "Switch_Port": "0P64",

[0054] "Switch_CDFPID": "0xA0"

[0055] },

[0056] {

[0057] "Switch_Port": "0P80",

[0058] "Switch_CDFPID": "0xA1"

[0059] }, ......

[0061] {

[0062] "Switch_Port": "7P128",

[0063] "Switch_CDFPID": "0xD9"

[0064] }

[0066] }

[0067] Correspondingly, the first mapping relationship is as follows:

[0068] {

[0069] "Mapping1":

[0070] { ​

[0071] "Switch_CDFPID": "0xA1",

[0072] "Switch_PortID": "0P80",

[0073] "VenderID": "0x0A",

[0074] "DeviceID": "0x01"

[0075] },

[0076] {

[0077] "Switch_CDFPID": "0xA2",

[0078] "Switch_PortID": "0P96",

[0079] "VenderID": "0x0B",

[0080] "DeviceID": "0x02"

[0081] },

[0082] {

[0083] "Switch_CDFPID": "0xA6",

[0084] "Switch_PortID": "1P80",

[0085] "VenderID": "0x0C",

[0086] "DeviceID": "0x01"

[0087] }, ......

[0090] }

[0091] The first management controller can send the first mapping relationship to the second controllers 21 in each first computing resource module 2 through the Redfish interface.

[0092] In an exemplary embodiment, as shown in Figure 2 the interconnection resource module 1 further includes:

[0093] a plurality of first storage components 13, and the plurality of first storage components 13 are correspondingly connected to the plurality of first communication ports;

[0094] ​The first controller 11 is further configured to write the first identifier of the first communication port connected to the first storage component 13 into the first storage component 13.

[0095] In this embodiment, the interconnection resource module 1 is further provided with a plurality of first storage components 13, and the first storage components 13 are respectively connected to the communication links between the first communication port and the switching port ( Figure 2 only one first storage component 13 corresponding to each switching component 12 is shown as being connected to the communication link between the first communication port and the switching port) for storing the first identifier of the first communication port. Among them, the first storage component 13 can select an I / O expander, which has an interrupt output and a configuration register. Specifically, the SDA (data line) and SCL (clock line) pins of the I / O expander can be connected to the I2C bus, and the I2C address of the I / O expander can be set through the corresponding pins. After power-on reset, the registers of the I / O expander will be set to the default values. By writing to the configuration register through the I2C bus, the direction of the I / O pins (input or output) can be set. The output port register can be selected as the register for storing Switch_CDFPID, and the first controller 11 writes Switch_CDFPID into the selected register through the I2C bus. Through the above method, the Switch_CDFPID of each CDFP port can be stored in the corresponding first storage component 13, so as to achieve a unique identifier, which is convenient for accurately identifying and managing each CDFP port in the system.

[0096] In an exemplary embodiment, a plurality of second communication ports are correspondingly connected to a plurality of slots in the first computing resource module 2, and the slots are used for installing local acceleration cards;

[0097] The second controller 21 is further configured to sequentially scan the plurality of second communication ports, identify the first identifiers stored in the first storage components 13 connected to the plurality of second communication ports, establish a second mapping relationship between the third identifiers of the plurality of slots and the identified plurality of first identifiers, and determine the identifier information of the acceleration cards in the plurality of slots according to the first mapping relationship and the second mapping relationship.

[0098] In this embodiment, the first computing resource module 2 is provided with a plurality of slots, and a plurality of acceleration cards are respectively installed in the slots. The second controller 21 sequentially scans the I2C channels where the plurality of second communication ports of the first computing resource module 2 where it is located are located, so as to identify the first storage component 13 at the IO Box end. Specifically, the second controller 21 sends an address scan command through the I2C bus. When receiving the response signal returned by the first storage component 13 based on the address scan command, the second controller 21 sends a read command to read the register content of the first storage component 13, obtains the first identifier of the CDFP port stored therein, and establishes a second mapping relationship based on the third identifier of the slot connected to the second communication port and the first identifier, where the third identifier may be the acceleration card silk screen (Location).

[0099] Exemplarily, the second mapping relationship is as follows:

[0100] {

[0101] "Mapping2":

[0102] {

[0103] "Switch_CDFPID": "0xA1",

[0104] "Location":"Slot0"

[0105] },

[0106] {

[0107] "Switch_CDFPID": "0xA2",

[0108] "Location":"Slot5"

[0109] }, ......

[0112] }

[0113] Then, based on the first mapping relationship and the second mapping relationship, the final mapping relationship can be established as follows:

[0114] {

[0115] "Mapping3":

[0116] {

[0117] "Switch_CDFPID": "0xA1",

[0118] "Location":"Slot0", ​

[0119] "Switch_Port": "0P80",

[0120] "VenderID": "0x0A",

[0121] "DeviceID": "0x01"

[0122] },

[0123] {

[0124] "Switch_CDFPID": "0xA2",

[0125] "Location":"Slot5",

[0126] "Switch_Port": "0P96",

[0127] "VenderID": "0x0B",

[0128] "DeviceID": "0x02"

[0129] }, ......

[0132] }

[0133] Thus, the VenderID and DeviceID of the acceleration cards installed on each slot can be obtained. The second controller loads different processing logics according to the VenderID and DeviceID of the acceleration cards on different slots to implement the monitoring operation of different heterogeneous acceleration cards.

[0134] The acceleration card monitoring process adopting the above architecture refers to Figure 3 ​As shown, when the entire cabinet of the fusion architecture server is powered on, the first controller interacts with the firmware of each switching component in sequence, accesses the configuration space of the acceleration card devices under each switching port in sequence, and obtains identification information. The first controller aggregates the second identification of the switching ports, the first identification of the first communication ports, and the acceleration card identification information to obtain a first mapping relationship. The first controller sends the first mapping relationship to the second controllers of each second chassis through the communication interface. The second controllers scan the communication channels where each second communication port is located in sequence, identify the first storage component, read the first identification of the stored first communication port, and obtain a second mapping relationship after aggregation. The second controllers aggregate the first mapping relationship and the second mapping relationship to obtain the final mapping relationship, so as to obtain the identification information of the acceleration cards on each slot. The second controllers load different processing logics according to the identification information of the acceleration cards on different slots to implement the monitoring of different acceleration cards. The second controllers report the monitoring information to the first controller, and the first controller presents it after aggregation.

[0135] In an exemplary embodiment, please refer to Figure 4 , the interconnected resource module 1 includes a plurality of first communication ports, the first computing resource module 2 includes a plurality of second communication ports, the plurality of first communication ports and the plurality of second communication ports are correspondingly connected, and the plurality of first communication ports are correspondingly connected to the plurality of switching ports;

[0136] The device monitoring system of the server further includes:

[0137] A plurality of second storage components 3, and the second storage components 3 are arranged on the communication links of the first communication ports and the second communication ports;

[0138] The first controller 11 is further configured to establish a first mapping relationship according to the corresponding relationship between the first communication ports and the switching ports and the identification information of the acceleration cards connected under the switching ports, and write the identification information of the plurality of acceleration cards into the corresponding second storage components 3 according to the first mapping relationship;

[0139] The second controller 21 is further configured to scan the plurality of second communication ports in sequence to obtain the identification information stored in the second storage components 3 arranged on the communication links of the second communication ports.

[0140] In this embodiment, after the entire cabinet is powered on collaboratively, the first controller 11 interacts with the firmware of each switching component 12 through the I2C channel in sequence, accesses the configuration space of the acceleration card devices under each switching port in sequence, obtains the VenderID and DeviceID, and aggregates the second identification Switch_PortID of the switching ports of each switching component 12, the second identification Switch_CDFPID of the first switching port, and the corresponding relationship between the Vendor ID and the DeviceID in the acceleration card configuration space to obtain a first mapping relationship.

[0141] In this embodiment, the communication link between the first communication port and the second communication port includes a connector for connecting a CDFP cable. The second storage component 3 is the storage component in the CDFP connector, which can specifically be an EEPROM.

[0142] The first controller 11 sequentially scans the I2C channels where each first communication port on the IO Box is located according to the first mapping relationship, and writes the VendorID and DeviceID in each accelerator card configuration space into the EEPROM (Electrically Erasable Programmable Read-Only Memory) of the CDFP connector for transmission to the second controller 21. Specifically, the second controller 21 sequentially scans the I2C channels where the second communication ports of the first computing resource module 2 it is located in are located, and reads the VendorID and DeviceID of the accelerator card stored in the EEPROM of the CDFP connector.

[0143] In an exemplary embodiment, multiple first communication ports are each configured with an independent first identifier, and multiple switching ports are each configured with an independent second identifier;

[0144] The first controller 11 is further configured to pre-store the correspondence between the first communication ports and the switching ports established through the first identifier and the second identifier.

[0145] In this embodiment, the solution of configuring independent identifiers for the switching ports and the first communication ports and the solution of establishing the correspondence between the first communication ports and the switching ports based on the first identifier and the second identifier are as described in the above embodiment, and will not be elaborated herein.

[0146] In an exemplary embodiment, multiple second communication ports are correspondingly connected to multiple slots in the first computing resource module 2, and the slots are used for installing local accelerator cards;

[0147] The second controller 21 is further configured to establish a third mapping relationship between the third identifiers of the multiple slots and the obtained multiple identifier information, so as to determine the identifier information of the local accelerator cards in the multiple slots based on the third mapping relationship.

[0148] In this embodiment, taking one second communication port as an example, after the second controller 21 obtains the identifier information of the accelerator card stored in the EEPROM of the CDFP connector corresponding to this second communication port, it establishes a third mapping relationship between the third identifier of the slot corresponding to this second communication port and the identifier information. The second controller 21 loads different processing logics according to the VenderID and DeviceID of each accelerator card in the third mapping relationship to implement the monitoring operation of different accelerator cards.

[0149] In an exemplary embodiment, with reference to Figure 4 , the interconnection resource module 1 further includes:

[0150] A first expansion component 14, a first end of the first expansion component 14 is connected to the first controller 11, and a plurality of second ends of the first expansion component 14 are correspondingly connected to a plurality of first communication ports.

[0151] In this embodiment, in order to further reduce the occupation of the communication ports of the first controller 11, a first expansion component 14 is further provided in this embodiment. A first end of the first expansion component 14 is connected to a communication port of the first controller 11, and a plurality of second ends of the first expansion component 14 are correspondingly connected to a plurality of first communication ports of the interconnection resource module 1 one by one.

[0152] In an exemplary embodiment, the first controller 11 is specifically configured to interact with the firmware of a plurality of switching components 12 periodically, and obtain the identification information of the acceleration card by accessing the configuration space of the acceleration card connected under the switching port of the switching component 12.

[0153] In this embodiment, the first controller 11 periodically obtains the identification information of each acceleration card, so as to identify whether the configuration of the acceleration card has changed, thereby updating the first mapping relationship, so that the second controller 21 can timely adjust the monitoring mode of the changed acceleration card. This embodiment can timely detect changes in the acceleration card configuration (such as hot plugging operation, firmware update, etc.), and this dynamic adaptability ensures that the system can always accurately identify the status of the acceleration card during operation.

[0154] Of course, as another alternative embodiment, the identification information of the new acceleration card can also be obtained after receiving the trigger of the second controller 21. During the operation of the server, if there is a hot plug change of the acceleration card, the second controller 21 detects the change signal and triggers the first controller 11. In this embodiment, unnecessary resource waste can be avoided, and at the same time, the system's rapid response to changes is ensured.

[0155] The support for the hot plug operation of the acceleration card in the above two embodiments enables the system to dynamically adjust the configuration without restarting, improving the availability and user experience of the system.

[0156] In an exemplary embodiment, the second controller 21 is further configured to obtain and output the first monitoring information of the local acceleration card in the first computing resource module 2;

[0157] The first controller 11 is further configured to receive and prompt the first monitoring information.

[0158] In this embodiment, the first monitoring information is the monitoring data obtained by the second controller 21 (BMC) regarding the operating status of the local acceleration card. This information can be used to evaluate the health status, performance, and whether there are any abnormalities of the acceleration card, including but not limited to temperature information: the chip temperature of the acceleration card (such as the temperature of the GPU, FPGA, etc.), the status of the cooling system (such as the fan speed); power consumption information: the real-time power consumption data of the acceleration card, the voltage and current status of the power supply module; performance metrics: the computing performance of the acceleration card (such as FPS (Frames Per Second), TFLOPS (Tera FLoating-Point Operations Per Second), etc.), memory usage rate and bandwidth utilization; hardware status: the presence status of the acceleration card (whether it is plugged in or unplugged), hardware fault information (such as hardware errors, memory errors, etc.); operating status, the task status running on the acceleration card (such as task completion rate, task queue length), abnormal events in the System Event Log; firmware information: the firmware version of the acceleration card, the firmware operating status (such as whether an update is required); sensor data: environmental data collected by sensors (such as humidity, air pressure, etc.).

[0159] The first controller 11 receives and prompts the first monitoring information, and specifically can display the monitoring information in real time on the server management interface (such as the Web interface), or display the monitoring information through the console or command line tool (such as the IPMI (Intelligent Platform Management Interface) tool).

[0160] This embodiment obtains the operating status of the acceleration card in real time, discovers and processes faults in a timely manner, and can improve the reliability and stability of the system.

[0161] In an exemplary embodiment, the second controller 21 is further configured to determine and execute the adjustment operation corresponding to the full-box adjustment condition when the first monitoring information of the local acceleration card meets the full-box adjustment condition.

[0162] In this embodiment, there can be multiple whole-box adjustment conditions. For example, the first adjustment condition can be that the average temperature of the acceleration cards in the first computing resource module 2 exceeds the preset threshold. When the first monitoring information meets this adjustment condition, the optional adjustment operations include, but are not limited to, increasing the fan speed to enhance heat dissipation, reducing the power consumption of the acceleration cards (such as reducing the GPU frequency) to reduce heat generation, triggering an alarm to notify the operation and maintenance personnel to check the heat dissipation system, etc. The second adjustment condition can be that the total power consumption of the acceleration cards in the box exceeds the preset threshold (such as the maximum output power of the power supply module). When the first monitoring information meets this adjustment condition, the optional adjustment operations include, but are not limited to, dynamically adjusting the power consumption allocation of the acceleration cards, giving priority to ensuring the power consumption requirements of critical tasks, turning off some non-critical acceleration cards to reduce the total power consumption, triggering an alarm to notify the operation and maintenance personnel to check the status of the power supply module, etc. The third adjustment condition can be that the performance indicators of the acceleration cards (such as FPS, TFLOPS) are lower than the preset threshold. When the first monitoring information meets this adjustment condition, the optional adjustment operations include, but are not limited to, adjusting the operating frequency or voltage of the acceleration cards to optimize performance, reallocating the task load, migrating high-load tasks to acceleration cards with higher performance, triggering an alarm to notify the operation and maintenance personnel to optimize the task scheduling. The fourth adjustment condition can be that the hardware status of the acceleration cards is detected to be abnormal (such as hardware errors, memory errors). When the first monitoring information meets this adjustment condition, the optional adjustment operations include, but are not limited to, migrating the tasks of the faulty acceleration cards to other normal acceleration cards, triggering an alarm to notify the operation and maintenance personnel to perform hardware repair or replacement, and recording the fault information in the system event log (SEL).

[0163] Of course, in addition to including the above whole-box adjustment conditions, other adjustment conditions and corresponding adjustment operations can also be set according to the actual engineering needs, which are not specifically limited in this embodiment.

[0164] In an exemplary embodiment, referring to Figure 5 , the server further includes a second computing resource module 4;

[0165] The device monitoring system of the server further includes:

[0166] A third controller 41 disposed in the second computing resource module 4, configured to receive and prompt the first monitoring information.

[0167] The second computing resource module 4 in this embodiment is a Host Box, which includes a third communication port, a processor 42, and a third controller 41. The third controller 41 and the processor 42 communicate through LPC, the processor 42 and the third communication port communicate through PCIe, and the third controller 41 and the third communication port communicate through I2C. The third communication port in the second computing resource module 4 is correspondingly connected to the first communication port in the interconnection resource module 1 through a cable. The first monitoring information obtained by the second controller 21 can also be prompted through the third controller 41.

[0168] In an exemplary embodiment, the first controller 11 is further configured to obtain and output the second monitoring information corresponding to the interconnected resource module 1;

[0169] The third controller 41 is further configured to determine and execute an adjustment operation corresponding to the overall machine adjustment condition when the received first monitoring information and / or second monitoring information meets the overall machine adjustment condition.

[0170] In this embodiment, the first monitoring information is obtained by the second controller 21 in the first computing resource module 2, including the operating status of its local acceleration card (such as temperature, power consumption, performance metrics, hardware status, etc.). The second monitoring information is obtained by the first controller 11, and the monitoring information related to the interconnected resource module 1, for example: the status of the switching component 12: the temperature, power consumption, bandwidth utilization, etc. of the switching chip; the status of the I / O port: the connection status, data transfer rate, error rate, etc. of the I / O port; the status of the power supply module: the voltage, current, temperature, etc. of the power supply module; the environmental sensor data: the environmental data such as the temperature, humidity, air pressure, etc. inside the chassis.

[0171] The overall machine adjustment conditions and the corresponding adjustment operations include but are not limited to the following examples:

[0172] Temperature-related overall machine adjustment condition: The first monitoring information or the second monitoring information shows that the overall system temperature exceeds a preset threshold (such as 75 °C). Adjustment operation: Increase the rotation speed of all chassis fans; reduce the power consumption of the acceleration card and the switching component 12 to reduce heat generation; trigger an alarm to notify the operation and maintenance personnel to check the heat dissipation system.

[0173] Power consumption-related overall machine adjustment condition: The first monitoring information or the second monitoring information shows that the total system power consumption exceeds the maximum output power of the power supply module. Adjustment operation: Dynamically adjust the power consumption allocation of the acceleration card and the switching component 12 according to the task priority; migrate high-power tasks to other low-power devices or chassis; trigger an alarm to notify the operation and maintenance personnel to check the status of the power supply module.

[0174] Performance-related overall machine adjustment condition: The first monitoring information or the second monitoring information shows that the overall system performance is lower than a preset threshold (such as the overall computing task completion time is extended). Adjustment operation: Reallocate the task load, optimize the utilization rate of the acceleration card and the switching component 12, adjust the operating frequency or voltage of the acceleration card, optimize the performance, and trigger an alarm to notify the operation and maintenance personnel to optimize the task scheduling.

[0175] Hardware status-related overall machine adjustment condition: The first monitoring information or the second monitoring information shows that the hardware status is abnormal (such as a failure of the acceleration card or the switching component 12). Adjustment operation: Migrate the tasks of the faulty device to other normal devices, enable standby devices or ports to ensure the normal operation of the system, and trigger an alarm to notify the operation and maintenance personnel to perform hardware repair or replacement.

[0176] Of course, in addition to including the above-mentioned whole machine adjustment conditions, other adjustment conditions and corresponding adjustment operations can also be set according to the actual engineering needs, which are not specifically limited in this embodiment.

[0177] In an exemplary embodiment, the device monitoring system of the server further includes:

[0178] A network switch 5, the first network port of the network switch 5 is connected to the first controller 11, the second network port of the network switch 5 is connected to the second controller 21, and the third port of the network switch 5 is connected to the third controller 41.

[0179] In this embodiment, for the fusion architecture system aiming to improve the IO expansion ability, referring to Figure 5 , there is 1 interconnection resource module 1, 2 first computing resource modules 2 and 8 second computing resource modules 4 in the whole cabinet. One BMC module is designed in each chassis to be responsible for Box management. All BMCs are connected under the same network switch 5 to form an internal local area network. The BMC in the interconnection resource module 1 serves as the pooled management controller PSMC (i.e., the first controller 11 in this embodiment), and the PSMC can communicate with each device Box BMC (the second controller 21 and the third controller) through the network to implement the management and control functions of the entire system.

[0180] In a second aspect, the present invention also provides a server, including an interconnection resource module, a plurality of first computing resource modules, a plurality of second computing resource modules, and the device monitoring system of the server described in any one of the above embodiments;

[0181] The interconnection resource module includes a plurality of switching components and a plurality of first communication ports, the first computing resource module includes a plurality of acceleration cards and a plurality of second communication ports, the second computing resource module includes a processor and a third communication port, the interconnection resource module and the first computing resource module are connected through the first communication port and the second communication port, and the interconnection resource module and the second computing resource module are connected through the third communication port and the first communication port.

[0182] In an exemplary embodiment, the interconnection resource module includes a plurality of switching boards, and at least one switching component is provided on the switching board;

[0183] The interconnection resource module further includes a second expansion component, the first end of the second expansion component is connected to the first controller in the interconnection resource module, and a plurality of second ports of the second expansion component are correspondingly connected to the switching components on the plurality of switching boards.

[0184] In an exemplary embodiment, the first computing resource module further includes a plurality of signal retiming components. The first end of the signal retiming component is connected to the second communication port, and the second end of the signal retiming component is connected to the acceleration card.

[0185] In an exemplary embodiment, the first computing resource module further includes a third expansion component. The first end of the third expansion component is connected to the second controller, and a plurality of second ends of the third expansion component are correspondingly connected to a plurality of acceleration cards in the first computing resource module.

[0186] In an exemplary embodiment, the server further includes a heat dissipation component configured to perform a heat dissipation operation in response to a heat dissipation regulation instruction output by the first controller in the interconnection resource module or the second controller in the first computing resource module or the third controller in the second computing resource module.

[0187] The device monitoring system of the server includes:

[0188] A first controller disposed in the interconnection resource module and connected to a plurality of switching components in the interconnection resource module. The first controller is configured to interact with the firmware of the plurality of switching components to obtain the identification information of the acceleration cards connected under the switching ports of the switching components;

[0189] A second controller disposed in the first computing resource module and connected to a plurality of acceleration cards in the first computing resource module. The second controller is configured to determine the monitoring modes of the plurality of acceleration cards according to the identification information of the plurality of acceleration cards obtained by the first controller, and perform monitoring operations on the plurality of acceleration cards according to the monitoring modes of the plurality of acceleration cards.

[0190] In an exemplary embodiment, the interconnection resource module includes a plurality of first communication ports, the first computing resource module includes a plurality of second communication ports, the plurality of first communication ports and the plurality of second communication ports are correspondingly connected, and the plurality of first communication ports are correspondingly connected to a plurality of switching ports;

[0191] The first controller is further configured to establish and send a first mapping relationship according to the corresponding relationship between the first communication port and the switching port and the identification information of the acceleration card connected under the switching port;

[0192] The second controller is further configured to obtain the identification information of the local acceleration card according to the corresponding relationship between the second communication port and the first communication port and the first mapping relationship.

[0193] In an exemplary embodiment, each of the plurality of first communication ports is configured with an independent first identifier, and each of the plurality of switching ports is configured with an independent second identifier;

[0194] The first controller is further configured to pre-store the corresponding relationship between the first communication port and the switching port established through the first identifier and the second identifier.

[0195] In an exemplary embodiment, the interconnection resource module further includes:

[0196] A plurality of first storage components, and the plurality of first storage components are correspondingly connected to a plurality of first communication ports;

[0197] The first controller is further configured to write the first identifier of the first communication port connected to the first storage component into the first storage component.

[0198] In an exemplary embodiment, a plurality of second communication ports are correspondingly connected to a plurality of slots in the first computing resource module, and the slots are used for installing local acceleration cards;

[0199] The second controller is further configured to sequentially scan the plurality of second communication ports, identify the first identifiers stored in the first storage components connected to the plurality of second communication ports, establish a second mapping relationship between the third identifiers of the plurality of slots and the identified plurality of first identifiers, and determine the identification information of the acceleration cards in the plurality of slots according to the first mapping relationship and the second mapping relationship.

[0200] In an exemplary embodiment, the interconnection resource module includes a plurality of first communication ports, the first computing resource module includes a plurality of second communication ports, the plurality of first communication ports and the plurality of second communication ports are correspondingly connected, and the plurality of first communication ports are correspondingly connected to a plurality of switching ports;

[0201] The device monitoring system of the server further includes:

[0202] A plurality of second storage components, and the second storage components are arranged on the communication link between the first communication port and the second communication port;

[0203] The first controller is further configured to establish a first mapping relationship according to the correspondence between the first communication port and the switching port and the identification information of the acceleration card connected under the switching port, and write the identification information of the plurality of acceleration cards into the corresponding second storage components according to the first mapping relationship;

[0204] The second controller is further configured to sequentially scan the plurality of second communication ports and obtain the identification information stored in the second storage components arranged on the communication link of the second communication ports.

[0205] In an exemplary embodiment, each of the plurality of first communication ports is configured with an independent first identifier, and each of the plurality of switching ports is configured with an independent second identifier;

[0206] The first controller is further configured to pre-store the correspondence between the first communication port and the switching port established through the first identifier and the second identifier.

[0207] In an exemplary embodiment, a plurality of second communication ports are correspondingly connected to a plurality of slots in the first computing resource module, and the slots are used for installing local acceleration cards;

[0208] The second controller is further configured to establish a third mapping relationship between the third identifiers of the multiple slots and the obtained multiple identifier information, so as to determine the identifier information of the acceleration cards in the multiple slots based on the third mapping relationship.

[0209] In an exemplary embodiment, the interconnection resource module further includes:

[0210] A first expansion component, a first end of the first expansion component is connected to the first controller, and multiple second ends of the first expansion component are correspondingly connected to multiple first communication ports.

[0211] In an exemplary embodiment, the first controller is specifically configured to interact with the firmware of the multiple switching components periodically to obtain the identifier information of the acceleration cards connected under the switching ports of the switching components.

[0212] In an exemplary embodiment, the second controller is further configured to obtain and output first monitoring information of the multiple acceleration cards in the first computing resource module;

[0213] The first controller is further configured to receive and prompt the first monitoring information.

[0214] In an exemplary embodiment, the second controller is further configured to determine and execute an adjustment operation corresponding to the full-box adjustment condition when the first monitoring information of the multiple acceleration cards meets the full-box adjustment condition.

[0215] In an exemplary embodiment, the device monitoring system of the server further includes:

[0216] A third controller disposed in the second computing resource module, configured to receive and prompt the first monitoring information.

[0217] In an exemplary embodiment, the first controller is further configured to obtain and output second monitoring information corresponding to the interconnection resource module;

[0218] The third controller is further configured to determine and execute an adjustment operation corresponding to the whole-machine adjustment condition when the received first monitoring information and / or second monitoring information meets the whole-machine adjustment condition.

[0219] In an exemplary embodiment, the device monitoring system of the server further includes:

[0220] A network switch, a first network port of the network switch is connected to the first controller, a second network port of the network switch is connected to the second controller, and a third port of the network switch is connected to the third controller.

[0221] In a third aspect, the present invention further provides a method for monitoring devices of a server, which is applied to a first controller. The first controller is disposed in an interconnection resource module of the server, and the first controller is connected to at least one switching component in the interconnection resource module. The server further includes a first computing resource module. The method for monitoring devices of the server includes:

[0222] Interact with the firmware of at least one switching component;

[0223] Obtain the identification information of the acceleration card connected under the switching port of the switching component, so that the second controller disposed in the first computing resource module can determine the monitoring mode of the local acceleration card according to the identification information of the acceleration card obtained by the first controller, and perform monitoring operations on the local acceleration card according to the monitoring mode of the local acceleration card; the local acceleration card is the acceleration card disposed in the same first computing resource module as the second controller.

[0224] In a fourth aspect, the present invention further provides a method for monitoring devices of a server, which is applied to a second controller. The second controller is disposed in the first computing resource module of the server, and the second controller is connected to the local acceleration card. The server further includes an interconnection resource module. The method for monitoring devices of the server includes:

[0225] Obtain the identification information of the acceleration card connected under the switching port of the switching component obtained by the interaction between the first controller in the interconnection resource module and at least one switching component;

[0226] Determine the monitoring mode of the local acceleration card according to the identification information of the acceleration card obtained by the first controller;

[0227] Perform monitoring operations on the local acceleration card according to the monitoring mode of the local acceleration card; the local acceleration card is the acceleration card disposed in the same first computing resource module as the second controller.

[0228] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0229] The above has introduced in detail a server, its device monitoring system and method provided by the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and modifications can still be made to the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.

Claims

1. A server equipment monitoring system, characterized in that: The server includes an interconnection resource module and at least one first computing resource module, and the device monitoring system of the server includes: A first controller provided in the interconnected resource module is connected to at least one switching component in the interconnected resource module, and the first controller is configured to interact with firmware of at least one of the switching components to obtain identification information of an accelerator card connected to a switching port of the switching component; a second controller provided in the first computing resource module, connected to a local acceleration card, the second controller being configured to determine a monitoring mode of the local acceleration card according to identification information of the local acceleration card acquired by the first controller, and to perform a monitoring operation on the local acceleration card according to the monitoring mode of the local acceleration card, the local acceleration card being the acceleration card provided in the same first computing resource module as the second controller; The monitoring mode is determined based on an I2C slave address and a monitoring method corresponding to a type determined by identification information of the local accelerator card. The monitoring operation includes establishing an out-of-band monitoring channel of the local accelerator card according to the I2C slave address, and monitoring the local accelerator card according to the monitoring method.

2. The device monitoring system for a server according to claim 1, characterized in that: The interconnection resource module includes a plurality of first communication ports, the first computing resource module includes a plurality of second communication ports, the plurality of first communication ports are correspondingly connected to the plurality of second communication ports, and the plurality of first communication ports are correspondingly connected to the plurality of switch ports; The first controller is further configured to establish and send a first mapping relationship according to the correspondence between the first communication port and the switch port and the identification information of the accelerator card connected to the switch port; The second controller is further configured to obtain identification information of the local acceleration card according to a correspondence between the second communication port and the first communication port and the first mapping relationship.

3. The device monitoring system of the server according to claim 2, characterized in that: The plurality of first communication ports are each configured with an independent first identifier, and the plurality of switch ports are each configured with an independent second identifier; The first controller is further configured to pre-store a corresponding relationship between the first communication port and the switch port established by the first identifier and the second identifier.

4. The device monitoring system for a server according to claim 3, characterized in that: The interconnection resource module also includes: A plurality of first storage components, wherein the plurality of first storage components are correspondingly connected to the plurality of first communication ports; The first controller is further configured to write a first identifier of the first communication port connected to the first storage component into the first storage component.

5. The device monitoring system for a server according to claim 4, characterized in that: The plurality of the second communication ports are correspondingly connected to the plurality of slots in the first computing resource module, and the slots are used to install the local acceleration card; The second controller is also configured to scan the multiple second communication ports in sequence, identify the first identifier stored in the first storage component connected to the multiple second communication ports, establish a second mapping relationship between the third identifiers of the multiple slots and the identified multiple first identifiers, and determine the identification information of the local acceleration cards in the multiple slots based on the first mapping relationship and the second mapping relationship.

6. The device monitoring system for a server according to claim 1, characterized in that: The interconnection resource module includes a plurality of first communication ports, the first computing resource module includes a plurality of second communication ports, the plurality of first communication ports are correspondingly connected to the plurality of second communication ports, and the plurality of first communication ports are correspondingly connected to the plurality of switch ports; The equipment monitoring system of the server also includes: a plurality of second storage components, wherein the second storage components are arranged on the communication link between the first communication port and the second communication port; The first controller is further configured to establish a first mapping relationship according to the correspondence between the first communication port and the switch port and the identification information of the acceleration card connected to the switch port, and write the identification information of the plurality of acceleration cards into the corresponding second storage component according to the first mapping relationship; The second controller is also configured to scan the plurality of second communication ports in sequence, and obtain identification information stored in a second storage component on a communication link of the second communication port.

7. The device monitoring system for a server according to claim 6, characterized in that: The plurality of first communication ports are each configured with an independent first identifier, and the plurality of switch ports are each configured with an independent second identifier; The first controller is further configured to pre-store a corresponding relationship between the first communication port and the switch port established by the first identifier and the second identifier.

8. The device monitoring system for a server according to claim 7, characterized in that: The plurality of the second communication ports are correspondingly connected to the plurality of slots in the first computing resource module, and the slots are used to install the local acceleration card; The second controller is further configured to establish a third mapping relationship between third identifiers of the plurality of slots and the acquired plurality of identification information, so as to determine identification information of the local acceleration cards in the plurality of slots based on the third mapping relationship.

9. The device monitoring system for a server according to claim 6, characterized in that: The interconnection resource module also includes: A first expansion component, wherein a first end of the first expansion component is connected to the first controller, and a plurality of second ends of the first expansion component are correspondingly connected to a plurality of the first communication ports.

10. The equipment monitoring system for a server according to claim 1, characterized in that: The first controller is specifically configured to interact with the firmware of the plurality of switching components periodically to obtain identification information of the acceleration card connected to the switching port of the switching component.

11. The device monitoring system for a server according to any one of claims 1 to 9, characterized in that: The second controller is further configured to obtain and output first monitoring information of the local acceleration card in the first computing resource module; The first controller is further configured to receive and prompt the first monitoring information.

12. The device monitoring system for a server according to claim 11, characterized in that: The second controller is further configured to determine and execute an adjustment operation corresponding to the whole box adjustment condition when the first monitoring information of the local acceleration card meets the whole box adjustment condition.

13. The device monitoring system for a server according to claim 11, characterized in that: The server also includes a second computing resource module; The equipment monitoring system of the server also includes: A third controller disposed in the second computing resource module is configured to receive and prompt the first monitoring information.

14. The device monitoring system for a server according to claim 13, characterized in that: The first controller is further configured to obtain and output second monitoring information corresponding to the interconnected resource module; The third controller is further configured to determine and execute an adjustment operation corresponding to the whole machine adjustment condition when the received first monitoring information and / or second monitoring information meets the whole machine adjustment condition.

15. The device monitoring system for a server according to claim 13, characterized in that: The equipment monitoring system of the server also includes: A network switch, wherein a first network port of the network switch is connected to the first controller, a second network port of the network switch is connected to the second controller, and a third port of the network switch is connected to the third controller.

16. A server, characterized in that: A device monitoring system comprising an interconnected resource module, a plurality of first computing resource modules, a plurality of second computing resource modules, and a server as described in any one of claims 1 to 15; The interconnected resource module includes multiple switching components and multiple first communication ports, the first computing resource module includes multiple accelerator cards and multiple second communication ports, the second computing resource module includes a processor and a third communication port, the interconnected resource module and the first computing resource module are connected through the first communication port and the second communication port, and the interconnected resource module and the second computing resource module are connected through the third communication port and the first communication port.

17. The server according to claim 16, characterized in that The interconnection resource module includes a plurality of switch boards, and at least one switch component is disposed on the switch board; The interconnected resource module further includes a second extension component, a first end of the second extension component is connected to the first controller in the interconnected resource module, and a plurality of second ports of the second extension component are correspondingly connected to the switching components on the plurality of switching boards.

18. The server according to claim 16, characterized in that: The first computing resource module further includes a plurality of signal retiming components, a first end of the signal retiming component is connected to the second communication port, and a second end of the signal retiming component is connected to the acceleration card.

19. The server according to claim 18, characterized in that The first computing resource module also includes a third expansion component, a first end of the third expansion component is connected to the second controller, and multiple second ends of the third expansion component are correspondingly connected to multiple acceleration cards in the first computing resource module.

20. The server according to any one of claims 16 to 19, characterized in that: The server also includes a heat dissipation component, which is configured to respond to a heat dissipation control instruction output by the first controller in the interconnected resource module or the second controller in the first computing resource module or the third controller in the second computing resource module to perform a heat dissipation operation.

21. A device monitoring method for a server, characterized in that: Applied to a first controller, the first controller is arranged in an interconnection resource module of the server, the first controller is connected to at least one switching component in the interconnection resource module, the server further comprises a first computing resource module, and the device monitoring method of the server comprises: interacting with firmware of at least one of the switching components; Obtain identification information of an accelerator card connected to a switching port of the switching component, so that a second controller disposed in the first computing resource module determines a monitoring mode of a local accelerator card according to the identification information of the accelerator card obtained by the first controller, and performs a monitoring operation on the local accelerator card according to the monitoring mode of the local accelerator card; the local accelerator card is an accelerator card disposed in the same first computing resource module as the second controller; the monitoring mode is determined based on an I2C slave address and a monitoring mode corresponding to a type determined by the identification information of the local accelerator card, and the monitoring operation includes establishing an out-of-band monitoring channel of the local accelerator card according to the I2C slave address, and monitoring the local accelerator card according to the monitoring mode.

22. A device monitoring method for a server, characterized in that: Applied to a second controller, the second controller is provided in a first computing resource module of the server, the second controller is connected to a local acceleration card, the server further includes an interconnection resource module, and the device monitoring method of the server includes: Acquire identification information of an accelerator card connected to a switch port of a switch component obtained by interaction between a first controller in the interconnected resource module and firmware of at least one switch component; determining a monitoring mode of the local accelerator card according to the identification information of the accelerator card acquired by the first controller; The local accelerator card is monitored according to the monitoring mode of the local accelerator card; the local accelerator card is the accelerator card that is arranged in the same first computing resource module as the second controller; the monitoring mode is determined based on an I2C slave address and a monitoring mode corresponding to a type determined by identification information of the local accelerator card, and the monitoring operation includes establishing an out-of-band monitoring channel of the local accelerator card according to the I2C slave address, and monitoring the local accelerator card according to the monitoring mode.

Citation Information

Patent Citations

  • Multi-accelerator card heterogeneous server and resource link reconstruction method

    CN117687956A

  • Method and system for determining mapping relation, storage medium and electronic device

    CN117978811A