Method for managing a plurality of GPU boxes, product, device and storage medium

By acquiring the identification information of the GPU BOX and switching the bus link, the problem of managing multiple GPU BOXes in a dual-node server is solved, and effective management of different forms and connection methods is achieved.

WO2026086283A1PCT designated stage Publication Date: 2026-04-30INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
INSPUR SUZHOU INTELLIGENT TECH CO LTD
Filing Date
2025-07-07
Publication Date
2026-04-30

AI Technical Summary

Technical Problem

In dual-node servers, existing technologies cannot effectively manage multiple GPU boxes, especially with different numbers and configurations of ports and interconnection methods. They cannot determine the current configuration, leading to management difficulties.

Method used

By obtaining the identification information of the GPU BOX, determining the total number and connection method, and using a gate and/or bus arbiter to switch the bus link of the target communication port, channel management of multiple GPU BOXes can be achieved.

Benefits of technology

It enables efficient management of multiple GPU boxes in a dual-node server, applicable to different numbers and configurations of ports and interconnection methods, thus improving management efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025107326_30042026_PF_FP_ABST
    Figure CN2025107326_30042026_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of computers. Disclosed are a method and apparatus for managing a plurality of GPU BOXes, a device and a storage medium. The method comprises: acquiring identity identification information sent by second communication ports in a plurality of GPU BOXes connected to a plurality of first communication ports on a dual-node server panel; on the basis of the identity identification information, determining the total number of the GPU BOXes and connection types between the GPU BOXes and the first communication ports; on the basis of the total number and the connection types, determining the connection type between the current dual-node server and the plurality of GPU BOXes and, on the basis of the connection type, determining a target communication port from among the plurality of first communication ports; and, by means of a selector and / or a bus arbiter, switching a target bus link corresponding to the target communication port to the target communication port, so as to perform corresponding channel management on the corresponding GPU BOX by means of the target communication port. The present application manages the plurality of GPU BOXes by the dual-node server.
Need to check novelty before this filing date? Find Prior Art

Description

A method, product, device, and storage medium for managing multi-GPU boxes.

[0001] Cross-reference to related applications

[0002] This application claims priority to Chinese Patent Application No. 202411481608.0, filed on October 23, 2024, entitled "A Multi-GPU BOX Management Method, Apparatus, Device and Storage Medium", the entire contents of which are incorporated herein by reference. Technical Field

[0003] This application relates to the field of computer technology, and in particular to a multi-GPU BOX management method, apparatus, device, and storage medium. Background Technology

[0004] With the rapid development of AI (Artificial Intelligence) computing technology, the scale of AI servers that need to be deployed is also increasing, with high-density GPU BOX (Graphics Processing Unit Box) being the primary type. A GPU BOX contains a large number of GPUs (Graphics Processing Units) for model training, inference, and other tasks. Furthermore, the GPU BOX connects to the server host, enabling management and business interaction functions.

[0005] In physical implementation, the GPU BOX exposes a CDFP (Compact Data Center Fabric Port, a high-speed data transmission port) and connects to the CDFP port of the server host via a DAC (Direct Attach Cable). In addition to data services, the server host's BMC (Baseboard Management Controller) and the GPU BOX's internal BMC can also perform out-of-band management and communication. The BMC of the server host and the GPU BOX typically use an IPMB (Intelligent Platform Management Bus) channel, physically employing an I2C (Inter-Integrated Circuit bus).

[0006] Currently, in a single-node server, managing the GPU BOX can be achieved simply by adding an I2C link to the CDFP port connecting to the GPU BOX and connecting it to the BMC of the server motherboard. That is, a single GPU BOX only requires one IPMB channel for host management. However, to increase the density of server hosts within a single rack, two server motherboards are often placed in the same rack, creating a dual-node server. Currently, there is no specific configuration for connecting GPU BOXes to dual-node servers. Since dual-node servers have their own independent hardware and operating systems, managing multiple GPU BOXes from a dual-node server requires meeting various adaptation scenarios, such as one server with two GPU BOXes or one server with four GPU BOXes. Furthermore, the interconnection methods between the dual-node server and the GPU BOX may vary, such as a GPU BOX's service link being connected to two server nodes simultaneously or only to one server node, resulting in different management logic and paths. Summary of the Invention

[0007] In view of this, the purpose of this application is to provide a multi-GPU BOX management method, apparatus, device, and storage medium, which can solve the problem that when a dual-node server host is equipped with GPU BOXes of different port numbers and configurations, and when the host and GPU BOXes have different interconnection methods, the dual-node server cannot determine the current configuration and interconnection method, and therefore cannot manage multiple GPU BOXes. This enables the dual-node server to manage multiple GPU BOXes. The specific solution is as follows:

[0008] In a first aspect, this application discloses a multi-GPU BOX management method applied to a dual-node server, comprising:

[0009] Each GPU BOX connected to multiple first communication ports on the dual-node server panel obtains the identification information sent by the second communication port of each of the multiple GPU BOXes connected to the multiple first communication ports on the dual-node server panel; the identification information is used to uniquely identify the GPU BOX corresponding to the second communication port.

[0010] The total number of GPU BOXes and the connection method between each GPU BOX and the first communication port are determined based on multiple identity information.

[0011] The connection type between the current dual-node server and multiple GPU BOXes is determined based on the total number and connection method, and the target communication port is determined from multiple first communication ports based on the connection type.

[0012] The target bus link corresponding to the target communication port is switched to the target communication port through the swivel and / or bus arbitrator, so as to perform corresponding channel management on the corresponding GPU BOX through the target communication port.

[0013] In some embodiments, the connection type includes a first connection type, a second connection type, a third connection type, and a fourth connection type;

[0014] The four connection types are as follows: the first type is a dual-node server connecting two GPU boxes, with the communication port of each GPU box connected to the communication port of only one server node; the second type is a dual-node server connecting two GPU boxes, with the communication port of each GPU box connected to the communication ports of both server nodes; the third type is a dual-node server connecting four GPU boxes, with the communication port of each GPU box connected to the communication port of only one server node; and the fourth type is a dual-node server connecting four GPU boxes, with the communication port of each GPU box connected to the communication ports of both server nodes.

[0015] Before obtaining the identity information sent by the second communication ports of each of the multiple GPU BOXes connected to the multiple first communication ports on the dual-node server panel, the process also includes:

[0016] According to preset rules, target communication ports for channel management are selected from multiple first communication ports on the dual-node server panel, corresponding to the first connection type, second connection type, third connection type, and fourth connection type.

[0017] The number of target communication ports corresponding to the first and second connection types is 2, and the number of target communication ports corresponding to the third and fourth connection types is 4. The management channel of each target communication port corresponding to the first and third connection types comes from a uniquely determined server node, while the management channel of each target communication port corresponding to the second and fourth connection types comes from a non-unique server node. The server node is either the first server node or the second server node in a dual-node server.

[0018] In some embodiments, the connection type between the current dual-node server and multiple GPU BOXes is determined based on the total number and connection method, including:

[0019] When the total number is 2, and the connection method is that multiple second communication ports in each GPU BOX are only connected to the communication ports of the first server node or the second server node, the connection type between the current dual-node server and multiple GPU BOX is determined to be the first connection type.

[0020] In some embodiments, the connection type between the current dual-node server and multiple GPU BOXes is determined based on the total number and connection method, including:

[0021] When the total number is 2, and the connection method is that the second communication port in each GPU BOX is connected to the communication ports of both the first server node and the second server node, the connection type between the current dual-node server and multiple GPU BOXes is determined to be the second connection type.

[0022] In some embodiments, the connection type between the current dual-node server and multiple GPU BOXes is determined based on the total number and connection method, including:

[0023] When the total number is 4, and the connection method is that the second communication port in each GPU BOX is only connected to the communication port of the first server node or the second server node, the connection type between the current dual-node server and multiple GPU BOXes is determined to be the third connection type.

[0024] In some embodiments, the connection type between the current dual-node server and multiple GPU BOXes is determined based on the total number and connection method, including:

[0025] When the total number is 4, and the connection method is that the second communication port in each GPU BOX is connected to the communication ports of the first server node and the second server node, the connection type between the current dual-node server and multiple GPU BOXes is determined to be the fourth connection type.

[0026] In some embodiments, switching the target bus link corresponding to the target communication port to the target communication port via a selector includes:

[0027] Enable the corresponding target bus link, and switch the target bus link corresponding to the target communication port to the target communication port through the selector.

[0028] In some embodiments, switching the target bus link corresponding to the target communication port to the target communication port via a selector and / or a bus arbiter includes:

[0029] Enable the corresponding target bus link, and switch the target bus link corresponding to the target communication port to the target communication port through a selector and / or bus arbitrator.

[0030] In some embodiments, the multi-GPU BOX management method further includes:

[0031] Detect whether the data fields on the target bus link have been completely transmitted;

[0032] If the data fields on the target bus link have been transmitted, the target bus link is released and another server node preempts the target bus link.

[0033] The target bus link is routed to the target communication port via another server node, so that the corresponding GPU BOX can be managed through the other server node and the target communication port.

[0034] In some embodiments, the multi-GPU BOX management method further includes:

[0035] A preset number of signal pins are reserved in the second communication port of each GPU BOX.

[0036] In some embodiments, the identification information sent by each of the second communication ports in multiple GPU BOXes connected to multiple first communication ports on a dual-node server panel is obtained, including:

[0037] The system obtains and sends the identity information of each second communication port in multiple GPU BOXes connected to multiple first communication ports on the dual-node server panel after performing resistor pull-up and pull-down processing on a preset number of signal pins.

[0038] In some embodiments, determining the total number of GPU BOXes and the connection method between each GPU BOX and the first communication port based on multiple identity information includes:

[0039] Based on the order of multiple first communication ports, sort the multiple identity information corresponding to the GPU BOX to obtain sorted identity information;

[0040] The total number of GPU BOXes and the connection method between each GPU BOX and the first communication port are determined based on the sorted identification information.

[0041] Secondly, this application discloses a computer program product, including a computer program / instruction, which, when executed by a processor, implements the aforementioned multi-GPU BOX management method.

[0042] Thirdly, this application discloses an electronic device, including a processor and a memory; wherein the processor implements the aforementioned multi-GPU BOX management method when executing a computer program stored in the memory.

[0043] Fourthly, this application discloses a computer non-volatile readable storage medium for storing computer programs; wherein, when the computer program is executed by a processor, it implements the aforementioned multi-GPU BOX management method.

[0044] As can be seen, this application is applied to a dual-node server. First, it obtains the identity information sent by the second communication port of each of the multiple GPU BOXes connected to the multiple first communication ports on the dual-node server panel. Then, based on the multiple identity information, it determines the total number of GPU BOXes and the connection method between each GPU BOX and the first communication port. Based on the total number and connection method, it determines the connection type between the current dual-node server and the multiple GPU BOXes. Based on the connection type, it determines the target communication port from the multiple first communication ports. Then, it switches the target bus link corresponding to the target communication port to the target communication port through a selector and / or a bus arbiter, so as to perform corresponding channel management on the corresponding GPU BOX through the target communication port. This application, when implementing channel management of multiple GPU boxes by a dual-node server, determines the total number of GPU boxes and the connection method between each GPU box and the communication port on the server panel based on the identification information sent by each communication port in the multiple GPU boxes. Based on the total number and connection method, it determines the connection type between the dual-node server and the multiple GPU boxes. Finally, it switches the target bus link corresponding to the connection type to the corresponding target communication port through a selector and / or bus arbitrator. This method solves the problem that when a dual-node server host is paired with GPU boxes of different numbers and configurations, and when the host and GPU boxes have different interconnection methods, the dual-node server cannot determine the current configuration and interconnection method, thus failing to manage multiple GPU boxes. This enables dual-node server management of multiple GPU boxes and is applicable to various application scenarios with different numbers of GPU boxes, different numbers of GPU box interfaces, and different GPU box interconnection methods. It allows the dual-node server to be paired with GPU boxes of any configuration, number, and connection method, thus filling a gap in the field of dual-node servers paired with multiple GPU boxes. Attached Figure Description

[0045] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0046] Figure 1 is a flowchart of a multi-GPU BOX management method disclosed in this application;

[0047] Figure 2 is a schematic diagram of a connection method from the internal node of a dual-node server to the server panel port disclosed in this application, and different connection methods of the dual-node server to the GPU BOX;

[0048] Figure 3 is a schematic diagram of a connection method for a first connection type between a dual-node server and two GPU BOXes disclosed in this application;

[0049] Figure 4 is a schematic diagram of a second connection type between a dual-node server and two GPU BOXes disclosed in this application;

[0050] Figure 5 is a schematic diagram of a third connection type between a dual-node server and four GPU BOXes disclosed in this application;

[0051] Figure 6 is a schematic diagram of a fourth connection type between a dual-node server and four GPU BOXes disclosed in this application;

[0052] Figure 7 is a schematic diagram of the GPU BOX ID identified under different connection methods disclosed in this application;

[0053] Figure 8 is a flowchart of a specific multi-GPU BOX management process disclosed in this application;

[0054] Figure 9 is a schematic diagram of the connections and functional annotations of the modules in Figure 8 disclosed in this application;

[0055] Figure 10 is a schematic diagram of a multi-GPU BOX management device disclosed in this application;

[0056] Figure 11 is a structural diagram of an electronic device disclosed in this application. Detailed Implementation

[0057] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0058] This application discloses a multi-GPU BOX management method in some embodiments, applied to a dual-node server, as shown in Figure 1. The method includes:

[0059] Step S11: Obtain the identity information sent by each of the second communication ports in the multiple GPU BOXes connected to the multiple first communication ports on the dual-node server panel; the identity information is used to uniquely identify the GPU BOX corresponding to the second communication port.

[0060] It should be noted that the multi-GPU BOX management scheme proposed in this application is applied to a dual-node server. The dual-node server contains two server nodes, namely the first server node and the second server node. Furthermore, the dual-node server panel contains multiple communication ports (such as CDFP ports), which can be connected to the communication ports (such as CDFP ports) on the GPU BOX via a bus (such as the I2C bus) to perform corresponding channel management, such as IPMB channel management. Specifically, as shown in Figure 2, the dual-node server contains two server nodes, namely server node A and server node B. The dual-node server's panel contains eight CDFP ports. Server node A's internal data links correspond to ports 1, 2, 3, and 4, while server node B's internal data links correspond to ports 5, 6, 7, and 8. Figure 2 also illustrates two different connection methods between the GPU BOX and the dual-node server. Connection method A involves all CDFP ports of the GPU BOX connecting to the corresponding data link ports of server node A, such as GPU BOX ports 1 and 2 connecting to server node A ports 1 and 2, respectively. Connection method B involves all CDFP ports of the GPU BOX connecting to the corresponding data link ports of both server node A and server node B, such as GPU BOX port 1 connecting to server node A port 1, and GPU BOX port 2 connecting to server node B port 5.

[0061] In some embodiments, the communication ports in each GPU BOX send identification information to the corresponding communication ports on the dual-node server panel. This identification information is used to uniquely identify the GPU BOX corresponding to the communication port currently sending information, i.e., the GPU BOX ID (Identity document). Therefore, the dual-node server can receive the identification information sent by the second communication ports of each of the multiple GPU BOXes connected to the multiple first communication ports on the current dual-node server panel. It should be noted that the identification information sent by multiple communication ports from the same GPU BOX is the same, all identifying the same GPU BOX.

[0062] In some embodiments, before performing multi-GPU BOX management, a predetermined number of signal pins may be reserved in the second communication port of each GPU BOX. That is, a predetermined number of signal pins are reserved in advance in each communication port of the GPU BOX.

[0063] Correspondingly, the identification information sent by each second communication port in one of the multiple GPU BOXes connected to multiple first communication ports on the dual-node server panel is obtained. Specifically, this may include obtaining the identification information generated and sent by each second communication port in one of the multiple GPU BOXes connected to multiple first communication ports on the dual-node server panel after applying pull-up and pull-down resistors to a preset number of its signal pins. In some embodiments, a corresponding identification code can be pre-set for each GPU BOX as an identifier to distinguish it from other GPU BOXes. When a preset time period is reached, three preset signal pins are selected from the signal pins of the CDFP port of the GPU BOX, and the ID of the corresponding GPU BOX is generated by applying pull-up and pull-down resistors to these three signal pins. For example, the voltage level of the signal pins is changed by applying pull-up and pull-down resistors to generate a 010 voltage level signal, and this voltage level signal (i.e., 010) is used as the ID of the corresponding GPU BOX. In this way, when the communication port on the dual-node server panel receives the above-mentioned voltage level signal sent by each communication port in the GPU BOX, it can know which GPU BOX it comes from, the number of ports in each GPU BOX, and the connection method of the communication ports between them.

[0064] Step S12: Determine the total number of GPU BOXes and the connection method between each GPU BOX and the first communication port based on multiple identity information.

[0065] In some embodiments, after obtaining the identity information sent by each of the second communication ports in multiple GPU BOXes, all the received identity information can be sorted first. The sorting rule can be based on the arrangement order of multiple communication ports on the dual-node server panel, such as sorting first according to the communication port order of the first server node in the dual-node server, then sorting according to the communication port order of the second server node in the dual-node server, and then determining the total number of all GPU BOXes and the connection method between each GPU BOX and multiple first communication ports on the dual-node server panel according to the sorted identity information. For example, after receiving the identity information sent from each communication port in multiple GPU BOXes, it is sorted according to the order of the communication ports on the dual-node server panel, resulting in 001, 010, 011, 100, 001, 010, 011, 100. Through analysis, it can be seen that the total number of GPU BOXes is 4 (i.e., the number of different identity information), the total number of communication ports on the dual-node server panel is 8 (i.e., the total number of all identity information), and each GPU BOX has 2 communication ports, which are connected to the communication ports corresponding to the two server nodes (where 001, 010, 011, 100 correspond to the first server node in the dual-node server, and 001, 010, 011, 100 correspond to the second server node in the dual-node server).

[0066] Step S13: Determine the connection type between the current dual-node server and multiple GPU BOXes based on the total number and connection method, and determine the target communication port from multiple first communication ports based on the connection type.

[0067] In some embodiments, after determining the total number of GPU BOXes and the connection method between each GPU BOX and the first communication port, the connection type between the current dual-node server and the multiple GPU BOXes is further determined based on the total number and the connection method, and the target communication port for channel management is determined from the multiple first communication ports on the dual-node server panel based on the connection type.

[0068] The connection types specifically include a first connection type, a second connection type, a third connection type, and a fourth connection type. Specifically, the first connection type involves a dual-node server connecting two GPU boxes, with each GPU box's communication port connected to only one server node's communication port. For example, as shown in Figure 3, the dual-node server (including server node A and server node B) connects to two GPU boxes, GPU Box1 and GPU Box2, and each GPU box contains four communication ports. The four communication ports of GPU Box1 are connected to the four communication ports corresponding to server node A (i.e., ports 1, 2, 3, and 4), and the four communication ports of GPU Box2 are connected to the four communication ports corresponding to server node B (i.e., ports 5, 6, 7, and 8).

[0069] Specifically, the second connection type involves a dual-node server connecting two GPU boxes, with each GPU box's communication port connected to the communication ports of both server nodes. For example, as shown in Figure 4, the dual-node server (containing server node A and server node B) connects to two GPU boxes, GPU Box1 and GPU Box2, and each GPU box contains four communication ports. Ports 1 and 2 of GPU Box1 are connected to ports 1 and 2 of server node A, respectively, while ports 3 and 4 of GPU Box1 are connected to ports 5 and 6 of server node B, respectively. In other words, GPU Box2 is connected to the communication ports of both server nodes simultaneously. Furthermore, GPU Box2 is connected in the same way as GPU Box1, with ports 1 and 2 of GPU Box2 connected to ports 3 and 4 of server node A, respectively, while ports 3 and 4 of GPU Box1 are connected to ports 7 and 8 of server node B, respectively.

[0070] Specifically, the third connection type involves a dual-node server connecting four GPU boxes, with each GPU box's communication port connected only to the communication port of a single server node. For example, as shown in Figure 5, the dual-node server (containing server node A and server node B) connects to four GPU boxes: GPU Box1, GPU Box2, GPU Box3, and GPU Box4. Each GPU box contains two communication ports: ports 1 and 2 of GPU Box1 are connected to ports 1 and 2 of server node A, ports 1 and 2 of GPU Box2 are connected to ports 5 and 6 of server node B, ports 1 and 2 of GPU Box3 are connected to ports 3 and 4 of server node A, and ports 1 and 2 of GPU Box4 are connected to ports 7 and 8 of server node B.

[0071] Specifically, the fourth connection type involves a dual-node server connecting four GPU boxes, with each GPU box's communication port connected to the communication ports of both server nodes. For example, as shown in Figure 6, the dual-node server (including server node A and server node B) connects to four GPU boxes: GPU Box1, GPU Box2, GPU Box3, and GPU Box4. Each GPU box contains two communication ports. Ports 1 and 2 of GPU Box1 are connected to port 1 of server node A and port 5 of server node B, respectively. Ports 1 and 2 of GPU Box2 are connected to port 2 of server node A and port 6 of server node B, respectively. Ports 1 and 2 of GPU Box3 are connected to port 3 of server node A and port 7 of server node B, respectively. Ports 1 and 2 of GPU Box4 are connected to port 4 of server node A and port 8 of server node B, respectively.

[0072] In some embodiments, by analyzing and classifying the connection types of dual-node server hosts paired with GPU BOXes of different port numbers and configurations, as well as the connection methods of the host (i.e., server node) and GPU BOXes under different interconnection methods, it is possible to enable dual-node servers to manage multiple GPU BOXes. This is applicable to various application scenarios with different numbers of GPU BOXes, different numbers of GPU BOX interfaces, and different GPU BOX interconnection methods, and improves the efficiency of managing multiple GPU BOXes.

[0073] It is understandable that the corresponding identity information differs for different connection methods. Referring to Figure 7, when the received identity information, ordered by interface sequence, is 001, 001, 001, 001, 010, 010, 010, 010, it indicates that the total number of GPU BOXes is 2, the total number of communication ports on the dual-node server panel is 8, and each GPU BOX has 4 communication ports, connected to the 4 communication ports corresponding to a single server node. Therefore, the current connection type can be determined to be the first connection type (corresponding to the connection method in Figure 3). Similarly, when the received identity information, ordered by interface sequence, is 001, 001, 010, 010, 001, 001, 010, 010, it indicates that the total number of GPU BOXes is 2, the total number of communication ports on the dual-node server panel is 8, and each GPU BOX has 4 communication ports, connected to the 4 communication ports corresponding to a single server node. The GPU BOX has 4 communication ports, and each GPU BOX is connected to one of the two server nodes' corresponding communication ports. Therefore, the current connection type can be determined to be the second connection type (corresponding to the connection method in Figure 4). Similarly, when the received identity information is ordered by interface as 001, 001, 011, 011, 010, 010, 100, 100, we know that the total number of GPU BOXes is 4, the total number of communication ports on the dual-node server panel is 8, and each GPU BOX has 2 communication ports, connected to the two communication ports corresponding to a single server node. Therefore, the current connection type can be determined to be the third connection type (corresponding to the connection method in Figure 5). Similarly, when the received identity information is ordered by interface as 001, 010, 011, 100, 001, 010, 011, 100, we know that the total number of GPU BOXes is 4, the total number of communication ports on the dual-node server panel is 8, and each GPU BOX has 2 communication ports, connected to the two communication ports corresponding to a single server node. The BOX has two communication ports, and it is connected to the communication ports corresponding to the two server nodes. Therefore, it can be determined that the current connection type is the fourth connection type (corresponding to the connection method in Figure 6).

[0074] In addition, before obtaining the identification information sent by each of the second communication ports in the multiple GPU BOXes connected to the multiple first communication ports on the dual-node server panel, the process further includes: selecting target communication ports for channel management corresponding to the first connection type, second connection type, third connection type, and fourth connection type from the multiple first communication ports on the dual-node server panel according to preset rules; wherein, the number of target communication ports corresponding to the first connection type and the second connection type is 2, and the number of target communication ports corresponding to the third connection type and the fourth connection type is 4, and the management channel of each target communication port corresponding to the first connection type and the third connection type originates from a uniquely determined server node, while the management channel of each target communication port corresponding to the second connection type and the fourth connection type originates from a non-unique server node; the server node is either the first server node (i.e., server node A in Figure 2) or the second server node (i.e., server node B in Figure 2) in the dual-node server. It should be noted that, in this application, target communication ports for channel management are set for the server nodes and GPU BOXes corresponding to different connection types according to preset rules, wherein the number of target communication ports corresponding to the first connection type and the second connection type is 2, and the number of target communication ports corresponding to the third connection type and the fourth connection type is 4. For example, referring to Figure 3, when it is the first connection type, according to the preset rules, the path between CDFP port 1 of GPU BOX1 and CDFP port 1 of server node A is used as the IPMB management path, and CDFP port 1 of server node A is set as the target communication port (that is, the port selected to connect each GPU). (The first port in a group of CDFP ports of BOX is used as the management channel interface). Similarly, if CDFP port 5 of server node B is set as the target communication port, then the number of target communication ports corresponding to the first connection type is 2. When it is the second connection type, as shown in Figure 4, if CDFP ports 1 and 3 of server node A are set as target communication ports, then the number of target communication ports corresponding to the second connection type is 2. When it is the third connection type, as shown in Figure 5, if CDFP ports 1 and 3 of server node A, and CDFP ports 5 and 7 of server node B are set as target communication ports, then the number of target communication ports corresponding to the third connection type is 4. When it is the fourth connection type, as shown in Figure 6, if CDFP ports 1, 2, 3, and 4 of server node A are set as target communication ports, then the number of target communication ports corresponding to the fourth connection type is 4.

[0075] Additionally, it should be noted that the management channels for each target communication port corresponding to the first and third connection types originate from a uniquely determined server node. Referring to Figures 3 and 5, in Figure 3, the IPMB management channel of CDFP port 1 on the server panel originates from a uniquely determined server node A, and the IPMB management channel of CDFP port 5 on the server panel originates from a uniquely determined server node B; or in Figure 5, the IPMB management channels of CDFP ports 1 and 3 on the server panel originate from a uniquely determined server node A, and CDFP ports 5 and 7 on the server panel originate from a uniquely determined server node B.

[0076] Furthermore, the management channels for the target communication ports corresponding to the second and fourth connection types originate from different server nodes (they may originate from server node A or server node B). Referring to Figures 4 and 6, the management channels for the target communication ports corresponding to the third and fourth connection types may originate from server node A or server node B. For example, the IPMB management paths for CDFP ports 1 and 3 on the server panel in Figure 4 may originate from either server node A or server node B. The specific path selected needs to be determined through arbitration and / or strobing during real-time multi-GPU BOX management. Similarly, the IPMB management paths for CDFP ports 1, 2, 3, and 4 on the server panel in Figure 6 may originate from either server node A or server node B.

[0077] In some embodiments, the connection type between the current dual-node server and multiple GPU BOXes is determined based on the total number and connection method. Specifically, this may include: when the total number is 2, and the connection method is that multiple second communication ports in each GPU BOX are connected only to the communication ports of the first server node or the second server node, the connection type between the current dual-node server and multiple GPU BOXes is determined to be the first connection type. For example, referring to Figure 3, when the total number of GPU BOXes is 2, and the connection method is that the four CDFP communication ports in a single GPU BOX are connected only to the CDFP communication ports of server node A or server node B, the connection type between the current dual-node server and multiple GPU BOXes is determined to be the first connection type.

[0078] In some embodiments, the connection type between the current dual-node server and multiple GPU BOXes is determined based on the total number and connection method. Specifically, this may include: when the total number is 2, and the connection method is that the second communication port in each GPU BOX is connected to the communication ports of both the first server node and the second server node, the connection type between the current dual-node server and multiple GPU BOXes is determined to be the second connection type. For example, referring to Figure 4, when the total number of GPU BOXes is 2, and the connection method is that all four CDFP communication ports in a single GPU BOX are connected to the CDFP communication ports of both server node A and server node B, the connection type between the current dual-node server and multiple GPU BOXes is determined to be the second connection type.

[0079] In some embodiments, the connection type between the current dual-node server and multiple GPU BOXes is determined based on the total number and connection method. Specifically, this may include: when the total number is 4, and the connection method is that the second communication port in each GPU BOX is connected only to the communication port of the first server node or the second server node, the connection type between the current dual-node server and multiple GPU BOXes is determined to be a third connection type. For example, referring to Figure 5, when the total number of GPU BOXes is 4, and the connection method is that two CDFP communication ports in a single GPU BOX are connected to the CDFP communication ports of server node A or server node B, the connection type between the current dual-node server and multiple GPU BOXes is determined to be a third connection type.

[0080] In some embodiments, the connection type between the current dual-node server and multiple GPU BOXes is determined based on the total number and connection method. Specifically, this may include: when the total number is 4, and the connection method is that the second communication port in each GPU BOX is connected to the communication ports of both the first server node and the second server node, the connection type between the current dual-node server and multiple GPU BOXes is determined to be the fourth connection type. For example, referring to Figure 6, when the total number of GPU BOXes is 4, and the connection method is that the two CDFP communication ports in a single GPU BOX are connected to the CDFP communication ports of both server node A and server node B, the connection type between the current dual-node server and multiple GPU BOXes is determined to be the fourth connection type.

[0081] Step S14: Switch the target bus link corresponding to the target communication port to the target communication port through the selector and / or bus arbiter, so as to perform corresponding channel management on the corresponding GPU BOX through the target communication port.

[0082] In some embodiments, after determining the target communication port from a plurality of first communication ports based on the connection type, the target bus link corresponding to the target communication port can be switched to the target communication port through a selector and / or a bus arbiter, thereby performing corresponding channel management on the corresponding GPU BOX through the target communication port.

[0083] In some embodiments, when the connection type is the first connection type, the target bus link corresponding to the target communication port is switched to the target communication port through a gating switch. Specifically, this may include: enabling the corresponding target bus link and switching the target bus link corresponding to the target communication port to the target communication port through a gating switch. Specifically, referring to Figure 8, for the first connection type, the ID sequence of the GPU BOX can be sorted by a gating switch to obtain the truth value. Then, based on the sorted truth value, the current connection method is identified and the bus link is automatically switched. As shown in Figure 8, after determining the target communication port (i.e., CDFP ports 1 and 5 on the server panel), the target bus links of the corresponding CDFP ports (i.e., I2C_A1 and I2C_B1) are first enabled, and then the target bus links corresponding to the target communication port (i.e., I2C_A1 and I2C_B1) are switched to the target communication port (i.e., CDFP ports 1 and 5 on the server panel) through a gating switch (i.e., a gating switch). Additionally, as shown in Figure 9, the dashed lines in Figure 8 represent I2C links, the solid lines represent I2C links that are ultimately connected to the CDFP port, the dashed boxes indicate that the CDFP port is used as the IPMB management port under a certain connection method, and the solid boxes represent ordinary CDFP data interfaces, which will not be used as IPMB management ports under any connection method.

[0084] In some embodiments, when the connection type is the third connection type, the target bus link corresponding to the target communication port is switched to the target communication port via a selector. Specifically, this may include: enabling the corresponding target bus link and switching the target bus link corresponding to the target communication port to the target communication port via a selector. Specifically, referring to Figure 8, for the third connection type, similar to the first connection type described above, after determining the target communication port (i.e., CDFP ports 1, 3, 5, and 7 on the server panel), the target bus links of the corresponding CDFP ports (i.e., I2C_A1, I2C_A3, I2C_B1, and I2C_B3) are first enabled, and then the target bus links corresponding to the target communication port (i.e., I2C_A1, I2C_A3, I2C_B1, and I2C_B3) are switched to the target communication port (i.e., CDFP ports 1, 3, 5, and 7 on the server panel) via a selector switch (selector).

[0085] In some embodiments, when the connection type is the second connection type, the target bus link corresponding to the target communication port is switched to the target communication port through a gating device and / or a bus arbiter. Specifically, this may include: enabling the corresponding target bus link and switching the target bus link corresponding to the target communication port to the target communication port through a gating device and / or a bus arbiter. Specifically, referring to Figure 8, for the second connection type, after determining the target communication port (i.e., CDFP ports 1 and 3 on the server panel), the target bus links (i.e., I2C_A1 and I2C_A3) of the corresponding CDFP ports are first enabled. Then, the ID sequence of the GPU BOX is sorted through the gating device to obtain the truth value, and the current connection method is identified based on the sorted truth value. Next, the I2C link is directed to the corresponding arbiter through the gating device to arbitrate the bus source corresponding to the target communication port. Then, the target I2C link (i.e., I2C_A1 or I2C_A3) corresponding to the bus source is directed to the target communication port (i.e., CDFP ports 1 and 3 on the server panel).

[0086] In some embodiments, when the connection type is the fourth connection type, the target bus link corresponding to the target communication port is switched to the target communication port through a gating switch and / or a bus arbiter. Specifically, this may include: enabling the corresponding target bus link, and switching the target bus link corresponding to the target communication port to the target communication port through a gating switch and / or a bus arbiter. Specifically, referring to Figure 8, for the fourth connection type, after determining the target communication port (i.e., CDFP ports 1, 2, 3, and 4 on the server panel), the target bus links of the corresponding CDFP ports (i.e., I2C_A1, I2C_A2, I2C_A3, and I2C_A4) are first enabled, and then the GPU is switched through a gating switch (gating switch). The BOX ID sequence is sorted to obtain the truth value, and the current connection method is identified based on the sorted truth value. Then, the I2C links (i.e., I2C_A1 and I2C_A3) are directed to the corresponding arbitrators through the selectors. The arbitrators arbitrate the bus source corresponding to the target communication port, and then the target I2C link (i.e., I2C_A1 or I2C_A3) corresponding to the bus source is directed to the target communication port (i.e., CDFP port 1 or 3 on the server panel). At the same time, the I2C links (i.e., I2C_A2 and I2C_A4) are arbitrated through the arbitrators, and based on the bus source corresponding to the target communication port, the corresponding target I2C link (i.e., I2C_A2 or I2C_A4) is directed to the target communication port (i.e., CDFP port 2 or 4 on the server panel).

[0087] In some embodiments, data transmission on the target bus link may further include: detecting whether the data field on the target bus link has been completely transmitted; if the data field on the target bus link has been completely transmitted, releasing the target bus link and preempting the target bus link through another server node; and directing the target bus link to the target communication port through the other server node, so as to perform IPMB channel management on the corresponding GPU BOX through the other server node and the target communication port. For example, the bus arbiter in the dual-node server detects whether the data field on the target bus link (transmitted by server node A) has been completely transmitted. If the data field on the target bus link has been completely transmitted, releasing the target bus link (such as the bus after I2C_A1 through the selector) and preempting the target bus link through server node B, and then directing the target bus link (such as the bus after I2C_B1 through the selector) to the target communication port (such as CDFP port 1 on the server panel) through server node B, so as to perform IPMB channel management on the corresponding GPU BOX (GPU BOX1 in Figure 6) through server node B and the target communication port (i.e., CDFP port 1 on the server panel). It should be noted that the bus arbiter allows node A to gain management rights by default, that is, it prioritizes selecting the I2C link of server node A as the final output and directing it to the CDFP port of the server. However, it can also implement a preemption logic. After the data field transmission on the I2C link of server node A is completed, the corresponding I2C link is released. At this time, server node B can preempt the I2C link and then direct the I2C link of server node B to the CDFP port of the server, thereby realizing the management of the GPU BOX by server node B.

[0088] As can be seen, the embodiments of this application are applied to a dual-node server. First, the identity information sent by the second communication port of each of the multiple GPU BOXes connected to the multiple first communication ports on the dual-node server panel is obtained. Then, based on the multiple identity information, the total number of GPU BOXes and the connection method between each GPU BOX and the first communication port are determined. Based on the total number and the connection method, the connection type between the current dual-node server and the multiple GPU BOXes is determined. Based on the connection type, the target communication port is determined from the multiple first communication ports. Then, the target bus link corresponding to the target communication port is switched to the target communication port through a selector and / or a bus arbiter, so as to perform corresponding channel management on the corresponding GPU BOX through the target communication port. In this embodiment of the application, when implementing channel management of multiple GPU BOXes by a dual-node server, the total number of GPU BOXes and the connection method between each GPU BOX and the communication port on the server panel are determined based on the identification information sent by each communication port in the multiple GPU BOXes. Based on the total number and connection method, the connection type between the dual-node server and the multiple GPU BOXes is determined. Finally, a selector and / or bus arbiter switches the target bus link corresponding to the connection type to the corresponding target communication port. This method solves the problem that when a dual-node server host is paired with GPU BOXes of different port numbers and configurations, and when the host and GPU BOXes have different interconnection methods, the dual-node server cannot determine the current configuration and interconnection method, thus failing to manage multiple GPU BOXes. This enables dual-node server management of multiple GPU BOXes and is applicable to various application scenarios with different numbers of GPU BOXes, different numbers of GPU BOX interfaces, and different GPU BOX interconnection methods. It allows the dual-node server to be paired with GPU BOXes of any configuration, number, and connection method, thus filling a gap in the field of dual-node servers paired with multiple GPU BOXes.

[0089] This application discloses a specific multi-GPU BOX management method in some embodiments, applied to a dual-node server, as shown in Figure 10. The method includes:

[0090] Step S21: Obtain the identity information sent by each of the second communication ports in the multiple GPU BOXes connected to the multiple first communication ports on the dual-node server panel; the identity information is used to uniquely identify the GPU BOX corresponding to the second communication port.

[0091] Step S22: Sort the multiple identity information corresponding to the GPU BOX according to the order of multiple first communication ports to obtain sorted identity information.

[0092] In some embodiments, after obtaining the identity information sent by each of the second communication ports of multiple GPU BOXes connected to multiple first communication ports on the dual-node server panel, the identity information corresponding to the GPU BOXes is further sorted according to the order of each first communication port to obtain sorted identity information. For example, referring to Figure 7, Figure 7 shows the sorted identity information (i.e., the sorted GPU BOX ID) under four different connection methods. The first connection method is the identity information obtained by sorting all the identity information (i.e., GPU BOX ID) corresponding to the GPU BOXes from left to right and from top to bottom, namely 001, 001, 001, 001, 010, 010, 010, 010.

[0093] Step S23: Determine the total number of GPU BOXes and the connection method between each GPU BOX and the first communication port based on the sorted identification information.

[0094] Furthermore, based on the sorted identification information (i.e., 001, 001, 001, 001, 010, 010, 010), the total number of all GPU BOXes is determined to be 2, the total number of communication ports on the dual-node server panel is 8, and each GPU BOX has 4 communication ports, which are connected to the 4 communication ports corresponding to a single server node. Therefore, the current connection type can be determined to be the first connection type (corresponding to the connection method in Figure 3). Similarly, the connection methods corresponding to Figures 4 to 6 can be determined.

[0095] Step S24: Determine the connection type between the current dual-node server and multiple GPU BOXes based on the total number and connection method, and determine the target communication port from multiple first communication ports based on the connection type.

[0096] Step S25: Switch the target bus link corresponding to the target communication port to the target communication port through the selector and / or bus arbiter, so as to perform corresponding channel management on the corresponding GPU BOX through the target communication port.

[0097] For more detailed processing procedures of steps S21, S24, and S25, please refer to the corresponding content disclosed in the foregoing embodiments, which will not be repeated here.

[0098] As can be seen, in this embodiment, the identity information sent by each second communication port of each of the multiple GPU BOXes connected to the multiple first communication ports on the dual-node server panel is firstly obtained. Then, the identity information corresponding to the GPU BOXes is sorted according to the order of the multiple first communication ports to obtain sorted identity information. Based on the sorted identity information, the total number of GPU BOXes and the connection method between each GPU BOX and the first communication port are determined. Then, based on the total number and the connection method, the connection type between the current dual-node server and the multiple GPU BOXes is determined. Based on the connection type, the target communication port is determined from the multiple first communication ports. Finally, the target bus link corresponding to the target communication port is switched to the target communication port through a selector and / or a bus arbiter, thereby performing corresponding channel management on the corresponding GPU BOXes through the target communication port. This application embodiment sorts the identification information sent by each communication port in the received multiple GPU BOXes, and determines the total number of multiple GPU BOXes and the connection method between each GPU BOX and multiple communication ports on the server panel based on the sorted identification information. In this way, the configuration and interconnection method of the current dual-node server and multiple GPU BOXes can be accurately identified, thereby realizing the management of multiple GPU BOXes by the dual-node server.

[0099] Furthermore, some embodiments of this application also disclose an electronic device. FIG11 is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content in the figure should not be considered as any limitation on the scope of use of this application.

[0100] Figure 11 is a schematic diagram of the structure of an electronic device 20 provided in some embodiments of this application. Specifically, the electronic device 20 may include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 stores a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the multi-GPU BOX management method disclosed in any of the foregoing embodiments. Furthermore, in some embodiments, the electronic device 20 may specifically be an electronic computer.

[0101] In some embodiments, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 25 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.

[0102] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored thereon can include operating system 221, computer program 222, etc., and the storage method can be temporary storage or permanent storage.

[0103] The operating system 221 is used to manage and control the various hardware devices on the electronic device 20 and the computer program 222, which may be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program capable of performing the multi-GPU BOX management method executed by the electronic device 20 as disclosed in any of the foregoing embodiments, the computer program 222 may further include a computer program capable of performing other specific tasks.

[0104] Furthermore, this application also discloses a computer-nonvolatile readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned multi-GPU BOX management method. Specific steps of this method can be found in the corresponding content disclosed in the foregoing embodiments, and will not be repeated here.

[0105] Furthermore, some embodiments of this application also disclose a computer program product, including a computer program / instructions, which, when executed by a processor, implement the steps of the multi-GPU BOX management method disclosed above.

[0106] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.

[0107] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0108] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0109] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.

[0110] The above provides a detailed description of a multi-GPU BOX management method, apparatus, device, and storage medium provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A multi-GPU BOX management method, characterized in that, Applied to dual-node servers, including: Each GPU BOX connected to multiple first communication ports on the dual-node server panel acquires its identity information sent from its second communication port; the identity information is configured to uniquely identify the GPU BOX corresponding to the second communication port. The total number of GPU BOXes and the connection method between each GPU BOX and the first communication port are determined based on multiple identity information. Based on the total number and the connection method, the connection type between the current dual-node server and the multiple GPU BOXes is determined, and the target communication port is determined from the multiple first communication ports based on the connection type; The target bus link corresponding to the target communication port is switched to the target communication port through a selector and / or a bus arbiter, so as to perform corresponding channel management on the corresponding GPU BOX through the target communication port.

2. The multi-GPU BOX management method according to claim 1, characterized in that, The connection types include a first connection type, a second connection type, a third connection type, and a fourth connection type; The first connection type involves a dual-node server connecting two GPU boxes, with each GPU box's communication port connected only to the communication port of a single server node. The second connection type involves a dual-node server connecting two GPU boxes, with each GPU box's communication port connected to the communication ports of both server nodes. The third connection type involves a dual-node server connecting four GPU boxes, with each GPU box's communication port connected only to the communication port of a single server node. The fourth connection type involves a dual-node server connecting four GPU boxes, with each GPU box's communication port connected to the communication ports of both server nodes.

3. The multi-GPU BOX management method according to claim 2, characterized in that, Before obtaining the identity information sent by each of the second communication ports in the multiple GPU BOXes connected to the multiple first communication ports on the dual-node server panel, the method further includes: According to preset rules, the target communication ports configured for channel management are selected from the multiple first communication ports on the dual-node server panel, corresponding to the first connection type, the second connection type, the third connection type, and the fourth connection type. Wherein, the number of target communication ports corresponding to the first connection type and the second connection type is 2, the number of target communication ports corresponding to the third connection type and the fourth connection type is 4, and the management channel of each target communication port corresponding to the first connection type and the third connection type originates from a uniquely determined server node, while the management channel of each target communication port corresponding to the second connection type and the fourth connection type originates from a non-unique server node; the server node is the first server node or the second server node in the dual-node server.

4. The multi-GPU BOX management method according to claim 3, characterized in that, The process of determining the connection type between the current dual-node server and the multiple GPU BOXes based on the total number and the connection method includes: When the total number is 2, and the connection method is that multiple second communication ports in each GPU BOX are only connected to the communication ports of the first server node or the second server node, the connection type between the current dual-node server and the multiple GPU BOX is determined to be the first connection type.

5. The multi-GPU BOX management method according to claim 3, characterized in that, The process of determining the connection type between the current dual-node server and the multiple GPU BOXes based on the total number and the connection method includes: When the total number is 2, and the connection method is that the second communication port in each GPU BOX is connected to the communication ports of both the first server node and the second server node, the connection type between the current dual-node server and the multiple GPU BOXes is determined to be the second connection type.

6. The multi-GPU BOX management method according to claim 3, characterized in that, The process of determining the connection type between the current dual-node server and the multiple GPU BOXes based on the total number and the connection method includes: When the total number is 4, and the connection method is that the second communication port in each GPU BOX is only connected to the communication port of the first server node or the second server node, the connection type between the current dual-node server and the multiple GPU BOXes is determined to be the third connection type.

7. The multi-GPU BOX management method according to claim 3, characterized in that, The process of determining the connection type between the current dual-node server and the multiple GPU BOXes based on the total number and the connection method includes: When the total number is 4, and the connection method is that the second communication port in each GPU BOX is connected to the communication ports of both the first server node and the second server node, the connection type between the current dual-node server and the multiple GPU BOXes is determined to be the fourth connection type.

8. The multi-GPU BOX management method according to claim 4 or 6, characterized in that, Switching the target bus link corresponding to the target communication port to the target communication port via a selector includes: Enable the corresponding target bus link, and switch the target bus link corresponding to the target communication port to the target communication port through a selector.

9. The multi-GPU BOX management method according to claim 5 or 7, characterized in that, The step of switching the target bus link corresponding to the target communication port to the target communication port through a selector and / or a bus arbiter includes: Enable the corresponding target bus link, and switch the target bus link corresponding to the target communication port to the target communication port through a selector and / or a bus arbiter.

10. The multi-GPU BOX management method according to claim 9, characterized in that, Also includes: Detect whether the data fields on the target bus link have been completely transmitted; If the data field on the target bus link has been transmitted, the target bus link is released, and another server node preempts the target bus link. The target bus link is routed to the target communication port via the other server node, so that the corresponding GPU BOX can be managed through the other server node and the target communication port.

11. The multi-GPU BOX management method according to claim 1, characterized in that, Also includes: A preset number of signal pins are reserved in the second communication port of each GPU BOX.

12. The multi-GPU BOX management method according to claim 11, characterized in that, The step of acquiring the identity information sent by each of the second communication ports in the multiple GPU BOXes connected to the multiple first communication ports on the dual-node server panel includes: The identity information generated and sent by each of the second communication ports in the multiple GPU BOXes connected to the multiple first communication ports on the dual-node server panel is obtained by applying resistor pull-up and pull-down processing to a preset number of signal pins.

13. The multi-GPU BOX management method according to claim 1, characterized in that, The step of determining the total number of GPU BOXes and the connection method between each GPU BOX and the first communication port based on multiple identity information includes: Based on the order of the multiple first communication ports, the multiple identity information corresponding to the GPU BOX is sorted to obtain sorted identity information; Based on the sorted identification information, the total number of the multiple GPU BOXes and the connection method between each GPU BOX and the first communication port are determined.

14. The multi-GPU BOX management method according to claim 13, characterized in that, The sorting of multiple identity information corresponding to the GPU BOX based on the order of multiple first communication ports to obtain sorted identity information includes: The multiple identity information corresponding to the GPU BOX is sorted according to the arrangement order of the multiple first communication ports on the dual-node server panel to obtain the sorted identity information.

15. The multi-GPU BOX management method according to claim 13, characterized in that, The sorting of multiple identity information corresponding to the GPU BOX based on the order of multiple first communication ports to obtain sorted identity information includes: The identification information corresponding to the GPU BOX is sorted from left to right and from top to bottom to obtain the sorted identification information.

16. The multi-GPU BOX management method according to claim 1, characterized in that, The different connection methods between the GPU BOX and the first communication port correspond to different identity information.

17. The multi-GPU BOX management method according to claim 1, characterized in that, The dual-node server includes two server nodes, namely a first server node and a second server node. The panel of the dual-node server includes multiple communication ports, which are configured to connect to the communication interface of the GPU BOX via a bus.

18. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the multi-GPU BOX management method as described in any one of claims 1 to 17.

19. An electronic device, characterized in that, It includes a processor and a memory; wherein, when the processor executes a computer program stored in the memory, it implements the multi-GPU BOX management method as described in any one of claims 1 to 17.

20. A computer-defined non-volatile readable storage medium, characterized in that, It is configured to store a computer program; wherein, when the computer program is executed by a processor, it implements the multi-GPU BOX management method as described in any one of claims 1 to 17.

Citation Information

Patent Citations

  • Model training method and device and cluster system

    CN111327692A

  • Intelligent server system with multiple heterogeneous nodes

    CN115454633A

  • Resource scheduling method, computer equipment, storage medium and program product

    CN118394533A

  • Multi-GPU BOX management method and device, equipment and storage medium

    CN119011330A

  • Virtual target port aggregation

    US20170338977A1