Method for hot standby switching of the entire network system, electronic device, medium and product

By implementing fast resource mapping and data switching of failed host devices in the PCIe switch system, the business interruption problem caused by fixed port mapping of PCIe switch is solved, and the reliability and fault tolerance of the system are improved.

CN119906663BActive Publication Date: 2025-05-27INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510391417.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-05-27
Estimated Expiration
2045-03-31

AI Technical Summary

Technical Problem

The port mapping of PCIe switches is relatively fixed, making it difficult to quickly switch data to backup devices, resulting in service interruptions.

Method used

By determining whether there is a target host device in the faulty state among multiple host devices, deleting the faulty device, generating fault information, updating the routing table of the PCIe switch, and starting the hot standby switch function of the entire machine, mapping the resources of the faulty host device to the backup host device in the normal state.

Benefits of technology

It realizes rapid resource mapping and data switching of PCIe switches in case of failure, avoids business interruption, and improves system reliability and fault tolerance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119906663B_ABST
    Figure CN119906663B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for hot standby switching of an entire network system, an electronic device, a medium, and a product, which relates to the field of computer technologies and includes: when there is a first target host device in a fault state among multiple host devices, deleting the first target host device and generating corresponding fault information, receiving a first topology configuration file of a new graphics processing unit updated by a multi-host control processing unit based on the fault information, updating a routing table of a PCIe switch based on the first topology configuration file, and mapping the resources of the graphics processing unit of the first target host device to a second target host device in a normal state, solving problems such as relatively fixed port mapping of the PCIe switch and difficulty in quickly switching data to a standby device, and realizing flexible connection of each GPU and continuous update of the routing table through the PCIe switch to achieve efficient data interaction and resource mapping among each GPU.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technologies, and in particular, to a method for switching the whole machine hot standby of a network system, an electronic device, a medium, and a product. Background Art

[0002] With the growth of demands, it has become a research hotspot to efficiently interconnect more GPUs (Graphics Processing Units). Due to problems such as high latency and limited transmission semantics in Ethernet interconnection, it cannot meet the requirements of some high-performance computing scenarios. Therefore, it has become a trend to build a multi-card supernode through PCIe (Peripheral Component Interconnect Express) Switch interconnection to achieve a more direct data transmission path.

[0003] In related technologies, in a PCIe interconnection system, a scheme using NTB (Non-Transparent Bridge) to implement multi-host communication is usually adopted. However, multi-host communication needs to be implemented through a VS (Virtual Switch) in the same PCIe Switch. At the same time, the port mapping and device recognition mechanisms of the PCIe Switch are relatively fixed, and it is difficult to quickly switch data traffic to a standby device, resulting in service interruption, which urgently needs to be solved. Summary of the Invention

[0004] The present invention provides a method for switching the whole machine hot standby of a network system, an electronic device, a medium, and a product, so as to at least solve the problem in related technologies that the port mapping of a PCIe switch is relatively fixed, and it is difficult to quickly switch data to a standby device, resulting in service interruption.

[0005] The present invention provides a method for switching the whole machine hot standby of a network system. The network system includes multiple host devices, each host device is equipped with at least one graphics processing unit, and all graphics processing units are connected to a PCIe switch. The method includes the following steps:

[0006] Determine whether there is a first target host device in a fault state among the multiple host devices;

[0007] If there is a first target host device in the fault state among the multiple host devices, delete the first target host device, generate fault information of the first target host device, and send the fault information to a multi-host control processing unit through a preset management interface;

[0008] Receive a first topology configuration file of a new graphics processing unit updated by the multi-host control processing unit based on the fault information, update the routing table of the PCIe switch based on the first topology configuration file, and after the routing table of the PCIe switch is updated, initiate the whole-machine hot standby switching function of the multi-host to map the resources of the graphics processing unit of the first target host device to a second target host device in a normal state.

[0009] The present invention also provides a whole-machine hot standby switching device for a network system. The network system includes multiple host devices, each host device is equipped with at least one graphics processing unit, and all graphics processing units are connected to a PCIe switch, including:

[0010] A judgment module, configured to judge whether there is a first target host device in a fault state among the multiple host devices;

[0011] A generation module, configured to, if there is a first target host device in the fault state among the multiple host devices, delete the first target host device, generate fault information of the first target host device, and send the fault information to the multi-host control processing unit through a preset management interface;

[0012] A mapping module, configured to receive a first topology configuration file of a new graphics processing unit updated by the multi-host control processing unit based on the fault information, update the routing table of the PCIe switch based on the first topology configuration file, and after the routing table of the PCIe switch is updated, initiate the whole-machine hot standby switching function of the multi-host to map the resources of the graphics processing unit of the first target host device to a second target host device in a normal state.

[0013] The present invention also provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor executes the program to implement the whole-machine hot standby switching method of the network system as described in the above embodiment.

[0014] The present invention also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any one of the above whole-machine hot standby switching methods of the network system are implemented.

[0015] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of any one of the above whole-machine hot standby switching methods of the network system are implemented.

[0016] With the present invention, when there is a first target host device in a failed state among multiple host devices, the first target host device is deleted, and corresponding fault information is generated. A first topology configuration file of a new graphics processing unit updated based on the fault information is received, so as to update the routing table of the PCIe switch based on the first topology configuration file, and map the resources of the graphics processing unit of the first target host device to a second target host device in a normal state. This solves problems such as the relatively fixed port mapping of the PCIe switch and the difficulty in quickly switching data to a standby device. Through the PCIe switch, flexible connection of each GPU and continuous update of the routing table are achieved, so as to realize efficient data interaction and resource mapping among GPUs. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] To more clearly illustrate the embodiments of the present invention, the following will briefly introduce the drawings required for the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0018] Figure 1 Schematic diagram of the whole-machine hot standby switching method in the related art;

[0019] Figure 2 Flowchart of the whole-machine hot standby switching method of a network system provided by an embodiment of the present invention;

[0020] Figure 3 Partial schematic diagram of the GPU interconnection topology in a PCIe interconnection system provided by an embodiment of the present invention;

[0021] Figure 4 Schematic diagram of the software topology management system architecture provided by an embodiment of the present invention;

[0022] Figure 5 Schematic diagram of the GPU topology configuration file before deleting a faulty host provided by an embodiment of the present invention;

[0023] Figure 6 Schematic diagram of the GPU topology configuration file after deleting a faulty host provided by an embodiment of the present invention;

[0024] Figure 7 Schematic diagram of the GPU topology configuration file after adding a host provided by an embodiment of the present invention;

[0025] Figure 8 Schematic diagram of the update of the GPU topology configuration file after adding a host provided by an embodiment of the present invention;

[0026] Figure 9A block diagram example of a whole - machine hot - standby switching device for a network system according to an embodiment of the present invention;

[0027] Figure 10 A schematic structural diagram of an electronic device according to an embodiment of the present invention. Detailed implementation manners

[0028] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the protection scope of the present invention.

[0029] It should be noted that in the description of the present invention, the terms "include", "comprise" or any other variation thereof are intended to cover a non - exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present invention are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0030] In order to enable those skilled in the art of the present technology to better understand the solution of the present invention, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.

[0031] An embodiment of the present invention provides a whole - machine hot - standby switching method for a network system. The method will be described in detail in combination with the execution process of the whole - machine hot - standby switching method for the network system.

[0032] Before introducing the embodiments of the present invention, first introduce the specific implementation methods of related technologies, mainly a solution for implementing multi - host communication using a non - transparent bridge (NTB) in a high - speed peripheral component interconnect (PCIe) interconnect system, such as Figure 1As shown in the figure, (1) The upstream port 501 of the first host domain is used to send a first message to the routing module 502. The routing module 502 is used to respond to receiving the first message sent by the upstream port 501 of the first host domain, and confirm whether the port of the first message is mounted with an NTB. If the port of the first message is mounted with an NTB, the first message is sent to the non-transparent bridge virtual side 504 through the centralized switching module 503; (3) The centralized switching module 503 is used to forward the first message from the routing module 502 to the non-transparent bridge virtual side 504. The non-transparent bridge virtual side 504 is used to forward the first message to the conversion module 505. The non-transparent bridge virtual side 504 is a virtual endpoint; (4) The conversion module 505 is used to convert the address and ID (Identifier) information of the first message from the upstream port 501 of the first host domain to the downstream port 507 of the second host domain and the upstream port 508 of the second host domain, generate a second message, and forward the second message through the routing module 502 and the centralized switching module 503 to the non-transparent bridge link side 506. The non-transparent bridge link side 506 is the endpoint device physically connected to the downstream port 507 of the second host domain in the upstream port 501 of the first host domain; (5) The routing module 502 and the centralized switching module 503 are also used to forward the second message from the conversion module 505 to the non-transparent bridge link side 506. The non-transparent bridge link side 506 is used to receive the second message, thereby achieving the purpose of multi-host communication.

[0033] However, there are some drawbacks in the above solution, mainly manifested as follows: (1) Implementing multi-host device communication is achieved through VS in the same PCIe switch, and it is difficult to expand to the case of multiple PCIe switches; (2) In this technical solution, in fact, the communication is achieved through the upstream ports of two host devices. By this method to achieve GPU communication, if it involves the whole machine hot standby switching, problems such as too long re-enumeration time, difficult to guarantee signal integrity, and poor compatibility will occur for the host device. Therefore, based on the above existing problems, the embodiments of the present invention innovatively propose a whole machine hot standby switching method based on a PCIe interconnection network system to achieve flexible connection between GPUs, effectively avoiding the disadvantages of traditional interconnection methods during topology switching. Among them, the management solution implemented by the present invention can also be flexibly expanded to a larger PCIe interconnected supernode system. The following will be described in detail by specific embodiments.

[0034] Specifically, Figure 2 FIG. is a schematic flowchart of a whole machine hot standby switching method for a network system provided by an embodiment of the present invention. The network system includes multiple host devices, each host device is equipped with at least one graphics processing unit, and all graphics processing units are connected to a PCIe switch.

[0035] As Figure 2As shown, the method for hot standby switching of the entire network system includes the following steps:

[0036] In step S201, it is determined whether there is a first target host device in a fault state among multiple host devices.

[0037] According to an embodiment of the present invention, before determining whether there is a first target host device in a fault state among multiple host devices, it further includes: initializing the network system, and after the network system is initialized, performing resource enumeration and address allocation on each graphics processing unit in the multiple host devices to make each graphics processing unit match a corresponding identity identifier; generating graphics processing unit topology information based on each graphics processing unit after identity identifier matching, and performing topology management on each graphics processing unit according to the graphics processing unit topology information.

[0038] Specifically, to achieve flexible connection between GPUs and implement an efficient and flexible method for hot standby switching of the entire machine in a PCIe interconnection network system, the embodiment of the present invention adopts a reasonable hardware topology design and software management scheme to achieve rapid switching and recovery of the network system in case of a fault, thereby improving the reliability and fault tolerance of the system.

[0039] Specifically, first, the hardware topology design of the embodiment of the present invention is introduced, as Figure 3 shown (a partial connection schematic diagram of the present invention). The network system (such as a supernode system) of the embodiment of the present invention includes multiple host devices HOST. Each host device can be an AI (Artificial Intelligence) server. Each host device is equipped with at least one graphics processing unit GPU. All GPUs are connected to a PCIe switch (i.e., PCIe Switch). The present invention takes 4 host devices, with each host device equipped with 8 GPU nodes as an example for discussion. That is, the 4 host devices can include host device 1 (HOST1), host device 2 (HOST2), host device 3 (HOST3), and host device 4 (HOST4). Each host device is connected to 8 GPUs. For example, HOST1: GPU1 - GPU8, HOST2: GPU9 - GPU16, HOST3: GPU17 - GPU24, HOST4: GPU25 - GPU32, a total of 32 GPUs. This shows that the 32 GPU nodes in the embodiment of the present invention are connected through a PCIe Switch, thereby achieving efficient full interconnection.

[0040] Further, during the hardware topology design process, to ensure the normal operation of the network system and the correct connection of hardware devices, first, the network system needs to be initialized; second, after initializing the network system, resource enumeration and address allocation are performed on each GPU in each host device to make each GPU match the corresponding identity identifier, so as to ensure that each GPU has a unique identifier and a correct connection relationship in the network system; finally, GPU topology information is generated based on each GPU after identity identifier matching, and topology management is performed on each GPU according to the GPU topology information, so as to ensure that the GPU resources in the network system can be correctly managed and accessed.

[0041] Thus, based on the above steps of network system initialization, resource enumeration and address allocation for GPUs, and topology management for each GPU, the efficiency and reliability of the network system can be ensured, and at the same time, it is ensured that the network system can operate normally and support subsequent hot standby switching and fault recovery.

[0042] According to an embodiment of the present invention, initializing the network system includes: connecting multiple host devices based on a PCIe switch to construct a target network interconnection architecture according to the multiple connected host devices; starting a baseboard management controller and a multi-host control processing unit, and configuring the port mapping of the PCIe switch to obtain a routing table of the PCIe switch; after the network system is started, initializing the routing table of the PCIe switch.

[0043] Specifically, during the process of initializing the network system, first, multiple host devices and the GPUs of each host device are connected based on a PCIe Switch. The GPUs are connected to the PCIe signal enhancement card PCIe Retimer Card inside the host device through the PCIe Switch inside the host device, and then through CDFP (400GigabitForm-factorPluggable, a standard module interface supporting a data transmission rate of 400 Gbps) cables, the PCIe signal enhancement card PCIe Retimer Card inside the host device is connected to the switch box PCIe Switch BOX. Inside the switch box PCIeSwitch BOX, it may be forwarded through multiple levels of PCIe Switch chips. At the same time, through CDFP cables, the PCIe Switch BOX can also be connected to the PCIe signal enhancement card PCIe Retimer Card inside another host device, and then connected to the PCIe Retimer Card through the PCIe Switch inside another host device, so as to realize the interconnection of GPUs across host devices.

[0044] In an embodiment of the present invention, 4 hosts and 8 GPUs of each host are interconnected through a PCIe Switch to construct a target network interconnection architecture through multiple interconnected host devices, forming an efficient high-performance interconnection network, that is, a PCIe Fabric network (a high-performance interconnection network based on PCIe), so as to ensure that the GPUs inside each host are fully interconnected through the PCIe Switch and can communicate through the PCIe Fabric network; secondly, enter the startup system management function to start the BMC (Baseboard Management Controller) and mCPU (Multi-host Control Processing Unit). Among them, the BMC is responsible for hardware monitoring, fault detection, and system management functions through IPMI (Intelligent Platform Management Interface) or SMBus (System Management Bus). The mCPU is management software running on the host device, which can communicate with the BMC through the Redfish API (Application Programming Interface) and is responsible for managing GPU topology, resource allocation, and hot standby switching; then, configure the port mapping of the PCIe Switch to obtain the routing table of the PCIe switch to ensure that each GPU can communicate through the correct port. When the network system starts, initialize the routing table of the PCIe Switch to ensure that data packets can be correctly routed to the target GPU.

[0045] Thus, based on the network initialization, it can lay a foundation for the normal operation of the network system to ensure that the hardware devices, management software, and network architecture can be correctly configured and started, ensure that each GPU device can communicate through the correct port, and data packets can be correctly routed to the target GPU.

[0046] According to an embodiment of the present invention, resource enumeration and address assignment are performed on each graphics processing unit in multiple host devices to make each graphics processing unit match the corresponding identity identifier, including: when multiple host devices start, scan at least one graphics processing unit configured corresponding to each host device and allocate resources to at least one graphics processing unit configured corresponding to each host device; use each host device to allocate local addresses to at least one graphics processing unit configured corresponding to each host device, and based on the address assignment result, manage and access the resources of at least one graphics processing unit configured corresponding to each host device through each host device.

[0047] Specifically, when multiple host devices are started, each host device will perform resource enumeration, that is, the host device will scan at least one GPU configured corresponding to it. For example, host device 1 will scan 8 GPUs connected to it and allocate hardware resources for each GPU. For example, memory address space, I / O (Input / Output) ports, etc. will be allocated for each GPU to ensure that the GPU can work properly, that is, the GPU can be correctly accessed and used. At the same time, each host device will also allocate a local address (Local Address) for at least one GPU configured corresponding to it. This Local Address is only valid inside the host device. Then the host device manages and accesses the internal GPU hardware resources based on the Local Address. For example, read or write the video memory of the GPU.

[0048] For example, assume that there are 4 host devices HOST1, HOST2, HOST3, and HOST4 in the network system. Each host device is connected to 8 GPUs (for example, HOST1: GPU1 - GPU8, HOST2: GPU9 - GPU16, HOST3: GPU17 - GPU24, HOST4: GPU25 - GPU32), a total of 32 GPUs. When multiple host devices are started, each host device scans and identifies the 8 GPUs connected to it and allocates a Local Address for each GPU. For example, HOST1 allocates Local Address 0x1000 for GPU1 and Local Address 0x2000 for GPU2. HOST2 allocates Local Address 0x3000 for GPU9 and Local Address 0x4000 for GPU10. Thus, HOST1 can access the resources of GPU1 through Local Address 0x1000.

[0049] It should be noted that the Local Address is usually implemented based on the memory - mapped I / O mechanism. The host device maps the registers or video memory of the GPU into the address space of each GPU of the host device. The host device accesses the GPU hardware resources by reading and writing these local addresses. Among them, in the PCIe interconnection system, each GPU has a unique address space, that is, the Local Address is the unique identifier of the GPU inside the host device.

[0050] Thus, in the network system based on PCIe interconnection, the host device manages and accesses the GPU hardware resources by allocating a Local Address for each GPU, thereby ensuring efficient communication between the host device and the GPU.

[0051] According to an embodiment of the present invention, after local addresses are assigned to at least one graphics processing unit configured for each host device, the method further includes: determining an access graphics processing unit and an accessed graphics processing unit in each host device; and accessing resources of the accessed graphics processing unit by using the access graphics processing unit based on the local address of the accessed graphics processing unit.

[0052] Specifically, in the GPU cluster inside the host device, data exchange is also required between each GPU. This exchange is also implemented based on the Local Address, that is, each GPU accesses the hardware resources of the target GPU through the Local Address of the target GPU. That is to say, first, the access GPU and the accessed GPU in each host device are determined, and then based on the Local Address of the accessed GPU, the access GPU accesses the resources of the accessed GPU.

[0053] For example, among GPU1 - GPU8 in HOST1, if GPU1 and GPU2 need to exchange data, at this time, GPU1 can access the resources of GPU2 through the Local Address 0x2000 of GPU2.

[0054] Therefore, based on the above discussion, whether the host device accesses the GPU resources configured correspondingly, or the GPUs inside the host device access each other, it is achieved through the Local Address. The GPUs inside the host device access the GPU hardware resources through the Local Address, thereby ensuring efficient communication between the GPUs inside the host device.

[0055] According to an embodiment of the present invention, after resource enumeration and address assignment are performed on the graphics processing units of multiple host devices, the method further includes: determining the graphics processing units of the access host device and the accessed host device; and managing and accessing resources of the graphics processing unit of the accessed host device by using the access host device based on the global address of the graphics processing unit of the accessed host device.

[0056] Specifically, in the PCIe Fabric network constructed above, each host device is also connected to other host devices and the GPUs corresponding to other host devices through a PCIe Switch to ensure a fully interconnected topology.

[0057] Specifically, in the PCIe Fabric network, first, the BMC and mCPU usually assign a globally unique Global Address to each GPU. The Global Address is valid throughout the system and supports cross-host data routing. Among them, the Global Address can include a Domain ID: identifying the domain of the PCIe Switch, a Bus Number: identifying the PCIe bus, a Device Number: identifying the GPU device, and a Function Number: identifying the functional unit of the GPU, etc. Secondly, determine the GPUs of the access host device and the accessed host device. Based on the routing table of the PCIe Switch, map the Global Address to the GPU of the accessed host device. When the GPU changes (such as adding or removing a device), dynamically update the routing table of the PCIe Switch. Among them, the process of updating the routing table of the PCIe Switch is usually completed by the BMC or mCPU through a management interface (such as the Redfish API). Thirdly, if GPU1 of HOST1 needs to send data to GPU9 of HOST2, GPU1 sends the data packet to the connected PCIe Switch. The PCIe Switch looks up the routing table according to the Global Address of the GPU of the accessed host device, performs cross-host transmission, and determines the next hop node of the data packet. The data packet is forwarded through multiple PCIe Switches and finally reaches GPU9. Thus, according to the Global Address of the GPU of the accessed host device, the resources of the GPU of the accessed host device are managed and accessed by using the access host device, realizing cross-host data routing and ensuring the correct routing of data at the same time.

[0058] For example, based on the 4 hosts (HOST1, HOST2, HOST3, HOST4) in the above example, each host is connected to 8 GPUs. When multiple host devices are started, a Global Address is assigned to each GPU. For example, GPU1:0:1:0:0, GPU2:0:1:1:0, GPU9:1:2:0:0, GPU10:1:2:1:0. The routing table of the PCIe Switch is configured in each PCIe Switch to map the Global Address to the port of the specific PCIe Switch. When GPU1 in HOST1 needs to transfer data to GPU10 in HOST2, GPU1 sends the data packet to the PCIeSwitch inside HOST1 to which it is connected. The data packet includes the Global Address 1:2:1:0 of GPU10. The PCIe Switch inside HOST1 looks up the routing table based on the Global Address 1:2:1:0 of GPU10. The routing table determines the next hop node of the data packet. The data packet is transmitted across hosts through the PCIe Fabric network. After being forwarded by multiple PCIe Switches, each PCIe Switch will look up the routing table based on the Global Address 1:2:1:0 of GPU10 to determine the next hop node of the data packet. Finally, it reaches the PCIe Switch of HOST2 where GPU10 is located. The PCIe Switch of HOST2 looks up the routing table based on the Global Address 1:2:1:0 of GPU10 and determines that the data packet needs to be transmitted to GPU10. The PCIe Switch of HOST2 forwards the data packet to the PCIe port of GPU10. GPU10 receives the data packet through its PCIe port and can correctly access the resources of GPU10.

[0059] Thus, by assigning a Global Address to each GPU and configuring the routing table in the PCIe Switch, efficient data routing can be achieved. This mechanism supports GPU communication across hosts and can dynamically respond to device changes in the system to ensure efficient data transmission and system reliability.

[0060] Furthermore, during the process of implementing cross-host transmission, it is necessary to convert the Local Address in the host device to the Global Address in the PCIe Fabric network based on the PCIe address mapping mechanism, so as to achieve cross-host data routing.

[0061] Specifically, first, at each PCIe port where the GPU is located, an address mapping table is designed to convert the Local Address into the Global Address. Each Local Address corresponds to a Global Address. This address mapping table is usually initialized by system management software such as BMC or mCPU when the network system starts up. Secondly, based on the Local Address and Global Address assigned to each GPU, the mapping relationship between the Local Address and the Global Address is written into the address mapping table. For example, Local Address 0x1000 → Global Address 0:1:0:0, Local Address 0x2000 → Global Address 0:1:1:0. Thirdly, when the host device needs to access a certain GPU, a data packet is sent using the Local Address, and the address mapping mechanism at the PCIe port converts the Local Address into the Global Address. The data packet is routed to the target GPU corresponding to the Global Address through the PCIe Fabric network according to the Global Address. Or, when a GPU needs to access another target GPU, a data packet is sent using the Global Address of the target GPU. The PCIe Switch looks up the routing table according to the Global Address of the target GPU and forwards the data packet to the target GPU. At the same time, when the GPUs in the network system change (such as adding or removing devices), the address mapping table needs to be dynamically updated to ensure that the data packets can be correctly routed to the new GPUs.

[0062] Thus, by designing the PCIe address mapping mechanism to convert the Local Address in the host device into the global Global Address, cross-host data routing can be achieved. This mechanism ensures the efficient management and access of all GPU resources in the system and supports the reliable operation of large-scale GPU cluster systems.

[0063] According to an embodiment of the present invention, topology management of each graphics processing unit is performed according to the graphics processing unit topology information, including: sending the graphics processing unit topology information to a multi-host control processing unit; receiving a second topology configuration file generated by the multi-host control processing unit based on the graphics processing unit topology information, and updating the routing table of the PCIe switch according to the second topology configuration file.

[0064] According to an embodiment of the present invention, after receiving the second topology configuration file generated by the multi-host control processing unit based on the GPU topology information, it further includes: at the PCIe port of each GPU, updating the address mapping table of each GPU and associating the local address and the global address of each GPU.

[0065] Specifically, to implement the whole-machine hot standby switching process of the multi-host device, the embodiments of the present invention need to design the interaction interface between the mCPU management software and the BMC and the software topology management architecture. The embodiments of the present invention adopt the GPU topology (GPUtopo) interface as the interaction interface between the mCPU and the BMC, which is used to transfer GPU topology information and routing configuration. The core of the GPUtopo interface is to realize the communication between the mCPU and the BMC through the Redfish API.

[0066] Specifically, the GPU topology information is stored in the JSON (JavaScript Object Notation) format and may include the following fields:

[0067] Id: The unique identifier (integer type) of the GPU;

[0068] Type: The type of the GPU (string type), which is divided into Local (local GPU) and Remote (remote GPU);

[0069] LocalAddrStart: The starting position of the local address of the GPU in the host device (integer type);

[0070] LocalAddrSize: The size of the local address space of the GPU (integer type);

[0071] MappingValid: Whether the address mapping is valid (integer type, 0 means invalid (illegal), 1 means valid (legal));

[0072] MappingFlag: Whether the address mapping has been set to the PCIe Switch (integer type, 0 means not set, 1 means set);

[0073] GlobalDomain: The PCIe Switch domain ID where the GPU is located (integer type);

[0074] GlobalAddrStart: The starting position of the global address of the GPU in the PCIe Fabric network (integer type).

[0075] The specific information is shown in Table 1:

[0076] Table 1

[0077]

[0078] The software topology management architecture is as follows Figure 4 shown: During the startup or operation of the network system, first, the basic input / output system collects the hardware asset information of the network system and transfers it to the baseboard management controller through H2B (Host to BMC, host device to baseboard management controller), and obtains the Local Address of each host device, and the LocalAddressTrap local address capture mechanism monitors and processes specific or abnormal events; secondly, the multi-host control processing unit sends a request to the baseboard management controller through the Redfish API in the LAN (Local Area Network), and the baseboard management controller obtains the GPU topology information and sends the obtained GPU topology information to the multi-host control processing unit; thirdly, the multi-host control processing unit sets the GPU routing topology information and sends the set GPU routing topology information to the high-speed peripheral component switch PCIe Switch. At the same time, the multi-host control processing unit generates a second topology configuration file based on the received GPU topology information and sends the second topology configuration file to the baseboard management controller through the Redfish API. The baseboard management controller updates the routing table of the PCIe Switch according to the second topology configuration file and returns the update result to the multi-host control processing unit; finally, the baseboard management controller sets the global address capture mechanism Global AddressTrap through MCTP (Management ComponentTransport Protocol) over SMBus (System Management Bus) and configures the PCIe Switch DLUT (DirectLookup Table, direct lookup table register) to send to each host device to obtain the global GPU topology information, and sends the global GPU topology information to the operating system client through KCS (Keyboard Controller Style) to achieve address mapping and routing decision-making.

[0079] For example, the second topology configuration file can be expressed as:

[0080] {

[0081] "GPU Topo Info": [{

[0082] "Id": 0,

[0083] "Type": "Local",

[0084] “LocalAddrStart”: 30580167147520,

[0085] “LocalAddrSize”: 36

[0086] “MappingValid”: 1

[0087] “MappingFlag”: 1

[0088] “GlobalDomain”: 16

[0089] “GlobalAddrStart”: 5634997092352

[0090] }, {

[0091] “Id”: 1,

[0092] “Type”: “Remote”,

[0093] “LocalAddrStart”: 30511447670784,

[0094] “LocalAddrSize”: 36

[0095] “MappingValid”: 1

[0096] “MappingFlag”: 1

[0097] “GlobalDomain”: 32

[0098] “GlobalAddrStart”: 5634997092352

[0099] }

[0101] }

[0102] Thus, when a certain host fails, the mCPU can notify the BMC to update the GPU topology information through the Redfish API. The BMC updates the routing table of the PCIe Switch according to the updated new topology information, so as to realize mapping the GPU resources of the faulty host device to the standby host. Through the Redfish API, the mCPU and the BMC can achieve efficient communication and data exchange.

[0103] ​Furthermore, in GPU topology management, it is also necessary to update the address mapping table according to the GPU topology information to ensure the correct association of the Local Address and Global Address of each GPU.

[0104] Specifically, after the mCPU generates the second topology configuration file based on the GPU topology information, at the PCIe port of each GPU, update the address mapping table of each GPU and associate the Local Address and Global Address of each GPU. After associating the Local Address and Global Address of each GPU, the BMC updates the routing table of the PCIe Switch according to the second topology configuration file, so as to ensure that data packets can be correctly routed to the target GPU.

[0105] For example, taking the above four host devices (HOST1, HOST2, HOST3, HOST4) as an example, first, the mCPU management software sends a request to the BMC through the Redfish API. The BMC obtains the GPU topology information of the current system and sends the GPU topology information to the mCPU. Secondly, the mCPU generates the second topology configuration file based on the received GPU topology information. Thirdly, at the PCIe port where each GPU is located, update the address mapping table of each GPU and associate the Local Address and Global Address, for example:

[0106] GPU1: Local Address 0x1000 → Global Address 0:1:0:0

[0107] GPU2: Local Address 0x2000 → Global Address 0:1:1:0

[0108] Finally, the BMC updates the routing table of the PCIe Switch according to the second topology configuration file to ensure that data packets can be correctly routed to the target GPU.

[0109] Thus, through the interaction interface design between the mCPU management software and the BMC, the management and hot standby switching of GPU topology information are realized through the Redfish API. The interface design includes functions such as obtaining GPU topology information, setting GPU topology information, and hot standby switching, supporting dynamic update of the routing table of the PCIe Switch, ensuring the efficiency and reliability of the system, and at the same time ensuring that data packets can be correctly routed to the target GPU based on the address mapping update, whether inside the same host or across hosts, so as to ensure that the GPU resources in the system can be correctly managed and accessed.

[0110] In summary, based on the above hardware topology design and software topology management architecture design of the network system, the BMC monitors the status of hardware devices (such as GPUs, host devices, PCIe Switches, etc.) in the network system in real time through sensors or monitoring interfaces (such as IPMI (Intelligent Platform Management Interface) or SMBus), and then based on the real-time monitoring, determines whether there is a first target host device in a fault state among multiple host devices, so as to send the fault information to the mCPU in a timely manner, enabling the mCPU to update the GPU topology configuration according to the fault information of the first target host device. The following will be described in combination with specific embodiments.

[0111] In step S202, if there is a first target host device in a fault state among multiple host devices, the first target host device is deleted, the fault information of the first target host device is generated, and the fault information is sent to the multi-host control processing unit through a preset management interface.

[0112] The preset management interface can be selected by those skilled in the art according to actual test requirements and is not specifically limited herein.

[0113] Specifically, first, a fault detection is performed. If there is a first target host device in a fault state among multiple host devices, that is, the BMC detects that a certain host device or node has a fault (such as a hardware fault, communication interruption, etc.), the first target host device is deleted, the fault information of the first target host device is generated, and a fault reminder can be generated at the same time. Then, the BMC sends the fault information and the fault reminder to the mCPU through a preset management interface, such as the Redfish API or other management interfaces. Among them, the fault detection is the trigger condition for the hot standby switching management, ensuring that the system can detect faults in a timely manner and start the switching process. Secondly, the mCPU generates a first topology configuration file for the new GPU according to the fault information of the first target host device, and sends the first topology configuration file of the new GPU to the BMC, so that subsequent resource mapping operations can be performed according to the first topology configuration file of the new GPU.

[0114] In step S203, receive the first topology configuration file of the new graphics processing unit updated by the multi-host control processing unit based on the fault information, update the routing table of the PCIe switch based on the first topology configuration file, and after the routing table of the PCIe switch is updated, start the whole machine hot standby switching function of the multi-host, and map the resources of the graphics processing unit of the first target host device to the second target host device in a normal state.

[0115] Specifically, after the BMC receives the first topology configuration file of the new GPU updated by the mCPU based on the fault information, first, the routing table of the PCIe Switch is updated based on the first topology configuration file through a management interface (such as I2C (Inter-Integrated Circuit), SMBus, etc.) to ensure that the updated routing table of the PCIe Switch is consistent with the new topology. After the routing table of the PCIe Switch is updated, when it is necessary to map the resources of the GPU of the first target host device to the second target host device in a normal state, the whole-machine hot standby switching function of the multi-host is started at this time, and the PCIeSwitch maps the resources of the GPU of the first target host device to the second target host device in a normal state according to the updated routing table of the PCIe Switch.

[0116] Thus, through fault detection, topology update, routing table update, and data rerouting, the continuity and reliability of the network system in case of faults are ensured, and the resources of the GPU of the host device in a fault state can be mapped to the host device in a normal state, thereby realizing the whole-machine hot standby switching function based on the PCIe interconnection network system.

[0117] Furthermore, if the BMC detects that a certain GPU inside the host is in a fault state, it is also necessary to generate the fault information and fault reminder of the GPU in the fault state, and send the fault information and fault reminder of the GPU in the fault state to the mCPU. The mCPU generates a new GPU topology file according to the fault information of the GPU in the fault state, and sends the new GPU topology file to the BMC. The BMC updates the routing table of the PCIe Switch based on the new GPU topology file, and after the routing table of the PCIeSwitch is updated, the PCIe Switch maps the resources of the GPU in the fault state to other GPUs in the normal state within the host device according to the updated routing table of the PCIe Switch.

[0118] Thus, through fault detection, topology update, routing table update, and data rerouting, not only can the resource mapping between multiple host devices be realized, but also the resource mapping of different GPUs within the same host device can be realized.

[0119] According to an embodiment of the present invention, updating the routing table of the PCIe switch based on the first topology configuration file includes: obtaining the first value corresponding to the first target field set by the multi-host control processing unit, and adjusting the second value of the second target field according to the first value; updating the routing table of the PCIe switch based on the first value and the second value.

[0120] According to an embodiment of the present invention, adjusting the second value of the second target field according to the first value includes: determining whether the first value is a first preset value; if the first value is the first preset value, determining that there is a resource of the graphics processing unit to be mapped currently, otherwise, determining that there is no resource of the graphics processing unit to be mapped currently; if there is a resource of the graphics processing unit to be mapped currently, determining whether the resource of the graphics processing unit to be mapped has been mapped; if the resource of the graphics processing unit to be mapped has been mapped, adjusting the second value of the second target field to the first preset value, if the resource of the graphics processing unit to be mapped has not been mapped or the resource of the graphics processing unit to be mapped has been deleted, adjusting the second value of the second target field to the second preset value.

[0121] Wherein, both the first preset value and the second preset value can be defined by those skilled in the art according to actual monitoring requirements, or obtained through a limited number of computer simulations, and no specific limitation is made here.

[0122] Specifically, during the process of updating the routing table of the PCIe switch based on the first topology configuration file, the mCPU confirms whether a certain placeholder needs to be mapped to the remote GPU by setting the first target field MappingValid. Therefore, during the operation of the system, the mCPU can dynamically update the MappingValid field as needed to ensure the correctness of the address mapping. For example, if the first value corresponding to the MappingValid set by the mCPU is the first preset value 1, it is determined that there is a resource of the GPU to be mapped currently, and the resource of the GPU to be mapped is legal, and at the same time, the resource of the GPU to be mapped is mapped to the remote GPU. If the first value corresponding to the MappingValid set by the mCPU is not the first preset value, for example, the first value is 0, it is determined that there is no resource of the GPU to be mapped currently or the resource of the GPU to be mapped is illegal, and at this time, it is not necessary to map the resource of the GPU to be mapped to the remote GPU, that is, the resource of the GPU to be mapped has not been mapped to the remote GPU.

[0123] Further, if there is currently a resource of the GPU to be mapped, it is further determined whether the resource of the GPU to be mapped has been mapped, that is, whether the resource of the GPU to be mapped has been set to the PCIe Switch. If the resource of the GPU to be mapped has been mapped, that is, the resource of the GPU to be mapped has been set to the PCIe Switch, the second value of the second target field MappingFlag is adjusted to the first preset value 1. If the resource of the GPU to be mapped has not been mapped or the resource of the GPU to be mapped is deleted, that is, the resource of the GPU to be mapped has not been set to the PCIe Switch, the second value of MappingFlag is adjusted to the second preset value 0. The setting conditions of MappingValid and MappingFlag are shown in Table 2:

[0124] Table 2

[0125]

[0126] For example, the mCPU sets the MappingValid field to confirm the address mapping status of each Placeholder. For example:

[0127] GPU1: MappingValid = 1 (mapped to the remote GPU (legal))

[0128] GPU2: MappingValid = 0 (not mapped to the remote GPU (illegal))

[0129] The BMC sets the MappingFlag field to confirm the status of each mapping. For example:

[0130] GPU1: MappingFlag = 1 (the mapping has been set to the PCIe Switch)

[0131] GPU2: MappingFlag = 0 (the mapping has not been set to the PCIe Switch)

[0132] When a failure occurs in HOST2, the mCPU updates the MappingValid field and maps the GPU resources of HOST2 to HOST4. For example:

[0133] GPU9: MappingValid = 1 (mapped to the remote GPU)

[0134] The BMC updates the MappingFlag field to ensure the correct update of the routing table of the PCIe Switch. For example:

[0135] GPU9: MappingFlag = 1 (The mapping has been set in the PCIeSwitch)

[0136] Thus, by setting MappingValid to confirm whether a certain Placeholder needs to be mapped to a remote GPU, and by setting MappingFlag to confirm whether a certain mapping has been set in the PCIe Switch, through the collaborative work of these two fields, efficient address mapping and routing table update can be achieved to ensure the correctness and reliability of the whole-machine hot standby switching.

[0137] According to an embodiment of the present invention, after deleting the first target host device, it further includes: determining whether a new target host device needs to be added; if a new target host device needs to be added, starting the new target host device, and obtaining parameter information of each graphics processing unit in the new target host device; integrating the parameter information into the global configuration file of the network system, and receiving a third topology configuration file of each graphics processing unit in the new target host device updated based on the parameter information, so as to update the routing table of the PCIe switch based on the third topology configuration file.

[0138] According to an embodiment of the present invention, before determining whether a new target host device needs to be added, it further includes: determining whether the first target host device has been repaired; if the first target host device has been repaired, using the repaired first target host device as the new target host device, and re-adding it to the target network interconnection architecture through the PCIe switch.

[0139] Specifically, when it is detected that there is a first target host device in a fault state among multiple host devices, the first target host device can be deleted at this time. For example, deleting HOST3. Before deletion, the GPU topology configuration files of each host device saved in the mCPU and BMC are as Figure 5 shown. If the first target host device is illegal and needs to be deleted, at this time, the GPU topology configuration files of each host device are updated, as Figure 6 shown. For example, the MappingValid in HOST1, HOST2, and HOST4 is updated from "MappingValid = 1" to "MappingValid = 0", and after deleting the first target host device, as Figure 7As shown by HOST1, HOST2, and HOST4 in , update the "MappingFlag": 1 in HOST1, HOST2, and HOST4 to "MappingFlag": 0, and call the Redfish API provided by the BMC to set the GPU topology and synchronize it to the BMC. The BMC can then update the routing table of the PCIe Switch according to the changes in the updated GPU topology configuration file to ensure that data packets can be correctly routed to the remaining GPU devices.

[0140] Further, after deleting the first target host device and updating the routing table of the PCIe Switch, there are 3 hosts (HOST1, HOST2, HOST4) in the network system at this time, with each host connected to 8 GPUs, for a total of 24 GPUs. If it is necessary to add the deleted first target host device HOST3, the GPU resources of HOST3 can be re-incorporated into the network system. However, before incorporation, it is necessary to determine whether the first target host device has been repaired. First, if the first target host device has been repaired, the repaired first target host device is used as the new target host device, that is, HOST3, and it is re-added to the target network interconnection architecture through the PCIe switch. And when a new target host device needs to be added, start the new target host device. After the new target host device is started, the management software of the PCIe Switch obtains the parameter information of each GPU in the new target host device through the interface provided by the BMC to obtain the GPU topology. The management software of the PCIe Switch updates the parameter information of each GPU of each GPU node in the new target host device, including Global Address, connection relationship, etc., and the remaining hosts (HOST1, HOST2, HOST4) update the GPU topology configuration file correspondingly, as Figure 6 and Figure 7 shown, and Figure 6 update the Mapping flag in HOST1, HOST2, and HOST4 in from "Mapping flag=1" to Figure 7 "Mapping flag=0" in , and Figure 6The "Id": 1, "Type": "Remote", "LocalInfo": Filled, "MappingValid": 1, "MappingFlag": 1, "GlobalInfo": Filled, "Id": 2, "Type": "Remote", "Local Info": Filled, "MappingValid": 1, "MappingFlag": 1, "GlobalInfo": Filled, and "Id": 3, "Type": "Remote", "Local Info": Filled, "MappingValid": 1, "MappingFlag": 1, "GlobalInfo": Filled in the original HOST3 (i.e., delete) are updated respectively to Figure 7 "Id": 1, "Type": "Remote", "Local Info": Filled, "MappingValid": 0, "MappingFlag": 0, "GlobalInfo": Filled, "Id": 2, "Type": "Remote", "Local Info": Filled, "MappingValid": 0, "MappingFlag": 0, "GlobalInfo": Filled, and "Id": 3, "Type": "Remote", "Local Info": Filled, "MappingValid": 0, "MappingFlag": 0, "GlobalInfo": Filled.

[0141] Secondly, the resources of the GPUs in the new target host device (HOST3) are re-integrated into the global configuration file of the network system, and the third topology configuration file of each GPU in the new target host device updated based on the parameter information is received. The mCPU management software calls the Redfish API provided by the BMC to set the GPU topology, and synchronizes the third topology configuration file of each GPU in the new target host device to the BMC, as Figure 8As shown, update the MappingValid in HOST1, HOST2, and HOST4 from "MappingValid = 0" to "MappingValid = 1", and based on the previous step, update HOST3 to "Id": 1, "Type": "Remote", "Local Info": Filled, "MappingValid": 1, "MappingFlag": 0, "GlobalInfo": Filled, "Id": 2, "Type": "Remote", "Local Info": Filled, "MappingValid": 1, "MappingFlag": 0, "GlobalInfo": Filled, and "Id": 3, "Type": "Remote", "Local Info": Filled, "MappingValid": 1, "MappingFlag": 0, "GlobalInfo": Filled. Thus, the BMC updates the routing table of the PCIe switch according to the new third topology configuration file, ensuring that data packets can be correctly routed to the newly added GPU devices.

[0142] Thus, by adding new host devices, the overall computing power of the entire network system can be improved, supporting larger-scale parallel computing, suitable for high-concurrency tasks such as deep learning and scientific computing. Also, the new host devices can serve as backup nodes to take over tasks when existing nodes fail, improving the fault tolerance of the system. At the same time, the new host devices can share computing and storage resources with other nodes, improving resource utilization.

[0143] In summary, based on the discussion of the above embodiments, the present invention can achieve the following beneficial effects:

[0144] (1) By reasonably configuring the PCIe Switch, flexible connections between various GPU cards are achieved, ensuring efficient data interaction between multiple GPU cards, significantly improving the internal data transmission efficiency of the system, reducing latency, and enhancing the overall computing performance;

[0145] (2) Support for network systems interconnected by PCIe Switch with 64 cards or even more, making this technical solution applicable to large-scale data centers and high-performance computing environments, providing higher scalability and flexibility;

[0146] (3) When there are changes in the GPU devices in the system (such as adding or removing devices), dynamically update the routing table of the PCIeSwitch. The update process is completed by the BMC or mCPU through the management interface (such as the Redfish API) to ensure the continuity and reliability of the system and enhance the fault tolerance of the system;

[0147] (4) Each host reserves resources for all GPU nodes in the system at the resource enumeration node and generates a Local Address within this host for accessing the GPU, so as to ensure that each GPU can be correctly accessed and used, avoid address conflicts and resource contention, optimize the resource configuration of the system, and improve resource utilization;

[0148] (5) Through a reasonable PCIe address mapping mechanism, convert the Local Address in the host domain into a Global Address. Even in the event of a host failure, it can quickly switch to a standby host and continue to provide services to achieve data routing within the entire system.

[0149] According to the whole-machine hot standby switching method of the network system proposed by the embodiment of the present invention, when there is a first target host device in a fault state among multiple host devices, delete the first target host device and generate corresponding fault information. Receive the first topology configuration file of the new graphics processing unit updated by the multi-host control processing unit based on the fault information, so as to update the routing table of the PCIe switch based on the first topology configuration file, and map the resources of the graphics processing unit of the first target host device to the second target host device in a normal state, solving problems such as the relatively fixed port mapping of the PCIe switch and the difficulty in quickly switching data to the standby device. Through the PCIe switch, flexible connection of each GPU and continuous update of the routing table are realized to achieve efficient data interaction and resource mapping among each GPU.

[0150] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.

[0151] The embodiment of the invention also provides a whole-machine hot standby switching device for a network system. The network system includes multiple host devices, each host device is equipped with at least one graphics processing unit, and all graphics processing units are connected to the PCIe switch.

[0152] Figure 9 It is a block diagram example of the whole-machine hot standby switching device for the network system of the embodiment of the present invention.

[0153] As shown Figure 9 in the figure, the whole-machine hot standby switching device 10 of the network system includes: a judgment module 100, a generation module 200, and a mapping module 300.

[0154] Among them, the judgment module 100 is used to judge whether there is a first target host device in a fault state among multiple host devices;

[0155] The generation module 200 is used to, if there is a first target host device in a fault state among multiple host devices, delete the first target host device, generate fault information of the first target host device, and send the fault information to the multi-host control processing unit through a preset management interface;

[0156] The mapping module 300 is used to receive the first topology configuration file of the new graphics processing unit updated by the multi-host control processing unit based on the fault information, update the routing table of the PCIe switch based on the first topology configuration file, and after the routing table of the PCIe switch is updated, start the whole-machine hot standby switching function of the multi-host, and map the resources of the graphics processing unit of the first target host device to the second target host device in a normal state.

[0157] According to an embodiment of the present invention, before judging whether there is a first target host device in a fault state among multiple host devices, the judgment module 100 further includes:

[0158] An initialization unit is used to initialize the network system, and after the network system is initialized, perform resource enumeration and address assignment on each graphics processing unit among multiple host devices, so that each graphics processing unit matches a corresponding identity identifier;

[0159] A topology management unit is used to generate graphics processing unit topology information based on each graphics processing unit after identity identifier matching, and perform topology management on each graphics processing unit according to the graphics processing unit topology information.

[0160] According to an embodiment of the present invention, the initialization unit includes:

[0161] A construction subunit is used to connect multiple host devices based on the PCIe switch to construct a target network interconnection architecture according to the connected multiple host devices;

[0162] A configuration subunit is used to start the baseboard management controller and the multi-host control processing unit, and configure the port mapping of the PCIe switch to obtain the routing table of the PCIe switch;

[0163] An initialization subunit is used to initialize the routing table of the PCIe switch after the network system is started.

[0164] According to an embodiment of the present invention, a topology management unit includes:

[0165] A sending subunit, configured to send graphics processing unit topology information to a host control processing unit;

[0166] A first update subunit, configured to receive a second topology configuration file generated by a multi-host control processing unit based on the graphics processing unit topology information, and update the routing table of a PCIe switch according to the second topology configuration file.

[0167] According to an embodiment of the present invention, after receiving the second topology configuration file generated by a multi-host control processing unit based on the graphics processing unit topology information, the update subunit further includes:

[0168] An update sub-component, configured to update the address mapping table of each graphics processing unit at the PCIe port of each graphics processing unit, and associate the local address and the global address of each graphics processing unit.

[0169] According to an embodiment of the present invention, an initialization unit includes:

[0170] An allocation subunit, configured to, when multiple host devices are started, scan at least one graphics processing unit configured corresponding to each host device, and allocate resources to the at least one graphics processing unit configured corresponding to each host device;

[0171] A first management access subunit, configured to use each host device to allocate a local address to the at least one graphics processing unit configured corresponding to each host device, and manage and access the resources of the at least one graphics processing unit configured corresponding to each host device based on the address allocation result.

[0172] According to an embodiment of the present invention, after using each host device to allocate a local address to the at least one graphics processing unit configured corresponding to each host device, the management access subunit further includes:

[0173] A determination sub-component, configured to determine the graphics processing unit for accessing and the graphics processing unit for being accessed in each host device;

[0174] An access sub-component, configured to access the resources of the graphics processing unit for being accessed by using the graphics processing unit for accessing based on the local address of the graphics processing unit for being accessed.

[0175] According to an embodiment of the present invention, after resource enumeration and address allocation are performed on each graphics processing unit in multiple host devices, the initialization unit further includes:

[0176] A determination subunit, configured to determine the graphics processing units of the access host device and the host device for being accessed;

[0177] The second management access subunit is used to manage and access the resources of the graphics processing unit of the accessed host device based on the global address of the graphics processing unit of the accessed host device.

[0178] According to an embodiment of the present invention, the mapping module 300 includes:

[0179] The first acquisition unit is used to acquire the first value corresponding to the first target field set by the multi-host control processing unit, and adjust the second value of the second target field according to the first value;

[0180] The first update unit is used to update the routing table of the PCIe switch based on the first value and the second value.

[0181] According to an embodiment of the present invention, the acquisition unit includes:

[0182] The first judgment subunit is used to judge whether the first value is the first preset value;

[0183] The determination subunit is used to determine that there is a resource of the graphics processing unit to be mapped currently if the first value is the first preset value, otherwise, determine that there is no resource of the graphics processing unit to be mapped currently;

[0184] The second judgment subunit is used to judge whether the resource of the graphics processing unit to be mapped has been mapped if there is a resource of the graphics processing unit to be mapped currently;

[0185] The adjustment subunit is used to adjust the second value of the second target field to the first preset value if the resource of the graphics processing unit to be mapped has been mapped, and adjust the second value of the second target field to the second preset value if the resource of the graphics processing unit to be mapped has not been mapped or the resource of the graphics processing unit to be mapped is deleted.

[0186] According to an embodiment of the present invention, after deleting the first target host device, the generation module 200 further includes:

[0187] The judgment unit is used to judge whether a new target host device needs to be added;

[0188] The second acquisition unit is used to start a new target host device and acquire the parameter information of each graphics processing unit in the new target host device if a new target host device needs to be added;

[0189] The second update unit is used to integrate the parameter information into the global configuration file of the network system, and receive the third topology configuration file of each graphics processing unit in the new target host device updated based on the parameter information, so as to update the routing table of the PCIe switch based on the third topology configuration file.

[0190] According to an embodiment of the present invention, before determining whether a new target host device needs to be added, the determination unit further includes:

[0191] A third determination subunit, configured to determine whether the first target host device has completed repair;

[0192] A second update subunit, configured to, if the first target host device has completed repair, use the first target host device that has completed repair as the new target host device, and re-add it to the target network interconnection architecture through a PCIe switch.

[0193] In summary, for the description of the features in the corresponding embodiments of the whole-machine hot standby switching device of the network system, reference can be made to the relevant descriptions in the corresponding embodiments of the whole-machine hot standby switching method of the network system, which will not be elaborated here one by one.

[0194] An embodiment of the present invention further provides an electronic device, which may include:

[0195] A memory 1001, a processor 1002, and a computer program stored on the memory 1001 and executable on the processor 1002.

[0196] When the processor 1002 executes the program, it implements the whole-machine hot standby switching method of the network system provided in the above embodiment.

[0197] Further, the electronic device further includes:

[0198] A communication interface 1003, configured for communication between the memory 1001 and the processor 1002.

[0199] The memory 1001 is used to store a computer program executable on the processor 1002.

[0200] The memory 1001 may include a high-speed RAM memory, and may also include a non-volatile memory, such as at least one disk memory.

[0201] If the memory 1001, the processor 1002, and the communication interface 1003 are implemented independently, the communication interface 1003, the memory 1001, and the processor 1002 can be interconnected through a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 10 only a thick line is used to represent it in Figure 10 , but it does not mean that there is only one bus or one type of bus.

[0202] Optionally, in specific implementation, if the memory 1001, the processor 1002, and the communication interface 1003 are integrated on a chip, the memory 1001, the processor 1002, and the communication interface 1003 can communicate with each other through an internal interface.

[0203] The processor 1002 may be a Central Processing Unit (CPU), or an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention.

[0204] Embodiments of the present invention also provide a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any of the above embodiments of the whole-machine hot standby switching method of the network system when running.

[0205] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media that can store computer programs such as a USB flash drive, a Read-Only Memory (ROM), a Random Access Memory (RAM), a mobile hard disk, a magnetic disk, or an optical disc.

[0206] Embodiments of the present invention also provide a computer program product. The above computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above embodiments of the whole-machine hot standby switching method of the network system.

[0207] Those skilled in the art may further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0208] The above has introduced in detail a method for hot standby switching of a whole network system provided by the present invention. Specific examples are used herein to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.

Claims

1. A method for hot standby switching of a network system, characterized in that: The network system includes a plurality of host devices, each of which is equipped with at least one graphics processing unit, and all graphics processing units are connected to a PCIe switch, and includes the following steps: Determining whether there is a first target host device in a fault state among the multiple host devices; If there is a first target host device in the fault state among the multiple host devices, deleting the first target host device, generating fault information of the first target host device, and sending the fault information to the multi-host control processing unit through a preset management interface; receiving a first topology configuration file of a new graphics processing unit updated by the multi-host control processing unit based on the fault information, updating a routing table of the PCIe switch based on the first topology configuration file, and after the routing table of the PCIe switch is updated, starting a hot standby switching function of the multi-host, mapping resources of the graphics processing unit of the first target host device to a second target host device in a normal state; Among them, before determining whether there is a first target host device in a faulty state among the multiple host devices, it also includes: initializing the network system, and after the network system is initialized, performing resource enumeration and address allocation for each graphics processing unit in the multiple host devices, so that each graphics processing unit matches the corresponding identity identifier; generating graphics processing unit topology information for each graphics processing unit after the identity identifier is matched, and performing topology management on each graphics processing unit according to the graphics processing unit topology information.

2. The whole machine hot standby switching method of the network system according to claim 1, characterized in that: The initializing the network system includes: Connecting the plurality of host devices based on the PCIe switch to construct a target network interconnection architecture according to the connected plurality of host devices; Starting a baseboard management controller and a multi-host control processing unit, and configuring the port mapping of the PCIe switch to obtain a routing table of the PCIe switch; After the network system is started, the routing table of the PCIe switch is initialized.

3. The whole machine hot standby switching method of the network system according to claim 1, characterized in that: The performing topology management on each graphics processing unit according to the graphics processing unit topology information includes: Sending the graphics processing unit topology information to the multi-host control processing unit; A second topology configuration file generated by the multi-host control processing unit based on the graphics processing unit topology information is received, and a routing table of the PCIe switch is updated according to the second topology configuration file.

4. The whole machine hot standby switching method of the network system according to claim 3 is characterized in that: After receiving a second topology configuration file generated by the multi-host control processing unit based on the graphics processing unit topology information, the method further includes: At the PCIe port of each graphics processing unit, the address mapping table of each graphics processing unit is updated, and the local address and the global address of each graphics processing unit are associated.

5. The whole machine hot standby switching method of the network system according to claim 1, characterized in that: The performing resource enumeration and address allocation on each graphics processing unit in the plurality of host devices so that each graphics processing unit matches a corresponding identity identifier includes: When the plurality of host devices are started, scanning at least one graphics processing unit of the corresponding configuration through each of the host devices, and allocating resources to the at least one graphics processing unit of the corresponding configuration; Each host device is used to allocate a local address to at least one graphics processing unit of the corresponding configuration, and based on the address allocation result, resources of the at least one graphics processing unit of the corresponding configuration are managed and accessed through each host device.

6. The whole machine hot standby switching method of the network system according to claim 5, characterized in that: After allocating a local address to at least one graphics processing unit of corresponding configuration by using each host device, the method further includes: Determining an accessing graphics processing unit and an accessed graphics processing unit in each of the host devices; Based on the local address of the accessed graphics processing unit, the resources of the accessed graphics processing unit are accessed by the accessing graphics processing unit.

7. The whole machine hot standby switching method of the network system according to claim 5, characterized in that: After performing resource enumeration and address allocation on each graphics processing unit in the plurality of host devices, the method further includes: Determining graphics processing units of an accessing host device and an accessed host device; Based on the global address of the graphics processing unit of the accessed host device, the resources of the graphics processing unit of the accessed host device are managed and accessed by the accessing host device.

8. The whole machine hot standby switching method of the network system according to claim 1, characterized in that: The updating of the routing table of the PCIe switch based on the first topology configuration file includes: Acquire a first value corresponding to a first target field set by the multi-host control processing unit, and adjust a second value of a second target field according to the first value; A routing table of the PCIe switch is updated based on the first value and the second value.

9. The whole machine hot standby switching method of the network system according to claim 8, characterized in that: The adjusting the second value of the second target field according to the first value includes: Determining whether the first value is a first preset value; If the first value is a first preset value, it is determined that there are currently resources of the graphics processing unit to be mapped; otherwise, it is determined that there are currently no resources of the graphics processing unit to be mapped; If the resources of the graphics processing unit to be mapped currently exist, determining whether the resources of the graphics processing unit to be mapped have been mapped; If the resources of the graphics processing unit to be mapped have been mapped, the second value of the second target field is adjusted to the first preset value; if the resources of the graphics processing unit to be mapped have not been mapped or the resources of the graphics processing unit to be mapped are deleted, the second value of the second target field is adjusted to the second preset value.

10. The whole machine hot standby switching method of the network system according to claim 1, characterized in that: After deleting the first target host device, the method further includes: Determine whether a new target host device needs to be added; If the new target host device needs to be added, the new target host device is started, and parameter information of each graphics processing unit in the new target host device is obtained; The parameter information is integrated into a global configuration file of the network system, and a third topology configuration file of each graphics processing unit in a new target host device updated based on the parameter information is received, so as to update a routing table of the PCIe switch based on the third topology configuration file.

11. The whole machine hot standby switching method of the network system according to claim 10, characterized in that: Before determining whether a new target host device needs to be added, the following steps are also included: Determining whether the first target host device has been repaired; If the first target host device has been repaired, the first target host device that has been repaired is used as the new target host device and is re-added to the target network interconnection architecture through the PCIe switch.

12. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the whole-machine hot standby switching method of the network system as claimed in any one of claims 1 to 11.

13. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the whole-machine hot standby switching method of the network system according to any one of claims 1 to 11.

14. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the whole-machine hot standby switching method of the network system as claimed in any one of claims 1 to 11 is implemented.

Citation Information

Patent Citations

  • Semantic intelligent information publishing and subscribing method based on P2P technology

    CN103412883A

  • Method and device for determining network topology and computer storage medium

    CN112751714A