Computing system, control method and device thereof, equipment, medium and product
Through hardware decoupling and flexible resource allocation, the problem of limited acceleration computing resource expansion capabilities in traditional multi-card systems is solved, high-density accelerated computing device deployment and large-scale computing power expansion are achieved, and computing performance and resource utilization are improved.
Patent Information
- Application Number
- CN202510897025.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-06-30
AI Technical Summary
The acceleration computing resource expansion capability of traditional multi-card systems is limited. The tightly coupled design leads to poor resource expansion, making it difficult to adapt to dynamic task requirements, and the acceleration card occupies a large space, which limits the number and density of acceleration cards in the server.
The general computing resources, high-performance switching units and accelerated computing resources are hardware decoupled. The system is divided into host equipment, first switch and accelerated computing equipment. Through the uplink and downlink port connection of the switching controller, flexible allocation and high-density deployment of accelerated computing resources are realized, the physical link binding between the motherboard and the acceleration card is unbundled, and the board layout in various forms is supported, and the resource pool configuration is optimized.
It breaks through the bottleneck of accelerated computing resources expansion, realizes high-density accelerated computing device deployment, supports single host to expand large-scale computing power and multi-host sharing large-scale computing resource pools, improving computing performance and resource utilization.
Smart Images

Figure CN120406680A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to a computing system, a control method, a device, a device, a medium and a product thereof. Background Art
[0002] With the development of artificial intelligence technology, the demand for computing power has increased significantly. In related technologies, the computing power of a single server is expanded by deploying acceleration cards inside the server, or a multi-machine multi-card cluster is constructed by interconnecting multiple servers. However, traditional multi-card systems usually adopt a tightly coupled design, resulting in a bottleneck in the expansion ability of accelerated computing resources. Summary of the Invention
[0003] The present invention provides a computing system, a control method, a device, a device, a medium and a product thereof, so as to at least solve the problem that the expansion ability of accelerated computing resources in a multi-card system reaches a bottleneck in related technologies.
[0004] The present invention provides a computing system, including: a host device, a first switch, and an accelerated computing device; Wherein, the first switch includes a plurality of interconnected switching controllers; The motherboard in the host device is connected to the upstream port of the switching controller through the first port and the first bus on the chassis panel of the host device; The accelerated computing device includes a first board and a cable adapter board. Both the first board and the cable adapter board are provided with a first slot for installing an acceleration card. The first connector of the first board is correspondingly connected to the downstream port of the switching controller through the second port and the second bus on the chassis panel of the accelerated computing device, and the connection point of the cable adapter board is correspondingly connected to the downstream port of the switching controller through the second port and the second bus.
[0005] The present invention also provides a control method for a computing system, which is applied to a first system controller of a first switch, including: Scanning the upstream ports of the switching controllers of the first switch to determine a first topology between the switching controllers and the host device; Scanning the downstream ports of the switching controllers to determine a second topology between the switching controllers and the accelerated computing device; Performing link reconstruction between the host device and the accelerated computing device according to the first topology and the second topology; Wherein, the motherboard in the host device is connected to the upstream port of the switching controller through the first port and a cable on the chassis panel of the host device; The acceleration computing device includes a first board and a cable adapter board. Both the first board and the cable adapter board are provided with first slots for installing acceleration cards. The first connector of the first board is connected to the downstream port of the switching controller through the second port on the chassis panel of the acceleration computing device and a cable. The cable adapter board is connected to the second connector of the first board through a cable.
[0006] The present invention also provides a control device for a computing system, including: A scanning module, configured to scan the upstream port of the switching controller of the first switch to determine the first topology between the switching controller and the host device; scan the downstream port of the switching controller to determine the second topology between the switching controller and the acceleration computing device; A control module, configured to perform link reconstruction between the host device and the acceleration computing device according to the first topology and the second topology; Wherein, the main board in the host device is connected to the upstream port of the switching controller through the first port on the chassis panel of the host device and a cable; The acceleration computing device includes a first board and a cable adapter board. Both the first board and the cable adapter board are provided with first slots for installing acceleration cards. The first connector of the first board is connected to the downstream port of the switching controller through the second port on the chassis panel of the acceleration computing device and a cable. The cable adapter board is connected to the second connector of the first board through a cable.
[0007] The present invention also provides an electronic device, including: a memory for storing a computer program; a processor for implementing the steps of any one of the above control methods of the computing system when executing the computer program.
[0008] The present invention also provides a non-volatile storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any one of the above control methods of the computing system are implemented.
[0009] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of any one of the above control methods of the computing system are implemented.
[0010] With the present invention, since the general computing resources, high-performance switching units, and acceleration computing resources are decoupled in hardware according to different logical functions, the system is divided into a host device, a first switch, and an acceleration computing device according to different logical functions. The first switch includes a plurality of interconnected switching controllers. The upstream ports of the switching controllers are connected to the first ports on the chassis panel of the host device, and the downstream ports of the switching controllers are connected to the second ports on the chassis panel of the acceleration computing device, thereby breaking through the bottleneck of the expansion ability of the acceleration computing resources caused by the limited space of the server chassis. At the same time, the present invention releases the binding relationship at the physical link level between the motherboard and the acceleration card, can flexibly adjust the port configuration and resource allocation path of the switching controller, and finely divides the shared resource pool to achieve the on-demand allocation of the acceleration computing resources. Further, for the problem that the acceleration card occupies a large space in the expansion of the acceleration computing resources, a first board card and a cable adapter board are arranged in the acceleration computing device, and both are provided with first slots for installing the acceleration card, so as to flexibly arrange the acceleration card in the chassis of the acceleration computing device in two forms of board cards and leave space for devices such as a relatively large power supply in the chassis, thereby enabling more acceleration cards to be further deployed in the limited chassis space, realizing a high-density deployment of the acceleration computing device. Therefore, the present invention solves the problem that the expansion ability of the acceleration computing resources in a multi-card system reaches a bottleneck, and also lays a foundation for a single host to expand large-scale computing power and build a multi-host shared large-scale computing resource pool. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] To more clearly illustrate the embodiments of the present invention, the accompanying drawings required for use in the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0012] Figure 1 Structural schematic diagram of the first computing system provided by an embodiment of the present invention; Figure 2 Structural schematic diagram of the second computing system provided by an embodiment of the present invention; Figure 3 Structural schematic diagram of the third computing system provided by an embodiment of the present invention; Figure 4 Structural schematic diagram of the fourth computing system provided by an embodiment of the present invention; Figure 5 Structural schematic diagram of the first cabinet provided by an embodiment of the present invention; Figure 6 Structural schematic diagram of the second cabinet provided by an embodiment of the present invention; Figure 7Schematic diagram of the front window structure of a chassis of an acceleration computing device provided by an embodiment of the present invention; Figure 8 Schematic diagram of the rear window structure of a chassis of an acceleration computing device provided by an embodiment of the present invention; Figure 9 Exploded view of the chassis structure of an acceleration computing device provided by an embodiment of the present invention; Figure 10 Schematic diagram of the board connection of the first acceleration computing device provided by an embodiment of the present invention; Figure 11 Schematic diagram of the board connection of the second acceleration computing device provided by an embodiment of the present invention; Figure 12 Schematic diagram of the structure of a first switch provided by an embodiment of the present invention; Figure 13 Schematic diagram of the interconnection architecture of a first switch provided by an embodiment of the present invention; Figure 14 Schematic diagram of the structure of a host device provided by an embodiment of the present invention; Figure 15 Schematic diagram of the board connection of a host device provided by an embodiment of the present invention; Figure 16 Schematic diagram of dynamic resource allocation provided by an embodiment of the present invention; Figure 17 Power on / off control architecture diagram of a computing system provided by an embodiment of the present invention; Figure 18 Schematic diagram of the first power supply architecture provided by an embodiment of the present invention; Figure 19 Schematic diagram of the second power supply architecture provided by an embodiment of the present invention; Figure 20 Schematic diagram of the system management architecture of an acceleration computing device provided by an embodiment of the present invention; Figure 21 Circuit diagram of a second board provided by an embodiment of the present invention; Figure 22 Schematic diagram of the structure of the fifth computing system provided by an embodiment of the present invention; Among them, 100 is the host device, 101 is the main board, 102 is the high-speed signal module of the host device, 103 is the hard disk module, 104 is the hard disk, 105 is the hard disk backplane, 106 is the fan module of the host device, 107 is the power module of the host device, 108 is the host management board, 109 is the host signal transfer board, 110 is the chassis panel of the host device, 111 is the first port, 200 is the first switch, 201 is the switch controller, 202 is the tray, 203 is the external port of the first switch, 204 is the internal interconnection port of the first switch, 205 is the first system controller, 206 is the management board, 207 is the base board, 208 is the second bracket of the first switch, 208 is the uplink port, 209 is the downlink port, 300 is the acceleration computing device, 301 is the acceleration card, 302 is the power module of the acceleration computing device, 304 is the second port, 305 is the first board, 3051 is the gold finger slot, 3052 is the second connector, 306 is the second board, 3061 is the gold finger, 3062 is the third connector, 307 is the cable transfer board, 3071 is the solder joint, 308 is the fan management board, 309 is the fan wall of the acceleration computing device, 310 is the power base board, 311 is the first bracket, 312 is the acceleration card fixing cross beam, 313 is the input / output board of the acceleration card, 314 is the network board of the acceleration card, 315 is the upper cover of the chassis of the acceleration computing device, 316 is the chassis of the acceleration computing device, 317 is the chassis panel of the acceleration computing device, 318 is the first bus, 319 is the second bus, 320 is the first slot, 321 is the fourth connector, 322 is the signal relay device. Detailed implementation manners
[0013] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0014] It should be noted that in the description of the present invention, the terms "include", "comprise" or any other variation thereof are intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present invention are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0015] In order to enable those skilled in the art of the present technology to better understand the solution of the present invention, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0016] Some key terms used in the embodiments of the present invention will be explained here first.
[0017] With the rapid development of fields such as artificial intelligence, large-scale deep learning, and high-performance computing (HPC), the number of model parameters and the amount of training data have increased exponentially. The computing power and communication efficiency of traditional single-card or small-scale multi-card systems can no longer meet the real-time training requirements of models with hundreds of billions of parameters. Single-machine multi-card systems, due to their high integration, low-latency communication, and easy manageability, have become one of the core infrastructures for training ultra-large-scale models. However, traditional single-machine multi-card architectures usually adopt a tightly coupled design, suffering from problems such as resource contention, limited scalability, and poor fault isolation. There is an urgent need to achieve flexible allocation and efficient collaboration of computing power, storage, and network through a decoupled architecture.
[0018] Traditional single-machine multi-card systems deploy a motherboard and multiple accelerator cards in one server. Through the interconnection between the motherboard and the accelerator cards and the interconnection between the accelerator cards, low-latency communication between multiple accelerator cards is achieved. However, this tightly coupled design makes it difficult to change the board topology once it is determined, and it is difficult to adapt to dynamic task requirements, such as the flexible switching of hybrid parallel training (data parallelism, model parallelism). A single-point failure may affect global tasks, and there is a lack of dynamic resource migration ability.
[0019] In addition, accelerator cards are usually large in size. For example, the double-width graphics processing unit (GPU) widely used in high-performance computing and artificial intelligence tasks requires two expansion slot widths of the high-speed serial computer expansion bus (Peripheral Component Interconnect Express, PCIe), and needs to be equipped with a larger fan and heat sink, which poses higher requirements for the chassis space and motherboard layout. As a result, usually only 4 to 8 accelerator cards can be deployed in one server. Although more accelerator cards can be accommodated by replacing the chassis with a larger size, this requires replacing the large cabinet chassis in the data center and expanding the space of the data center, which will bring huge costs.
[0020] To solve the problem that the expansion ability of the accelerated computing resources brought by the tightly coupled design of the traditional multi-card system reaches a bottleneck, the computing system, its control method, device, equipment, medium and product provided by the embodiments of the present invention decouple the general computing resources, high-performance switching unit and accelerated computing resources in terms of hardware according to different logical functions. The system is divided into a host device, a first switch and an accelerated computing device according to different logical functions. The first switch includes a plurality of interconnected switch controllers. The upstream port of the switch controller is connected to the first port of the chassis panel of the host device, and the downstream port of the switch controller is connected to the second port of the chassis panel of the accelerated computing device, thereby breaking through the bottleneck of the expansion ability of the accelerated computing resources caused by the limited space of the server chassis. At the same time, the present invention releases the binding relationship at the physical link level between the main board and the acceleration card, can flexibly adjust the port configuration and resource allocation path of the switch controller, and finely divides the shared resource pool to realize the on-demand allocation of the accelerated computing resources. Further, for the problem that the acceleration card occupies a large space in the expansion of the accelerated computing resources, a first board and a cable adapter board are arranged in the accelerated computing device, and both are provided with a first slot for installing the acceleration card, so as to flexibly arrange the acceleration card in the chassis of the accelerated computing device in two forms and leave space for devices such as a relatively large power supply in the chassis, so that more acceleration cards can be further deployed in the limited chassis space, realizing a high-density deployment of the accelerated computing device. Therefore, the present invention solves the problem that the expansion ability of the accelerated computing resources of the multi-card system reaches a bottleneck, and also lays a foundation for a single host to expand large-scale computing power and build a multi-host shared large-scale computing resource pool.
[0021] Figure 1 It is a schematic structural diagram of the first computing system provided by the embodiments of the present invention; Figure 2 It is a schematic structural diagram of the second computing system provided by the embodiments of the present invention; Figure 3 It is a schematic structural diagram of the third computing system provided by the embodiments of the present invention; Figure 4 It is a schematic structural diagram of the fourth computing system provided by the embodiments of the present invention; Figure 5 It is a schematic structural diagram of the first cabinet provided by the embodiments of the present invention; Figure 6 It is a schematic structural diagram of the second cabinet provided by the embodiments of the present invention.
[0022] Such as Figure 1As shown in the figure, the computing system provided by the embodiment of the present invention may include: a host device, a first switch, and an accelerated computing device; wherein, the first switch includes a plurality of interconnected switching controllers; the motherboard in the host device is connected to the upstream port of the switching controller through the first port and the first bus of the chassis panel of the host device; the accelerated computing device includes a first board and a cable adapter board, both the first board and the cable adapter board are provided with a first slot for installing an acceleration card, the first connector of the first board is correspondingly connected to the downstream port of the switching controller through the second port and the second bus of the chassis panel of the accelerated computing device, and the connection point of the cable adapter board is correspondingly connected to the downstream port of the switching controller through the second port and the second bus.
[0023] In the embodiment of the present invention, the accelerated computing device may include one or more acceleration cards, and the types of the acceleration cards may include but are not limited to a Graphics Processing Unit (GPU), a Field-Programmable Gate Array (FPGA), etc., which are installed on a Peripheral Component Interconnect Express (PCIe) slot in the accelerated computing device. Multiple acceleration cards may form an accelerated computing resource pool, or different types of acceleration cards may form a heterogeneous computing acceleration resource pool.
[0024] In the embodiment of the present invention, the switching controller (PCIe Switch) may adopt a fifth-generation PCIe switch, such as PEX89144.
[0025] As Figure 1 shown, taking a single-machine 32-card cluster as an example, 8 switching controllers may be adopted, encapsulated in the first switch, to construct a 2*4 Peripheral Component Interconnect Express (PCIe) Fabric topology. Figure 1 Among them, a single switching controller may support 144 fifth-generation PCIe links (PCIe GEN5 lanes), that is, 9 groups of x16 PCIe lanes. 4 groups of x16 lanes in each switching controller are set to the Fabric mode for Fabric interconnection between 8 switching controllers to form a PCIe Fabric network; 1 group of x16 lanes is set to the host mode for connecting the central processing unit of the host device; 4 groups of x16 lanes are set to the device mode for connecting acceleration cards.
[0026] Through such as Figure 1The structure of the computing system shown. Two central processing units (CPU 0, CPU 1) in a host device form a dual - way host. Through 8 switch controllers, a Fabric network is formed, which can connect 32 accelerator cards (GPU 0 - GPU 31). At the same time, through the PCIe Fabric network, the accelerator cards can also communicate with each other with a maximum of 3 hops (each passing through a switch controller is a hop). For example, for the accelerator cards under switch controller 0 and switch controller 2 to communicate point - to - point, since there is no Fabric link between switch controller 0 and switch controller 2, it needs to jump through switch controller 1, and the communication path needs to pass through switch controller 0, switch controller 1, and switch controller 2, and the communication hop count between the accelerator cards is 3 hops.
[0027] Regarding the accelerator cards connected to the same switch controller as a group, the accelerator cards within the group can aggregate the communication bandwidth of 512GB / s through the internal processing of the switch controller. The accelerator cards in different groups communicate through the Fabric interconnection bus between switch controllers, and the aggregated communication bandwidth is also 512GB / s. At the same time, the accelerator cards between different groups can also use a bridge to expand the interconnection, thereby further increasing the interconnection communication bandwidth.
[0028] In practical applications, in a 2 * 4 Fabric topology, for the communication between accelerator cards under two switch controllers without a Fabric link, due to the large number of hops, there are problems such as high software scheduling development difficulty and high communication latency.
[0029] To solve the above problems, while keeping the external IO ports of the switch controllers unchanged, the Fabric topology between the 8 switch controllers is modified. Figure 2 It is a 2 * 2 fully - interconnected topology for grouping. One group consists of 4 switch controllers and 16 accelerator cards, ensuring that the aggregated communication bandwidth of 512GB / s remains unchanged. The 32 cards are divided into two groups of 16 cards at the software level to achieve high - speed point - to - point communication with low latency and easy implementation within a single group. Figure 3 To change the 4 Fabric - mode x16 lanes inside the switch controller to 8 Fabric - mode x8 lanes, according to Figure 3 the interconnection path in it can achieve a 2 * 4 full interconnection. In the full - interconnection topology, there is a Fabric link between every two switch controllers. Therefore, for point - to - point communication between accelerator cards under any two switch controllers, it only needs to pass through two switch controllers, thus reducing the communication hop count to at most 2 hops and reducing the communication latency.
[0030] The computing system provided by the embodiments of the present invention, compared with the traditional 8-card AI server architecture, quadruples the number of acceleration cards connected to the host, and can achieve southbound interconnection between cards without requiring the acceleration cards themselves to support the inter-card interconnection protocol. The southbound interconnection scale of the acceleration cards quadruples, and the southbound interconnection bandwidth can also reach 512 GB / s.
[0031] At the same time, the 32-card system interconnection architecture also supports horizontal (Scale Out) networking expansion. Compared with the traditional AI server group, the number of hosts required for a thousand-card cluster is reduced to one-fourth, and the number of network cards, switches, etc. used also decreases accordingly.
[0032] Whether it is a traditional 8-card AI server or the single-machine 32-card decoupled system provided by the embodiments of the present invention, the hosts in their interconnection topologies are different PCIe buses (PCIe Bus) of the central processor of a single host device. Substantially, their upstream is a single host. As Figure 4 shown, the embodiments of the present invention also provide a multi-level multi-card decoupled system, which may include 8 host devices. The 8 host devices share an accelerated computing resource pool, which is essentially different from the interconnection topology of the single-machine 32 cards. Multi-host sharing can use an accelerated computing resource pool to process multiple tasks, improve the utilization rate of IO resources, achieve dynamic resource allocation and load balancing, realize the fusion of multi-platform processor computing power and the collaborative scheduling of heterogeneous acceleration resources, and alleviate the performance expansion bottleneck problem of the data center. Through Figure 4 the topology design described above, each of the 8 host devices outputs a group of x16 PCIe signals, and through a Fabric network composed of 8 switching controllers, 32 acceleration cards (x16 PCIe lanes interfaces) can be connected. The 8 host devices can use any acceleration card in the accelerated computing resource pool through management software; at the same time, through the PCIe Fabric network, mutual communication between up to 3 hops among the 32 acceleration cards can also be achieved. The acceleration cards connected to the same switching controller are regarded as a group. The acceleration cards within the group aggregate the communication bandwidth of 512 GB / s through the internal processing of the switching controller; the acceleration cards in different groups communicate through the Fabric interconnection bus between the switching controllers, and the aggregated communication bandwidth is also 512 GB / s. At the same time, the acceleration cards in different groups can also use a bridge for interconnection expansion, thereby further increasing the interconnection communication bandwidth.
[0033] To reduce the number of communication hops and communication latency, Figure 4 the multi-machine multi-card decoupled system shown in Figure 2 and Figure 3 can also apply the PCIe Fabric network between the switching controllers shown in
[0034] Therefore, the computing system provided by the embodiments of the present invention is different from traditional servers. It has a hardware decoupling design for the host device, the first switch, and the acceleration computing device, solving the problem of rigid IO resources in traditional servers and the inability to flexibly allocate them. At the same time, by separately deploying the acceleration card in the chassis of the acceleration computing device, the number and density of acceleration cards for accelerating computing resources in a single-machine multi-card cluster and a multi-card cluster are expanded, effectively solving the problems of small computing power density and non-concentration of computing power resources. Moreover, it does not rely on a private protocol and only needs to use the PCIe protocol to complete the point-to-point interconnection between multiple cards.
[0035] In the embodiments of the present invention, the host device, the first switch, and the acceleration computing device are vertically installed in the cabinet, and the first switch is located between the host device and the acceleration computing device. To shorten the uplink and downlink communication links of the first switch, the first switch is installed between the host device and the acceleration computing device in the cabinet.
[0036] Based on the above-mentioned single-machine multi-card computing system architecture, its cabinet layout can be as Figure 5 shown. According to different logical functions, it can be divided into: a general computing resource pool, a high-performance switching unit, an acceleration computing resource pool, etc. Among them, the general computing resource pool is composed of general computing units in the host device 100, mainly providing the computing function of the system; the acceleration computing resource pool is composed of heterogeneous computing units in the acceleration computing device 300, mainly providing acceleration computing functions for the entire computing system to achieve higher computing performance, parallelism, and energy efficiency ratio; the high-performance switching unit is a high-speed data interconnection module formed by the switching controller 201 in the first switch 200. It is closely combined with the basic software and is connected to the general computing unit and the acceleration computing resource pool through high-speed cables, realizing the physical decoupling of general computing and heterogeneous acceleration computing. At the same time, it realizes the efficient configuration management of heterogeneous acceleration cards 301 and the cross-node direct sharing of resources, improving the access efficiency of heterogeneous resources and providing a highly reliable and high-performance data interconnection network for the entire system. The system supports hybrid air-cooling and liquid-cooling, and the liquid-cooling adopts a negative-pressure liquid-cooling method. In addition, the cabinet can also include wire winding arranged in the top 1U space, as well as a data network, a task network, and a management network."U" is the unit of the server chassis, and 1U is equal to 4.445 centimeters (1.75 inches). In the data center, 4U chassis servers are widely used.
[0037] As Figure 6 shown, for a multi-machine multi-card computing system, multiple host devices 100 can be adjacent and vertically installed at the bottom of the cabinet.
[0038] The computing system provided by the embodiment of the present invention can solve problems such as the performance expansion of the host system in the current data center being limited by the system interconnect bandwidth, the performance mismatch between levels of storage, and the low utilization rate of I / O resources. Through the high-performance switching of the bus inside the first switch 200, the bottleneck of the scale expansion of the accelerator card 301 in the traditional server can be broken through, and a larger-scale accelerated computing cluster can be expanded.
[0039] Based on the above embodiment, the embodiment of the present invention continues to describe the structure of the accelerated computing device 300.
[0040] Figure 7 It is a schematic front window structure diagram of the chassis of an accelerated computing device provided by the embodiment of the present invention; Figure 8 It is a schematic rear window structure diagram of the chassis of an accelerated computing device provided by the embodiment of the present invention; Figure 9 It is an exploded view of the chassis structure of an accelerated computing device provided by the embodiment of the present invention; Figure 10 It is a schematic diagram of the board connection of the first accelerated computing device provided by the embodiment of the present invention; Figure 11 It is a schematic diagram of the board connection of the second accelerated computing device provided by the embodiment of the present invention.
[0041] In the computing system introduced in the above embodiment, an accelerated computing cluster of 32 accelerator cards 301 can be formed. The 32 accelerator cards 301 can be deployed in two accelerated computing devices 300, and the accelerated computing device 300 can adopt a 4U chassis.
[0042] In the embodiment of the present invention, a first board 305 can be provided with multiple first slots for installing multiple accelerator cards 301. A cable transfer board 307 can be provided with one or more first slots.
[0043] The cable transfer board 307 (Paddle card) is an intermediate printed circuit board (PCB) between the connector and the cable. A first slot can be deployed thereon and connected to the first board 305 through a cable to achieve a flexible layout inside the chassis of the accelerated computing device 300.
[0044] To further save the chassis space of the accelerated computing device 300, the connection point of the cable transfer board 307 can be a solder joint 3071, and the other end of the first bus 318 is connected to the second connector 3052 of the first board 305, that is, using the solder joint 3071 to replace the connector and reducing the space occupied by the connector. The second connector 3052 can adopt a multi-channel input / output connector (Multi-Channel I / O, MCIO).
[0045] An embodiment of the present invention may further include a second board 306. The second board 306 is provided with a signal relay device 322, and the signal relay device 322 is disposed between the first connector of the first board 305 and the second port 304 on the panel of the chassis 316 of the acceleration computing device 300. The signal relay device 322 may adopt a high-speed signal retimer. In the embodiment of the present invention, the retimer may adopt a retimer chip. One retimer chip can be used to provide one or more communication channels, which can be respectively used for signal relaying of high-speed data signals, clock signals, etc. In the embodiment of the present invention, one second board 306 may be provided with one or more signal relay devices 322. In some alternative embodiments of the embodiment of the present invention, one second board 306 may be provided with four retimer chips for use by four acceleration cards 301.
[0046] As Figure 7 shown, in the front window of the chassis 316 of the acceleration computing device 300, a fan wall 309 of the acceleration computing device 300 can be vertically installed. The fan wall 309 may include a plurality of fans (10 8056 fans can be adopted).
[0047] As Figure 9 shown, the first board 305 and the cable adapter board 307 can be arranged in a stepped manner in the chassis 316 of the acceleration computing device 300 to install the acceleration cards 301 vertically in two layers perpendicular to the bottom surface of the chassis 316 of the acceleration computing device 300. Thus, the acceleration cards 301 installed in the front row of the chassis and the acceleration cards 301 installed in the rear row of the chassis are at different heights, so as to ensure their respective heat dissipation requirements.
[0048] In the embodiment of the present invention, when the acceleration computing device 300 includes a second board 306, the gold fingers of the second board 306 can be inserted into the gold finger slots 3051 of the first board 305; the second board 306 is provided with a signal relay device 322, and the signal relay device 322 is disposed between the first connector of the first board 305 and the second port 304 on the panel of the chassis 316 of the acceleration computing device 300.
[0049] As Figure 9 shown, the acceleration computing device 300 may include two first boards 305. One first board 305 may be provided with two gold finger slots 3051 for inserting the gold fingers of the second board 306 and three first slots 320 for installing the acceleration cards 301.
[0050] In an embodiment of the present invention, the second board 306 may also be connected to a cable adapter board 307 to directly connect to the first slot 320 on the cable adapter board 307. Specifically, the acceleration computing device 300 may further include a cable adapter board 307 whose solder joint 3071 is connected to the third connector 3062 of the second board 306. The third connector 3062 may employ a Multi-Channel I / O (MCIO) connector.
[0051] In an embodiment of the present invention, the chassis 316 of the acceleration computing device 300 may include a first air duct and a second air duct arranged vertically; a first bracket 311 is provided on one side of the chassis 316 of the acceleration computing device 300 close to the air outlet; corresponding to the first air duct, the first board 305 is disposed above the first bracket 311; corresponding to the second air duct, the cable adapter board 307 is disposed at the bottom of the chassis 316 of the acceleration computing device 300 close to the air inlet side. The first air duct and the second air duct may be isolated by a plastic air duct housing. Then, in the chassis 316 of the acceleration computing device 300 as shown in Figure 9 Figure 5, ten cable adapter boards 307 are disposed at the bottom of the chassis close to the front window of the chassis. After the acceleration card 301 is installed, the acceleration card 301 occupies the lower 3U space of the chassis. The first board 305 is lifted by the first bracket 311, and the first board 305, the second board 306, and the acceleration card 301 installed on the first board 305 occupy the upper 3U space of the chassis. Then, at the air inlet of the front window of the chassis, cold air is blown through the first air duct to the first board 305, the second board 306, and the acceleration card 301 installed on the first board 305 in the upper 3U space at the rear side of the chassis, and through the second air duct to the cable adapter board 307 and the acceleration card 301 installed thereon in the lower 3U space at the front side of the chassis.
[0052] In an embodiment of the present invention, corresponding to the second air duct, the power module 302 of the acceleration computing device 300 may be disposed below the first bracket 311. In addition, a power module 302 may also be provided above the first bracket 311 and between the acceleration cards 301 inserted in the two groups of first boards 305.
[0053] In addition, in the chassis 316 of the acceleration computing device 300, the fan management board 308 can be arranged at the rear side of the fan wall 309. An acceleration card fixing cross beam 312 can be arranged in the middle of the chassis 316 of the acceleration computing device 300 to fix the acceleration cards 301 on the 10 front-row cable adapter boards 307. The chassis 316 of the acceleration computing device 300 can also include an input / output board card 313 and a network board card 314. The input / output board card 313 provides network interfaces, various buttons, and indicator lights. The circuit traces on the network board card 314 can be used for the network interfaces of the acceleration cards 301 to connect to the outside of the chassis 316 of the acceleration computing device 300. A power supply base plate 310 is provided at the bottom of the chassis 316 of the acceleration computing device 300, and a third system controller is provided on the power supply base plate 310 for monitoring and managing the acceleration computing device 300 and communicating with the host device 100 and the first switch 200. The chassis 316 of the acceleration computing device 300 is provided with a chassis upper cover 315, which is used to be buckled on the upper part of the chassis 316 of the acceleration computing device 300 after the components inside the chassis are installed.
[0054] Thus, the acceleration computing device 300 provided by the embodiment of the present invention can install 10 fan walls 309 of 8056 in a 4U chassis, install 10 acceleration cards 301 in the front row, install 6 acceleration cards 301 in the rear row, and the two first board cards 305 on the left and right in the rear row are carriers of the second board card 306, providing PCIe signal routing. Thus, the acceleration computing device 300 provided by the embodiment of the present invention can install 16 double-width full-height three-quarter-length GPU cards in a 4U chassis. The 16 GPU cards can be connected to the switching controller 201 through 16 retimer chips to output 16 groups of PCIe x16 lanes signals, and there is no fixed connection sequence.
[0055] Such as Figure 8As shown, the rear window of the chassis of the acceleration computing device 300, i.e., the chassis panel 317 of the acceleration computing device 300, may be provided with a plurality of second ports 304. The second ports 304 may adopt 400G pluggable connectors. The 16 groups of high-speed links in the acceleration computing device 300 are accessed through four second circuit boards 306 near the left and right sides of the chassis panel 317 of the acceleration computing device 300. Among them, 8 groups of high-speed links are connected to 8 cable adapter boards 307 through third connectors 3062 inside the circuit board on the second circuit board 306, and the other 8 groups of high-speed links are respectively connected to 2 first circuit boards 305 through the gold fingers 3061. Each first circuit board 305 accesses 4 groups of high-speed links. Among these, 3 groups of high-speed links are respectively connected to 3 acceleration cards 301 through the first slots 320 of the first circuit board 305, and the other 1 group of high-speed links is connected to 1 cable adapter board 307 through the second connector 3052 of the first circuit board 305. Therefore, a total of 10 cable adapter boards 307 access 10 groups of high-speed links, and 10 acceleration cards 301 can be expanded at the front window of the chassis 316 of the acceleration computing device 300; the first circuit boards 305 access a total of 6 groups of high-speed links, and 6 acceleration cards 301 can be expanded at the rear window of the chassis 3 of the acceleration computing device 300. The whole machine can support 16 acceleration cards 301 in total.
[0056] As Figure 10 shown, in the acceleration computing device 300, the second circuit board 306 is plugged on the gold finger slot 3051 of the first circuit board 305 through the gold finger 3061. A plurality of signal relay devices 322 are provided on the second circuit board 306. The first slots 320 are provided on both the first circuit board 305 and the cable adapter board 307. The second connector 3052 of the first circuit board 305 is correspondingly connected to the solder joint 371 of the cable adapter board 307 through the first bus 318. The third connector 3062 of the second circuit board 306 is correspondingly connected to the solder joint 371 of the cable adapter board 307 through the second bus 319. The fourth connector 321 of the second circuit board 306 is connected to the downstream port 209 of the first switch 200 through a high-speed cable to connect to the host device side.
[0057] The high-speed link topology of the acceleration computing device 300 provided by the embodiment of the present invention may be as Figure 11 shown. A total of 16 high-speed signal retimers are included in the acceleration computing device 300. Its upstream is connected to the downstream port of the first switch 200 through a high-speed interface, and its downstream is connected to the first slots 320 on the first circuit board 305 and the cable adapter board 307 respectively through paths such as high-speed cables and gold fingers 3061 through a high-speed interface (denoted as the third connector 3062).
[0058] The embodiment of the present invention continues to describe the structure of the first switching device.
[0059] Figure 12 Schematic diagram of the structure of a first switch provided by an embodiment of the present invention; Figure 13 Schematic diagram of the interconnection architecture of a first switch provided by an embodiment of the present invention.
[0060] In an embodiment of the present invention, the switching controller 201 in the first switch 200 is a key bridge for realizing the hardware decoupling of the computing system, the software and hardware reconstruction, and the on-demand combination of resources. The core of the first switch 200 is implemented by 4 groups of high-performance switching single boards, which are mainly responsible for completing the high-performance data interconnection and transmission of the entire system. The management part in the first switch 200 includes functions such as high-performance data link topology switching management, high-performance switching unit power-on / off management, and whole cabinet power-on / off management, and is the core module of the computing system.
[0061] As Figure 12 shown, through system structure optimization, the extreme utilization of compact space is realized, and a switching module interface with ultra-high performance and a large number of ports is provided, that is, Figure 12 the external port 203 of the first switch 200 shown. A first switch 200 can include 40 external ports 203. The topology supports interconnections of 1*2, 2*2, 2*3, and 2*4 (switching controller 201), and can meet the interconnection topologies and various service scenario application requirements of the general computing resource pool and the acceleration computing resource pool of the whole cabinet, and realize full-interconnection non-blocking data transmission. Multiple topology schemes in the system architecture are implemented in the first switch 200. The external ports 203 of the first switch 200 can all be uplink / downlink multiplexed. In the computing system, it can be divided into an uplink port 208 for connecting the first port 111 of the chassis panel 110 of the host device 100 and a downlink port 209 for connecting the second port 304 of the chassis panel 317 of the acceleration computing device 300.
[0062] The first switch 200 supports multi-host shared I / O (Multi-host shared I / O); supports device hot add / remove (Device hot add / remove); real-time topology management and status monitoring, and dynamic allocation of I / O resources. There is a first system controller 205 in the first switch 200 to manage and monitor the status of each part, and is responsible for communicating with the heterogeneous computing acceleration resource pool and the general computing resource pool. The first system controller 205 can adopt a baseboard management controller (BMC).
[0063] Through the modular design of key components such as a high-performance switching board, a management board 206 (located on the substrate 207), fans, and power supplies, the flexibility of the overall machine structure, the universality of components, the ease of use of the system, and the maintainability are improved. It can be quickly expanded and a large-scale large artificial intelligence acceleration computing cluster can be quickly built to cope with the computing power challenges required for training huge artificial intelligence models.
[0064] Figure 12 In [description], the switching board can be respectively installed in two trays 202, and the two trays 202 are installed in two layers in the chassis of the first switch 200. The outer side of the tray 202 is the external port 203 of the first switch 200, which can be used as an uplink port 208 or a downlink port 209. The inner side of the tray 202 is the internal interconnection port 204 of the first switch 200, and the interconnection between the switching controllers 201 can be realized through the interconnection of the internal interconnection port 204.
[0065] The BMC of the first switch 200 (denoted as the first system controller 205) is provided on the substrate 207. A second bracket 208 can be provided at the rear side of the chassis of the first switch 200, and the power module and fan of the first switch 200 are installed at the lower part of the second bracket 208.
[0066] In the embodiment of the present invention, the switching connection ports of the switching controller 201 are respectively connected to all other switching controllers 201 in the first switch 200 through the internal interconnection port 204, and the communication hop count is reduced through the full interconnection between the switching controllers 201.
[0067] In the embodiment of the present invention, if the number of switching connection ports of the switching controller 201 is less than the total number of all other switching controllers 201 in the first switch 200, one switching connection port of the switching controller 201 is at least connected to multiple other switching controllers 201 in the first switch 200 through two cables respectively. That is, it can be Figure 12 The internal interconnection port 204 of the first switch 200 shown in the figure is used for disconnection to realize the full interconnection between a larger number of switching controllers 201.
[0068] The high-speed interface connected by the switching controller 201 (i.e., the external port 203 of the first switch 200) can be a 400G pluggable connector (Compact 400G Form-factor Pluggable, CDFP).
[0069] The embodiment of the present invention continues to describe the structure of the host device 100.
[0070] Figure 14 It is a schematic structural diagram of a host device provided by an embodiment of the present invention; Figure 15 It is a schematic diagram of board card connection of a host device provided by an embodiment of the present invention.
[0071] In an embodiment of the present invention, the host device 100 may include multiple mainboards 101, and at least one central processing unit is provided on each mainboard 101. For example, one mainboard 101 may be provided with one CPU to achieve the minimum modular design of the mainboard 101, and different CPU platforms and models can be compatible by replacing the mainboard 101.
[0072] The general computing resource pool formed by the host device 100 realizes the extreme utilization of compact space through the optimized design of the system structure, provides ultra-high density artificial intelligence training computing power, ultra-high memory bandwidth and storage capacity, a single node supports large-scale artificial intelligence training computing tasks, and covers the application requirements of mainstream models and business scenarios. Through the system modular design, the flexibility of the whole machine structure, the interoperability of components, the ease of use and maintainability of the system are improved, and a large-scale artificial intelligence acceleration computing cluster can be built through rapid expansion to cope with the computing power challenges required for training a huge amount of artificial intelligence models.
[0073] As Figure 14 shown, the host device 100 may be a dual-single-way 2U machine. 103 and 104 are front-mounted hard disk modules and hard disks, which can provide an ultra-large local storage space. 105 is the hard disk backplane. 101 are two single-way mainboards. 106 is the fan module of the host device 100, and 6 6056 cooling fans can be used. 108 is the host management board integrating the master and slave second system controllers. 107 is the power module of the host device 100, and two general redundant power supplies can be used. 102 is the high-speed signal module of the host device 100 at the rear. The rear window of the chassis of the host device 100, that is, the chassis panel 110 of the host device 100, is provided with a first port 111, which may include, for example, 8 groups of x16 PCIe high-speed interfaces. Each part of the host device 100 is modularly designed, and the single-way mainboard 101 is compatible with different platforms and different models of CPUs.
[0074] The host device 100 provided by the embodiment of the present invention may have 2 CPUs and a maximum of 24 RDIMM DDR5 memories. As Figure 15As shown in the figure, the dual-single motherboard can be managed by two second system controllers (Second System Controller 0 and Second System Controller 1) on the host management board 108. Each can output PCIe signals through the high-speed interfaces on the motherboard 101, and is connected to the first port 111 on the chassis panel 110 of the host device 100 through a cable and the host signal transfer board 109. The first port 111 can adopt a high-speed interconnect interface, such as a 400G pluggable connector. Usually, a single motherboard can output 4 groups of x16 PCIe signals, so the dual-single motherboard can output a maximum of 8 groups of x16 PCIe signals. In addition, the dual-single motherboard shares a set of Universal Serial Bus (USB) and Video Graphics Array (VGA), supports two groups of OCP3.0 network cards (RJ45), and supports the Multihost function. After the 8 groups of x16 PCIe high-speed interfaces are connected to the first port 111 through the x16 host signal transfer board 109, they are connected to the first switch 200 through high-speed interconnect cables.
[0075] The embodiments of the present invention continue to describe the implementation of software-defined link reconstruction of the computing system.
[0076] Figure 16 It is a circuit diagram of a second board card provided by an embodiment of the present invention; Figure 17 It is a power-on and power-off control architecture diagram of a computing system provided by an embodiment of the present invention; Figure 18 It is a schematic diagram of the first power supply architecture provided by an embodiment of the present invention; Figure 19 It is a schematic diagram of the second power supply architecture provided by an embodiment of the present invention; Figure 20 It is a schematic diagram of the system management architecture of an acceleration computing device provided by an embodiment of the present invention; Figure 21 It is a circuit diagram of a second board card provided by an embodiment of the present invention.
[0077] The computing system provided by the embodiments of the present invention may further include a first system controller 205 disposed in the first switch 200; the first system controller 205 may be used to scan the upstream port 208 of the switch controller 201 and the downstream port 209 of the switch controller 201 to determine the first topology between the switch controller 201 and the host device 100 and the second topology between the switch controller 201 and the acceleration computing device 300.
[0078] The first system controller 205 can also be used to perform hot removal between the host device 100 and the corresponding switch controller 201 according to the first topology and the second topology after receiving a resource release request sent by the host device 100, and initialize the acceleration computing device 300 corresponding to the resource release request through the switch controller 201, so as to allocate the acceleration computing device 300 to the corresponding host device 100 according to the resource expansion request.
[0079] Based on the management topology of the computing system provided in the embodiments of the present invention, the decoupled system of the multi-host shared acceleration computing resource pool can dynamically allocate resources. As Figure 16 shown, among 8 host devices 100 (host 0 to host 7), taking host 1 releasing the acceleration card 3011 and allocating it to host 2 as an example. Host 1 processes and ends the process related to the acceleration card 3011, and the general computing resource pool initiates a dynamic adjustment request for the acceleration computing resources to the first system controller 205 in the first switch 200. The first system controller 205 hot-removes the acceleration card 3011 through the hot removal capability of the switch controller 201. The system controller of host 2 sends a request to the first switch 200 to obtain the physical location of the new device. After the first system controller 205 in the first switch 200 confirms the physical location of the device, it sends a reset signal to the heterogeneous computing unit in the acceleration computing resource pool and performs link training, so that the device operating state returns to the default value. The first system controller 205 reallocates the device to host 2 through the Fabric path. Host 2 can see the newly added device without being aware of the service, that is, the dynamic switching of heterogeneous computing power resources is completed.
[0080] The computing system provided in the embodiments of the present invention can achieve cross-node and multi-host sharing through the method of dynamic resource allocation, allocate resources on demand and apply them elastically, maximize the release of heterogeneous computing power, and achieve performance optimization in different application scenarios.
[0081] In terms of system monitoring and management, the computing system provided in the embodiments of the present invention can be divided into two layers. One layer is independently managed within the chassis 316 of the host device 100, the first switch 200, and the acceleration computing device 300. Each of the three chassis has its own system controller, which independently manages conventional items such as fan control and temperature monitoring. The second layer is the system-level management. As Figure 17 shown, the system controllers of the three parts of the host device 100, the first switch 200, and the acceleration computing device 300 are connected together through an integrated circuit bus and a network, and are redundant to each other.
[0082] Each system controller reads the port ID, encapsulates the basic information of the unit into a message, and sends it to the management system interconnection network. After each system manager obtains and parses the message, it accesses the system controller at the other end to obtain the port connection information and establishes a port and management IP mapping topology.
[0083] Each system controller in the computing system uses a dedicated auto-discovery network protocol to broadcast its own management address, device assets and other information in the network. The system controller obtains all BMC IP information and provides an IP control visualization page according to the topology of the interconnection network.
[0084] Based on the identified system topology relationship, the system controller will develop a whole-machine system topology view in the Web interface, reflect the association of each unit topology, and display the detailed information of each unit in a visualization page. The detailed information of the unit includes node type, power-on state, asset information, overall health status, etc. After the above topology identification, based on the management topology, functions such as system asset management, coordinated power-on and power-off, dynamic power consumption management, and fault management are realized.
[0085] The computing system provided by the embodiment of the present invention may further include a second system controller disposed in the host device 100 and a third system controller disposed in the acceleration computing device 300. The first system controller 205 is further configured to, after receiving a power-on signal, send a power-on signal to the third system controller to enable the third system controller to control the acceleration computing device 300 to power on, then control the first switch 200 to power on, and then send a power-on signal to the second system controller to enable the second system controller to control the host device 100 to power on.
[0086] The power-on and power-off control architecture of the computing system provided by the embodiment of the present invention is as Figure 17 shown. Compared with a general server, it is composed of multiple host devices 100, acceleration computing devices 300, first switches 200, etc. The reset logic of the entire system is relatively complex, and its control logic is: the host device 100, acceleration computing device 300, and first switch 200 are first prepared to power on separately. During the power-on process, the power-on timing of the entire system link is controlled by the system controller. First, the acceleration computing device 300 is powered on, then the first switch 200 is powered on, and finally the host device 100 is powered on to ensure the correctness of the timing link of the entire system. In case of an emergency, the system controller performs a reset restart operation or a power-off operation on the corresponding link to ensure the stability and reliability of other links.
[0087] As Figure 17As shown, the first system controller 205, the second system controller, and the third system controller can communicate based on Ethernet and the second switch. During the power-on process, the first system controller 205 sends a power-on signal to the third system controller through the second switch, so that the third system controller controls the power-on of each component in the acceleration computing device 300 through the third controller. Then, the first system controller 205 controls the power-on of each component in the first switch 200 through the first controller. Finally, the first system controller 205 sends a power-on signal to the second system controller through the second switch, so that the second system controller controls the power-on of each component of the host device 100 through the second controller. During the power-on process, the second controller monitors the platform reset signal (PLTRST) sent by the central processing unit and sends a peripheral reset signal (PERST) to the second system controller and the expansion card. The second controller also sends the peripheral reset signal to the first controller, and the first controller sends the peripheral reset signal to the third controller to reset each component in the computing system.
[0088] As Figure 18 shown, the power supply architecture adopted by the computing system according to the embodiment of the present invention can be based on a common redundant power supply (CRPS).
[0089] As Figure 19 shown, the power supply architecture adopted by the computing system according to the embodiment of the present invention can also be implemented based on a high-voltage direct current power supply (HVDC).
[0090] Then the power interfaces in each chassis of the computing system can be compatible with two power modules. The high-voltage direct current power supply can be used to solve the technical problem of multi-level transformation of the alternating current uninterruptible power supply (UPS) power supply link topology in the power supply architecture, extremely reduce the power supply link, and improve the power conversion efficiency. The high-voltage direct current power supply can provide a maximum power of 9000W per unit, and the input is directly connected to the 380V DC bus.
[0091] There are a large number of acceleration cards 301 in the acceleration computing resource pool, but not all acceleration cards 301 need to work at full power at any time. In the embodiment of the present invention, as shown, the third system controller can monitor the power consumption of each part through the integrated circuit bus, control the output of the power module 302 (which can include power module 0 and power module 1) of the acceleration computing device 300 through the power management bus (PMBUS). At the same time, it can obtain work tasks through two paths, namely the integrated circuit bus or the network. Through algorithm optimization, scattered tasks are merged, and the frequent wake-up of multiple acceleration cards 301 is reduced. By dynamically adjusting the working state and energy consumption of system resources, the overall system power consumption is minimized, "energy supply on demand" is realized, and the energy consumption per unit of computing power is reduced.
[0092] Taking the second board card 306 of the acceleration computing device 300 as an example, as shown, the second board card 306 is provided with four high-speed signal retimers (high-speed signal retimer 0, high-speed signal retimer 1, high-speed signal retimer 2, high-speed signal retimer 3), and four fourth connectors 321 (i.e., the fourth connectors 0, fourth connectors 1, fourth connectors 2, and fourth connectors 3 shown). Each fourth connector 321 introduces two sets of integrated circuit bus signals from the host device 100. One set is used for BMCI2C for CDFP ID address recognition, and the other set of mCPU I2C is used to manage the connected acceleration card 301 through the acceleration card 301 controller (mCPU).
[0093] In , x = 0, 1, 2, 3.
[0094] The high-speed signal retimer can also be provided with debugging pins connected to a debugger. The debugging pins can be a 3-pin header. Each fourth connector 321 can separately input a reset signal (PERST_CDFP0, PERST_CDFP1, PERST_CDFP2, PERST_CDFP3) and a power-on enable signal PWR_EN, which are used to control the individual power-on and reset (reset) of the acceleration card 301 for each path of the host device 100.
[0095] As shown, the reset of the high-speed signal retimer (PERSTx) can be controlled by the reset signal (PERSTx_CDFP) output by the host device 100 and the reset signal (PERSTx_CPLD) output by the third controller.
[0096] Each fourth connector 321 can be provided with a set of indicator lights. The red light can indicate that the data transmission fails and is always on, which is controlled by the third system controller. The green light can indicate that there is data transmission and flashes, and is always on when there is no data transmission, which is controlled by the GPIO of the high-speed signal retimer.
[0097] Each high-speed signal retimer can be equipped with a configuration memory for storing the configuration information of the high-speed signal retimer, such as the configuration memory 0, configuration memory 1, configuration memory 2, and configuration memory 3 shown. The configuration memory can adopt an electrically erasable programmable read-only memory (EEPROM).
[0098] In an embodiment of the present invention, the second board 306 may further be provided with a first bus arbiter; the input ends of the first bus arbiter are respectively connected to the status management controller and the accelerator card 301 controller of the acceleration computing device 300, and the output end of the first bus arbiter is connected to the status signal pin of the high-speed signal retimer.
[0099] As shown, the first bus arbiter may adopt PCA9641, which is used to access the integrated circuit bus signals (Retimerx_SMBUS) of the third system controller and the integrated circuit bus signals (mCPU_SMBUSx) of the accelerator card 301 controller to the high-speed signal retimer.
[0100] In an embodiment of the present invention, the second board 306 may further be provided with a second bus arbiter; the input end of the second bus arbiter is connected to the status signal pin of the third system controller, and the output end of the second bus arbiter is respectively connected to the status signal pins of multiple high-speed signal retimers.
[0101] As shown by Retimer0_SMBUS, Retimer1_SMBUS, Retimer2_SMBUS, Retimer3_SMBUS, the second bus arbiter may adopt PCA9548, which is used for the third system controller to monitor the status of each high-speed signal retimer.
[0102] In an embodiment of the present invention, the second board 306 may further be provided with a third bus arbiter; the input end of the third bus arbiter is connected to the status signal pin of the third system controller, and the output end of the third bus arbiter is respectively connected to the signal status pins of the configuration memories of the high-speed signal retimers.
[0103] As shown by EEPROM0_SMBUS, EEPROM1_SMBUS, EEPROM2_SMBUS, EEPROM3_SMBUS, the third bus arbiter may adopt PCA9546, which is used for the third system controller to monitor the status of each configuration memory.
[0104] As shown, on one second board 306, two high-speed signal retimers are connected through two third connectors 3062 (such as The third connectors 0 and 1 shown, and the cable connection cable adapter board 307. High-speed signal retimers 0 and 1 respectively transmit high-speed data signals x16 PCIe in and x16 PCIe out, as well as status signals (SMBUS0, SMBUS1) and clock signals (GPU0_CLK, GPU1_CLK) between the third connectors 0 and 1.
[0105] High-speed signal retimers 2 and 3 are connected to the accelerator card 301 installed on the first board 305 through the gold fingers 3061 of the second board 306. High-speed data signals x16 PCIe in and x16 PCIe out, as well as status signals (SMBUS2, SMBUS2) and clock signals (GPU2_CLK, GPU3_CLK) are transmitted between high-speed signal retimers 2 and 3 and the accelerator card 301 installed on the first board 305.
[0106] In the embodiment of the present invention, each accelerator card 301 is separately powered, and the power-on and power-off between different accelerator cards 301 do not interfere with each other. When an interconnection relationship is established between the host device 100 and the acceleration computing device, the power-on and power-off of the corresponding accelerator card 301 can be controlled by the enable signal of the host device 100. When the host device 100 issues a power-on enable signal (PWR enable) for a certain accelerator card 301, the third controller issues a 12V power enable signal (P12V_PWR_EN), a 3.3V standby power enable signal (P3V3_STBY_PWR_EN), and a 3.3V power enable signal (P3V3_PWR_EN) for the accelerator card 301 to supply power to the accelerator card 301. When the enable signal of any host device 100 is valid, the key operation of the power-on / off signal input device (PWR Button) of the acceleration computing device 300 will be ignored by the third controller; at the same time, the power-on enable switch (PWR_EN Button) corresponding to the accelerator card 301 on the host device side on the management interface (BMC WEB) of the third system controller will be in a prohibited operation state. Both the third controller and the third system controller will receive the accelerator card 301 reset (GPU_PRSNT) signal from the first slot 320. When the GPU_PRSNT of a certain first slot is pulled high, it indicates that the accelerator card 301 has been inserted into the first slot 320, and the enable signal from the host device side will be valid. Otherwise, the third system controller will save the fault data and notify the third controller to ignore the enable signal from the host device side.
[0107] When the enable signal of the host device is invalid, the power-on and power-off operations of the corresponding accelerator card 301 can be implemented through the BMC WEB interface, or the power-on and power-off control of all accelerator cards 301 can be achieved by pressing the button of the power-on and power-off signal input device on the rear IO board. When the PWRBTN button on the BMC WEB interface is pressed, the third system controller will send a control signal (BMC_GPUx_WEB_BTN) to the third controller, telling the third controller which accelerator card 301 needs to be powered on. The third controller will then supply the 12V, 3V3_STBY, and 3V3 power required by the accelerator card 301. When the button of the power-on and power-off signal input device on the rear IO board is pressed, the power-on and power-off signal input device will send a signal to the third controller. If some accelerator cards 301 are in the powered-on state, the button can be used to power off all of them with one key; if all accelerator cards 301 are in the powered-off state, the button can be used to power on all accelerator cards 301 with one key.
[0108] The third controller receives the PWR Button signal from the rear IO board, the GPU power supply enable signals from the Host and BMC respectively. At the same time, the BMC receives the GPU_PRSNT signal. When the third controller receives the EN signal but does not receive the PRSNT signal of this slot, the third controller will not issue PWR_EN, and the third system controller will record the event.
[0109] The third system controller receives the power-on and power-off enable signal of the accelerator card 301 from the host device on one hand, and also receives the power-on and power-off enable signal of the accelerator card 301 output by the third controller, and the VR output P3V3_GPUx_PWRGD GPU power-on state signal. Through the internal logic judgment of the third system controller, the power supply state of the accelerator card 301 is displayed on the BMC WEB interface.
[0110] In addition, the CDFP_EN signal of the host device is also connected to the first board 305 through the gold finger 3061, and then connected to the third system controller / third controller of the power supply backplane 310 through a cable. The signals do not affect each other, realizing the independent power-on and power-off control of the accelerator card 301.
[0111] The first board 305 is connected to the power supply backplane 310 through a 2X6 power connector to obtain 12V, 12V_GPUx (x = 0, 1, 2, the same below), and 3V3_STBY power. Three power (VR) chips are placed on each first board 305 to generate 3V3_GPUx. 3V3_STBY_GPUx is also branched out by the switch chip on the first board 305. Each 3V3_GPUx and 3V3_STBY_GPUx has its own enable signal, which is connected to the third controller on the power supply backplane 310 through an x8 MCIO connector, and the third controller controls the individual power-on and power-off of each accelerator card 301.
[0112] As shown, three I2C signals come from the third system controller on the power supply backplane 310, which are respectively used for ID recognition and clock management of the accelerator card 301, as well as managing high-speed signal retimers, connecting sensors and field replaceable units (FRUs). A temperature detector is provided on the first board 305. The temperature detector can adopt TMP75. Five temperature detectors can be provided on the first board 305 to monitor the temperature, and the board configuration information is obtained through 1 field replaceable unit.
[0113] The computing system provided by the embodiment of the present invention may further include a first clock generator; the output end of the first clock generator is respectively connected to the clock pin of the high-speed signal retimer and the clock pin of the accelerator card 301 in the first slot 320.
[0114] A first clock generator may also be provided on the first board 305. The first clock generator can adopt CK440 to provide an asynchronous clock for the entire system. CK440 emits 8 groups of clock signals, which are connected to the gold finger 3061 and are designed for compatibility (Colay) with the synchronous clock sent from the host device end on the second board 306.
[0115] As shown, the 100M CLK output by each group of fourth connectors 321 has unreliable signal quality due to too long trace length, and a reserved design is made. The clock is provided by CK440 on the first board 305.
[0116] As shown, the clock signal (CLK) output by the fourth connector 321 passes through a clock buffer to provide clock signals Retimerx_CLK and GPUx_CLK for the high-speed signal retimer and the accelerator card 301.
[0117] It is a schematic structural diagram of the fifth computing system provided by the embodiment of the present invention.
[0118] In the above embodiments, a single-machine 32-card cluster or a multi-host 32-card cluster can be constructed based on 8 switching controllers 201.
[0119] Based on the above embodiments, the computing system provided by the embodiments of the present invention may further include a third switch; the third switch is connected to the network interfaces of at least two acceleration computing devices 300.
[0120] The third switch may include a parameter plane access switch and a parameter plane aggregation switch. As shown, for the 64-card system acceleration card 301, an acceleration card 301 with an integrated network interface is selected. Two single-machine 32-card computing systems are used, and a communication link between the two 32-card computing systems is established through the parameter plane access switch and the parameter plane aggregation switch. The bandwidth from the acceleration card 301 to the parameter plane access switch is 100G, and the bandwidth from the parameter plane access switch to the parameter plane aggregation switch is 400G. 100G * 32 = 400G * 8, and there is no bandwidth loss in the communication link between the two single-machine 32-card decoupled systems. By introducing more parameter plane network switches according to this method, large-scale thousands-of-card or even tens-of-thousands-of-card systems can be constructed.
[0121] Thus, the computing system provided by the embodiments of the present invention can use an acceleration card 301 with an integrated acceleration core and network interface, connect to an Ethernet switch through optical fibers, and can perform horizontal expansion (scale out) networking to expand a thousands-of-card or tens-of-thousands-of-card cluster.
[0122] The embodiments of the present invention provide a control method for a computing system, which is applied to a first system controller of a first switch and may include: scanning the upstream ports of the switching controllers of the first switch to determine a first topology between the switching controllers and the host devices; scanning the downstream ports of the switching controllers to determine a second topology between the switching controllers and the acceleration computing devices; and performing link reconstruction between the host devices and the acceleration computing devices according to the first topology and the second topology.
[0123] Among them, the main board in the host device is connected to the upstream port of the switching controller through a first port and a cable on the chassis panel of the host device; the acceleration computing device includes a first board and a cable adapter board. Both the first board and the cable adapter board are provided with first slots for installing acceleration cards. A first connector of the first board is connected to the downstream port of the switching controller through a second port and a cable on the chassis panel of the acceleration computing device, and the cable adapter board is connected to a second connector of the first board through a cable.
[0124] The control method of the computing system provided by the embodiments of the present invention may refer to the introduction of the above computing system embodiments.
[0125] An embodiment of the present invention further provides a control device for a computing system, which may include: a scanning module, configured to scan the upstream port of the switching controller of the first switch to determine the first topology between the switching controller and the host device; scan the downstream port of the switching controller to determine the second topology between the switching controller and the acceleration computing device; and a control module, configured to perform link reconstruction between the host device and the acceleration computing device according to the first topology and the second topology.
[0126] Wherein, the main board in the host device is connected to the upstream port of the switching controller through the first port and cable of the chassis panel of the host device; the acceleration computing device includes a first board and a cable adapter board, both the first board and the cable adapter board are provided with first slots for installing acceleration cards, the first connector of the first board is connected to the downstream port of the switching controller through the second port and cable of the chassis panel of the acceleration computing device, and the cable adapter board is connected to the second connector of the first board through a cable.
[0127] The control device of the computing system provided by the embodiment of the present invention may refer to the introduction of the above computing system embodiment.
[0128] An embodiment of the present invention further provides an electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above embodiments of the control method of the computing system.
[0129] An embodiment of the present invention further provides a non-volatile storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps in any one of the above embodiments of the control method of the computing system when running.
[0130] In an exemplary embodiment, the above non-volatile storage medium may include, but is not limited to: various media such as a USB flash drive, a read-only memory (ROM for short), a random access memory (RAM for short), a mobile hard disk, a magnetic disk, or an optical disc that can store computer programs.
[0131] An embodiment of the present invention further provides a computer program product, the above computer program product includes a computer program, and when the computer program is executed by a processor, the steps in any one of the above embodiments of the control method of the computing system are implemented.
[0132] An embodiment of the present invention further provides another computer program product, including a non-volatile non-volatile storage medium, the non-volatile non-volatile storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any one of the above embodiments of the control method of the computing system are implemented.
[0133] Those skilled in the art may further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0134] The above has introduced in detail a computing system and its control method, device, equipment, medium, and product provided by the present invention. Specific examples are used herein to illustrate the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principles of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the present invention.
Claims
1. A computing system, characterized in that, Including: A host device, a first switch, and an acceleration computing device; Wherein, the first switch includes a plurality of interconnected switching controllers; The motherboard in the host device is connected to the upstream port of the switching controller through the first port and the first bus of the chassis panel of the host device; The acceleration computing device includes a first board and a cable adapter board. Both the first board and the cable adapter board are provided with a first slot for installing an acceleration card. The first connector of the first board is correspondingly connected to the downstream port of the switching controller through the second port and the second bus of the chassis panel of the acceleration computing device. The connection point of the cable adapter board is correspondingly connected to the downstream port of the switching controller through the second port and the second bus.
2. The computing system according to claim 1, wherein The first board and the cable adapter board are arranged in a stepped manner in the chassis of the acceleration computing device to install the acceleration card in two layers perpendicular to the bottom surface of the chassis of the acceleration computing device.
3. The computing system according to claim 2, wherein The acceleration computing device further includes a second board, and the gold fingers of the second board are inserted into the gold finger slots of the first board; The second board is provided with a signal relay device, and the signal relay device is arranged between the first connector of the first board and the second port of the chassis panel of the acceleration computing device.
4. The computing system according to claim 3, wherein The acceleration computing device further includes the cable adapter board whose connection point is connected to the third connector of the second board.
5. The computing system according to claim 2, wherein The chassis of the acceleration computing device includes a first air duct and a second air duct arranged vertically; A first bracket is provided on one side of the chassis of the acceleration computing device close to the air outlet; Corresponding to the first air duct, the first board is arranged above the first bracket; Corresponding to the second air duct, the cable adapter board is arranged at the bottom of the chassis of the acceleration computing device close to the air inlet side.
6. The computing system according to claim 5, wherein Corresponding to the second air duct, the power module of the acceleration computing device is arranged below the first bracket.
7. The computing system according to claim 1, wherein The connection point of the cable adapter board is a solder joint, and the other end of the first bus is connected to the second connector of the first board.
8. The computing system according to claim 1, wherein The host device includes a plurality of motherboards, and each motherboard is provided with at least one central processing unit.
9. The computing system according to claim 1, wherein The switching connection ports of the switching controller are respectively connected to all other switching controllers in the first switch.
10. The computing system according to claim 9, wherein The number of switching connection ports of the switching controller is less than the total number of all other switching controllers in the first switch. One switching connection port of the switching controller is connected to multiple other switching controllers in the first switch through at least two cables.
11. The computing system according to claim 1, wherein The host device, the first switch, and the acceleration computing device are vertically installed in the cabinet, and the first switch is located between the host device and the acceleration computing device.
12. The computing system according to claim 1, wherein It further includes a first system controller arranged in the first switch; The first system controller is used to scan the upstream port and the downstream port of the switching controller to determine the first topology between the switching controller and the host device and the second topology between the switching controller and the acceleration computing device.
13. The computing system according to claim 12, wherein The first system controller is further configured to, after receiving a resource release request sent by the host device, perform a hot removal between the host device and the corresponding switch controller according to the first topology and the second topology, and initialize the acceleration computing device corresponding to the resource release request through the switch controller, so as to allocate the acceleration computing device to the corresponding host device according to a resource expansion request.
14. The computing system according to claim 12, wherein It further includes a second system controller disposed in the host device and a third system controller disposed in the acceleration computing device; The first system controller is further configured to, after receiving a power-on signal, send a power-on signal to the third system controller to enable the third system controller to control the acceleration computing device to power on, then control the first switch to power on, and then send a power-on signal to the second system controller to enable the second system controller to control the host device to power on.
15. The computing system according to claim 1, wherein It further includes a third switch; The third switch is connected to network interfaces of at least two of the acceleration computing devices.
16. A control method for a computing system, characterized in that, The first system controller applied to the first switch includes: Scanning the upstream port of the switch controller of the first switch to determine a first topology between the switch controller and the host device; Scanning the downstream port of the switch controller to determine a second topology between the switch controller and the acceleration computing device; Performing link reconstruction between the host device and the acceleration computing device according to the first topology and the second topology; Wherein, the main board in the host device is connected to the upstream port of the switch controller through a first port on the chassis panel of the host device and a cable; The acceleration computing device includes a first board and a cable adapter board. Both the first board and the cable adapter board are provided with a first slot for installing an acceleration card. A first connector of the first board is connected to the downstream port of the switch controller through a second port on the chassis panel of the acceleration computing device and a cable, and the cable adapter board is connected to a second connector of the first board through a cable.
17. A control device of a computing system, characterized in that, It includes: A scanning module, configured to scan the upstream port of the switch controller of the first switch to determine a first topology between the switch controller and the host device; Scanning the downstream port of the switch controller to determine a second topology between the switch controller and the acceleration computing device; A control module, configured to perform link reconstruction between the host device and the acceleration computing device according to the first topology and the second topology; Wherein, the main board in the host device is connected to the upstream port of the switch controller through a first port on the chassis panel of the host device and a cable; The acceleration computing device includes a first board and a cable adapter board. Both the first board and the cable adapter board are provided with a first slot for installing an acceleration card. A first connector of the first board is connected to the downstream port of the switch controller through a second port on the chassis panel of the acceleration computing device and a cable, and the cable adapter board is connected to a second connector of the first board through a cable.
18. An electronic device, characterized in that: It includes: A memory, configured to store a computer program; A processor for implementing the steps of the control method of the computing system as described in claim 16 when executing the computer program.
19. A non-volatile storage medium, characterized in that: A computer program is stored in the non-volatile storage medium, wherein the computer program implements the steps of the control method of the computing system as described in claim 16 when executed by a processor.
20. A computer program product comprising a computer program, characterized in that The computer program implements the steps of the control method of the computing system as described in claim 16 when executed by a processor.
Citation Information
Patent Citations
Efficient heat dissipation air duct structure of communication tower and working method of efficient heat dissipation air duct structure
CN114980664A
Mainboard and computing device
CN115708040A
Memory system, memory resource adjusting method and device, electronic equipment and medium
CN116225177A
Distributed resource management method, device, system and equipment and storage medium
CN117472596A
Multi-accelerator card heterogeneous server and resource link reconstruction method
CN117687956A
Cited By
System time sequence management method and system and electronic equipment
CN120704472A
Network card, server mainboard and server
CN120729824A
Electronic equipment and operation and maintenance method
CN120743837A