An accelerator card and control method thereof and an acceleration computing system

By introducing a storage controller and storage components into the accelerator card, the local storage space can be directly expanded, solving the problem of excessively long access paths for extended memory by the accelerator card and improving the utilization rate of computing resources and the efficiency of distributed computing.

CN120492381BActive Publication Date: 2025-10-21LANGCHAO ELECTRONIC INFORMATION IND CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510983939.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-17
Publication Date
2025-10-21
Estimated Expiration
2045-07-17

AI Technical Summary

Technical Problem

The excessively long path when the accelerator card accesses extended memory leads to increased latency and limited extended memory capacity, affecting the performance of distributed computing.

Method used

An accelerator card is provided, including a processor core, a storage controller, and a storage component. The storage component is directly mounted on the storage controller to expand the local storage space, and an address mapping relationship is established between the processor core and the storage component to enable direct read and write operations.

Benefits of technology

It improves the utilization rate of single computing resources, reduces the communication volume between accelerator cards, breaks through the communication bottleneck, and improves the efficiency of distributed computing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120492381B_ABST
    Figure CN120492381B_ABST
Patent Text Reader

Abstract

The application discloses an acceleration card, a control method thereof and an acceleration computing system, and relates to the technical field of computers. The acceleration card comprises a processor core, a storage controller, a first connector and a storage component. The processor core of the acceleration card can mount the storage component based on the local storage controller to expand the local storage space, and the acceleration card has both computing power performance and storage performance. The processor core is configured to access the storage component through the storage controller via the first connector, establish a first mapping relationship between the address space of the processor core and the address of the storage component, perform a read-write task on the storage component based on the first mapping relationship, and respond to the read-write task by the storage controller to perform a read-write operation on the storage component. Therefore, the utilization rate of a single computing power resource can be improved in distributed computing, the communication volume between acceleration cards can be reduced, the communication bottleneck can be broken through, and the distributed computing efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to an accelerator card, a control method thereof, and an accelerated computing system. Background Art

[0002] With the development of artificial intelligence (AI) technology, the demand for computing power has increased significantly. By connecting accelerator cards to server hosts and interconnecting them, large-scale accelerated computing clusters can be built to achieve distributed computing, solving the computational challenges of large-scale AI models. However, high-computing accelerator cards often have limited storage resources and require the use of memory expansion cards as extended memory for the accelerator cards. However, when the accelerator card accesses this extended memory, a memory copy operation must be performed by the central processing unit (CPU). This lengthens the access path, increases latency, and limits the capacity of the extended memory, leading to bottlenecks in distributed computing performance. Summary of the Invention

[0003] The present invention provides an accelerator card, a control method thereof, and an accelerated computing system, so as to at least solve the problem in the related art that the path for the accelerator card to access the extended memory is too long.

[0004] The present invention provides an accelerator card, comprising: a processor core, a storage controller, a first connector, and a storage component;

[0005] Wherein, the storage controller is provided between the processor core and the first connector, and the first connector is also connected to the storage component;

[0006] The processor core is configured to access the storage component via the first connector through the storage controller, establish a first mapping relationship between the address space of the processor core and the address of the storage component, and perform read and write tasks on the storage component based on the first mapping relationship;

[0007] The storage controller is used to respond to the read and write tasks to perform read and write operations on the storage component.

[0008] The present invention also provides an accelerated computing system, comprising a plurality of interconnected acceleration cards;

[0009] The accelerator card includes: a processor core, a storage controller, a first connector and a storage component;

[0010] Wherein, the storage controller is provided between the processor core and the first connector, and the first connector is also connected to the storage component;

[0011] The processor core is configured to access the storage component through the storage controller, establish a first mapping relationship between the address space of the processor core and the address of the storage component, and perform read and write tasks on the storage component based on the first mapping relationship;

[0012] The storage controller is used to respond to the read and write tasks to perform read and write operations on the storage component.

[0013] The present invention also provides a control method for an accelerator card, which is applied to a processor core of the accelerator card, comprising:

[0014] Accessing the storage component through a storage controller to obtain storage resource information of the storage component;

[0015] Establishing a first mapping relationship between the address space of the processor core and the address of the storage component according to the storage resource information;

[0016] Executing a read and write task on the storage component based on the first mapping relationship;

[0017] The storage controller is arranged between the processor core and the first connector of the acceleration card, and the first connector is also connected to the storage component.

[0018] Through the present invention, an accelerator card including a processor core, a storage controller, a first connector, and a storage component is provided. The storage controller is arranged between the processor core and the first connector, and the first connector is also connected to the storage component. Based on this, the processor core of the accelerator card can mount the storage component based on the local storage controller to expand the local storage space, thereby realizing an accelerator card with both computing power performance and storage performance. The processor core is used to access the storage component through the storage controller via the first connector, establish a first mapping relationship between the address space of the processor core and the address of the storage component, and perform read and write tasks on the storage component based on the first mapping relationship. The storage controller responds to the read and write tasks to perform read and write operations on the storage component, thereby improving the utilization rate of a single computing power resource in distributed computing, reducing the communication volume between accelerator cards, breaking through the communication bottleneck, and improving the efficiency of distributed computing. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the embodiments of the present invention, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0020] Figure 1 A schematic diagram of the structure of an accelerator card provided in an embodiment of the present invention;

[0021] Figure 2 A schematic diagram of the structure of a computing processor provided in an embodiment of the present invention;

[0022] Figure 3 A schematic diagram of the structure of interconnection between accelerator cards provided by an embodiment of the present invention;

[0023] Among them, 100 is an accelerator card, 101 is a computing processor, 102 is a memory particle, 103 is a first non-volatile storage device, 104 is a first connector, 105 is a third connector, 106 is an optical module, 107 is a management controller, 108 is a power module, and 109 is a first gold finger interface; 200 is a server host, 300 is an Ethernet switch, and 400 is a unified address management server. DETAILED DESCRIPTION

[0024] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.

[0025] It should be noted that, in the description of the present invention, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. The terms "first," "second," etc., in the present invention are used to distinguish similar objects, and are not used to describe a particular order or precedence.

[0026] In order to enable those skilled in the art to better understand the solutions of the present invention, the present invention is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0027] Modern data centers are employing an increasing number of accelerator cards and processor cores for cloud computing and artificial intelligence (AI) computing. However, the average memory size and bandwidth allocated to each processor core have not increased accordingly. However, AI model computations require more accelerator cards and memory units to support them. The computing power of a single accelerator card is limited, necessitating collaborative computations involving multiple accelerator cards. This involves interconnecting and communicating between accelerator busses and transmitting intermediate computational cache data. The memory accessible by a single accelerator card is limited, so further expansion of accelerator card memory capacity is needed to enable greater memory sharing between accelerator cards.

[0028] Compute Express Link (CXL) is a high-speed serial protocol that allows fast and reliable data transfer between components within a computer system. It aims to address bottlenecks in high-performance computing, including memory capacity, memory bandwidth, and I / O latency. CXL also enables memory expansion and sharing, and can communicate with peripherals such as computing accelerators (GPUs and FPGAs), providing faster and more flexible data exchange and processing.

[0029] In related technologies, memory expansion cards based on the Computer Express Memory Interconnect Type 3 (CXL Type 3) protocol and constructed from memory chips in the form of gold fingers for the high-speed Peripheral Component Interconnect Express (PCIe) bus can provide extended memory for accelerator cards. However, access to this extended memory requires a memory copy operation within the central processing unit (CPU), or through the CPU's internal root complex (RC) controller or PCIe switch. This results in a long access path, high latency, and limited expansion capacity.

[0030] In order to solve the problem of the path process of the accelerator card accessing the extended memory when using a memory expansion card as the extended memory of the accelerator card, an embodiment of the present invention provides an accelerator card including a processor core, a storage controller, a first connector and a storage component. The storage controller is arranged between the processor core and the first connector, and the first connector is also connected to the storage component. Based on this, the processor core of the accelerator card can mount the storage component based on the local storage controller to expand the local storage space, thereby realizing an accelerator card with both computing power performance and storage performance. The processor core is used to access the storage component through the storage controller through the first connector, establish a first mapping relationship between the address space of the processor core and the address of the storage component, and perform read and write tasks on the storage component based on the first mapping relationship. The storage controller responds to the read and write tasks to perform read and write operations on the storage component, thereby improving the utilization rate of a single computing power resource in distributed computing, reducing the communication volume between accelerator cards, breaking through the communication bottleneck, and improving the efficiency of distributed computing.

[0031] Figure 1 A schematic diagram of the structure of an accelerator card provided in an embodiment of the present invention; Figure 2 A schematic diagram of the structure of a computing processor provided in an embodiment of the present invention.

[0032] The acceleration card provided by an embodiment of the present invention may include: a processor core, a storage controller, a first connector and a storage component; wherein the storage controller is arranged between the processor core and the first connector, and the first connector is also connected to the storage component; the processor core is used to access the storage component through the first connector via the storage controller, establish a first mapping relationship between the address space of the processor core and the address of the storage component, and perform read and write tasks on the storage component based on the first mapping relationship; the storage controller is used to respond to the read and write tasks to perform read and write operations on the storage component.

[0033] In the embodiment of the present invention, the accelerator card may refer to a graphics processing unit (GPU).

[0034] The external interfaces of an accelerator card typically include a first interface for connecting to a server host and an inter-card interconnect connector for connecting to other accelerator cards. The first interface is typically a PCIe interface, plugged into the first slot (PCIe slot) of the server host using a gold finger. The inter-card interconnect connector can be a high-speed NVLink inter-card interconnect connector. Different accelerator cards in the same server can be directly interconnected through the inter-card interconnect connector, and accelerator cards in different servers can be interconnected using the inter-card interconnect connector and the NVLink-compatible switch controller.

[0035] In this embodiment of the present invention, the first connector for connecting to the storage component can be the inter-card interconnect connector of the accelerator card, and the storage controller is the cross-board forwarding control module corresponding to the first connector. In other words, one or more inter-card interconnect connectors of the accelerator card can be configured as the first connector for connecting to the storage component, and the cross-board forwarding control module in the accelerator card, which originally controls the inter-card interconnect connector to implement the inter-card interconnection function, can be configured as the storage controller.

[0036] In other optional implementations of the embodiments of the present invention, if the inter-card interconnect connector of the accelerator card is used as the first connector for connecting the storage component, a storage controller can also be deployed separately based on the hardware resources of the computing processor of the accelerator card, and the connection between the inter-card interconnect connector and the cross-board forwarding control module can be changed to a connection with the storage controller.

[0037] In an embodiment of the present invention, the first connector for connecting the storage component can also use the first pin of the first interface of the accelerator card; the first interface is the interface used by the accelerator card to connect to the server host; the storage controller is connected to the storage component via the first pin, the onboard wiring of the server host, and the first slot of the server host, and the first slot is used to install the storage component. The first interface is the interface used by the accelerator card to connect to the server host. If the number of channels of the first interface is sufficient, the first interface can be configured as two lines, one of which is still used to connect to the server host, and the other is used to connect to the first slot via the onboard wiring on the server host. The storage component is installed in the first slot, and the storage server is connected to the storage component via the first connector.

[0038] Therefore, by utilizing the external interface resources of the acceleration card to configure a first connector for connecting a storage component and implementing a storage controller based on the computing processor of the acceleration card, the problem of low computing power utilization caused by the computing power of the acceleration card not being matched with memory resources can be solved.

[0039] In an embodiment of the present invention, the processor core can also be obtained by programming the logic circuit of the first controller to implement an accelerator card. The first controller can be a programmable controller, such as a field programmable gate array (FPGA), or other types of programmable controllers.

[0040] The logic circuit based on the programmable controller provides hardware resources. According to the open source code of the accelerator card, an accelerator card instance can be built to serve as the processor core. Other hardware resources on the first controller, such as serial-to-parallel conversion channels (Serdes), computer express interconnect controller (CXL IP) resources, storage controller resources, etc., can also be used to mount larger-capacity storage components, so that high-computing-power accelerator card instances can be directly mounted on the extended memory, thereby solving the problem of low computing power utilization caused by the lack of matching memory resources for the computing power of the accelerator card in related technologies.

[0041] like Figure 1 As shown, the processor core and the memory controller may be provided in the computing processor 101 .

[0042] The computing processor 101 may include one or more processor cores (e.g. Figure 2 Processor core 1, processor core 2, ..., processor core N) are shown.

[0043] If the computing processor 101 includes multiple processor cores, the multiple processor cores can share the storage resources of the storage component, or the storage resources can be divided into storage resources corresponding to different processor cores, or the storage component can include both shared storage resources and storage resources unique to the processor cores.

[0044] like Figure 1 As shown, the accelerator card 100 may include a first gold finger interface 109. The first gold finger interface 109 may be a 16-lane fifth generation PCIe (PCIe Gen5 x16) gold finger interface. The accelerator card 100 is inserted into a PCIe slot of the server host 200 via the first gold finger interface 109. If the first connector 104 is implemented using the pins of the first interface, it is also implemented using the pins of the first gold finger interface 109.

[0045] In the embodiment of the present invention, the connector of the accelerator card 100 used for interconnection between cards may be referred to as a second connector.

[0046] The second connector may include a third connector 105 for interconnecting different acceleration cards 100 in the same server. Different acceleration cards 100 in the same server may be interconnected through the third connector 105 and cables.

[0047] The third connector 105 may be an NVLink connector, so different accelerator cards 100 in the same server can be directly connected via the third connector 105 .

[0048] The third connector 105 can also be a multi-channel input / output (MCIO) connector, and the cable can be a PCIe bus. The third connector 105 of the accelerator card 100 can then be connected to the switch controller via a cable, enabling interconnection between the accelerator cards 100 and between the accelerator card 100 and the server host 200 via the switch controller.

[0049] In other optional implementations of the present invention, at least one thread of at least one processor core runs a root complex (RC) bridge driver to configure the storage component as a bus endpoint (EP) device. This allows the processor core to actively access the storage component of another accelerator card 100 through a direct connection between accelerator cards 100, without going through the server host 200, switch controller, or adapter card. This further shortens the access path and reduces access latency. The third connector 105 utilizes a multi-channel input / output (MCIO) connector. Each MCIOx4 interface contains four pairs of high-speed SerDes (Serializer-Deserializer) channels, each supporting 50 Gb / s bandwidth. This provides a unidirectional bandwidth of 200 Gb / s and a bidirectional bandwidth of 400 Gb / s, meeting the high-bandwidth communication requirements of direct inter-card connections. Any two cards within a node can then communicate directly via multi-channel link technology (MC-link), facilitating system integration without requiring switches or adapter cards, and offering high bandwidth and low latency.

[0050] like Figure 1 As shown, the second connector of the accelerator card 100 for interconnecting cards can also include an optical module 106. The accelerator card 100 can access the unified address management server 400 through the optical module 106 to report local storage resources, accept the unified address management server 400 to allocate a super-node unified address space for multiple accelerator cards 100, and initialize the unified address translation table (UATT) based on the allocated super-node unified address space and the address space of the local processor core, which is recorded as the second mapping relationship. Therefore, based on the global unified address, it is possible to realize the interconnection between accelerator cards 100 in the same server and the interconnection between accelerator cards 100 across servers, and the server host 200 can also access the accelerator card 100 based on the global unified address. The optical module 106 can adopt a 400G optical fiber network module, through Figure 2 The optical module port shown is connected to a fiber optic network.

[0051] In an embodiment of the present invention, the optical module 106 can also adopt the inter-card interconnection connector of the accelerator card, and the inter-card interconnection connector can be configured as the optical module 106. One or more inter-card interconnection connectors of the accelerator card 100 can be configured as the optical module 106, and the cross-board forwarding control module originally used to control the inter-card interconnection connector to realize the inter-card interconnection function in the accelerator card can be configured as the super-node forwarding module provided in the embodiment of the present invention or the super-node forwarding module can be separately deployed based on the hardware resources of the computing processor 101 of the accelerator card 100, so as to achieve a larger communication bandwidth when interconnecting across server cards.

[0052] In an embodiment of the present invention, at least one thread of at least one processor core of the accelerator card 100 runs an RC bridge driver. Different accelerator cards 100 located in the same server can be directly interconnected through the first connector 104 and a cable. Different accelerator cards 100 across servers can be point-to-point interconnected through the optical module 106, or many-to-many interconnected can be achieved through the optical module 106 and an Ethernet switch.

[0053] In an embodiment of the present invention, the processor core performs read and write tasks on the storage component based on the first mapping relationship, which may include: the processor core reports storage resource information of the accelerator card 100 to the unified address management server 400, and receives the super node unified address space allocated by the unified address management server 400 to the accelerator card 100, initializes a second mapping relationship between the super node unified address space and the address space of the processor core, and performs read and write tasks on the storage component based on the second mapping relationship and the first mapping relationship.

[0054] like Figure 1 As shown, the accelerator card 100 may also include a management controller 107, which is connected to the processor core of the accelerator card 100 and is used to monitor the status of the accelerator card's components. The management controller 107 can be connected to the computing processor 101, the power module 108, and other components of the accelerator card 100 to monitor the operating status of each component and perform fault detection, fault recording, and fault reporting. The management controller 107 can be a microcontroller (MCU) that uses sensors to collect information such as voltage, current, and temperature of components on the accelerator card 100.

[0055] The accelerator card provided by an embodiment of the present invention includes a processor core, a storage controller, a first connector and a storage component. The storage controller is arranged between the processor core and the first connector, and the first connector is also connected to the storage component. Based on this, the processor core of the accelerator card can mount the storage component based on the local storage controller to expand the local storage space, thereby realizing an accelerator card with both computing power performance and storage performance. The processor core is used to access the storage component through the storage controller via the first connector, establish a first mapping relationship between the address space of the processor core and the address of the storage component, and perform read and write tasks on the storage component based on the first mapping relationship. The storage controller responds to the read and write tasks to perform read and write operations on the storage component, thereby improving the utilization rate of a single computing power resource in distributed computing, reducing the communication volume between accelerator cards, breaking through the communication bottleneck, and improving the efficiency of distributed computing.

[0056] Based on the above embodiments, Figure 1As shown, in the accelerator card 100 provided in the embodiment of the present invention, the storage controller may include a memory access controller ( Figure 1 (not shown), the storage component may be a memory particle 102.

[0057] The type of memory particle 102 can be Double Data Rate (DDR) memory, and the packaging form used can be a Dual Inline Memory Module (DIMM) form, concentrating multiple memory particles 102 on a circuit board, or a Registered DIMM (RDIMM), an Unbuffered DIMM (UDIMM), a Mini Dual In-line Memory Module (Mini-DIMM) and other packaging forms. Memory particles 102 of other types and other packaging forms can also be used.

[0058] like Figure 2 As shown, the memory access controller is provided in the computing processor 101 and is used to connect to the memory chip 102. The memory access controller is used to parse and forward access request information and access response information to the storage component. The access request information may include the access type and access address.

[0059] To further expand the storage resources of the accelerator card 100, in the accelerator card 100 provided in an embodiment of the present invention, the storage controller may include a memory access controller and a storage access controller; the storage component includes a memory chip 102 and a first non-volatile storage device 103; the memory access controller is also connected to the memory chip 102 through a corresponding first connector 104; the first end of the storage access controller is connected to the memory access controller; and the second end of the storage access controller is connected to the first non-volatile storage device 103 through the corresponding first connector 104.

[0060] The memory access controller then parses and forwards access request information and access response information for the storage component, which may include: the memory access controller parses the access request information, and if the access address in the access request information hits the memory, reads the corresponding access response information from the memory cell 102 and returns it to the memory access controller; if the access address in the access request information does not hit the memory, forwards the access request information to the storage access controller, and caches the response data in the access response information returned by the storage access controller in the memory cell 102. The storage access controller is configured to read the corresponding response data from the first non-volatile storage device 103 according to the access address in the access request information sent by the memory access controller, and returns it. This allows the first non-volatile storage device 103 to expand storage resources for the accelerator card 100 and improve the processing efficiency of access tasks based on the memory-storage architecture.

[0061] In an embodiment of the present invention, the first non-volatile storage device 103 may be, but is not limited to, a solid-state drive (such as a non-volatile memory express solid-state drive NVMe SSD), a flash memory (NAND Flash Memory), an electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, EEPROM), etc. Other types of non-volatile storage devices may also be used.

[0062] The accelerator card 100 provided in an embodiment of the present invention can use the first connector 104 to directly expand the first non-volatile storage device 103, such as an NVMe solid-state drive, thereby reducing the coupling between the accelerator card 100 and the server when using the extended memory, reducing the latency of the processor core of the accelerator card 100 accessing the extended memory, and increasing the local storage capacity of the accelerator card 100.

[0063] In the above embodiment, the storage component can be used as a shared storage for multiple processor cores. Figure 2 As shown, if the memory particles 102 serve as shared memory for multiple processor cores, the storage controller may also include a shared memory controller disposed between the processing core and the memory access controller. The shared memory controller is used to execute a first data movement task between the server host 200 and the acceleration card 100 where it is located and / or a second data movement task between different processor cores within the acceleration card 100.

[0064] The shared memory controller can implement shared memory based on the Computer Express Memory Interconnect (CXL.mem) protocol and, as a DMA controller, can mount an extended DDR4 DIMM controller. The DMA controller and memory can cache part of the data in the NVMe solid-state drive. The DMA controller is responsible for data movement between the host and the accelerator card 100, data movement within the accelerator card 100, and RDMA data movement over the network.

[0065] Based on the above embodiments, Figure 2 As shown, the accelerator card 100 may further include a port routing forwarding module, an inter-board forwarding module, and a second connector; the port routing forwarding module is disposed between the processor core and the storage controller, and is further connected to the inter-board forwarding module; the inter-board forwarding module is connected to another accelerator card 100 via the second connector, and is used to forward data packets between the accelerator card 100 where it is located and the other accelerator card 100; the port routing forwarding module is used to perform data packet format conversion between the accelerator card 100 where it is located and the other accelerator card 100.

[0066] In an embodiment of the present invention, the port routing forwarding module and the cross-board forwarding module may be provided in the computing processor 101 and implemented based on the logic circuit programming of the first accelerator card.

[0067] The port routing forwarding module can be used to convert data packets based on a unified address for accessing the memory of other local acceleration cards 100 into data packets based on a port ID and an offset address, and send the data packets to the inter-board forwarding module.

[0068] In the embodiment of the present invention, the cross-board forwarding module may be an intra-node forwarding module; the intra-node forwarding module is connected to another acceleration card 100 of the server via a second connector and a cable.

[0069] In this embodiment of the present invention, the intra-node forwarding module, namely the intra-node forwarding Media Access Control Address (MAC) module, can be used to forward data packets that access the memory of other local accelerator cards 100 based on the port ID. For example, two interconnected accelerator cards 100 in a server can access the memory of another accelerator card 100 through the local intra-node forwarding module.

[0070] In the embodiment of the present invention, the cross-board forwarding module may be a super-node forwarding module; the super-node forwarding module is connected to another acceleration card 100 of another server via a second connector.

[0071] In the embodiment of the present invention, the supernode forwarding module, ie, the supernode forwarding MAC module, may be used to send and receive data packets from the optical module 106 accessing the local memory.

[0072] like Figure 2 As shown, in this embodiment of the present invention, the second connector corresponding to the supernode forwarding module can be an optical module 106. The accelerator card 100 also includes a remote direct memory access protocol stack module located between the supernode forwarding module and the storage controller. The optical module 106 can be described in the above embodiment. In this embodiment of the present invention, the remote direct memory access protocol stack module can be used to run the Remote Direct Memory Access (RDMA) protocol over Ethernet to process direct memory access (DMA) requests from the network access accelerator card 100's extended memory.

[0073] In an embodiment of the present invention, in order to achieve many-to-many interconnection of accelerator cards across servers, the super-node forwarding module is connected to another accelerator card of another server through a second connector, which may include: the super-node forwarding module is connected to an Ethernet switch through the optical module 106 of the accelerator card where it is located, and is connected to the optical module 106 of another accelerator card of another server through the Ethernet switch.

[0074] The accelerator card 100 provided by the embodiment of the present invention may further include an arbitrator provided between the port routing and forwarding module and the storage controller; the arbitrator is used to arbitrate multiple access tasks to the storage component.

[0075] In an embodiment of the present invention, the cross-board forwarding module may include an intra-node forwarding module and a super-node forwarding module; the intra-node forwarding module is connected to another accelerator card 100 in the server via a corresponding second connector and a cable; the super-node forwarding module is connected to another accelerator card 100 in another server via a corresponding second connector; the accelerator card 100 also includes a remote direct memory access protocol stack module, which is arranged between the super-node forwarding module and the arbitrator; the type of access task to the storage component may be: at least one of the server host 200's access task to the storage component, the other accelerator card 100 in the server's access task to the storage component, and the other accelerator card 100 in another server's access task to the storage component.

[0076] The accelerator card 100 provided in an embodiment of the present invention may further include a control register module connected to the processor core, the remote direct memory access protocol stack module, the arbitrator, and the storage controller respectively; the control register module is used by the processor core to manage remote direct memory access registers of the remote direct memory access protocol stack module, the arbitrator, and the storage controller.

[0077] like Figure 2As shown, the control register module and the port routing forwarding module can be mounted on the system bus, thereby shortening the communication path with the processor core.

[0078] When traditional accelerator cards communicate with resources outside the server where they are located, they need to go through the RDMA network card and the PCIe switch (PCIe Switch) for point-to-point communication. The accelerator card 100 provided in the embodiment of the present invention has an optical module 106, which can be connected to the RDMA over Converged Ethernet version 2 (RoCEv2) network or standard Ethernet to communicate with other computing nodes or storage nodes. The computing processor 101 of the accelerator card 100 does not need to go through the PCIe switch and then connect to the RDMA network card in the traditional solution. Using the RoCE protocol stack based on the fiber optic network, remote initialization access to the NVMe solid-state drive can be achieved. The NVMe solid-state drive serves as an extended memory accessible to the processor core of the accelerator card 100, greatly improving the shared memory capacity and reducing the degree of data parallel splitting and data communication volume. Based on the 400G fiber optic network and the address management server, unified memory management of multiple nodes is realized, such as reporting the local memory size of each node during the power-on self-test. The 400G fiber optic network of each accelerator card 100 also serves as a redundant backup channel for the direct connection channel between cards. When any two direct connection channels fail or are severely congested, it can provide connection communication and traffic diversion.

[0079] In summary, the accelerator card 100 provided in an embodiment of the present invention can adopt the optical module 106 and 400G optical network interconnection, support heterogeneous memory expansion and the processor core of the accelerator card 100 to mount extended memory and NMVe solid-state memory; the accelerator card 100 may include a computing processor 101, a first connector 104, a power management module, a management controller 107, a 400G optical module 106 and optical cable, PCIe gold finger and open memory interface (Open Memory Interface, OMI) memory module or DDR DIMM memory stick, etc.

[0080] If the processor core is implemented using the first controller programming, four accelerator cards 100 can be installed in a single server. Each accelerator card 100 has a gold finger interface that supports the PCIe Gen5 x16 standard and four MCIO x4 interfaces that support the PCIe Gen5 standard. The four card gold fingers are inserted into the server host 200 PCIe slots, and any two cards are connected via a high-speed MCIO x4 cable. The first MCIO x4 interface of each accelerator card 100 is connected to an NVMe solid-state drive via a converter. Each MCIO x4 interface contains four pairs of high-speed SerDes channels, each supporting a bandwidth of 50 Gb / s. This results in a unidirectional bandwidth of 200 Gb / s and a bidirectional bandwidth of 400 Gb / s for each MCIO x4 interface, which can meet the high-bandwidth communication requirements of direct connections between cards.

[0081] Figure 3 A schematic diagram of the structure of interconnection between acceleration cards provided by an embodiment of the present invention.

[0082] The following example uses a dual-core server in which four accelerator cards 100 (accelerator card 1, accelerator card 2, accelerator card 3, and accelerator card 4) are placed. The server host 200 of the dual-core server includes a central processing unit 0 and a central processing unit 1, and the accelerator cards 100 are interconnected. The connection relationship is as follows: Figure 3 As shown, the data flow of the accelerator card 100's direct communication between cards and access to local extended memory, access to extended memory on other instance cards in the node, and remote access to extended memory is described below.

[0083] The accelerator card 100 is powered on, and the server system loads the driver configuration file to complete initialization. The processor core 1 on each accelerator card 100 starts a thread to load the RC bridge driver, completing the link initialization with the extended memory EP device. The inter-card multi-channel (MC-link) physical layer is trained and connected.

[0084] The server system initiates the MAC layer initialization and transport layer self-test communication test of the MC-link link between the boards in the node according to the default configuration board interconnection topology.

[0085] Each server node reports the self-test results and storage resource size, and the unified address management server 400 starts the unified address management allocation application to allocate a super-node unified address space to each server node. Each node server initializes the unified address translation table (UATT).

[0086] After each server node obtains the unified address space, it allocates a unified address space to the storage components of each accelerator card 100 and initializes the base address mapping table (BAMT). The BAMT uses a 64-bit address, with the lower 48 bits all 0 and omitted. Table 1 shows an example of address bit mapping for bits [55:48] of four accelerator cards 100.

[0087] Table 1

[0088]

[0089] The unified address [62:56] indicates the ID of different server nodes, and

[63] distinguishes between unified addresses or local addresses.

[0090] The extended memory mounted under the accelerator card 100 can be accessed by the local central processor in the server node, the local accelerator card 100 in the node, other accelerator cards 100 in the node, and other nodes through the optical fiber network.

[0091] At the same time, the present invention proposes a multi-channel interconnection solution between the acceleration cards 100 within the node and a cross-node interconnection topology, which can not only realize the horizontal expansion and vertical expansion of the computing power cluster, but also enable direct communication between devices within the node. The direct connection channel and the network channel are redundant to each other, thereby improving the reliability of the computing system.

[0092] An embodiment of the present invention further provides an accelerated computing system, comprising a plurality of interconnected acceleration cards.

[0093] The accelerator card includes: a processor core, a storage controller, a first connector and a storage component; wherein the storage controller is arranged between the processor core and the first connector, and the first connector is also connected to the storage component; the processor core is used to access the storage component through the storage controller, establish a first mapping relationship between the address space of the processor core and the address of the storage component, and perform read and write tasks on the storage component based on the first mapping relationship; the storage controller is used to respond to read and write tasks to perform read and write operations on the storage component.

[0094] The computing system provided by the embodiment of the present invention can refer to the introduction of the above embodiment.

[0095] In the accelerated computing system provided in an embodiment of the present invention, an accelerator card includes a processor core, a storage controller, a first connector, and a storage component. The storage controller is disposed between the processor core and the first connector, and the first connector is also connected to the storage component. Based on this, the processor core of the accelerator card can mount the storage component based on the local storage controller to expand the local storage space, thereby realizing an accelerator card with both computing power and storage performance. The processor core is configured to access the storage component through the first connector via the storage controller, establish a first mapping relationship between the address space of the processor core and the address of the storage component, and execute read and write tasks on the storage component based on the first mapping relationship. The storage controller responds to the read and write tasks to execute read and write operations on the storage component. Functional modules such as memory expansion and sharing controller, RDMA protocol stack, and MC-link controller are implemented on the accelerator card, which has the functions of expanding the memory space accessible by the GPU, memory sharing between hosts and devices / between devices, and cross-node interconnection communication. The system can utilize the rich high-speed SerDes, CXL IP resources and NVMe controller IP of FPGA chips to mount DDR DIMM memory modules and NVMe SSD solid-state drives, and realize TB-level GPU direct mounting of extended memory and direct data communication between GPU instances through MC-link; thus, it can increase GPU shared memory, improve the utilization rate of single GPU computing resources, reduce data communication volume between GPUs, reduce communication bottlenecks and training and inference time of large models, and effectively reduce the data center cost of AI computing power.

[0096] An embodiment of the present invention also provides a control method for an accelerator card, which is applied to the processor core of the accelerator card and may include: accessing a storage component through a storage controller to obtain storage resource information of the storage component; establishing a first mapping relationship between the address space of the processor core and the address of the storage component based on the storage resource information; executing read and write tasks on the storage component based on the first mapping relationship; the storage controller is arranged between the processor core and the first connector of the accelerator card, and the first connector is also connected to the storage component.

[0097] The control method of the accelerator card provided in the embodiment of the present invention can refer to the introduction of the above embodiment.

[0098] An embodiment of the present invention further provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps of any of the above-mentioned embodiments of the control method for an accelerator card.

[0099] An embodiment of the present invention further provides a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps of any of the above-mentioned embodiments of the control method for an accelerator card when running.

[0100] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0101] An embodiment of the present invention further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of any of the above-mentioned embodiments of the control method for an accelerator card are implemented.

[0102] An embodiment of the present invention further provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any of the above-mentioned accelerator card control method embodiments are implemented.

[0103] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.

[0104] The above is a detailed introduction to an accelerator card, a control method thereof, and an accelerated computing system provided by the present invention. This article uses specific examples to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only intended to help understand the method and core concept of the present invention. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the scope of protection of the present invention.

Claims

1. An accelerator card, characterized in that: The accelerator card is a graphics processor, comprising: a processor core, a storage controller, a first connector, a storage component, a port routing forwarding module, an inter-board forwarding module, and a second connector; Wherein, the storage controller is provided between the processor core and the first connector, and the first connector is also connected to the storage component; The processor core is configured to access the storage component via the first connector through the storage controller, establish a first mapping relationship between the address space of the processor core and the address of the storage component, and perform read and write tasks on the storage component based on the first mapping relationship; The storage controller is used to respond to the read and write tasks to perform read and write operations on the storage component; The port routing forwarding module is provided between the processor core and the storage controller, and the port routing forwarding module is also connected to the cross-board forwarding module; The inter-board forwarding module is connected to another accelerator card via the second connector, and is used to forward memory access data packets between the accelerator card where the module is located and the other accelerator card; The port routing and forwarding module is used to perform data packet format conversion between the accelerator card where it is located and another accelerator card; The cross-board forwarding module includes at least an intra-node forwarding module; the intra-node forwarding module is connected to another accelerator card of the server via the second connector and a cable; The processor core performs a read and write task on the storage component based on the first mapping relationship, including: The processor core reports the storage resource information of the acceleration card to the unified address management server, and receives the super node unified address space allocated by the unified address management server to the acceleration card, initializes the second mapping relationship between the super node unified address space and the address space of the processor core, and performs read and write tasks on the storage component based on the second mapping relationship and the first mapping relationship.

2. The accelerator card according to claim 1, wherein: At least one thread of at least one of the processor cores runs a root complex bridge driver to configure the storage component as a bus endpoint device.

3. The accelerator card according to claim 1, wherein: The first connector is an inter-card interconnection connector of the acceleration card, and the storage controller is an inter-board forwarding control module corresponding to the first connector.

4. The accelerator card according to claim 1, wherein: The first connector is a first pin in the first interface of the accelerator card; The first interface is an interface used by the accelerator card to connect to the server host; The storage controller is connected to the storage component via the first pin, on-board wiring of the server host, and a first slot of the server host, where the first slot is used to install the storage component.

5. The accelerator card according to claim 1, wherein: The processor core is obtained by programming the logic circuit of the first controller to implement an acceleration card.

6. The accelerator card according to claim 1, wherein: The storage controller is a memory access controller, and the storage component is a memory particle.

7. The accelerator card according to claim 1, wherein: The storage controller includes a memory access controller and a storage access controller; the storage component includes a memory particle and a first non-volatile storage device; The memory access controller is also connected to the memory chip via a corresponding first connector; The first end of the storage access controller is connected to the memory access controller; The second end of the storage access controller is connected to the first non-volatile storage device through the corresponding first connector.

8. The accelerator card according to claim 6 or 7, characterized in that: The memory particles are shared memories of multiple processor cores; The storage controller also includes a shared memory controller arranged between the processor core and the memory access controller, and the shared memory controller is used to execute a first data movement task between the server host and the acceleration card and / or a second data movement task between different processor cores in the acceleration card.

9. The accelerator card according to claim 1, wherein: The cross-board forwarding module also includes a super node forwarding module; The super node forwarding module is connected to another acceleration card of another server through the second connector.

10. The accelerator card according to claim 9, wherein: The second connector corresponding to the supernode forwarding module is an optical module; The acceleration card further includes a remote direct memory access protocol stack module arranged between the super node forwarding module and the storage controller.

11. The accelerator card according to claim 10, wherein: The super node forwarding module is connected to another accelerator card of another server through the second connector, including: The supernode forwarding module is connected to an Ethernet switch via the optical module of the acceleration card where it is located, and is connected to the optical module of another acceleration card of another server via the Ethernet switch.

12. The accelerator card according to claim 1, wherein: It also includes an arbitrator provided between the port routing forwarding module and the storage controller; The arbitrator is used to arbitrate multiple access tasks to the storage component.

13. The accelerator card according to claim 12, wherein: The cross-board forwarding module includes an intra-node forwarding module and a super-node forwarding module; The intra-node forwarding module is connected to another accelerator card in the server via the corresponding second connector and cable; The super node forwarding module is connected to another accelerator card of another server through the corresponding second connector; The accelerator card further includes a remote direct memory access protocol stack module, which is provided between the supernode forwarding module and the arbitrator; The type of the access task to the storage component is at least one of: an access task to the storage component by a server host, an access task to the storage component by another accelerator card of the server, and an access task to the storage component by another accelerator card of another server.

14. The accelerator card according to claim 13, wherein: It also includes a control register module connected to the processor core, the remote direct memory access protocol stack module, the arbitrator and the storage controller respectively; The control register module is used by the processor core to manage remote direct memory access registers of the remote direct memory access protocol stack module, the arbitrator, and the storage controller.

15. The accelerator card according to claim 1, wherein: Also includes a management controller; The management controller is connected to the processor core and is used to monitor the component status of the accelerator card.

16. An accelerated computing system, characterized in that: It includes a plurality of interconnected accelerator cards; the accelerator card is a graphics processor; The accelerator card includes: a processor core, a storage controller, a first connector, a storage component, a port routing forwarding module, an inter-board forwarding module, and a second connector; Wherein, the storage controller is provided between the processor core and the first connector, and the first connector is also connected to the storage component; The processor core is configured to access the storage component through the storage controller, establish a first mapping relationship between the address space of the processor core and the address of the storage component, and perform read and write tasks on the storage component based on the first mapping relationship; The storage controller is used to respond to the read and write tasks to perform read and write operations on the storage component; The port routing forwarding module is provided between the processor core and the storage controller, and the port routing forwarding module is also connected to the cross-board forwarding module; The inter-board forwarding module is connected to another accelerator card via the second connector, and is used to forward memory access data packets between the accelerator card where the module is located and the other accelerator card; The port routing and forwarding module is used to perform data packet format conversion between the accelerator card where it is located and another accelerator card; The cross-board forwarding module includes at least an intra-node forwarding module; the intra-node forwarding module is connected to another accelerator card of the server via the second connector and a cable; The processor core performs a read and write task on the storage component based on the first mapping relationship, including: The processor core reports the storage resource information of the acceleration card to the unified address management server, and receives the super node unified address space allocated by the unified address management server to the acceleration card, initializes the second mapping relationship between the super node unified address space and the address space of the processor core, and performs read and write tasks on the storage component based on the second mapping relationship and the first mapping relationship.

17. A control method for an accelerator card, characterized in that: Processor cores used in accelerator cards include: Accessing the storage component through the storage controller to obtain storage resource information of the storage component; Establishing a first mapping relationship between the address space of the processor core and the address of the storage component according to the storage resource information; Executing a read and write task on the storage component based on the first mapping relationship includes: the processor core reporting storage resource information of the accelerator card to a unified address management server, receiving a supernode unified address space allocated by the unified address management server to the accelerator card, initializing a second mapping relationship between the supernode unified address space and the address space of the processor core, and executing the read and write task on the storage component based on the second mapping relationship and the first mapping relationship; Among them, the accelerator card is a graphics processor; The storage controller is provided between the processor core and the first connector of the accelerator card, and the first connector is also connected to the storage component; The acceleration card further includes a port routing forwarding module, an inter-board forwarding module and a second connector; The port routing forwarding module is provided between the processor core and the storage controller, and the port routing forwarding module is also connected to the cross-board forwarding module; The inter-board forwarding module is connected to another accelerator card via the second connector, and is used to forward memory access data packets between the accelerator card where the module is located and the other accelerator card; The port routing and forwarding module is used to perform data packet format conversion between the accelerator card where it is located and another accelerator card; The cross-board forwarding module includes at least an intra-node forwarding module; the intra-node forwarding module is connected to another acceleration card of the server via the second connector and a cable.

Citation Information

Patent Citations

  • Data processing system and method and computer system

    CN119046211A

  • Heterogeneous computing system, server and data center

    CN223092421U