Computing device, server, data processing method, and storage medium

Through CXL technology, the direct access and memory sharing between GPU and main memory is realized, which solves the problem of traditional GPU memory bottleneck and improves the memory expansion and performance of AI computing devices.

WO2025138849A1PCT designated stage expired Publication Date: 2025-07-03INSPUR SUZHOU INTELLIGENT TECH CO LTD

Patent Information

Application Number
PCT/CN2024/111106
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-29
Filing Date
2024-08-09
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Traditional central processing units (CPUs) and graphics processing units (GPUs) face memory bottlenecks when handling complex AI tasks and cannot meet the computing and storage needs of large-scale and complex AI workloads.

Method used

Compute Express Link (CXL) high-speed interconnection technology is adopted to realize direct access between GPU and main memory through PCIe switch, share memory resources, form a unified memory architecture, and expand the memory capacity of the GPU.

Benefits of technology

It improves the memory access efficiency of the GPU, reduces data transmission delay and replication overhead, saves hardware costs, and optimizes the performance of high-performance computing applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024111106_03072025_PF_FP_ABST
    Figure CN2024111106_03072025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of computers, and provides a computing device, a server, a data processing method, and a storage medium. The computing device comprises a central processing unit (CPU), an accelerator, and a first high-speed serial computer extension bus standard PCIe switch; the accelerator is connected to a first downlink port of the first PCIe switch; an uplink port of the first PCIe switch is connected to the CPU; each port of the first PCIe switch supports a compute express link (CXL) protocol; the accelerator is configured to perform an access operation on a main memory on the basis of the CXL protocol.
Need to check novelty before this filing date? Find Prior Art

Description

Computing device, server, data processing method and storage medium

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to the Chinese patent application filed with the China Patent Office on December 29, 2023, with application number 202311850460.9, and application name “Computing device, server, data processing method and storage medium”, all contents of which are incorporated by reference into this application. Technical Field

[0003] The present application relates to a computing device, a server, a data processing method and a storage medium. Background Art

[0004] In recent years, the widespread application of artificial intelligence (AI) has brought many technical challenges, especially in meeting the demands of machine learning and deep learning. Traditional central processing units (CPUs) and graphics processing units (GPUs) can face performance bottlenecks when handling complex AI tasks. These bottlenecks aren't caused by computing power, but by the memory requirements of accelerators like GPUs.

[0005] Therefore, how to provide an architecture or method for expanding accelerator memory becomes a technical problem that needs to be solved urgently.

[0006] Summary of the Invention

[0007] In a first aspect, according to an embodiment of the present application, there is provided a computing device comprising: a central processing unit (CPU), an accelerator, and a first high-speed serial computer expansion bus standard PCIe switch;

[0008] The accelerator is connected to the first downstream port of the first PCIe switch;

[0009] The uplink port of the first PCIe switch is connected to the CPU;

[0010] Each port of the first PCIe switch supports the Compute Express Link (CXL) protocol;

[0011] The accelerator is configured to perform access operations on the main memory based on the CXL protocol.

[0012] In a second aspect, according to an embodiment of the present application, a server is further provided, comprising: a computing device as described in any one of the first aspects above.

[0013] In a third aspect, according to an embodiment of the present application, there is further provided a data processing method, based on any computing device according to the first aspect, comprising:

[0014] The accelerator sends a first data request message to the main memory based on the CXL protocol;

[0015] The main memory sends the first task data to the accelerator through the first PCIe switch in response to the first data request message.

[0016] In a fourth aspect, according to an embodiment of the present application, a non-transitory computer-readable storage medium is also provided, on which computer-readable instructions are stored. When the computer-readable instructions are executed by a processor, any data processing method as described in the third aspect above is implemented.

[0017] The details of one or more embodiments of the present application are set forth in the accompanying drawings and the description below. Other features and advantages of the present application will become apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in this application or related technologies, the following is a brief introduction to the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0019] FIG1 is a schematic diagram of the structure of a computing device provided in an embodiment of the present application;

[0020] FIG2 is a schematic diagram of a basic architecture of a CPU and a GPU using a PCIe / CXL interface to enable the GPU to use the main memory according to an embodiment of the present application;

[0021] FIG3 is a schematic diagram of a basic architecture for implementing integrated memory between a CPU and an FPGA according to an embodiment of the present application;

[0022] FIG4 is a schematic diagram of implementing GPU access to integrated memory using a Switch with CXL functionality according to an embodiment of the present application;

[0023] FIG5 is a schematic diagram of a direct memory access architecture provided in an embodiment of the present application;

[0024] FIG6 is an expanded schematic diagram of a direct memory access architecture provided in an embodiment of the present application;

[0025] FIG7 is a schematic diagram of a memory pool expansion architecture according to an embodiment of the present application;

[0026] FIG8 is a second schematic diagram of a memory pool expansion architecture provided in an embodiment of the present application;

[0027] FIG9 is a schematic diagram showing the connection of eight GPUs, GPU0-GPU7, provided in an embodiment of the present application;

[0028] FIG10 is a schematic diagram of the structure of a server provided in an embodiment of the present application;

[0029] FIG11 is a flow chart of a data processing method according to an embodiment of the present application. DETAILED DESCRIPTION

[0030] To make the objectives, technical solutions, and advantages of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments of this application, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this application.

[0031] The terms "first", "second", etc. in this application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the terms used in this way are interchangeable where appropriate, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first" and "second" are generally of the same type, and do not limit the number of objects, for example, the first object can be one or more. In addition, "or" in this application represents at least one of the connected objects. For example, "A or B" covers three options, namely, Option 1: including A but not including B; Option 2: including B but not including A; Option 3: including both A and B. The character " / " generally indicates that the objects associated before and after are in an "or" relationship.

[0032] The term "indication" in this application can be either a direct indication (or explicit indication) or an indirect indication (or implicit indication). A direct indication can be understood as the sender explicitly informing the receiver of specific information, the operation to be performed, or the requested result, etc. in the instruction sent; an indirect indication can be understood as the receiver determining the corresponding information based on the instruction sent by the sender, or making a judgment and determining the operation to be performed or the requested result, etc. based on the judgment result.

[0033] In recent years, the widespread application of artificial intelligence (AI) has brought many technical challenges, especially in meeting the needs of machine learning and deep learning. Traditional central processing units (CPUs) and graphics processing units (GPUs) can face performance bottlenecks when handling complex AI tasks, as these tasks typically require extensive computing resources and memory capacity.

[0034] For AI applications, existing hardware can generally handle some tasks, but it may face performance and memory limitations when handling large-scale and complex AI workloads. The increasing size and complexity of AI models leads to greater computing and storage requirements, which places higher demands on hardware. In addition to computing efficiency challenges, AI training also faces issues of memory capacity and bandwidth. Deep learning models typically have a large number of parameters, requiring a large amount of memory capacity to store and process data. Furthermore, due to the massively parallel computing requirements of deep learning models, high memory bandwidth becomes crucial.

[0035] To address these challenges, many hardware manufacturers are continuously launching new solutions, including GPUs with higher memory capacity and dedicated AI accelerators. In addition, high-bandwidth memory technologies such as High Bandwidth Memory (HBM) and Graphics Double Data Rate (GDDR) are also widely adopted.

[0036] The amount of computation required for AI training is increasing significantly every year. The future bottleneck for AI training will not be limited by computing power, but by GPU memory. Therefore, developing an architecture or method to expand GPU memory has become a pressing technical challenge.

[0037] The computing device, server, data processing method and storage medium of the present application are described below in conjunction with Figures 1 to 11.

[0038] The computing device provided in this application implements a GPU extended memory architecture to solve the problem of GPU encountering a memory wall.

[0039] This application uses CXL (Compute Express Link) high-speed interconnect technology to provide a scale-out memory architecture. Scale-out memory is a method of expanding the available memory capacity in a computing system. When a computing system needs to process large amounts of data or perform memory-intensive tasks, the memory capacity of a single node may become insufficient. Scale-out memory aims to expand the available memory capacity by connecting multiple computing nodes (usually multiple computing devices or servers) together to form a large cluster that shares each other's memory resources.

[0040] The following is an introduction to CXL technology. CXL (Compute Express Link) is a high-speed interconnect technology designed to address the challenges of connecting memory and accelerators in data centers and computing systems. CXL is an open standard promoted by a consortium of computer hardware manufacturers. CXL technology was originally developed to enable shared memory between CPUs and Acceleration Function Units (AFUs), thereby enabling memory interconnection between processors (such as CPUs, ASICs, and FPGAs).

[0041] CXL 3.0 features include peer-to-peer messaging between peripheral devices, providing a direct memory access (DMA) architecture between peripheral devices without going through the CPU. This architecture, combined with enhanced hardware consistency mechanisms, allows peripheral memory areas to be shared with multiple host CPUs simultaneously. CXL 3.0 is the latest version of CXL technology, improving and expanding upon previous versions. Key features of CXL 3.0 include:

[0042] High bandwidth and low latency: CXL 3.0 provides high-bandwidth and low-latency data transfer capabilities, which can transfer data from memory to accelerators or other processing units more quickly, thereby improving system performance.

[0043] Memory Expansion: CXL 3.0 supports memory expansion, allowing multiple devices to share physical memory, thereby expanding the available memory capacity. This is very useful for processing large datasets and memory-intensive tasks.

[0044] Computational acceleration: CXL 3.0 supports efficient connections to computational accelerators (such as GPUs and FPGAs), enabling them to better collaborate with the main processor and memory to accelerate computing tasks.

[0045] Compatibility: CXL 3.0 is compatible with PCI Express (PCIe) and Memory CXL interconnect standards, allowing for smooth upgrades and migrations of existing hardware and software.

[0046] CXL defines three types of peripheral device applications:

[0047] Type 1: Operates through the CXL.io and CXL.cache protocols and is suitable for specialized accelerators that don't have dedicated memory. For example, some smart network cards or video accelerators can now access the host CPU's memory through the CXL protocol, sharing this memory with these peripherals for use as a cache. The basic concept of CXL Type 1 devices is similar to the architecture of a CPU with built-in graphics devices that can share the system's main memory. In both cases, peripherals can use main memory, eliminating the need for dedicated memory. CXL allows any PCIe peripheral device that supports the CXL.io and CXL.cache protocols to use the system's main memory.

[0048] Type 2: Operates through the CXL.io, CXL.cache, and CXL.mem protocols and is designed for general-purpose accelerators with built-in high-performance memory (GDDR or HBM memory), such as GPU cards, or FPGA-based or ASIC-based accelerator cards. Using the CXL protocol, it provides bidirectional memory sharing between these peripheral devices and the host CPU, allowing both peripheral devices to access the host CPU's memory and vice versa. By sharing the memory of the host CPU and peripheral devices, CXL Type 2 applications dynamically allocate memory resources between them, thereby improving overall system memory utilization.

[0049] Type 3: Operates via the CXL.io and CXL.mem protocols. This type of device is a memory expansion card based on dynamic random access memory (DRAM) or storage class memory (SCM). The host CPU can access the DRAM or non-volatile SCM memory on this memory expansion card through the CXL protocol.

[0050] FIG1 is a schematic diagram of the structure of a computing device provided in an embodiment of the present application. As shown in FIG1 , the computing device includes: a central processing unit (CPU) 110, an accelerator 120, and a first high-speed serial computer expansion bus standard (Peripheral Component Interconnect Express, PCIe) switch 130;

[0051] The accelerator 120 is connected to the first downstream port of the first PCIe switch 130;

[0052] The uplink port of the first PCIe switch 130 is connected to the CPU 110;

[0053] Each port of the first PCIe switch 130 operates in a computing express link CXL mode or a PCIe mode;

[0054] The accelerator is configured to perform access operations on the main memory based on the CXL protocol.

[0055] Here, the CXL protocol specifically refers to the CXL.io and CXL.cache protocols.

[0056] Optionally, the accelerator may be a GPU or other heterogeneous acceleration device, such as a Field Programmable Gate Array (FPGA) or an Application Specific Integrated Circuit (ASIC).

[0057] It is understood that the accelerator is connected to the CPU via the first PCIe switch, so that the accelerator can directly access the main memory of the system based on the CXL protocol. The access operation refers to writing data or reading data.

[0058] Each port of the first PCIe switch supports the Compute Express Link (CXL) protocol and operates in a Compute Express Link (CXL) mode or a PCIe mode, that is, each port of the first PCIe switch is a PCIe / CXL interface.

[0059] The following explanation uses the GPU as an example. Figure 2 shows a basic architecture diagram of the CPU and GPU using the PCIe / CXL interface to enable GPU access to main memory, as provided in an embodiment of the present application. As shown in Figure 2, this is a CXL Type 1 application. The GPU can access main memory through the CXL Type 1 application, which provides several advantages:

[0060] Shared resources: The GPU and host CPU share the same memory space and can directly access data in the main memory, avoiding the additional copying of data from the main memory to the GPU dedicated memory.

[0061] Cost savings: GPUs can directly expand their memory configuration through CXL.

[0062] Cache acceleration: Using main memory as cache memory can speed up GPU data access and improve performance.

[0063] The introduction of CXL technology will further promote the collaboration between GPU and host CPU, provide more efficient memory access and sharing, and thus optimize high-performance computing applications.

[0064] In an embodiment of the present application, the accelerator is connected to the CPU through a first PCIe switch, so that the accelerator can access the main memory based on the CXL protocol, thereby allowing the accelerator to directly operate on the main memory, expanding the memory available to the accelerator, providing more efficient memory access and sharing, and accelerating data access and improving performance.

[0065] In some embodiments, the computing device further comprises: a memory expansion unit,

[0066] The memory expansion unit is connected to the second downstream port of the first PCIe switch;

[0067] The local memory of the memory extension unit and the main memory constitute an integrated memory;

[0068] The accelerator is configured to access the unified memory based on the CXL protocol.

[0069] It can be understood that the memory expansion unit in this embodiment is connected to the second downstream port of the first PCIe switch, and the first downstream port of the first PCIe switch is connected to the accelerator, and the upstream port of the first PCIe switch is connected to the CPU, so point-to-point communication can be achieved between the accelerator and the memory expansion unit, the memory expansion unit can access the main memory, and accordingly, the CPU can also access the local memory of the memory expansion unit.

[0070] The local memory of the memory extension unit and the main memory form an integrated memory, and the accelerator can access the integrated memory based on the CXL protocol.

[0071] The following is an introduction to Converged Memory.

[0072] Converged Memory is a computing concept in which different types of memory technologies are consolidated or combined into a single memory pool or architecture. This approach aims to address the limitations and challenges of the traditional memory hierarchy, in which different memory types, such as DRAM, SRAM (Static Random Access Memory), and NAND (Not AND) flash memory, are used for specific purposes such as main memory, cache, and storage.

[0073] The idea behind unified memory is to create a unified memory system that can deliver better performance, energy efficiency, and simplified memory management. By combining multiple memory technologies, data can be shared and moved more efficiently between different levels of the memory hierarchy, reducing the need for data transfers between different memory types and potentially reducing latency.

[0074] Optionally, the memory extension unit includes at least one first processing unit having an independent memory.

[0075] Optionally, the first processing unit includes any one or combination of a field programmable gate array (FPGA), a complex programmable logic device (CPLD), a programmable logic device (PLD), an application specific integrated circuit (ASIC), a generic array logic (GAL), a system on chip (SOC), a software defined infrastructure (SDI) device and an artificial intelligence (AI) device.

[0076] The following explanation is given by taking the FPGA as the first processing unit. Figure 3 is a schematic diagram of the basic architecture for implementing integrated memory between the CPU and FPGA provided in an embodiment of the present application. As shown in Figure 3, it is a typical CXL Type 2 application, and bidirectional memory sharing between the FPGA and the host CPU. This allows devices such as FPGA to have high-performance independent memory (such as GDDR or HBM memory) and share memory with the CPU to achieve more efficient data transmission and processing. Through the CXL communication architecture, the CPU and FPGA can achieve tighter integration and collaboration, forming a integrated memory (Converged Memory) from the CPU's Host Memory and the FPGA's Optional Memory to improve the overall performance and flexibility of the system. Technology for integrating and sharing the memory resources of FPGA and CPU. Traditionally, FPGA and CPU have independent memory spaces, and data must be explicitly copied or transferred between them. Converged Memory technology enables the sharing of memory resources at the hardware and software levels, allowing the FPGA and CPU to access the same physical memory, thereby achieving more efficient data transmission and sharing.

[0077] Converged Memory technology brings the following benefits:

[0078] Data Sharing: The FPGA and CPU can directly share the same memory space without copying data. This allows data to be transferred more efficiently between the FPGA and CPU, reducing data transmission latency and copying overhead.

[0079] Flexibility: FPGA can access the CPU's memory space, allowing it to directly process data stored by the CPU. This provides greater flexibility and possibilities for collaborative work between FPGA and CPU.

[0080] Memory management: With Converged Memory, memory management can be more unified, simplifying the transfer and management of data between the FPGA and CPU, and improving the overall performance of the system.

[0081] Saving resources: Since the same memory resources are shared, the physical memory capacity required in the system can be reduced, saving hardware costs.

[0082] This application utilizes a switch (or switch) with CXL functionality to simultaneously implement the integrated memory function of the CPU and FPGA and the GPU's direct access to the same shared memory space between the CPU and FPGA. Figure 4 is a schematic diagram of an embodiment of this application using a switch with CXL functionality to implement GPU access to integrated memory. As shown in Figure 4, the upstream port of the switch is connected to the CPU root complex, and the downstream port is connected to the GPU and FPGA. The GPU can use the switch to allocate the FPGA's memory to make up for the GPU's own insufficient local memory. In addition, the physical layer of the CXL3.0 protocol uses the PCIe 6.0 interface. PCIe 6.0 is the latest version of the PCIe (Peripheral Component Interconnect Express) bus, and its transmission speed is 16GT / s (Gigabits per second). In other words, the speed of PCIe 6.0 is 16 Gigabits per second, which has a faster data transmission rate than previous versions (such as PCIe 5.0 and PCIe 4.0). It allows high-speed memory access and sharing between the host CPU and peripheral devices through the CXL protocol.

[0083] The features of the architecture in Figure 4 can be used extensively to implement the function of extending the memory pool.

[0084] The computing device provided in an embodiment of the present application also includes a memory expansion unit, which is connected to the second downstream port of the first PCIe switch; the local memory of the memory expansion unit and the main memory constitute an integrated memory, so that the accelerator can access the memory space shared by the CPU and the memory expansion unit, reducing the delay in data transmission and the overhead of copying. Since the CPU and the memory expansion unit share the same memory resources, the physical memory capacity required in the system can be reduced, saving hardware costs, memory management can be more unified, and the overall performance of the system is improved.

[0085] In some embodiments, the memory expansion unit includes a memory expansion board.

[0086] Specifically, the memory expansion board is connected to the second downstream port of the first PCIe switch;

[0087] The local memory of the memory expansion board and the main memory form an integrated memory;

[0088] The accelerator is configured to access the unified memory based on the CXL protocol. The CXL protocol includes the CXL.io and CXL.mem protocols. Through the CXL.io and CXL.mem protocols, the local memory of the memory expansion board and the main memory form a unified memory, allowing the accelerator to access the local memory on the memory expansion board.

[0089] Optionally, the memory expansion board is equipped with dynamic random access memory DRAM or storage class memory SCM.

[0090] In the computing device provided in the embodiment of the present application, the memory expansion unit can be a memory expansion board, and the accelerator can access the integrated memory composed of the local memory of the memory expansion board and the main memory, thereby expanding the memory available to the accelerator, providing more efficient memory access and sharing, and accelerating data access and improving performance.

[0091] In some embodiments, the computing device further includes: a Non-Volatile Memory Express (NVMe) solid state drive (SSD), the NVMe solid state drive being connected to the third downstream port of the first PCIe switch;

[0092] The accelerator is configured to access the NVMe solid-state drive based on the CXL protocol.

[0093] Specifically, the PCIe Switch supports point-to-point (P2P) communication between NVMe SSDs and GPUs, allowing data to be transferred directly between the NVMe SSD and GPU without involving the host CPU. This direct communication path can significantly reduce data transfer latency and CPU overhead, thereby improving overall system performance under certain workloads. In traditional PCIe configurations, data transfer between NVMe SSDs and GPUs requires sending data from the NVMe SSD to the host memory and then from the host memory to the GPU's memory. This process involves multiple jumps, adding additional latency. However, through the P2P communication of the PCIe Switch, the NVMe SSD and GPU can exchange data directly through the PCIe Switch without involving the host memory or CPU. This P2P communication is particularly suitable for tasks that require frequent data exchange between the NVMe SSD and GPU, such as data-intensive workloads such as artificial intelligence, machine learning, and high-performance computing. By enabling direct communication, the P2P function of the PCIe Switch can enhance the overall efficiency and performance of data-intensive applications and reduce data movement bottlenecks.

[0094] In some embodiments, the computing device further includes: a network interface controller (NIC), the NIC being connected to the fourth downstream port of the first PCIe switch; and the accelerator being configured to perform data exchange with the NIC based on a CXL protocol.

[0095] Specifically, the PCIe Switch hardware itself has peer-to-peer communication capabilities and supports the shortest path transmission between peers. The accelerator can directly transmit data to the Ethernet through the Switch P2P function and the NIC. The shortest path transmission only passes through the PCIe Switch and does not cause any CPU burden, reducing the waste of system resources.

[0096] Figure 5 is a schematic diagram of a direct memory access architecture provided in an embodiment of the present application. As shown in Figure 5, the architecture includes: a motherboard (MB) and AI Compliance.

[0097] Among them, the motherboard MB includes a CPU, 4 groups of PCIe switches (only 2 groups are shown in Figure 5), 8 groups of FPGAs (only 4 groups are shown in Figure 5), and 2 groups of MCIO (Mini Cool Edge IO) connectors (only one group is shown in Figure 5).

[0098] The CPU must be an X86 platform and support the CXL function.

[0099] 4 sets of PCIe Switches, including PCIe Switch1, Switch2, Switch3, and Switch4.

[0100] The PCIe switch downstream ports must support operation in either CXL or PCIe mode. They are suitable for data-intensive workloads such as AI. This model is Broadcom's Atlas 3 series PCIe Switch.

[0101] Eight FPGAs, shown in Figure 5, are connected to the downstream ports of the PCIe switch, enabling high-speed, low-latency, and highly efficient data transmission and communication between the CPU and the FPGAs. Each FPGA has multiple independent x16 lane PCIe endpoints. Each independent PCIe endpoint can be considered a PCIe device. The computing device has a built-in 8-channel DMA (Direct Memory Access) controller that supports DDR5 and LPDDR5 interfaces, as well as CXL.

[0102] Two sets of MCIO x16 connectors for memory expansion (scale-up).

[0103] In Figure 5, UP stands for UPstream, DP stands for Downstream, F stands for Fabric port, and EP stands for Endpoint. The root complex is a key component of the PCI Express (PCIe) bus architecture. It is a logical node used to manage the entire PCIe system. The root complex is typically implemented by a CPU, a computing device group, or an FPGA. In a PCIe system, each device must be connected to a PCIe bus. The root complex, as the starting point and central node of the PCIe bus architecture, is responsible for managing all PCIe devices and endpoints, including allocating and managing bus bandwidth, controlling transmission, and routing data.

[0104] Each set of PCIe Switch downstream ports is equipped with dual GPUs and dual NICs. This configuration is used in the training phase of machine learning, where the model needs to use a large-scale training data set for optimization and learning. The training phase has high requirements for network bandwidth speed because a large amount of training data needs to be transferred from storage devices (such as cloud storage or local storage) to the training server or device. The amount of data transferred during training is usually very large, especially in distributed training, where data needs to be exchanged frequently between multiple devices. Therefore, high-speed, low-latency network bandwidth is very important and can significantly affect training speed and efficiency. In large-scale machine learning training, high-bandwidth network connections and specialized network architectures are usually used to support large-scale data transmission.

[0105] The machine learning training phase requires high network bandwidth speeds. When designing and deploying a machine learning system, it is necessary to consider network bandwidth requirements based on the specific application scenario and data scale to ensure system stability and efficiency.

[0106] The PCIe switch hardware in Figure 5 inherently supports peer-to-peer communication, supporting the shortest path transmission between peers. For example, in Figure 5, the GPU can use the switch's P2P functionality to directly transmit data to the Ethernet network with the NIC. This shortest path transmission only passes through the PCIe switch and does not impose any CPU burden, thus reducing system resource waste. In Figure 5, in Figure 5, the PCIe switch also supports peer-to-peer (P2P) communication between the NVMe SSD and the GPU, allowing data to be transferred directly between the NVMe solid-state drive (NVMe SSD) and the graphics processing unit (GPU) without involving the host CPU. This direct communication path can significantly reduce data transfer latency and CPU overhead, thereby improving overall system performance under certain workloads. In traditional PCIe configurations, data transfer between an NVMe SSD and a GPU requires first sending the data from the NVMe SSD to the host memory and then from the host memory to the GPU memory. This process involves multiple hops and adds additional latency. However, through the PCIe switch's P2P communication, the NVMe SSD and GPU can exchange data directly through the PCIe switch without involving the host memory or CPU. This P2P communication is particularly suitable for tasks that require frequent data exchange between NVMe SSDs and GPUs, such as data-intensive workloads such as artificial intelligence, machine learning, and high-performance computing. By enabling direct communication, the PCIe Switch's P2P functionality can enhance the overall efficiency and performance of data-intensive applications and reduce data movement bottlenecks.

[0107] Figure 5 (a) illustrates a CXL Type 1 application, which enables the GPU to directly access the system's main memory. Figure 5 (b) illustrates a configuration that allows the GPU and FPGA to directly communicate peer-to-peer (P2P) with each other, allowing them to access each other's memory without the need for the host CPU. Specifically, when the GPU needs to access the FPGA's memory, it can directly transfer data to the FPGA's memory location via the PCIe switch. Similarly, when the FPGA needs to access the GPU's memory, it can directly read data from the GPU's memory via the PCIe switch. This direct communication reduces data transmission latency and CPU involvement, improving memory access efficiency and performance. This type of P2P memory access between the GPU and FPGA is often suitable for applications requiring efficient data exchange and sharing, such as those in artificial intelligence, machine learning, and high-performance computing. Enabling memory access between the GPU and FPGA via the PCIe switch provides a high-bandwidth, low-latency communication channel, further enhancing the system's computing and data processing capabilities.

[0108] Furthermore, as shown in Figure 5, a converged memory architecture can be implemented between the FPGA and main memory through the application of CXL Type 2, enabling bidirectional memory sharing between the FPGA acceleration device and the host CPU. This allows devices like the FPGA, which have high-performance independent memory (such as GDDR or HBM memory), to share memory with the CPU for more efficient data transfer and processing. This allows the CPU's main memory (host memory) and the FPGA's optional memory (optional memory) to form a converged memory, creating a massive memory pool.

[0109] In some embodiments, the computing device further comprises: a first MCIO connector,

[0110] The first MCIO connector is connected to the Fabric port of the first PCIe switch and is used to be connected to the second MCIO connector of another computing device.

[0111] The main function of the Fabric port is to support transmission between PCIe switches, and it has I / O sharing functions and DMA with non-blocking and linear acceleration features.

[0112] Figure 6 is an expanded schematic diagram of the direct memory access architecture provided in an embodiment of the present application. As shown in Figure 6, ⑤ is the process of dynamically adjusting and balancing workloads between multiple GPU devices by implementing point-to-point transmission between the GPUs of the two systems through the characteristics of Fabric Link. When performing parallel computing, the workload may be distributed to multiple GPUs to accelerate processing, but the performance and resources between different GPUs may vary, so dynamic adjustment is required to ensure optimal performance and efficiency. As shown in Figure 6, point-to-point communication between the FPGA of System 1 and the FPGA of System 2 can also be achieved through Fabric Link and CXL applications, and the memory pools of the two systems can be integrated. Such a design can achieve efficient data exchange and sharing, thereby improving the overall performance and efficiency of the system. This means that the FPGAs of the two systems can directly access each other's memory, achieve sharing and share memory resources, and thus more effectively manage and utilize memory.

[0113] This application also provides another architecture for extending memory.

[0114] In some embodiments, the memory expansion unit includes a memory pool including at least one second processing unit having independent memory and at least one second PCIe switch;

[0115] The downlink port of the second PCIe switch is connected to the first Endpoint port of the second processing unit, and the second Endpoint port of the second processing unit is connected to the fifth downlink port of the first PCIe switch.

[0116] FIG7 is a schematic diagram of a memory pool expansion architecture provided in an embodiment of the present application. The memory pool expansion architecture includes:

[0117] 1. MB (Mother Board)

[0118] The CPU must be an X86 platform and support the CXL function.

[0119] 2. AI Compliance, including:

[0120] 4 sets of PCIe Switch, PCIe Switch1, Switch2, Switch3, Switch4);

[0121] Each set of PCIe Switch downstream ports is equipped with dual GPUs and dual NICs;

[0122] Each PCIe Switch is paired with four NVMe SSDs, and the entire system has a total of 16 NVMe SSD storage units.

[0123] 3. Memory Pool, including:

[0124] 4 groups of FPGAs (such as FPGA0, FPGA1, FPGA2, and FPGA3 connected to the downstream ports of the PCIe Switch);

[0125] 2 sets of PCIe Switch (such as Switch5, Switch6);

[0126] Two sets of MCIO x16 connectors for FPGA scale-up.

[0127] As shown in Figure 7, the FPGA in this architecture has three sets of PCIe Endpoints. This architecture uses the downstream ports of PCIe Switch0, PCIe Switch1, PCIe Switch2, and PCIe Switch3 to connect to one set of Endpoint ports on FPGA0, FPGA1, FPGA2, and FPGA3, respectively. The downstream ports of PCIe Switch5 and PCIe Switch6 are used to connect to two sets of Endpoints on FPGA0, FPGA1, FPGA2, and FPGA3, respectively. PCIe Switch5 and PCIe Switch6 each have a Fabric Port connected to the MCIO connector, allowing system connections via cables, serving as a bridge for memory scale-up.

[0128] In an embodiment of the present application, the memory expansion unit can be a memory pool, which includes at least one second processing unit with independent memory and at least one second PCIe switch. The accelerator can access the local memory of the second processing unit in the memory pool, as well as the integrated memory composed of the local memory of the first processing unit and the main memory, thereby expanding the memory available to the accelerator, providing more efficient memory access and sharing, accelerating data access, and improving performance.

[0129] Optionally, the second processing unit includes any one or combination of a field programmable gate array FPGA, a complex programmable logic device CPLD, a programmable logic device PLD, an application specific integrated circuit ASIC, a general array logic GAL, a system on chip SOC, a software defined architecture SDI device and an artificial intelligence AI device.

[0130] Optionally, the memory pool further includes a third MCIO connector, which is connected to the Fabric port of the second PCIe switch and is configured to be connected to a fourth MCIO connector of another computing device.

[0131] Figure 8 is a second schematic diagram of the memory pool expansion architecture provided by an embodiment of the present application. As shown in Figure 8, by Daisy-chaining four computing devices (or systems) through MCIO, a large memory pool can be formed. The four computing devices will share memory expansion and dynamically allocate memory. At this point, the four host systems can form a star-linked topology. When a system has a large memory demand, it can be allocated to any system through the Fabric port, achieving dynamic resource allocation, thereby increasing the computing power of the system's computing units and maximizing resource optimization.

[0132] Optionally, the FPGA is configured to divide internal resources of the FPGA into different regions using a dynamic partitioning technique, or to use a direct memory access (DMA) controller to implement dynamic memory allocation and data transmission.

[0133] FPGAs can implement memory allocation, which refers to dynamically allocating and managing memory resources within the FPGA so that different modules or subsystems can share and use this memory. FPGA memory allocation is implemented in the following ways:

[0134] a) Dynamic Partitioning: FPGAs can use dynamic partitioning to divide their internal resources into different regions. Part of this region can be used as the FPGA main memory area. The remaining area can be partitioned and allocated to different GPUs or CPUs at runtime based on system requirements.

[0135] b) DMA Controller: FPGAs can use a DMA (Direct Memory Access) controller to dynamically allocate and transfer memory. The DMA controller can transfer data from the FPGA's memory area to other devices or modules, and also transfer data from other devices or modules to the FPGA's memory area, enabling dynamic memory sharing and utilization.

[0136] The FPGA used in this application's architecture features Multi-Channel DMAIP for PCI Express, primarily composed of H2DDM (Host-to-Device Data Mover) and D2HDM (Device-to-Host Data Mover) modules. It also provides DMA bypass functionality for the host, enabling PIO read and write operations to device memory.

[0137] The MCDMA engine operates on software DMA queues, used to transfer data between the local FPGA and the host. Each queue element is a software descriptor written by the driver / software. The hardware reads the queue descriptors and executes them. The hardware can support up to 2K DMA channels. For each channel, a separate queue is used for read and write DMA operations.

[0138] The H2DDM module transfers data from the host memory to the local memory through the PCIe hardware IP and the Avalon-MM Write Master / Avalon-ST Source interface.

[0139] The D2HDM module transfers data from device memory to host memory. It receives data from user logic through the Avalon-MM Read Master / Avalon-ST Sink interface and generates Mem Wr TLPs based on descriptor information (such as PCIe address (destination), data size, and MPS value), moves the data to the host, and transfers it to the receive buffer of the host memory.

[0140] In some embodiments, the accelerator is a graphics processing unit (GPU).

[0141] Figure 9 is a schematic diagram of the connection of eight GPUs, GPU0-GPU7, provided in an embodiment of the present application. This connection is called a Starlink topology, and is interconnected through PCIe Switch 0, PCIe Switch 1, PCIe Switch 2, and PCIe Switch 3. PCIe Switch 0, PCIe Switch 1, PCIe Switch 2, and PCIe Switch 3 are not shown in Figure 9. PCIe Switch 0's downstream ports are all connected to GPU 0 and GPU 1. The PCIe Switch hardware itself has peer-to-peer communication capabilities, supporting the shortest path transmission between peers. GPU 0 and GPU 1 can directly share their own memory through peer-to-peer communication. GPU 2 and GPU 3 are the downstream ports of PCIe Switch 1, GPU 4 and GPU 5 are the downstream ports of PCIe Switch 2, and GPU 6 and GPU 7 are the downstream ports of PCIe Switch 3. All of them can share their memory through the peer-to-peer function. Therefore, as shown in Figure 9, each group has an interconnection path. These paths allow the eight groups of GPUs to share memory. GPU memory sharing refers to the process of sharing the same block of memory between multiple computing units (usually threads or threads) when performing general computing on the GPU. In a GPU, the GPU usually has a large number of computing cores and can execute multiple computing units simultaneously to accelerate processing.

[0142] GPU memory sharing is achieved by using shared memory or global memory in the GPU. These memory areas can be accessed and operated by different computing units, allowing them to share data during the computation process. This design avoids unnecessary data copying and transfer, thereby improving computational efficiency and performance.

[0143] However, in GPU memory sharing, special attention must be paid to synchronization and race conditions. This is because multiple computing units accessing shared memory simultaneously may lead to data races and inconsistencies. To ensure data correctness, synchronization mechanisms (such as mutexes and semaphores) are required to control access to shared memory and ensure data consistency.

[0144] GPU memory sharing is widely used in many GPU computing applications, such as machine learning, deep learning, scientific computing, image processing, etc. By properly designing and managing memory sharing, the computing efficiency of the GPU can be maximized, resulting in faster and more efficient operations.

[0145] The computing device provided by the embodiments of the present application has the following beneficial effects: 1) Shared resources: The GPU and the host CPU share the same memory space and can directly access data in the main memory, avoiding the additional copying of data from the main memory to the GPU's dedicated memory. 2) Cost savings: The GPU and CPU can directly expand the configured memory capacity through CXL. 3) Cache acceleration: Using the main memory as cache memory can accelerate GPU data access and improve performance. 4) The introduction of CXL technology will further promote the collaborative work between the GPU and the host CPU, provide more efficient memory access and sharing, and thus optimize high-performance computing applications. 5) Using the CXL protocol, the memory between the host CPU and the accelerator can be dynamically allocated. This means that when an accelerator needs more memory space to process a specific task, it can request more memory from the host CPU, and the host CPU can also allocate part of the memory to the accelerator, sharing memory resources in an optimal manner. 6) Implementing a memory pool through the CXL protocol can significantly improve the overall performance of the system. The accelerator can directly access the host CPU's memory, avoiding the tedious data transmission and copying process, reducing the time for calculation and data processing, and thus speeding up the system's operation.

[0146] Figure 10 is a schematic diagram of the structure of the server provided in an embodiment of the present application. As shown in Figure 10, the server 1010 includes a computing device 1020. For an understanding of the computing device, reference can be made to the description in the previous embodiment, which will not be repeated here.

[0147] For example, the computing devices provided in the above embodiments of the present application can be used on AI servers to solve the problem of insufficient memory in the computing unit to meet the needs of increasingly complex and large AI models.

[0148] FIG11 is a flow chart of a data processing method provided in an embodiment of the present application. As shown in FIG11 , the data processing method includes:

[0149] Step 1110: The accelerator sends a first data request message to the main memory based on the CXL protocol.

[0150] In this step, the accelerator is connected to the CPU through the first PCIe switch, so that the accelerator directly accesses the system's main memory based on the CXL protocol. Therefore, when there is a need, the accelerator sends a first data request message to the main memory based on the CXL protocol. The first data request message is used to request the first task data stored in the main memory.

[0151] Optionally, the CXL protocol here specifically refers to CXL.io and CXL.cache protocols.

[0152] Step 1120 : The main memory sends the first task data to the accelerator through the first PCIe switch in response to the first data request message.

[0153] In this step, the main memory receives a first data request message sent by the accelerator. The first data request message is used to request the first task data stored in the main memory. Therefore, the main memory responds to the first data request message, obtains the first task data, and sends the first task data to the accelerator through the first PCIe switch.

[0154] In an embodiment of the present application, the accelerator obtains data directly from the main memory based on the CXL protocol, which expands the memory available to the accelerator, provides more efficient memory access and sharing, and can accelerate data access and improve performance.

[0155] In some embodiments, the method further comprises:

[0156] The accelerator performs a computing task based on the first task data to generate first result data;

[0157] The accelerator receives the second data request message sent by the CPU, and sends the first result data to the main memory.

[0158] It is understandable that the accelerator is connected to the CPU via the first PCIe switch, so that the accelerator can access the main memory based on the CXL protocol. After receiving the first task data, the accelerator performs the corresponding computing task and generates first result data. The first result data can be stored in the main memory for use by the CPU or other devices when performing other computing tasks. The accelerator can also store the first result data in the accelerator's local memory. When necessary, the CPU actively obtains the first result data from the accelerator. The accelerator receives the second data request message sent by the CPU and sends the first result data to the main memory, so that the CPU can also use the accelerator's memory, thereby providing more efficient memory access and sharing.

[0159] In an embodiment of the present application, the accelerator performs a computing task based on the first task data, generates first result data, and sends the first result data to the main memory. The CPU can also use the computing results of the accelerator or the data stored in the memory, thereby providing more efficient memory access and sharing.

[0160] In some embodiments, the method further comprises:

[0161] The accelerator sends a third data request message to the memory extension unit based on the CXL protocol;

[0162] The memory expansion unit sends the second task data to the accelerator through the first PCIe switch in response to the third data request message.

[0163] The operation of the method of this embodiment requires the presence of a memory expansion unit in the computing device. The memory expansion unit is connected to the second downstream port of the first PCIe switch; the local memory of the memory expansion unit and the main memory constitute an integrated memory; the accelerator is configured to perform access operations on the integrated memory based on the CXL protocol. It can be understood that the accelerator can access the integrated memory based on the CXL protocol, that is, the accelerator can obtain data from the local memory of the memory expansion unit. Specifically, the accelerator sends a third data request message to the memory expansion unit based on the CXL protocol. The memory expansion unit responds to the third data request message and sends the second task data to the accelerator through the first PCIe switch.

[0164] In an embodiment of the present application, the local memory of the memory expansion unit and the main memory constitute an integrated memory. The accelerator can access the integrated memory based on the CXL protocol. Data can be shared and moved more efficiently between different levels of the memory hierarchy, reducing the data transmission requirements between different memory types and potentially reducing latency.

[0165] In some embodiments, the method further comprises:

[0166] The accelerator stores the third task data to the memory expansion unit through the first PCIe switch.

[0167] The accelerator can access the integrated memory based on the CXL protocol, that is, the accelerator can store data in the local memory of the memory expansion unit through the CXL protocol. Specifically, the accelerator stores the third task data in the memory expansion unit through the first PCIe switch.

[0168] In an embodiment of the present application, the accelerator can access the memory space shared by the CPU and the memory expansion unit, reducing the delay in data transmission and the overhead of copying. Since the CPU and the memory expansion unit share the same memory resources, the physical memory capacity required in the system can be reduced, saving hardware costs, and memory management can be more unified, thereby improving the overall performance of the system.

[0169] On the other hand, the present application also provides a computer-readable instruction product, which includes computer-readable instructions. The computer-readable instructions can be stored on a non-transitory computer-readable storage medium. When the computer-readable instructions are executed by the processor, the computer can execute the above-mentioned data processing method embodiments, which will not be repeated here.

[0170] On the other hand, the present application also provides a non-transitory computer-readable storage medium having computer-readable instructions stored thereon. When the computer-readable instructions are executed by a processor, the above-mentioned data processing method embodiments are implemented, which will not be repeated here.

[0171] It should be noted that each implementation method of the present application can be freely combined, the order can be changed, or it can be executed separately, and does not need to rely on or depend on a fixed execution order.

[0172] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0173] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the relevant technology, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.

[0174] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A computing device, characterized in that, including: a central processing unit CPU, an accelerator, and a first Peripheral Component Interconnect Express (PCIe) switch; the accelerator is connected to a first downstream port of the first PCIe switch; an upstream port of the first PCIe switch is connected to the CPU; each port of the first PCIe switch supports the Compute Express Link (CXL) protocol; and the accelerator is configured to perform access operations on the main memory based on the CXL protocol.

2. The computing device according to claim 1, wherein It further includes: a memory expansion unit, the memory expansion unit is connected to a second downstream port of the first PCIe switch; the local memory of the memory expansion unit and the main memory form a unified memory; and the accelerator is configured to perform access operations on the unified memory based on the CXL protocol.

3. The computing device according to claim 2, wherein The memory expansion unit includes at least one first processing unit with independent memory.

4. The computing device according to claim 3, wherein The first processing unit includes any one or a combination of a Field Programmable Gate Array (FPGA), a Complex Programmable Logic Device (CPLD), a Programmable Logic Device (PLD), an Application Specific Integrated Circuit (ASIC), a Generic Array Logic (GAL), a System on Chip (SOC), a Software Defined Infrastructure (SDI) device, and an Artificial Intelligence (AI) device.

5. The computing device according to claim 2, wherein The memory expansion unit includes a memory expansion board.

6. The computing device according to claim 5, wherein The memory expansion board is equipped with a Dynamic Random Access Memory (DRAM) or a Storage Class Memory (SCM).

7. The computing device according to claim 1, wherein It further includes: a Non-Volatile Memory Host Controller Interface Specification (NVMe) solid state drive, the NVMe solid state drive is connected to a third downstream port of the first PCIe switch; and the accelerator is configured to perform access operations on the NVMe solid state drive based on the CXL protocol.

8. The computing device according to claim 1, wherein It further includes: a Network Interface Controller (NIC), the NIC is connected to a fourth downstream port of the first PCIe switch; and the accelerator is configured to perform data interaction with the NIC based on the CXL protocol.

9. The computing device according to claim 1, wherein It further includes: a first Multi-Chassis Input / Output (MCIO) connector, the first MCIO connector is connected to the Fabric port of the first PCIe switch for connection to a second MCIO connector of other computing devices.

10. The computing device according to claim 2, wherein The memory expansion unit includes a memory pool, the memory pool includes at least one second processing unit with independent memory and at least one second PCIe switch; and a downstream port of the second PCIe switch is connected to a first Endpoint port of the second processing unit, a second Endpoint port of the second processing unit is connected to a fifth downstream port of the first PCIe switch.

11. The computing device according to claim 10, wherein The memory pool further includes a third MCIO connector, the third MCIO connector is connected to the Fabric port of the second PCIe switch for connection to a fourth MCIO connector of other computing devices.

12. The computing device according to claim 10, wherein The second processing unit includes any one or a combination of a Field Programmable Gate Array (FPGA), a Complex Programmable Logic Device (CPLD), a Programmable Logic Device (PLD), an Application Specific Integrated Circuit (ASIC), a Generic Array Logic (GAL), a System on Chip (SOC), a Software Defined Infrastructure (SDI) device, and an Artificial Intelligence (AI) device.

13. The computing device according to claim 12, wherein The FPGA is used to divide the internal resources of the FPGA into different regions using dynamic partitioning technology, or to implement dynamic allocation and data transfer of memory using a direct memory access (DMA) controller.

14. The computing device according to any one of claims 1-13, characterized in that, The accelerator is a graphics processing unit (GPU).

15. A server, characterized in that, It includes the computing device according to any one of claims 1 to 14.

16. A data processing method, characterized in that Based on the computing device according to any one of claims 1 to 14, it includes: The accelerator sends a first data request message to the main memory based on the CXL protocol; and In response to the first data request message, the main memory sends first task data to the accelerator through the first PCIe switch.

17. The data processing method according to claim 16, wherein The method further includes: The accelerator executes a computing task based on the first task data to generate first result data; and The accelerator receives a second data request message sent by the CPU and sends the first result data to the main memory.

18. The data processing method according to claim 16, wherein The method further includes: The accelerator sends a third data request message to the memory expansion unit based on the CXL protocol; and In response to the third data request message, the memory expansion unit sends second task data to the accelerator through the first PCIe switch.

19. The data processing method according to claim 18, wherein The method further includes: The accelerator stores third task data to the memory expansion unit through the first PCIe switch.

20. A non-transitory computer-readable storage medium having computer-readable instructions stored thereon, characterized in that, When the computer-readable instructions are executed by a processor, they implement the data processing method according to any one of claims 16 to 19.

Citation Information

Patent Citations

  • Server and data center

    CN115964315A

  • Access acceleration system of storage device

    CN115994107A

  • CXL data transmission board card and method for controlling data transmission

    CN116501681A

  • Computing device, server, data processing method and storage medium

    CN117493237A

  • Integrated storage / processing devices, systems and methods for performing big data analytics

    US20140129753A1

Cited By

  • Distributed system, data processing method, equipment, medium and program product

    CN120578509A

  • Processor, processor communication method and computer equipment

    CN120723682A