Proximity data processing system for extra-large mixture-of-experts artificial intelligence model

MoNDE addresses memory and computation inefficiencies in MoE LLMs by using a CXL-enabled memory system to transfer activation data and optimize load balancing, enhancing performance and reducing latency.

WO2025143339A1PCT designated stage expired Publication Date: 2025-07-03SEOUL NATIONAL UNIVERSITY R&DB FOUNDATION

Patent Information

Application Number
PCT/KR2023/022047
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-29
Filing Date
2023-12-29
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Transformer-based large-scale language models (LLMs) face challenges with increased computation requirements, memory constraints, inter-device data transfer overhead, and dynamic token routing inefficiencies, particularly in Mixture-of-Experts (MoE) models, leading to performance bottlenecks and hardware inefficiencies.

Method used

The MoNDE (Mixture-of-Near-Data-Experts) system employs a CXL-enabled memory system with an NDP core unit to process MoE feed-forward layers, transferring activation data rather than expert parameters, leveraging high internal memory bandwidth and optimizing load balancing between GPU and NDP devices.

Benefits of technology

MoNDE significantly reduces inference latency by up to 2.11x and improves computational efficiency, addressing memory constraints and data transfer bottlenecks in capacity-constrained environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2023022047_03072025_PF_FP_ABST
    Figure KR2023022047_03072025_PF_FP_ABST
Patent Text Reader

Abstract

A mixture-of-experts artificial intelligence model according to the present invention comprises: a CXL control unit that manages a CXL protocol; an NDP control unit that includes a memory mapping register set for communicating with a host; a memory unit composed of a plurality of memory chips that provide a memory capacity to the host; and an NDP core unit that processes MoEs existing in the memory unit, wherein the NDP core unit transmits only hot experts to a GPU accelerator while memory-based calculation is performed on cold experts that exist in the memory unit.
Need to check novelty before this filing date? Find Prior Art

Description

Proximity data processing system for large-scale expert-mixed AI models

[0001] The present invention relates to a Mixture-of-Experts (MoE) artificial intelligence model for efficiently scaling a Large Language Model (LLM).

[0002] Transformer-based large-scale language models (LLMs) have demonstrated impressive performance in a variety of natural language processing (NLP) tasks, including question answering, machine translation, and code generation. This outstanding model performance can be attributed to the unprecedented model size, which has steadily increased over the years, as illustrated in Figure 1a.

[0003] Thanks to advances in modern computing systems, the size of LLMs has increased by approximately a million times over the past five years, enabling significant qualitative improvements.

[0004] However, large-scale models require increased computation, making it increasingly difficult for LLM to meet the strict latency constraints of inference tasks.

[0005] Meanwhile, Mixture of Experts (MoE) has attracted attention as a way to achieve high model quality for various computer vision and language modeling tasks while significantly scaling up model size at a fixed computational cost.

[0006] However, MoE models are notoriously parameter-inefficient, and several challenges remain to be addressed for a more practical deployment for inference:

[0007] First, there is the challenge of memory requirements.

[0008] The parameter size of MoE LLM increases asymptotically-linearly with the number of experts and can easily exceed the total GPU memory capacity of a multi-GPU computing node, as shown in Fig. 1b.

[0009] For example, the T5-Large requires approximately 3 GB of memory, while the Switch Transformers-Large, a 128-bit MoE version of the T5-Large, requires approximately 100 GB (34 times) of memory. In other words, applying MoE to LLM significantly increases memory requirements.

[0010] Next, there are challenges associated with moving data over slow interconnect links.

[0011] Modern deep learning frameworks provide parameter offloading and fetching between CPU and GPU devices to alleviate memory capacity issues.

[0012] Previously proposed MoE-specific offloading strategies offload sparse MoE parameters to the CPU while keeping the remaining dense parameters on the GPU. However, due to the significant inter-device data transfer overhead caused by low Peripheral Component Interconnect Express (PCIe) bandwidth, these parameter offloading techniques are not suitable for MoE inference. Despite bandwidth improvements in PCIe Gen 5 and 6 technologies, a significant gap still exists between inter-device and intra-device bandwidth.

[0013] Finally, there are challenges related to dynamic token routing.

[0014] Dynamic token routing and expert load imbalance are inherent characteristics of the MoE model, which substantially vary the workload characteristics of individual expert modules. The widely adopted batch matrix multiplication primitive, which performs a fixed-shape set of general matrix multiplications (GEMMs), inherently exhibits poor computational efficiency when handling sparse MoE computations. In particular, experts with few routed tokens perform GEMMs on large, thin matrices, which cannot saturate the compute capabilities of widely used GPUs, resulting in low compute utilization and poor memory-constrained performance.

[0015] To address the aforementioned inefficiencies, the present invention proposes MoNDE (Mixture-of-Near-Data-Experts), a Near-Data Processing (NDP) solution implemented on a PCIe-connected memory device. As the size of expert parameters expands with the number of experts, offloading sparse expert parameters to backup memory for storage becomes an inevitable issue in resource-constrained environments.

[0016] The technical task of the present invention is to solve the above-described problem, and to provide an expert mixed artificial intelligence model capable of resolving expert bias and dynamic load imbalance.

[0017] In addition, the technical task of the present invention is to resolve a bottleneck that may occur in parameter transmission.

[0018] In addition, the technical task of the present invention is to minimize parameter transmission in the expert calculation process.

[0019] The expert mixed artificial intelligence model according to the present invention is characterized by including a CXL control unit for managing a CXL protocol, an NDP control unit having a memory mapping register set for communicating with a host, a memory unit comprising a plurality of memory chips for providing memory capacity to the host, and an NDP core unit for processing MoE experts present in the memory unit.

[0020] In one embodiment, the NDP core unit is characterized in that it transfers only hot experts to the GPU accelerator while memory-based computations are performed on cold experts residing in the memory unit.

[0021] In one embodiment, the CXL control unit is characterized in that it offloads an NDP command using a CXL message and transmits the offloaded NDP command to the NDP control unit.

[0022] In one embodiment, the NDP control unit is characterized in that when the NDP command arrives, the NDP control unit generates a memory request and loads data into the NDP core unit tile by tile.

[0023] In one embodiment, the NDP control unit is characterized in that, after the memory calculation is completed, it generates a completion signal indicating that the calculation is completed by setting a memory-mapped completion register.

[0024] According to the MoNDE method proposed in the present invention, most expert parameter transfers are replaced with relatively inexpensive input token activation data between GPUs and memory devices.

[0025] That is, the MoE task dispatched using the activation data sent from the GPU is efficiently performed by the MoNDE NDP engine, and then the output activation data is sent back to the GPU for subsequent transformer tasks (i.e., tasks of the attention layer).

[0026] In this way, replacing large expert parameter transmissions with much smaller active data sets significantly reduces inference latency.

[0027] Additionally, expert tasks with few allocated tokens prioritize memory bandwidth over computational power, so the MoNDE NDP engine can handle them efficiently by leveraging high on-device memory bandwidth with a low logic budget.

[0028] According to the present invention, a single GPU-MoNDE device pair can improve MoE inference performance by up to 2.11 times compared to prior art enabling inference in capacity-constrained settings.

[0029] Figure 1 is a graph showing the current trend of LLM size.

[0030] Figure 2 is a conceptual diagram illustrating an overview of the Transformer block and MoE FFN layer with E = 3 experts and top 2 routing.

[0031] Figure 3 is a graph showing expert activation data of the MoE layer in NLLB-MoE.

[0032] Figure 4 is a graph showing a single expert FFN analysis of a switch transformer.

[0033] Figure 5 is a block diagram showing an overview of the MoNDE system according to the present invention.

[0034] Figure 6 is a block diagram showing the MoNDE NDP core architecture according to the present invention.

[0035] Figure 7 is a conceptual diagram comparing PMove and AMove.

[0036] Figure 8 is a conceptual diagram illustrating expert parameter movement and analysis of active data via PCIe.

[0037] Figure 9 is a flowchart showing the workflow of MoNDE according to the present invention.

[0038] Figure 10 is a conceptual diagram showing a comparison of workflows according to MoNDE optimization.

[0039] Figure 11 is a graph showing the end-to-end latency of a large-scale MoE model.

[0040] Figure 12 is a graph showing speed-up GPM.

[0041] Figure 13 is a graph showing a comparison of latency of switch transformers.

[0042] Figure 14 is a graph comparing the performance of MoNDE according to the prior art and the present invention.

[0043] Figure 15 is a graph showing AMove latency.

[0044] [Project ID]1711193550

[0045] [Assignment Number] 2021-0-00863-003

[0046] [Ministry Name] Ministry of Science and ICT

[0047] [Research Management Specialist Agency] Information and Communications Technology Planning and Evaluation Institute

[0048] [Research Project Name] Development of New Concept PIM Semiconductor Leading Technology (R&D)

[0049] [Research Project Title] Development of an Intelligent In-Memory Error Correction Device for High-Reliability Memory

[0050] [Contribution rate] 100 / 100

[0051] [Main Research Institution] Seoul National University Industry-Academic Cooperation Foundation

[0052] [Research Period] April 1, 2021 - December 31, 2024

[0053]

[0054] Hereinafter, embodiments of the present invention will be described in detail with reference to the attached drawings so that those skilled in the art can easily implement the present invention. However, the present invention may be implemented in various different forms and is not limited to the embodiments disclosed below. Furthermore, in order to clearly disclose the present invention in the drawings, parts unrelated to the present invention have been omitted, and identical or similar symbols in the drawings represent identical or similar components.

[0055] The purpose and effects of the present invention can be naturally understood or made clearer by the following description, and the purpose and effects of the present invention are not limited to the following description alone.

[0056] The purpose, features, and advantages of the present invention will become more apparent through the following detailed description. Furthermore, in describing the present invention, detailed descriptions of known technologies related to the present invention will be omitted if they are deemed to unnecessarily obscure the gist of the invention.

[0057] Hereinafter, an embodiment of the present invention will be described in detail with reference to the attached drawings.

[0058] First, we explain the Transformer-based model.

[0059] The Transformer model consists of an encoder and a decoder, each of which can be configured as a stack of Transformer blocks. Each Transformer block can be configured with Attention and Feed-forward layers, as shown in Figure 2.

[0060] At this point, the encoder and decoder can process the input data in different ways. Specifically, the encoder module can ingest the entire batched token sequence at once and encode it into an intermediate representation that summarizes the sequence context. When the encoder is used independently, the encoder output is fed to the language model head for various downstream tasks, such as sentiment analysis and question answering.

[0061] When used with a decoder, the output embedding for the last input token is used as input to the decoder for token generation, and the keys and values ​​generated inside the encoder's transformer block are passed to the decoder to enable encoder-decoder cross-multi-heading.

[0062] BERT (BiDirectional Encoder Representations from Transformers) is an example of an encoder-based Transformer model.

[0063] In contrast, the generative decoder module uses a single batch of tokens as input for each independent sequence for token generation.

[0064] The encoder's keys and values ​​are also passed as input to the cross-MHA. The generated output tokens are used as input to the decoder module to generate the next output token.

[0065] The decoder module continues to perform this iterative token generation until a predefined maximum number of tokens is reached or an end-of-sequence token is generated.

[0066] When the decoder is used independently, a summary step is added before token generation to allow the model to understand the input sequence context.

[0067] The output embedding for the last token in the input sequence is used as input to start the token generation step.

[0068] GPT (Generative Pre-trained Transformer) is an example of a decoder-based Transformer model. Some Transformer models, such as T5, use both encoder and decoder modules.

[0069] Table 1 describes the notation for model or data dimensions used in the following sections.

[0070]

[0071]

[0072] Below, we explain MoE.

[0073] Mixture-of-Experts (MoE) is an ensemble learning technique aimed at improving model performance. The primary motivation for using MoE is to increase model capacity without a proportional increase in computational complexity. Recent research has shown that Transformer models employing this technique (called MoE Transformers) achieve significant improvements over conventional high-density Transformers.

[0074] Figure 2 shows a high-level overview of the MoE Transformer. Here, MoE is applied to the Feedforward Network (FFN) of the Transformer block. The MoE FFN layer combines multiple copies of the FFN, called experts. A core component of MoE is the gating network, also known as the routing network.

[0075] The gating network determines the experts to whom the tokens are routed. For each input token, the gating function uses a softmax function to compute a probability distribution over the experts and generates a (token, expert) score map, which is used to route each token to the top k experts. The routed tokens are processed by each expert, and the expert outputs are combined to reconstruct the original token order and fed to the next Transformer block.

[0076] There are several variations of gating networks, each using different inputs to the softmax function. Sparsely Gated MoE uses two (d model , E) W, which is a weight matrix g Wow W noise trains. Each matrix contains d model -dim computes the score of each E expert for a token by multiplying the input embedding tokens. Then, before performing the softmax computation, only the top k values ​​are kept while the remaining scores are set to negative infinity (-Δ). This makes the final E-dim output vector sparse by setting the non-top k scores to 0.

[0077] Switch Transformers use a top-1 (i.e., k = 1) gating strategy (called switch routing) that routes input tokens to only a single expert, making the gating operation cheaper than the top-k gating operation. In practice, the gating operation is performed on B batches of S input tokens. Therefore, the gating network is (B S, d model ) and (d model , E) Perform matrix multiplication of matrices.

[0078] In the present invention, the characteristics of MoE transformers, which are more widely used, are analyzed using switch transformers.

[0079] Below, we describe CPU memory utilization for MoE LLM.

[0080] As LLM continues to scale to hundreds of billions of parameters, hosting the entire model parameters in GPU memory is unlikely to be a viable solution.

[0081] Meanwhile, new interconnect technologies such as Compute Express Link (CXL) can add up to tens of terabytes of memory capacity to CPUs. Consequently, significant efforts have been made recently to leverage large CPU memory capacities for LLM training and inference.

[0082] For example, Microsoft's DeepSpeed ​​and its parameter offloading implementation allow LLMs to run on limited GPU memory by offloading model parameters to CPU memory (or NVMe storage) and bringing them to the accelerator immediately when computation is needed.

[0083] Recent work also demonstrates a parameter offloading scheme for MoE inference, where dense parameters (e.g., those of the attention layer) are stored in the accelerator memory, and sparse but large MoE-specific FFN parameters are offloaded to CPU memory due to the memory capacity constraints of the accelerator.

[0084] Meanwhile, the following inefficiencies and problems exist in the process of executing MoE inference.

[0085] First, there is expert bias and dynamic load imbalance.

[0086] As described above, the gating network can determine which experts tokens are routed to. Therefore, the number of tokens each expert receives varies significantly across the MoE FFN layer.

[0087] Figure 3 shows the token distribution across experts when performing translation inference on sentence collocations from the FLORES-200 dataset using NLLB-MoE. As shown in Figure 3, only a limited number of experts are processing a large number of tokens, and these experts are defined as "hot experts."

[0088] In contrast, the majority of experts receive significantly fewer tokens, and are referred to as "cold experts."

[0089] The presence of hot experts is a widespread pattern across MoE Transformers, indicating that the workload (i.e., the compute-to-memory ratio) varies across experts. Hot experts have high workloads and, therefore, are computationally intensive because a larger number of routed tokens constitute a high computational load. Cold experts are allocated a relatively low computational load and are therefore memory-intensive. This imbalance can lead to hardware inefficiencies, as described below.

[0090] Typically, prior art techniques aim to balance the expert load by deleting tokens that impede hardware performance scaling at the risk of model quality degradation.

[0091] Second, there is a parameter transmission bottleneck.

[0092] As LLM scaling progresses through MoE, the limited GPU memory capacity will soon require parameters to be offloaded to auxiliary memory spaces such as CPU memory.

[0093] Therefore, modern deep learning frameworks have already begun to support parameter transfer between CPU and GPU memory as described above, allowing users and customers to run LLMs on billions of scales within GPU resources.

[0094] However, simply transferring the offloaded expert parameters to the GPU poses a major performance bottleneck when running MoE inference.

[0095] Figure 4 shows the performance analysis between expert parameter transfer, FFN calculation, and loopline analysis when the expert receives different numbers of tokens. Figure 4 shows that transferring a single expert parameter to the GPU takes significantly longer than FFN calculation (up to 27 times longer for a single routed token).

[0096] Furthermore, loopline analysis shows that for a small number of routed tokens, as is the case for most experts in the MoE layer, FFN computation is memory bandwidth limited, requiring less computational throughput from GPUs.

[0097] Third, the memory-bound nature of expert calculations.

[0098] For B parallel arrangement tokens, the specialized FFN is essentially a dimension (S, d model ) Х (d model , d ff ) and (S, d ff ) Х (d ff , d model ) performs a series of matrix multiplications. d model and d ff Assuming that is large enough, a large number of routed tokens (i.e., large S dimension) will result in a lot of reuse of input activation data and weight data.

[0099] This binds the above expert work to the calculation.

[0100] On the other hand, if the expert receives only a small number of tokens to compute, the S dimension is reduced and the reuse of input activation data and weight data is also reduced, limiting the memory bandwidth of FFN computation.

[0101] Therefore, most of the runtime is spent loading expert parameters to the GPU rather than performing the calculations.

[0102] Figure 4 demonstrates this by showing that a smaller number of routed tokens leads to significantly higher per-token computational latency and lower hardware efficiency. This suggests that per-token computations can vary by a factor of 1.2 to 46 for the evaluated cases. The imbalance in the MoE token distribution is considered unsympathetic by most experts, making the entire MoE operation memory-bound.

[0103] Fourth, there is the opportunity for short-range data MoE.

[0104] As described above, it appears that minimizing expert parameter transfer is the most important factor in reducing MoE LLM inference time when CPU-offloaded MoE expert parameters are present.

[0105] Accordingly, the present invention proposes a near-distance data processing (NDP) solution that can reduce data transfer through PCIe and provide cost-effective computation for memory-bound specialized computations. NDP devices can utilize high internal memory bandwidth and avoid bottlenecks caused by relatively slow PCIe links, making them an effective solution.

[0106] Additionally, it should be noted that many of the characteristics considered to be cold experts are memory-intensive. Meeting the computational demands of these MoE expert tasks can be achieved through a specialized memory-centric NDP design optimized for computing memory-intensive matrix operations.

[0107]

[0108] Below, MoNDE, an artificial intelligence model proposed in the present invention, is described.

[0109] In this invention, we propose MoNDE (Mixture-of-Near-Data-Experts), a CXL-enabled memory system with customized computing capabilities for processing MoE feed forward network (FFN) layers for deep learning neural network (DNN) inference.

[0110] MoNDE according to the present invention can store large expert parameters and process MoE FFN layers within CXL-based memory. Unlike conventional techniques that move large parameters to the GPU for computation, MoNDE takes a different approach, significantly reducing the burden of data movement by transferring activation data to the parameter location and performing expert computation on-site.

[0111]

[0112] - MoNDE device architecture

[0113] Figure 5 illustrates an overview of the MoNDE system. MoNDE can accommodate MoE FFN parameters by increasing memory capacity through the CXL interface, while simultaneously leveraging computing power to accelerate MoE LLM inference tasks.

[0114] As illustrated, the MoNDE memory may include at least one of a CXL controller (1), an NDP controller (2), an NDP core (3), and a device memory (4).

[0115] First, the CXL controller (1) can be configured as dedicated hardware responsible for managing the CXL protocol.

[0116] The CXL controller (1) can decompose incoming packets, extract memory requests initiated by the host, and reassemble responses from device memory into packets.

[0117] Additionally, the CXL controller (1) can receive NDP commands. That is, the CXL controller of MoNDE is enhanced in that it distinguishes between memory requests and NDP commands.

[0118] MoNDE can offload NDP commands using a CXL message called RwD (Request with Data). The RwD message contains a reserved opcode, indicating whether the accompanying data belongs to an NDP command. These NDP commands can be passed to MoNDE's NDP controller. Interface management by the MoNDE device controller will be described later.

[0119] Next, the NDP controller may be equipped with a set of memory-mapped registers for communicating with the host.

[0120] Specifically, when an NDP instruction arrives at the instruction register, the NDP controller can generate a memory request and load data into the NDP core on a tile-by-tile basis. Furthermore, the NDP controller can generate a completion signal by setting a memory-mapped completion register after the computation is complete.

[0121] Next, the NDP core (3) can process the MoE expert in the device memory (4).

[0122] The NDP core is a rectangular array of multiply-accumulate (MAC) processing elements (PEs) that can efficiently meet the changing computing demands of each expert depending on the input data dimension.

[0123] In particular, the input embedding dimension corresponding to the number of tokens routed to the expert can vary significantly, whereas for large-scale LLMs, d model and d ffConsidering that the dimension of is fixed and generally high (hundreds of thousands to tens of thousands), the NDP of MoNDE according to the present invention adopts an output-stationary systolic array design with small row and large column dimensions for high hardware utilization.

[0124] Scratchpad memory can store specialized parameters and active data read from memory and used in calculations. The read data is loaded into the systolic array buffer, and the final output data can be written back to DRAM. The expert FFN may include a special function unit (SFU) for active data functions (ReLU, GeLU).

[0125] Additionally, the present invention employs a series of SIMD instruction arrays that operate on different expert weight partitions and the entire input token active data. Furthermore, a MoNDE-specific hardware support device is designed to assist in providing data for MAC device input operands.

[0126] The Activation Processing Unit (APU) can preprocess token activation data exchanged between the GPU and the memory-mapped MoNDE activation data buffer within the APU. This is a single, serialized method that initially reduces the kernel call overhead of copying PCIe data between devices. Based on the data attachment address offset and size metadata for each expert's activation data, the APU can decode the input activation data tensor, split it into inputs for each expert, and populate the activation data scratchpad buffer.

[0127] Meanwhile, the reverse operation (i.e., concatenation) can be performed before the output expert active data is sent back to the GPU. The expert DMA unit tracks the offset addresses and sizes of static expert parameters within device memory. Once the FIFO expert workload buffer is filled with information requiring expert parameters, the DMA arbitration unit fetches the expert parameter data into the expert scratchpad buffer. This utilizes a parallel DRAM channel interface that can provide high-bandwidth expert parameter data to the NDP cores.

[0128] Next, the device memory (4) may be composed of a plurality of memory chips that provide sufficient memory capacity to the host and other peripheral devices.

[0129] In one embodiment, the device memory may be a DRAM module comprising multiple DDR5 chips. A memory controller connected to the DRAM may support access to multiple DRAM channels.

[0130] Compared to PCIe bandwidth, the internal memory bandwidth of MoNDE memory is several times higher. Assuming MoNDE memory is built with four channels of DDR5-3200, it can theoretically utilize a memory bandwidth of 102 GB / s. Compared to the bandwidth of x16 PCIe gen4 (up to 32 GB / s), the above configuration can provide data rates of more than three times. MoNDE NDP can efficiently execute professional calculations by leveraging this high internal memory bandwidth. If MoNDE NDP is not used, DRAM modules can be used as general memory expansion devices.

[0131]

[0132] - MoNDE optimization

[0133] 1) MoNDE's swap and offload strategy:

[0134] As described above, conventional LLM offloading frameworks required a parameter movement (PMove) strategy to transfer the corresponding expert parameters from CPU memory for MoE computation and remove other experts from the GPU when there is insufficient memory on the GPU.

[0135] In contrast, the present invention utilizes the AMove (Activation Movement) strategy to exchange active data generated in the Attention and FFN layers between the MoNDE memory device and the accelerator while processing most of the MoE expert tasks using the MoNDE NDP core.

[0136] Figure 7 shows an overview of the proposed AMove compared to PMove. Since the memory device according to the present invention already hosts the entire parameters, there is no need to send experts back to the CPU in PMove. However, if the MoE layer consists of hundreds or thousands of experts, computing a single MoE FFN layer is basically O(d model Х d ff Х E) parameters need to be moved. In comparison, AMove's data movement is O(B Х S Х d model )am.

[0137] More specifically, the following mathematical expressions 1 and 2 show more accurate overheads for parameter transmission of PMove and AMove, respectively.

[0138]

[0139]

[0140] d in Transformers ff is d model Since it scales proportionally, the amount of data movement of PMove is d model , which is expanded secondarily. Additionally, many MoE models consist of tens to thousands of experts (E) per MoE module.

[0141] dmodel , d ff As E is expected to continue to grow, the volume of data movement required for PMove will increase exponentially. Because PCIe bandwidth is limited, moving this large amount of data from memory devices to GPUs on demand exposes the DNN runtime to similarly high latency.

[0142] On the other hand, d model The data volume scaling of AMove is lower than that of PMove, and the data movement complexity is O(d) when B and S are small. model ) can be reduced to . This is generally the case for MoE inference.

[0143] Fig. 8 shows various d model We compare the batch sizes of 3072 tokens and single expert FFN parameters for the configuration. The line plot shows the larger d model It shows a monotonic and superlinear increase in the FFN-placement ratio for the MoE LLM. This means that as the MoE LLM grows, the amount of data for the MoE FFN parameters increases at a faster rate than the active data rate, making the AMove strategy a better choice than the PMove strategy used in existing AI frameworks.

[0144] Even if the AMove strategy cannot overlap token activation data movement and specialized computation, it reduces the data movement penalty by avoiding PMove and reducing communication volume, thereby reducing latency.

[0145] Additionally, AMove technology enables parallel processing of MoE experts on GPUs and MoNDE NDPs. This effect is further detailed below.

[0146] 2) GPU-MoNDE load balancing:

[0147] To efficiently and simultaneously utilize the basic accelerator and MoNDE NDP devices for MoE operation, we propose a mechanism that performs load balancing of computing devices and dispatches expert workloads to each hardware in a timely manner.

[0148] To achieve high computing performance, the decision to dispatch each expert to the appropriate system is crucial. For example, dispatching too many experts to the GPU incurs significant PMove overhead, which degrades inference performance. Conversely, dispatching too many experts to the MoNDE NDP and too few experts to the GPU can lead to computational inefficiency.

[0149] GPUs typically provide higher-performance computations than other computing devices, so assigning small workloads to them underutilizes valuable GPU resources. Maintaining a consistent balance between GPU and MoNDE NDP computation time is a key goal of the dispatching method.

[0150] With regard to load balancing, the algorithm according to the present invention is designed under two principles.

[0151] First, because PCIe parameter migration is the biggest bottleneck, the dispatch algorithm carefully controls the total number of experts transmitted over the slow PCIe link. Furthermore, considering the computational characteristics of hot and cold experts, compute-intensive hot experts are prioritized for dispatch to GPU accelerators, while memory-intensive cold experts are given lower priority.

[0152] The mathematical expression 3 below is used to find a balance point in dispatch decisions for GPU (requires PMove) and NDP (requires AMove) expert computations, taking into account the first principle described above.

[0153]

[0154] Specifically, in Equation 3, DataGPU and DataNDP represent the data volumes that must be read from the memory expansion device to dispatch experts to the GPU and NDP, respectively. That is, DataGPU corresponds to the data volume of the GPU expert that must be PMoved from the memory device to the GPU, and DataNDP corresponds to the data volume of the Near-Data expert that must be dispatched and used to the MoNDE NDP.

[0155] Additionally, in Equation 3, BWPCIe and BWMem represent the PCIe bandwidth and the internal memory bandwidth of the memory device, respectively, which can be verified in the hardware specifications of the system settings or measured using a profiling tool. By substituting the two bandwidth figures into Equation 3, the ratio between DataGPU and DataNDP can be determined.

[0156] For example, suppose that the algorithm according to the present invention performs load balancing for the NLLB-MoE model with 128 experts in each MoE layer. Also, suppose that the PCIe and internal memory device bandwidths are 24 GB / s and 256 GB / s. In this case, if we substitute the numbers into the above mathematical expression 3 and leave only the DataGPU term on the left side, we get DataGPU = 24 / 256 Х DataNDP _Х DataNDP.

[0157]

[0158] If the total number of active experts (i.e. experts who have received 0 or more tokens) is 88, the algorithm dispatches the top 8 experts with the most tokens (= 88Х1 1+10) to the GPU. The remaining 80 experts are placed in the MoNDE NDP.

[0159] According to Equation 3 above, GPU execution time and NDP execution time can be equalized. Since NDP experts are considered cold and memory-intensive, the NDP execution time is roughly calculated by the memory read time of the NDP expert parameter of size DataNDP. This explains the formula on the right side of Equation 3. Furthermore, since compute-intensive GPU systems can compute hot experts in a very short time, the total GPU execution latency of the load-balanced expert computation scheme is dominated by PMove time. Using the DataGPU and BWP CIe terms to approximate GPU execution time for PMove latency completes the formula on the left side of Equation 3.

[0160] The above method requires sorting experts based on the number of tokens routed in each inference iteration, which can be a burden on the CPU. Therefore, in this invention, we propose a dynamic dispatch threshold algorithm that achieves similar performance while reducing computational costs.

[0161] In the first inference iteration after the memory device is booted, the algorithm according to the present invention finds the top BWMem BWP CIe+BWMem ХEact expert (activated expert) with the most routed tokens based on the gating network results and initializes the dispatch threshold.

[0162] The threshold is maintained for a predefined number of inference iterations (e.g., 10), indicated by the update cycle, before being updated again. This method is effective because the token skewness characteristic among MoE experts is maintained across diverse data samples. To account for the varying computational intensity of different batch sizes, users can adjust the threshold by multiplying the above formula by a scaling factor.

[0163] Figure 9 illustrates the MoNDE dispatch engine, the CPU runtime software that manages the MoNDE dispatch algorithm. As an example, we describe the execution of the MoE layer using a single GPU, single MoNDE NDP example.

[0164] The MoE layer starts with a gating network whose weights reside on the GPU and processes the output activation tensor of the previous state layer to generate token-expert routing decisions. Next, the output activation tensor of the previous state layer (B Х S, d model ) is of size (1, d model ) can be split into individual tokens. In this case, the tokens can be grouped according to routing decisions.

[0165] The next series of tasks to make expert dispatch decisions can be managed by the MoNDE dispatch engine, which fetches token-expert gating decisions from the GPU and uses them to generate expert-token and expert-token-count maps on the CPU.

[0166] The above engine can group experts based on whether the number of tokens for each expert exceeds a dispatch threshold. This can be achieved through a simple element-wise larger value and subsequent masking operations.

[0167] Expert groups with a token count greater than the dispatch threshold are experts processed on the GPU, while other groups may include experts processed by the MoNDE NDP. Data movement and expert computation can be performed by the MoNDE dispatch engine runtime running on the host CPU based on this dispatch decision. Each GPU expert requires a PMove, after which each expert task can be executed on the GPU. The active output data of the GPU expert computation can remain on the GPU.

[0168] Meanwhile, NDP experts can request AMove before or after executing expert tasks on the NDP device to store the final output activation data on the GPU. Once all expert calculations are complete, the output tokens from all GPU and NDP experts collected on the GPU can be combined into a single MoE layer output activation data tensor.

[0169] In Figure 10, the workflow of the MoE Transformer block execution scheme related to parallel hardware streams is illustrated.

[0170] The GPU-MoNDE load balancing scheme proposed in the present invention can achieve the lowest latency due to reduced data movement, specialized processing units, and better load balancing.

[0171]

[0172] - MoNDE software stack

[0173] Below we describe the MoNDE software stack that enables computation near CXL-supported memory.

[0174] The MoNDE according to the present invention can provide a library that application programmers can utilize to achieve efficient MoE LLM inference.

[0175] The library can be implemented on top of the MoNDE device driver, which is responsible for CXL device memory management, NDP command generation, and NDP command offload.

[0176] 1) Memory management

[0177] All NDP kernel operands must be placed in MoNDE device memory before kernel execution, which causes problems when implementing the AMove operation.

[0178] Unlike expert parameters, the input active data size for a specific expert is dynamically determined based on the gating policy. This requires on-demand CXL memory allocation, which slows down the active data transfer speed.

[0179] Therefore, to minimize memory allocation overhead, the MoNDE library can integrate a special memory allocator that manages a pool of memory blocks within CXL memory. The allocator can pre-allocate memory blocks through the MoNDE device driver. The MoNDE device driver can internally track the starting address of each memory block. This starting address is retrieved later when generating NDP instructions.

[0180] 2) Execution flow

[0181] MoNDE according to the present invention can adopt a heterogeneous programming model in which the host executes a kernel and the NDP-offloaded kernel executes the kernel. MoE LLM inference using MoNDE begins with appropriately finding model parameters. Large-scale expert parameters can be stored in the enhanced MoNDE device memory, while other parameters can be stored in fast GPU memory. Input activation data for each expert is prepared after the MoE dispatch operation. The transmitted input activation data can be transferred from the GPU to the MoNDE memory device. After the transfer, the MoNDE device driver can generate an NDP instruction. An NDP instruction can consist of an opcode defining the operation and data specifying the physical address of each operand. Specifically, in the context of MoE FFN calculations, the physical address can be expressed using an offset and the size of the operand. The device driver can offload the NDP instruction to the MoNDE memory using a memory write request. The device driver can write the NDP instruction to the memory-mapped MoNDE instruction register. A memory write request can be converted into a CXL RwD message by the host CXL controller. The controller can examine the target address of the request. If the request targets the NDP instruction register, the controller can configure the opcode of the RwD message to one of the reserved values. This specific opcode is later examined by the MoNDE CXL controller to distinguish between standard memory requests and NDP instructions. The MoNDE device driver polls the completion register to determine whether the NDP kernel execution has completed.

[0182]

[0183] Next, the experimental setup workloads related to the present invention are described.

[0184] To evaluate the MoNDE proposed in the present invention, we use the pre-trained Switch Transformer and NLLB-MoE provided by Google and Meta in the Hugging Face repository.

[0185] The Switch Transformer model is based on the T5 model, which includes both encoder and decoder modules, but the feedforward layer is replaced with an MoE feedforward layer.

[0186] Specifically, experiments related to the present invention run 64- and 128-expert models with various batch sizes using the XSum dataset. The NLLB-MoE model is an MoE version of the NLLB-200 model, trained to translate 200 languages. Furthermore, the FLORES200 dataset, a many-to-many multilingual dataset, is used. The MoE parameters consist of weights for the experts and gating, while the non-MoE parameters include the remaining parameters, such as those of the attention and embedding layers. MoNDE can be applied to any encoder-only or decoder-only model that uses MoE, so it evaluates the performance of the encoder and decoder modules separately.

[0187] Table 2 below summarizes the evaluated workloads.

[0188]

[0189]

[0190] The software implementation is as follows.

[0191] To implement the MoNDE workflow, we use the Switch Transformers implementation provided by Hugging Face as a baseline. Because the original implementation was not well optimized for performance, we first modified the code to create a more robust baseline. For example, instead of excessively importing redundant experts as in the original code, we implemented wasteful parameter transfer, which only imports active experts to the GPU.

[0192] Additionally, to maximize overlap between asynchronous memory copies and other computational tasks, memory requests for parameter transfers are issued immediately after routing decisions are made in the dispatch phase. MoNDE operations are then implemented based on stronger criteria.

[0193] In particular, we implement a dropless and padding-free token routing algorithm similar to state-of-the-art MoE work. Furthermore, we implement active data movement by adding code to move token embedding activation data from the GPU to the CPU for expert computation and to move output token embedding activation data from the NDP to the GPU for attention computation on the GPU. Finally, we implement a MoNDE dispatch engine to offload expert computation to the NDP system.

[0194] The MoNDE NDP hardware models are as follows:

[0195] To evaluate the system-level performance of MoNDE, we implemented a cycle-level simulator. We specialized token routing statistics obtained by running the model on the context summarization task using the XSum dataset. Based on the token threshold and the number of tokens routed to each expert, the simulator outputs the cycles required for the MoNDE NDP to process the offloaded MoE expert computation.

[0196] In this invention, we obtain a detailed analysis of MoE operations and use the NVIDIA NSight profiler to isolate the latency required for MoE gating, dispatching, expert calculations, and combining operations. For NDP offload expert calculations, initially modeled as CPU calculations, only the latency of the expert calculations performed on the CPU is replaced with the latency of the NDP latency obtained from the simulator.

[0197] System level modeling is as follows.

[0198] Table 3 shows the system configuration for estimating end-to-end performance using MoNDE. We modeled the MoNDE system using a real machine and a cycle-level simulator. Since commercial CXL-supported memory is not yet available, we used CPU memory as backup memory for the MoE expert. The GPU performs the overall transformation work while retrieving expert parameters from the backup memory or offloading expert calculations to the CPU. The offloaded expert calculations are processed by the CPU, which calls the Intel MKL backend in the PyTorch function defining the expert calculation. Because we aim to replace expensive GPU resources with a less expensive NDP solution, we focused on single-GPU setups with or without MoNDE memory.

[0199]

[0200]

[0201] Below, the performance of MoNDE according to the present invention is described.

[0202] To demonstrate the effectiveness of MoNDE, we compare the following configurations. The ideal single GPU (ideal) assumes a GPU with infinite memory capacity, with all MoE and non-MoE layers residing in GPU memory. The remaining configurations offload MoE specialized storage to a memory expansion unit, which we model as CPU memory. MoE experts are computed using various compute units and policies.

[0203] Single GPU (GPM) support for PMove is a memory oversubscription scheme where active MoE experts are moved to the GPU for computation. CPUs with AMove support (CAM) leverage AMove to offload all expert computation to the CPU. GPU-CPU load balancing (C-LB) uses the MoNDE dispatch engine to jointly use the CPU and GPU for expert computation. GPU-MoNDE load balancing (MD-LB) replaces the CPU units in the C-LB scheme with MoNDE NDPs.

[0204] Figure 11 shows the end-to-end latency and throughput of MoE model inference.

[0205] GPUs with memory oversubscription suffer from performance degradation due to the large data movement penalty. Since PMove volume increases with model size, the performance degradation increases for larger models. CPU-only expert execution also suffers significantly due to its limited computational power. GPU-CPU load balancing also suffers from poor performance due to the excessive data movement required for load balancing decisions and the limited CPU computing power.

[0206] The proposed GPU-MoNDE load-balancing scheme improves upon the aforementioned issues and demonstrates superior performance across a wide range of models, achieving up to 11.1× and 7× performance improvements over the GPM scheme for the encoder and decoder, respectively. Furthermore, it achieves up to 10.7× and 5.1× performance improvements over the C-LB scheme for the encoder and decoder, respectively.

[0207] Meanwhile, the embedding dimension d model and describes the sensitivity to expert size E.

[0208] Figure 12 compares three Switch Transformers variants and demonstrates that the MoNDE scheme can provide greater performance benefits than the current memory oversubscription scheme. As the workload approaches compute limits, the encoder's returns diminish in large batches. Because many inference tasks involve small batches, the MoNDE scheme offers a performance advantage. Larger batches improve decoder performance because more experts are activated and MoNDE avoids the proportionally increasing PMove overhead.

[0209] Also, regarding the GPU-MoNDE load balancing issue, the practicality of the load balancing method of the MoNDE Dispatch Engine is explained below.

[0210] We compare static dispatching with a hard-coded threshold to determine whether an expert should be processed on the GPU or the MoNDE NDP. This static method dispatches experts with a token count greater than or equal to the threshold to the GPU, and the rest to the MoNDE NDP device. Furthermore, we vary the hard-coded threshold to several points between 0 and 64. Choosing a threshold of 0 is equivalent to GPM, and choosing a threshold of B×S is equivalent to CAM.

[0211] Figure 13 illustrates an evaluation of the Switch Transformer basic model. As illustrated in Figure 13, load balancing based on the MoNDE dispatch engine according to the present invention consistently outperforms the static threshold method across a variety of batch sizes. Furthermore, the MoNDE dispatch engine according to the present invention can establish appropriate load balancing points even under dynamic conditions, such as varying batch sizes.

[0212]

[0213] Meanwhile, Table 4 shows the area and power of the MoNDE NDP core assuming a 1 GHz clock frequency. In this invention, the MoNDE shrinker array is synthesized using Synopsys Design Compiler at a 28 nm technology node. Furthermore, the on-chip buffer is generated using a commercial memory compiler that uses the same technology. MoNDE utilizes a lightweight shrinker array to minimize the expansion of the CXL memory expander. The MoNDE NDP core occupies only a small portion of the PCIe card form factor. The PCIe card operates within a 75 watt power budget without an external power supply. The power consumption of the MoNDE NDP core remains within this range.

[0214] In conclusion, the features of the present invention are as follows.

[0215] -Multi-GPU systems.

[0216] Recent frameworks distribute expert parameters across multiple GPUs. However, multi-GPU systems are inefficient when delivering MoE models. In the gating network, each token activates only one or two experts, resulting in multiple GPUs with inactive experts remaining idle.

[0217] Figure 14 shows the performance of a single-GPU MoNDE system compared to a two-GPU system implemented with DeepSpeed. The evaluation is based on the NLLB-MoE model. The multi-GPU system performs better when multiple experts are active in the encoder module, but MoNDE maintains a similar latency level as the multi-GPU system.

[0218] -Multi-MoNDE scalability.

[0219] The present invention addresses the communication challenges of using multiple MoNDE devices for higher computational throughput. Specifically, we model a multi-MoNDE scenario using four NVIDIA 3090 GPUs connected via PCIe Gen3 ×16 lanes. Furthermore, we use the Switch Transformers baseline model for specialized parallel multi-GPU language modeling inference to model one-to-many and many-to-one AMove communications.

[0220] In this invention, we profiled the communication activity of a multi-GPU CPU-AM ​​(CAM) configuration, in which multiple GPUs simultaneously exchange token activation data with parallel CPU processes running on a single host CPU system. The evaluation method according to the invention generates PCIe traffic very similar to multi-name AMove traffic. Figure 15 shows the AMove latency of the MoE layer according to various batch sizes and the number of MoNDE devices connected to the PCIe bus.

[0221] Accordingly, we can see that communication latency increases superlinearly as the number of devices competing for the PCIe bus increases. This poses a potential threat to MoNDE scalability, as congestion on the PCIe bus can slow down one of the participating NDP devices, potentially causing a bottleneck if input activation data is not transmitted in a timely manner. Designing a scheduler for AMove to minimize local traffic contention can address this communication issue. Pipelining NDP computation by sequentially delivering a subset of token activation data can prevent NDP unit starvation.

[0222] -Cost-effectiveness.

[0223] Typically, multiple GPUs are used to fit LLM into GPU memory. This leads to significant hardware setup and maintenance costs. Since MoE expert parameters are inherently sparse, meaning they are only used conditionally, further reducing cost-effectiveness. The MoNDE system offers a cost-effective alternative by addressing data volume issues through a capacity-centric memory expansion solution that allows programs to utilize lightweight computing to avoid time-consuming data movement.

[0224] The present invention described above relates to one embodiment, which is merely an example. Those skilled in the art will appreciate that various modifications and equivalent other embodiments may be possible from this description. Therefore, the scope of the present invention is not limited by the above-described embodiment and the attached drawings.

Claims

1. CXL control unit that manages the CXL protocol; An NDP control unit having a set of memory-mapped registers for communicating with the host; A memory unit comprising a plurality of memory chips that provide memory capacity to the host; and Includes an NDP core section that processes MoE experts existing in the above memory section, The above NDP core part, An expert hybrid artificial intelligence model characterized in that only hot experts are transferred to a GPU accelerator while memory-based computation is performed on cold experts existing in the above memory unit.

2. In paragraph 1, The above CXL control unit, Offload NDP commands using CXL messages, An expert hybrid artificial intelligence model characterized by transmitting the offloaded NDP command to the NDP control unit.

3. In paragraph 1, The above NDP control unit, An expert hybrid artificial intelligence model characterized in that when the above NDP command arrives, a memory request is generated and data is loaded into the NDP core unit by tile.

4. In paragraph 3, The above NDP control unit, An expert hybrid artificial intelligence model characterized in that after a memory calculation is completed, a completion signal is generated to indicate that the calculation is completed by setting a memory-mapped completion register.

Citation Information

Patent Citations

  • Image generation using one or more neural networks

    US20220012858A1

  • Video recommendation with multi-gate mixture of experts soft actor critic

    US20220019878A1

  • Routing to expert subnetworks in mixture-of-experts neural networks

    WO2023147140A1

Cited By

  • Processor system and operation method for large-scale deep learning

    CN120448132A

  • Heterogeneous reasoning acceleration method and system for hybrid expert model

    CN121525859A