Sparse LLM data parallel processing system and method based on MLA
Through the sparse LLM data parallel processing system based on the MLA architecture, the video memory and communication bottleneck problems of large-scale language models in high concurrency scenarios are solved, and more efficient resource utilization and throughput are achieved, adapting to the dynamic load scheduling of sparse MoE models.
Patent Information
- Application Number
- CN202510378960.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-08
AI Technical Summary
The existing technology is difficult to effectively solve the problems of large-scale language models with excessive memory overhead, prominent communication bottlenecks, unbalanced expert network load and low KV cache reuse rate in high concurrency scenarios. Especially in sparse MoE models, traditional parallel strategies cannot adapt to high concurrency requirements.
The sparse LLM data parallel processing system based on the MLA architecture is adopted, and the input data is allocated to the data processing equipment with complete MLA modules through the data parallel forwarding component. Combined with hashing management and load-aware scheduling, the expert model load is dynamically adjusted, and the resource utilization and cache multiplexing rate are optimized.
It effectively avoids the memory redundancy of multi-card segmentation single-head Key/Value cache, improves inference throughput, adapts to higher concurrent requests, reduces response delay, and improves cluster resource utilization and throughput.
Smart Images

Figure CN120276850A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer information processing, and more particularly, to a data parallel processing system and method for large sparse LLMs based on MLA. Background Art
[0002] Large language model inference engines often receive a large number of user requests. Since large language model inference typically requires a large amount of computing resources, but the hardware resources are limited and cannot handle all requests simultaneously. In response to this, the large language model (LLM) DeepSeek-V3 adopts two major innovative architectures: the multi-head latent attention mechanism (MLA, Multi-head Latent Attention) and the DeepSeek MoE (DeepSeek Mixture of Experts) hybrid expert system. Its scale of 671B total parameters and 37B activation parameters poses new challenges to distributed inference. However, the continuous increase in model scale also brings higher hardware and computing overheads. Especially in the inference stage, deploying large models faces many challenges. First, the memory occupancy is huge. When the model parameters reach the scale of hundreds of billions, a single GPU (e.g., 80GB video memory) is difficult to accommodate all the parameters. In addition, intermediate results such as Key / Value caches (KV Cache) need to be saved during inference, and the video memory consumption grows almost linearly with the model scale. Second, there is an inference communication bottleneck. In a multi-GPU or multi-machine parallel environment, data exchange between nodes will significantly affect the inference speed. Especially when distributed communication operations such as All-Reduce, All-Gather, or Send-Receive need to be performed frequently, the network bandwidth becomes a limiting factor. Third, there is a high-concurrency and low-latency requirement. Many practical applications (such as online conversations, real-time translation) need to maintain a low response latency (Latency) under high-concurrency requests. If the parallel strategy is inappropriate or the communication is too frequent, the overall throughput of the system and the latency of a single request cannot be balanced. And with the rise of sparsification and mixture of experts (MoE), for the growing model parameters, a feasible idea is to adopt the "mixture of experts" (Mixture of Experts, MoE) structure, which only activates a small number of experts through dynamic routing, thereby theoretically improving the inference performance and saving computing. However, when this architecture is deployed distributively, there is a significant imbalance in the load of using the expert network, and it is difficult to achieve balanced use of the expert network.
[0003] As the order of magnitude of model parameters rapidly climbs from billions (10 9 ) to hundreds of billions (10 11 ) and even trillions (10 12), the model architecture has gradually shifted from dense (where all parameters are activated in each calculation, for example, in the traditional Transformer architecture, each input needs to be processed by all attention heads and feed-forward networks) to sparse (where only some parameters are dynamically activated, for example, in MoE (Mixture of Experts), the input is routed to a small number of expert networks according to a routing strategy, and the rest of the parameters remain "dormant"). As a result, the deployment difficulty and computational pressure faced during the inference phase have increased sharply.
[0004] Therefore, large language models often use multi-head attention (Multi-Head Attention, MHA) to capture and model context information. Traditional tensor parallelism strategies generally rely on "splitting multi-head attention along the head dimension". If the number of attention heads is sufficient, the computational and storage pressure can be effectively shared after splitting. However, in the multi-head latent attention (MLA, Multi-head Latent Attention) proposed by DeepSeek-V3, only one "latent head" is retained on the Key / Value side, that is, both Key and Value are in single-head form. Therefore, in the MLA scenario, since there are no multiple heads on the Key / Value side that can be split, continuing to split along the head dimension becomes meaningless. If tensor parallelism is forced, it often leads to the repeated storage of the Key / Value Cache in each parallel card, wasting video memory and incurring unnecessary communication synchronization overhead.
[0005] Therefore, the existing mainstream parallelism strategies for large-scale language model inference (such as simply relying on tensor parallelism TP and pipeline parallelism PP) are no longer sufficient to meet the requirements of ultra-large-scale or sparse (MoE) models in high-concurrency scenarios, specifically manifested in excessive video memory overhead, prominent communication bottlenecks, unbalanced hot-spot expert loads, inability to fully reuse KV caches, and lack of adaptive scheduling.
[0006] Therefore, it is expected that in the MLA architecture, the defect of video memory redundancy caused by splitting the single-head Key / Value cache among multiple computing devices (GPU cards) can be avoided, and the additional storage of duplicate information can be eliminated, enabling it to be applied to more scenarios with GPU card resources.
[0007] The above information disclosed in the background art section is only used to enhance the understanding of the background of the present disclosure, and thus it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0008] In view of this, the present disclosure provides a data processing system and method for large models, especially for LLMs based on the multi-head latent attention (MLA) architecture. Through data parallelism, the present disclosure effectively avoids the video memory redundancy caused by splitting a single-head Key / Value cache across multiple cards, while improving the inference throughput and adapting to a higher number of concurrent requests.
[0009] Other features and advantages of the present disclosure will become apparent from the following detailed description, or will be learned in part through the practice of the present disclosure.
[0010] According to one aspect of the present disclosure, a data parallel processing system for a sparse LLM based on MLA is provided, including: an input data splitting component for splitting input data; a data parallel forwarding component for parallelly distributing the sliced micro-batch data to data processing devices each having a complete MLA module; and a plurality of parallel data processing devices, each data processing device using a complete set of attention weights and cache structure based on the complete MLA module it owns, and performing attention calculation processing on the obtained micro-batch data locally using locally generated KV values.
[0011] The data parallel processing system for a sparse LLM based on MLA according to the present disclosure further includes: a request parsing component for obtaining token information of the input data for a user's input data; a subset splitting component for, based on a predetermined number of tokens as a base, taking each data processing request in the input data as a unit, and each time dividing out a token subset starting from the head of the token information of the received data processing request until the end of the token information, and generating hash values for all token subsets using a hash function to form one or more hash value sets of the data processing request; a cache status tracking component for, for each destination data processing device, caching the hash value set of each data processing request forwarded to it by the data parallel forwarding component and the load of the data processing requests being processed by each destination data processing device; a scheduling score calculation component for scoring each destination data processing device based on the matching degree between the hash value set corresponding to the current data processing request and the hash value sets corresponding to the data processing requests received by the cache status tracking component for each destination data processing device and the load of each destination data processing device, so that the data parallel forwarding component routes and forwards the current data processing request to the destination data processing device with the highest score or multiple destination data processing devices with the highest scores.
[0012] The data parallel processing system for MLA-based sparse LLM according to the present disclosure further includes: an expert model load monitoring component, which, for each expert model, in the inference stage, statistically monitors in real time the load of the data processing requests being processed by the expert model selected by the data parallel forwarding component; an expert model routing adjustment component, which, when the current load of any expert model exceeds a predetermined load threshold and is determined as a hot expert model, allocates reserved data processing devices for the hot expert model, or migrates the entire hot expert model from the current data processing device to the reserved data processing device, so that the parallel data processing device can parallelly forward the new input data requests that should be selected to the hot expert to the reserved data processing device.
[0013] The data parallel processing system for MLA-based sparse LLM according to the present disclosure, wherein the expert model routing adjustment component statistically monitors the current loads of all expert models, calculates the weights of the current loads of each expert model in the overall load, and allocates reserved data processing devices to a predetermined number of the highest-ranked expert models.
[0014] The data parallel processing system for MLA-based sparse LLM according to the present disclosure further includes: an expert model load monitoring component, which, for each expert model, in the inference stage, statistically monitors in real time the load of the data processing requests being processed by the expert model selected by the data parallel forwarding component; an expert model routing adjustment component, which, when the current load of any expert model exceeds a predetermined load threshold and is determined as a hot expert model, reallocates data processing devices with a larger capacity or more data processing devices for the hot expert model, and migrates the entire hot expert model from the current data processing device to the newly allocated and mapped data processing device, so that the parallel data processing device can parallelly forward the new input data requests that should be selected to the hot expert to the newly allocated and mapped data processing device.
[0015] According to another aspect of the present disclosure, there is also provided a data parallel processing method for MLA-based sparse LLM, including: splitting input data; parallelly allocating the segmented micro-batch data to data processing devices each having a complete MLA module; and through each data processing device, using a complete set of attention weights and cache structures possessed based on the complete MLA module it owns, performing attention calculation processing on the obtained micro-batch data locally using locally generated KV values.
[0016] The data parallel processing method for MLA-based sparse LLM according to the present disclosure further includes: before splitting the input data, for the user's input data, obtaining the token information of the input data; using a predetermined number of tokens as the base, taking each data processing request in the input data as a unit, each time dividing a token subset starting from the head of the token information of the received data processing request until the end of the token information is reached, and using a hash function to perform hashing on all token subsets to generate the hash value of each divided token subset, thereby forming one or more hash value sets of the data processing request; when the micro-batch data is parallelly allocated, for each destination data processing device, caching the hash value set of each data processing request forwarded by the data parallel forwarding component to it and the load of the data processing requests being processed by each destination data processing device; and based on the matching degree between the hash value set corresponding to the current data processing request and the hash value sets corresponding to the data processing requests received by each destination data processing device collected for each destination data processing device and the load of each destination data processing device, scoring each destination data processing device, so that the data parallel forwarding component routes and forwards the current data processing request to the destination data processing device with the highest score or one of the multiple highest scores.
[0017] The data parallel processing method for MLA-based sparse LLM according to the present disclosure further includes: for each expert model, during the inference stage, in real time, statistically calculating the load of the data processing requests being processed by the expert model selected by the data parallel forwarding component; and when the current load of any expert model exceeds a predetermined load threshold and is determined to be a hot expert model, allocating a reserved data processing device to the hot expert model, or migrating the entire hot expert model from the current data processing device to the reserved data processing device, so that the parallel data processing device can parallelly forward the new input data requests that should be selected to the hot expert to the reserved data processing device.
[0018] The data parallel processing method for the MLA architecture according to the present disclosure further includes: calculating the weight of the current load of each expert model in the overall load, and allocating reserved data processing devices to a predetermined number of the highest-ranked expert models.
[0019] The data parallel processing method based on MLA-based sparse LLM according to the present disclosure further includes: for each expert model, during the inference phase, the load of the data processing requests being processed by the expert model selected by the data parallel forwarding component is statistically calculated in real time; when the current load of any expert model exceeds a predetermined load threshold and is determined as a hot expert model, a data processing device with a larger capacity or more data processing devices is reallocated to the hot expert model, and the hot expert model as a whole is migrated from the current data processing device to the newly allocated and mapped data processing device, so that the parallel data processing device can parallelly forward the new input data requests that should be selected to the hot expert to the newly allocated and mapped data processing device.
[0020] In the data parallel processing system and method based on MLA-based sparse LLM according to the present disclosure, since each data processing device (GPU card or other computing card) only needs to store a set of MLA weights in the data parallel mode and does not need to store redundant information additionally, the video memory redundancy caused by multi-card splitting of a single-head Key / Value cache can be effectively avoided. At the same time, for scenarios with more GPU resources, only by expanding the parallel replicas in the DP dimension, the inference throughput can be improved simultaneously to adapt to a higher number of concurrent requests. Furthermore, due to the adoption of the load cache awareness means, the hashing management and monitoring of the user request context (Prompt or historical conversation content) are introduced, so that when the data parallel forwarding component distributes requests to each data parallel replica, it can preferentially select the GPU that already stores the same or similar context. More importantly, the dynamic expert parallelism of the present disclosure periodically (or in real time) monitors the load information of all experts during the inference phase, and adaptively adjusts the mapping relationship between experts and devices (or the routing of some Tokens) according to the actual distribution of expert loads, thereby improving the resource utilization rate of the cluster and alleviating the overload phenomenon of hot cards.
[0021] It should be understood that the above general description and the following detailed description are only exemplary and do not limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] By referring to the accompanying drawings and describing its exemplary embodiments in detail, the above and other objectives, features, and advantages of the present disclosure will become more apparent. The following described drawings are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0023] Figure 1 It is a block diagram of a first embodiment of a data parallel processing system based on MLA-based sparse LLM shown according to an exemplary embodiment.
[0024] Figure 2It is a block diagram of a second embodiment of a data parallel processing system for sparse LLM based on MLA shown according to an exemplary embodiment.
[0025] Figure 3 It is a block diagram of a third embodiment of a data parallel processing system for sparse LLM based on MLA shown according to an exemplary embodiment. Detailed implementation manners
[0026] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the concept of the example embodiments to those skilled in the art. Like reference numerals in the figures denote like or similar parts, and thus their repetitive description will be omitted.
[0027] In addition, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present disclosure. However, those skilled in the art will realize that the technical solutions of the present disclosure can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. can be adopted. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present disclosure.
[0028] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0029] The flowcharts shown in the drawings are only illustrative and do not necessarily include all the contents and operations / steps, nor are they necessarily executed in the order described. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined, so the actual execution order may change according to the actual situation.
[0030] It should be understood that although terms such as first, second, and third may be used herein to describe various components, these components should not be limited by these terms. These terms are used to distinguish one component from another. Thus, the first computing device discussed below can be referred to as the second computing device without departing from the teachings of the concept of the present disclosure. As used herein, the term "and / or" includes any one of the associated listed items and all combinations of one or more of them.
[0031] Those skilled in the art can understand that the accompanying drawings are only schematic diagrams of exemplary embodiments, and the modules or processes in the drawings are not necessarily essential for implementing the present disclosure. Therefore, they cannot be used to limit the protection scope of the present disclosure.
[0032] Figure 1 is a block diagram of a first embodiment of a data parallel processing system for sparse LLM based on MLA shown according to an exemplary embodiment. As Figure 1 shown, the data parallel processing system 100 for sparse LLM based on MLA includes: an input data splitting component 110, a data parallel forwarding component 120, and multiple parallel data processing devices 130. The input data splitting component 110 splits the input data. The data parallel forwarding component 120 distributes the split micro-batch data in parallel to the data processing devices each having a complete MLA module. Each data processing device 130 uses a complete set of attention weights and cache structures it has based on the complete MLA module it owns, and uses the locally generated KV values to perform attention calculation processing on the obtained micro-batch data locally. Since in multi-head latent attention (MLA), only one "latent head" is retained on the Key / Value side, that is, both Key and Value are in single-head form. Therefore, in the MLA scenario, the present disclosure deploys the MLA attention part to each data processing device, thereby enabling data parallelism.
[0033] In a data parallel architecture, since there is a complete Key / Value Cache locally, each GPU (or device) has a complete copy of the attention weights and cache structure, eliminating the need for Key / Value splitting and synchronization among multiple GPUs. This avoids the video memory duplication problem caused by having only a single head in the MLA and enables the local reuse of the Key / Value Cache in long sequence generation or multi-turn conversations, reducing frequent cross-card communication. Therefore, the input data is split through data parallelism, and different user requests or different batches of Token sequences are assigned to GPU replicas each with a complete MLA module. The particularity of traditional attention mechanisms exacerbates the distributed communication burden. Most traditional large models use the Transformer framework, and the multi-head attention (MHA) in it can usually be tensor-parallel split along the "head dimension". However, for the multi-head latent attention MLA, which has a single Key / Value head, continuing to use tensor parallelism to split along the head dimension will be ineffective, resulting in the Key / Value Cache having to be redundantly stored on all participating parallel cards, wasting video memory and causing frequent broadcasts / synchronizations. In contrast, the solution of the present disclosure does not require tensor parallel splitting along the "head dimension". Instead, after a certain GPU receives a request, it can directly perform attention calculations locally, maximizing the utilization of the Key / Value already stored locally. The data parallel processing system integrating MLA of the present disclosure effectively avoids the video memory redundancy caused by splitting the single-head Key / Value cache across multiple cards. Since each card only needs to store a set of MLA weights in the data parallel processing mode, there is no need to store duplicate information additionally. In addition, for scenarios with more GPU resources, only by expanding the parallel replicas in the data parallel dimension can the inference throughput be increased simultaneously to adapt to a higher number of concurrent requests. It should be noted that the MLA module can be implemented in an existing manner and will not be elaborated here.
[0034] Figure 2 It is a block diagram of a second embodiment of a data parallel processing system 200 for a sparse LLM based on MLA shown according to an exemplary embodiment. Compared with Figure 1 the data parallel processing system 100 for a sparse LLM based on MLA shown in a certain figure, the reference numerals of the same parts are similar to Figure 1 those in it, and the only difference is that they start with the number "2". Therefore, the description of the same parts adopts the description for Figure 1 it and will not be repeated here.
[0035] In high-concurrency inference scenarios, a large number of user requests arrive simultaneously. The low cache reuse efficiency in high-concurrency scenarios For online inference scenarios (such as chatbots and multi-turn dialogue systems), it is often necessary to reuse the KV cache among a large number of requests arriving simultaneously to reduce repeated calculations and the overhead during the prefill stage. However, without fine-grained scheduling aware of the cache (Cache-Aware), many requests may be scattered and executed on different nodes, resulting in the inability to reuse the cache with the same or similar context; even if there is available cache on the same node, it may not be effectively utilized due to unreasonable scheduling. When a large number of user requests arrive, if the scheduling strategy only considers the current load of each node or simple round-robin polling, it is very likely that the Key / Value Cache resources generated by the MLA module cannot be reused by subsequent requests with the same context. Especially during the prefill stage, if a large amount of existing context is repeatedly calculated, it will significantly lengthen the response latency and reduce the system throughput. Therefore, the data parallel processing system 200 based on MLA for sparse LLM of the present disclosure adopts a cache-aware scheduling strategy, introducing hash management and monitoring of the user request context (Prompt or historical dialogue content), so that when the data parallel forwarding component 220 distributes requests to each copy of the data parallelism, it can preferentially select the GPU that already stores the same or similar context.
[0036] Therefore, as Figure 2 shown, the data parallel processing system 200 based on MLA for sparse LLM includes: an input data splitting component 210, a data parallel forwarding component 220, a plurality of parallel data processing devices 230, a request parsing component 240, a subset splitting component 250, a cache status tracking component 260, and a scheduling score calculation component 270. First, after the system receives the high-concurrency input data from the user, the request parsing component 240 obtains the token information of the input data for the user's input data. As Figure 2 shown, in a large language model (LLM), a token is the basic unit of text, which can be a word, subword, or character. The model decomposes the input text into tokens for processing and understanding. The number of tokens determines the input length and computational complexity of the model, and the text data input by the user will first be converted into tokens and input into the model. The request parsing component 240 parses the current data processing request into one or more tokens. This facilitates the input data splitting component 210 to allocate different user requests or different batches of Token sequences to the GPU copies each having a complete MLA module.
[0037] Subsequently, the subset splitting component 250 uses a predetermined number of tokens as the basis and takes each data processing request in the input data as a unit. Each time, a token subset is divided from the head of the token information of the received data processing request until the end of the token information. Then, the hash function is used to hash all the token subsets to generate the hash value of each divided token subset, thereby forming one or more hash value sets of the data processing request. Usually, Prompt hashing calculation is performed to segment-hash all the Tokens in the user request. For example, a hash value is generated for every 32 Tokens to identify whether shorter prefixes or medium-length prefixes are the same.
[0038] Specifically, after receiving the user request, the context (Prompt) Tokens of the request are extracted and hashing calculation is performed (multiple lengths of sub-prefix hashes can be calculated at one time). Assume that prompt_tokens represents the Token sequence of the user request (with a length of L), and the system internally maintains an ordered list representing the sequentially increasing "truncation" boundaries (usually doubling, such as 64, 128, 256,..., and not exceeding the maximum input length allowed by the system). For each if then the first tokens are hashed once to obtain a hash value . Therefore, if a total of M valid hashes are generated, they can be recorded as where , and each hash value For example, in one segmentation method, the subset segmentation component 250 divides the tokens converted from the current data processing request according to the number of tokens per page of page attention, and draws the segmentation lines according to the rule of doubling each time. Suppose the number of tokens per page in the system is 32 (the base number), and the number of tokens in a request is 555. Then the segmentation lines are [32, 64, 128, 256, 512]. Usually, the small number of tokens remaining after 555 compared to 512 are ignored because this will not affect the accuracy of the final pre-filling or decoding. In this way, the first token subset is 32 tokens, the second token subset is 64 tokens, which includes the 32 tokens in the first token subset, and so on. In this way, the token subsets divided each time start from the 0th token of the request and are respectively tokens[0:32], tokens[0:64], tokens[0:128], tokens[0:256], tokens[0:512]. Tokens[0:i] represents the tokens from the 0th token to the ith token. For these token subsets (or token pages), 5 hash values can be converted through a fixed hash function. The advantage of this division is to maintain the overall linear complexity. For example, the complexity above is only 32 + 64 + 128 + 256 + 512 = 992 <= 2 * 555. Optionally, it is also possible to perform segmentation once per page, that is, with a predetermined number of tokens as the base number, and each time a token subset is divided starting from the head of the current data processing request. In two adjacent segmentations, the number of the token subset divided in the latter time increases by one base number compared to the number of the token subset divided in the former time. For example, 32, 64, 96, 128, 160,..., 512, 544. However, its complexity will reach n^2: 32 + 64 + 96 + 128 + 160 +... + 512 + 544 = 32 * (1 + 17) * 17 / 2. Here, 17 or n will increase linearly as the total number of tokens increases, so the complexity is quadratic.
[0039] The cache status tracking component 260 caches, for each destination data processing device, the set of hash values of each data processing request forwarded to it by the data parallel forwarding component and the load of the data processing requests being processed by each destination data processing device. Specifically, a "stored hash table" is maintained on each GPU (DP replica) to record the context hash values corresponding to those in its Key / Value Cache. Optionally, this "stored hash table" can be maintained at each data processing device (GPU (DP replica)) or deployed at the CPU. The cache status tracking component 260 collects and caches, for each destination data processing device, the set of hash values of each data processing request forwarded to it by the data parallel forwarding component 220 and the load of the data processing requests being processed by each destination data processing device. The cache status tracking component 260 obtains the load information and cache hash table of each GPU from the global data parallel forwarding component 220 or the distributed data parallel forwarding component 220. That is, when the data parallel forwarding component 220 forwards each data processing request, the cache status tracking component 260 obtains the ID of the data processing device that is the forwarding destination and caches the set of hash values of the data processing request being forwarded. Accordingly, it also knows how many data processing requests the data parallel forwarding component 220 has forwarded to each data processing request. That is, for each data processing device acting as a destination, a data structure is established to cache the relevant information of the data processing requests it receives. The data structure used for this information only needs to be able to record the hash values of the segmented token subsets of the data requests that have been received in the input order. For example, the cache status tracking component 260 uses a dict and a FIFO list to save the hash values of the existing requests from 0 to the dividing line for each data processing device used for data parallel inference. Therefore, the cache status tracking component 260 allocates memory of a predetermined length for each destination data processing device and, when the memory is full, replaces the hash values cached in the memory in a first-in, first-out manner.
[0040] The hash value in the dict here is somewhat like the index of the subset of tokens corresponding to the hash value. Or rather, the Dict is a mapping from this hash value to the number of occurrences of this hash value in the data processing device. The FIFO list is First In First Out (in fact, here First In First Out means that the oldest one is popped first). Since the number of pages is limited, after the cache is used up, new requests will push out the caches of old requests. So the FIFO list also has a size. Before the number of hashes stored reaches the maximum value of the FIFO list, the FIFO list will store all incoming requests' hash values according to the delimiter. Each request may store multiple hash values, which is determined by the number of delimiters, that is, the number of hash values, according to the length of the tokens. After the number of hashes stored reaches the maximum value of the FIFO list, the FIFO list will pop the oldest hash value each time and then put a new hash value in. Repeat this multiple times. The multiple hash values of a request will be input in ascending order of the number of pages. So there is also an order for the multiple hash values of the same request.
[0041] Subsequently, the scheduling score calculation component 270 scores each destination data processing device based on the matching degree between the hash value set corresponding to the current data processing request and the hash value sets corresponding to the data processing requests received by the cache status tracking component for each destination data processing device, and the load of each destination data processing device, so that the data parallel forwarding component 220 routes and forwards the current data processing request to the destination data processing device with the highest score or one of the multiple highest-scoring destination data processing devices. That is, when the data parallel forwarding component 220 selects a target GPU for a new request, it will comprehensively consider the real-time load of the GPU (such as the number of requests being processed currently, the number of Tokens to be generated) and the context hash matching degree with this request. If a certain GPU already contains the same hash, a large score increase can be given, making it more likely to be selected as the target node.
[0042] Specifically, the scheduling score calculation component 270 calculates the "comprehensive score" for each candidate GPU (in fact, all data processing devices 230 or GPU cards deployed distributively). The smaller the load, the higher the score, and the more cache hits, the higher the score. If there are no hits, the selection is made only based on the load information. Specifically, the score consists of two parts: the load balancing factor and the cache hit bonus (matching degree). The load balancing factor represents the number of requests that have been accumulated on the current node, which has a negative impact on the "comprehensive score", that is, the more requests, the lower the score. It can be expressed by the following mathematical formula: where represents the number of requests being processed by the d-th node. A weighted load metric can also be used, such as multiplying the number of requests by a coefficient of 0.8. Cache hit bonus means that if the current node has cached some or all of the prompt word hashes (indicating that the same context has been processed before), an additional bonus will be given to this node. If we let be the total hash of this request, let represent the set of hashes stored by node d, and let represent the "bonus weight" at the j-th segment of the corresponding hash, then the cache hit bonus can be written as: where is an indicator function, which means that if is in the cache of node d, the corresponding score will be added , otherwise no bonus is added.
[0043] Combining the above two items, the total score (Score) of node d can be expressed as: where is the current number of requests of the node, indicates whether the j-th segment hash exists in the cache of node d; represents the "bonus weight" of the j-th segment hash.
[0044] Subsequently, the data parallel forwarding component 220 routes and forwards the current data processing request to one destination data processing device with the highest score or multiple highest scores. That is, the request is assigned to the GPU with the highest comprehensive score, or in the case of multiple highest comprehensive scores, the request is assigned to the GPU with the highest matching degree or the highest load score. Selecting the GPU with the highest matching degree can reduce the amount of computation, and selecting the GPU with the highest load score means that the selected GPU has the smallest existing load and the most available computing power resources. After the selection and forwarding, the hash table of this GPU can be updated. As Figure 2 shown, the data processing device 230 feeds back cache information to the cache status tracking component 260. On the one hand, it can update the local hash table of this GPU, and it can also feed back the updated data to the cache status tracking component 260. If we let the set D represent all available node sets, then the node finally selected by the data parallel forwarding component 220 After the data processing request is forwarded, the data processing device 230 (GPU) side performs attention forward inference (especially the MLA part can directly use the existing Key / Value Cache). After completion, if there are still subsequent Token generations or subsequent similar requests, the cache state can be reused to continue accelerating.
[0045] Based on the above scoring mechanism, the scheduling strategy based on cache state tracking can, in a high-concurrency distributed inference scenario, the more requests received by a certain GPU, the lower its score, avoiding further task accumulation therein. At the same time, in the case where the same or similar context exists in a certain GPU node, the matching degree score is greatly increased, making it more likely to be selected as the destination data processing device, avoiding duplicate calculations in the entire system. Thereby effectively improving the KV cache reuse rate and reducing the duplicate computing power overhead in a large-scale cluster, while also preventing a certain node from being "overloaded alone" resulting in an increase in response latency. The present disclosure can particularly significantly improve the inference throughput and reduce the overall latency in multi-round dialogue and repeated context scenarios.
[0046] Figure 3 is a block diagram of a third embodiment of a data parallel processing system based on MLA-based sparse LLM shown according to an exemplary embodiment. Compared with Figure 2 the shown data parallel processing system 200 based on MLA-based sparse LLM, the reference numerals of the same parts are similar to Figure 2 those in, and the difference is only that it starts with the number "3". Therefore, the description of the same parts adopts the description for Figure 2 and will not be repeated here one by one.
[0047] The sparse LLM based on MLA is also a mixture of experts (MoE) model containing multiple sparse experts (Experts). The load of traditional MoE parallel is unbalanced. Due to its parameter sparseness in large models, the mixture of experts (MoE) can reduce the actual computing overhead while maintaining a large model capacity. However, when the scale of the MoE model continues to climb to tens of billions or hundreds of billions of parameters, the difficulty of traditional tensor parallel or pipeline parallel processing of the expert layer increases significantly. If tensor parallelism is used to split the weights within each expert, the communication volume increases dramatically with the increase of parallelism, and the actual computing load of a single card is smaller, resulting in a decrease in overall efficiency. The routing of requests by each expert during reasoning may be uneven, and it is very easy to have "hotspot experts", that is, the number of visits to a few experts far exceeds that of other experts, causing load imbalance and computing bottlenecks. That is, in the traditional mixture of experts (MoE) model, each expert is often fixedly deployed on a certain card (or a certain device group). In this way, when the request distribution is uneven, "hotspot experts" are prone to appear: some experts are frequently accessed, resulting in serious overload of the card where the hotspot experts are located, while other experts (or cards) are idle. This will cause problems such as a decrease in overall throughput and a sharp increase in long-tail latency. In addition, the traditional hybrid expert (MoE) lacks a dynamic scheduling strategy at the expert level. In the MoE model, although only a small number of experts need to be activated for a single reasoning, the actual business traffic is usually diverse and dynamic. Some time periods may be highly concentrated on a few hotspot experts, while other time periods are scattered on multiple experts. If the system only uses a fixed expert mapping (for example, "several experts per card"), when the request distribution changes, the load will be extremely unbalanced between experts, which will significantly reduce the system throughput and hardware utilization. The lack of "dynamic hybrid expert" scheduling for reasoning scenarios makes it difficult for large-scale MoEs to obtain stable and excellent reasoning performance in real applications. To this end, the MLA-based sparse LLM data parallel processing system 300 disclosed in the present invention proposes a "dynamic expert parallel" strategy, that is, by periodically (or in real time) monitoring the load information of all experts in the inference phase, and adaptively adjusting the mapping relationship from experts to devices (or the routing of some tokens) according to the actual distribution of expert loads, thereby improving the resource utilization of the cluster and alleviating the overload of hotspot cards.
[0048] Specifically, the data parallel processing system 300 for sparse LLM based on MLA includes: an input data splitting component 310, a data parallel forwarding component 320, multiple parallel data processing devices 330, a request parsing component 340, a subset splitting component 350, a cache status tracking component 360, a scheduling score calculation component 370, an expert model load monitoring component 380, and an expert model routing adjustment component 390.
[0049] As Figure 3 shown, for each expert model, the expert model load monitoring component 380 statistically calculates the load of data processing requests being processed by the expert model selected by the data parallel forwarding component in real time during the inference phase. In a distributed inference scenario, each input Token (or mini-batch of Tokens) is routed to several experts for calculation through a gating function. In this disclosure, gating operators such as MoEGate or other gating operators are used to determine the "Token → expert" routing, and the access volume or the number of Tokens processed by each expert is statistically calculated in real time. Other operators can also be used to determine the "Token → expert" routing, which belongs to common routing determination operators, so they will not be elaborated one by one. This disclosure can adopt a conventional method for determining the "Token → expert" routing, but the method itself is not the focus of this disclosure. In a specific implementation, the hidden states will output the expert number assigned to each token and the global count through the gating function, and then by combining the all_gather or all_to_all operations between multiple cards, the cumulative load information of all experts can be obtained at each node.
[0050] Specifically, if represents the number of Tokens processed by the i-th expert, then the "hotspot weight" can be obtained through normalization or other statistical methods: When the hotspot weights of some experts are much higher than the average level or exceed a predefined threshold, it is considered that there is a "hotspot" phenomenon for these experts, and they are determined as "hotspot experts", and additional resource migration or routing adjustment is required.
[0051] For "hotspot experts", task sharing can be achieved by using the data processing devices reserved by the distributed deployment data processing system. Specifically, when the current load of any expert model exceeds the predetermined load threshold and is determined to be a hotspot expert model, the expert model routing adjustment component 390 allocates reserved data processing devices to the hotspot expert model, or migrates the entire hotspot expert model from the current data processing device to the reserved data processing device, so that the parallel data processing device can forward the new input data requests that should be selected to the hotspot expert to the reserved data processing device in parallel. Generally speaking, a certain number of idle devices (or cards) are reserved for "potential hotspot experts" during deployment. When it is detected that the load of a certain expert exceeds the threshold, some of its requests are allocated to this reserved device for operation, or the parameters of this expert are also copied / migrated to the reserved device, achieving the effect of elastic expansion. This way of elastic expansion is also called "partial dynamic allocation", that is, the deployment locations of some expert models are elastic or flexible to meet more expert inferences. For example, the reserved expert pool is denoted as , if it is monitored that a certain expert exceeds the threshold τ (for example , where is the average load of all experts, and α>1 is the magnification factor), then let the idle expert device in be allocated to the hotspot expert e. Then, the routing table is updated, and the expert e is assigned to both the original node and the reserved node in the routing table (copying or sharding can be performed as needed, that is, an expert parallel strategy, where the tasks assigned to the same expert are parallelly assigned to multiple GPUs, or different parts of the same GPU), and during the gated output stage, Token routing is preferentially performed according to the "least idle copy". That is, between the GPU copies originally allocated to the hotspot expert and the reserved GPU copies, the loads are compared, and the newly arrived input requests are assigned to the GPU copy with the least load among the multiple GPU copies corresponding to the same expert.
[0052] Optionally, when the current load of any expert model exceeds a predetermined load threshold and is determined to be a hot expert model, the expert model routing adjustment component 390 reallocates a data processing device with a larger capacity or more data processing devices for the hot expert model, and migrates the entire hot expert model from the current data processing device to the newly allocated and mapped data processing device, so that the parallel data processing device can forward the new input data requests that should be selected to the hot expert to the newly allocated and mapped data processing device in parallel. This method is called "full dynamic allocation", that is, for all experts, overall dynamic allocation is performed. Specifically, in a distributed deployment data processing system, the entire system does not pre-fix any mapping relationship between experts and GPU replicas, but updates the "expert index → device mapping" based on real-time statistical load data. If the expert model routing adjustment component 390 determines that some experts are often under high load, more cards are allocated to them or re-routing adjustment is performed to achieve global load balancing. This strategy does not preset any fixed expert-device mapping, but continuously monitors during inference . Set a scheduling period T (for example, triggered once every 1000 - 5000 Tokens processed). For each expert e, calculate its current normalized load According to Sort and allocate additional resources to the top several experts, and recycle the device mapping of the lowest experts (if not necessary, only a single replica can be retained) to ensure overall load balance. After the scheduling is completed, each card broadcasts the updated routing table to ensure consistent address mapping for subsequent Tokens to see.
[0053] In summary, the present disclosure maintains a "routing table" for each parallel node (for example, each data processing device or GPU card), recording the distribution of all current experts. At the same time, for each expert e, the routing table gives the identifier of the device group d where it is located (for example, GPU number or process group ID). When a "hot expert" is detected, the expert model routing adjustment component 390 migrates (replicate / redistribute) the corresponding expert (or a part of its parameters) to a new idle node. Subsequently, the routing table is updated so that the gating function can distribute the computational load of this expert among multiple nodes during subsequent Token routing. In addition, in addition to the traditional "token → expert" routing, the present disclosure can also automatically add several "shared experts" to each Token on the "token → shared_expert" path, while not destroying the original sparsity, and distributing the computational load as much as possible.
[0054] The present disclosure can significantly alleviate the computing power imbalance problem caused by hot experts through the above-mentioned "dynamic expert parallelism". Compared with the traditional fixed expert allocation or the method of using tensor parallelism for experts, it can periodically or in real time disperse the calculations of overloaded experts to idle or reserved resources, avoid some cards from becoming bottlenecks, and achieve load balancing. Moreover, the present disclosure increases or decreases the expert replicas as needed. Especially in the "fully dynamic allocation" mode, it can flexibly adapt to the frequent changes in the request pattern and achieve elastic expansion. In addition, although the present disclosure will increase certain All-to-All or All-Gather operations, through the carefully designed routing table and local replication strategy, it ensures scalability in a large-scale cluster environment and realizes controllable communication overhead. Further, the present disclosure cannot scale the computing cards by using tensor parallelism because this will continuously reduce the overall efficiency, while using this strategy can expand the computing resources by hundreds of times and at the same time achieve a linear increase in efficiency, realizing scale scalability.
[0055] The following content is the test results and analysis of the present disclosure under different parallel strategies. The test mainly compared various parallel combination schemes such as the traditional TP8PP2 (2 machines, baseline) and the DP4TP8+EP8ETP4 (4 machines) proposed by the present disclosure, and measured and compared the inference performance indicators under different concurrency numbers, input lengths, and output lengths. The main measurement indicators are given in the table: - Concurrency number: The number of requests arriving simultaneously at the same moment; - Input length: The number of tokens in an input for one request (such as the length of the Prompt); - Output length: The maximum number of tokens that the model needs to generate; - TTFT (Time To First Token): The average latency (seconds) required from the arrival of the request to the generation of the first token; - TPOT (Time Per Output Token): The average time (seconds per token) required from the start of generation to each output token; - Total throughput (which can be understood as the total number of tokens generated or the cumulative processing volume during the measurement process); - TPS (Token Per Second): The number of tokens that can be generated per second on average, reflecting the overall throughput rate of the system.
[0056] It can be seen from the table that there are obvious differences in the inference performance indicators of the two main parallel schemes - TP8PP2 (2 machines, baseline) and DP4TP8+EP8ETP4 (4 machines) - under different concurrency scenarios: Single concurrency scenario (Concurrency=1) For TP8PP2 (2 machines), when concurrency = 1, input length = 128, and output length = 1024, TTFT ≈ 0.173 seconds, TPOT ≈ 0.025 seconds / token, and TPS ≈ 40.
[0057] Under the same conditions, DP4TP8+EP8ETP4 (4 machines) has TTFT≈0.193 seconds, TPOT≈0.032 seconds / token, and TPS≈31.25. It can be seen that the baseline solution (2 machines) has more advantages in first token latency and generation speed under single concurrency; the present disclosure (4 machines) has relatively greater communication and scheduling overheads due to the fact that it spans more machines, resulting in a disadvantage in the single request scenario.
[0058] High concurrency scenarios (Concurrency ≥ 32) When the number of concurrent connections increases, the disclosed solution (4 machines) shows obvious scalability advantages: Take concurrency = 32, input = 128, output = 1024 as an example: the baseline (2 machines) TPS ≈ 15.87; the disclosed (4 machines) TPS ≈ 25, which is an improvement of more than 50%.
[0059] When the number of concurrency continues to increase to 64, 128, and 256, the TPS of the present invention is always significantly higher than the baseline; and when the concurrency = 256, the TPS of the baseline is only ≈6.76, while the present invention can reach ≈14.49, and the gap is further widened.
[0060] This disclosure also supports concurrency = 512 (the baseline is not tested in this concurrency), which shows that it can still maintain a relatively impressive TPS (≈10.10) in large-scale concurrency scenarios.
[0061] The impact of input length on this disclosure When the input length increases from 128 to 2048, as the Prompt length increases, TTFT increases to varying degrees, and TPOT also increases; however, the overall TPS still maintains a high level under high concurrency.
[0062] When concurrency = 32, TPS ≈ 25 when input = 128 drops to TPS ≈ 21.28 when input = 2048; When the concurrency is higher, such as 64, 128, 256, etc., there is also a certain decrease, but the generation efficiency is still maintained at a higher level.
[0063] In summary, the mainstream parallel strategies for traditional large-scale language model inference (such as simply relying on tensor parallelism TP and pipeline parallelism PP) are difficult to handle ultra-large-scale or sparse (MoE) models in scenarios with high concurrency, excessive video memory overhead, prominent communication bottlenecks, unbalanced hot spot expert loads, inability to fully reuse KV caches, and lack of adaptive scheduling. The systems and methods of the present disclosure propose a DP+EP hybrid parallel inference method for trillion-scale sparse large language models, aiming to achieve significant optimizations in terms of resource utilization, communication efficiency, and inference latency.
[0064] Those skilled in the art can understand that the above-mentioned modules can be distributed in the device according to the description of the embodiments, or can be correspondingly changed in one or more devices that are uniquely different from this embodiment. The modules of the above embodiments can be combined into one module, or further split into multiple sub-modules.
[0065] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described here can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, USB flash drive, mobile hard disk, etc.) or on the network, including several instructions to enable a computing device (which can be a personal computer, server, mobile terminal, or network device, etc.) to execute the methods according to the embodiments of the present disclosure.
[0066] The exemplary embodiments of the present disclosure have been specifically shown and described above. It should be understood that the present disclosure is not limited to the detailed structures, setting methods, or implementation methods described here; on the contrary, the present disclosure is intended to cover various modifications and equivalent settings included within the spirit and scope of the appended claims.
Claims
1. A data parallel processing system for sparse LLM based on MLA, comprising: An input data segmentation component for segmenting input data; A data parallel forwarding component for parallelly distributing the segmented micro-batch data to data processing devices each having a complete MLA module; And Multiple parallel data processing devices, each of which uses a complete set of attention weights and cache structures based on the complete MLA module it owns, and performs attention calculation processing on the obtained micro-batch data locally using locally generated KV values.
2. The data parallel processing system for sparse LLM based on MLA according to claim 1, further comprising: A request parsing component for obtaining token information of the input data for the user's input data; A subset segmentation component, with a predetermined number of tokens as the base, taking each data processing request in the input data as a unit, dividing a token subset starting from the head of the token information received each time until the end of the token information, and generating hash values for all token subsets using a hash function to form one or more hash value sets of the data processing request; A cache status tracking component for caching, for each destination data processing device, the hash value set of each data processing request forwarded by the data parallel forwarding component to it and the load of the data processing requests being processed by each destination data processing device; A scheduling score calculation component for scoring each destination data processing device based on the matching degree between the hash value set corresponding to the current data processing request and the hash value sets corresponding to the data processing requests received by the cache status tracking component for each destination data processing device, and the load of each destination data processing device, so that the data parallel forwarding component routes and forwards the current data processing request to the destination data processing device with the highest score or multiple destination data processing devices with the highest scores.
3. The data parallel processing system for sparse LLM based on MLA according to claim 1 or 2, further comprising: An expert model load monitoring component for, for each expert model, statistically monitoring in real time during the inference phase the load of the data processing requests being processed by the expert model selected by the data parallel forwarding component; An expert model routing adjustment component for, when the current load of any expert model exceeds a predetermined load threshold and is determined to be a hot expert model, allocating reserved data processing devices to the hot expert model, or migrating the hot expert model as a whole from the current data processing device to the reserved data processing device, so that the parallel data processing devices can parallelly forward the new input data requests that should be selected to the hot expert to the reserved data processing device.
4. The data parallel processing system for sparse LLM based on MLA as claimed in claim 3, wherein the expert model routing adjustment component counts the current loads of all expert models, calculates the weights of the current loads of each expert model in the overall load, and allocates reserved data processing devices to a predetermined number of the highest-ranked expert models.
5. The data parallel processing system for sparse LLM based on MLA as claimed in claim 1 or 2, further comprising: An expert model load monitoring component, for each expert model, during the inference phase, in real time counts the load of the data processing requests being processed by the expert model selected by the data parallel forwarding component; An expert model routing adjustment component, when the current load of any expert model exceeds a predetermined load threshold and is determined to be a hot expert model, reallocates a data processing device with a larger capacity or more data processing devices to the hot expert model, and migrates the hot expert model as a whole from the current data processing device to the newly allocated and mapped data processing device, so that the parallel data processing device can parallelly forward the new input data requests that should be selected to the hot expert to the newly allocated and mapped data processing device.
6. A data parallel processing method for sparse LLM based on MLA, comprising: Segmenting the input data; Parallelly allocating the segmented micro-batch data to data processing devices each having a complete MLA module; And Using the local KV values generated by each data processing device with a complete set of attention weights and cache structure based on the complete MLA module it owns, to perform attention calculation processing on the obtained micro-batch data locally.
7. The data parallel processing method for sparse LLM based on MLA as claimed in claim 6, further comprising: Before segmenting the input data, for the user's input data, obtaining the token information of the input data; Taking a predetermined number of tokens as a base, taking each data processing request in the input data as a unit, each time dividing a token subset starting from the head of the token information of the received data processing request until the end of the token information, and using a hash function to generate hash values for all token subsets to form one or more hash value sets of the data processing request; When the micro-batch data is parallelly allocated, for each destination data processing device, caching the hash value set of each data processing request forwarded to it by the data parallel forwarding component and the load of the data processing requests being processed by each destination data processing device; And Score each destination data processing device based on the matching degree between the hash value set corresponding to the current data processing request and the hash value sets corresponding to the data processing requests received by each destination data processing device, as well as the load of each destination data processing device, so that the data parallel forwarding component can route and forward the current data processing request to the destination data processing device with the highest score or one of the multiple highest-scoring destination data processing devices.
8. The data parallel processing method for sparse LLM based on MLA according to claim 6 or 7 further includes: For each expert model, during the inference stage, in real-time, count the load of the data processing requests being processed by the expert model selected by the data parallel forwarding component; And When the current load of any expert model exceeds a predetermined load threshold and is determined as a hot expert model, allocate reserved data processing devices to the hot expert model, or migrate the entire hot expert model from the current data processing device to the reserved data processing device, so that the parallel data processing device can parallelly forward the newly input data requests that should be selected to the hot expert to the reserved data processing device.
9. The data parallel processing method for sparse LLM based on MLA according to claim 8 further includes: Calculate the weight of the current load of each expert model in the overall load, and allocate reserved data processing devices to a predetermined number of the highest-ranked expert models.
10. The data parallel processing method for sparse LLM based on MLA according to claim 6 or 7 further includes: For each expert model, during the inference stage, in real-time, count the load of the data processing requests being processed by the expert model selected by the data parallel forwarding component; When the current load of any expert model exceeds a predetermined load threshold and is determined as a hot expert model, reallocate data processing devices with a larger capacity or more data processing devices to the hot expert model, and migrate the entire hot expert model from the current data processing device to the newly allocated and mapped data processing devices, so that the parallel data processing device can parallelly forward the newly input data requests that should be selected to the hot expert to the newly allocated and mapped data processing devices.
Citation Information
Patent Citations
Load balancing method and device and electronic equipment
CN118819841A
Hybrid expert model distributed training method based on dynamic load balancing
CN118838711A
Model reasoning method and device and electronic equipment
CN119129746A
Cited By
Large language model indexing method and device, computer equipment and storage medium
CN122047460A