An Inference Running Acceleration Method for Processing Extremely Long Texts in Large Language Models

Through dynamic segmentation and scheduling attention heads, combined with asynchronous preloading and Key-Value cache management, the problem of computational load imbalance in extremely long text inference is solved, and the system throughput and stability is significantly improved.

CN119847763BActive Publication Date: 2025-05-30UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510315102.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-05-30
Estimated Expiration
2045-03-18

AI Technical Summary

Technical Problem

When large language models process extremely long text, the computational complexity of the self-attention mechanism leads to significant increase in inference delay and memory consumption, and existing parallel technologies are difficult to achieve load balancing, resulting in a decrease in computing efficiency.

Method used

By dynamically segmenting, reorganizing and scheduling the attention heads during the inference stage, a multi-head attention load balancing allocation strategy is established, and asynchronous preloading and Key-Value cache management are used to achieve load balancing of the inference process.

Benefits of technology

It effectively reduces the performance bottlenecks caused by unbalanced attention computing load, greatly improves the throughput and stability of large language models in extremely long text inference tasks, and does not require changing the model structure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119847763B_ABST
    Figure CN119847763B_ABST
Patent Text Reader

Abstract

The present invention discloses an inference operation acceleration method for processing extremely long texts in large language models, belonging to the technical field of inference optimization of large-scale pre-trained language models. Specifically: select a large language model and enable the sparse attention mode, input different texts, record the execution time of each attention head and the type of attention mode in different layers, and establish a statistical database; solve the multi-head attention load balancing allocation strategy of the large language model under actual texts; establish a weight index table by splitting the weight matrix; retrieve the weight sub-matrices corresponding to each attention head and load them to the corresponding GPU devices; realize the load balancing of the inference process by asynchronously preloading the MHA calculation weights and MLP calculation weights of adjacent layers and combining KV cache management. The present invention effectively avoids the problems of uneven multi-GPU load and resource idling by dynamically splitting, reorganizing, and scheduling attention heads in the inference stage, and significantly improves the system throughput of long sequence processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of inference optimization of large-scale pre-trained language models, and specifically relates to a method for accelerating the inference operation of processing extremely long texts for large language models. Background Art

[0002] In recent years, large language models (LLMs) based on the Transformer architecture, such as GPT, LLaMA, etc., have made breakthroughs in the field of natural language processing (NLP), becoming key milestones in the development of general artificial intelligence. As their application scenarios continue to expand to fields such as intelligent assistants, long document question answering, multi-turn conversations, automatic programming, etc., the length of the context sequence that the large language model needs to process in the inference stage has increased significantly, which poses higher requirements for the time efficiency and memory management of the computing system.

[0003] Currently, the computing systems of large language models have the following problems:

[0004] 1) Challenges of ultra-long sequence inference. The computational complexity of the self-attention mechanism in the Transformer model grows quadratically with the context length, resulting in a sharp increase in inference latency and a significant increase in memory consumption.

[0005] 2) Limitations of existing parallel technologies. The academic and industrial communities have proposed various parallel optimization strategies in the training stage, such as data parallelism, pipeline parallelism, tensor parallelism, sequence parallelism, etc., to accelerate model training and solve the problem of limited memory of a single GPU (graphics processing unit). However, these parallel optimization strategies are mainly designed for homogeneous attention mechanisms. When processing long texts, if heterogeneous and sparse attention patterns, such as retrieval heads, sparse heads, strided heads, etc., are enabled, it is easy to cause load imbalance among multiple GPUs and reduce the overall computational efficiency.

[0006] 3) Necessity of parallel optimization for long sequence inference. In long text application scenarios, the computational patterns of different attention heads often exhibit heterogeneous characteristics, that is, some heads need to calculate for global positions, such as the Full Attention Head, and some heads only need to focus on local or jumping information, such as the Sparse / Strided Attention Head, which will cause significant differences in computational overhead. Simply adopting tensor slicing or other homogeneous parallel strategies may lead to uneven allocation of GPU resources and affect the end-to-end latency performance.

[0007] Therefore, there is an urgent need for a parallel optimization method specifically for long text inference and supporting load balancing. Summary of the Invention

[0008] Aiming at the problems existing in the existing large language models in the scenario of long text reasoning, the present invention provides a method for accelerating the inference operation of large language models for processing extremely long texts. By dynamically splitting, reorganizing and scheduling attention heads in the inference stage, the problems of uneven multi-GPU load and resource idling are effectively avoided, and the system throughput of long sequence processing is significantly improved.

[0009] In order to achieve the above object, the technical method adopted by the present invention is as follows:

[0010] A method for accelerating the inference operation of large language models for processing extremely long texts, comprising the following steps:

[0011] Step 1, select a large language model, including layers of Transformer layers, and each self-attention layer of each layer includes attention heads, and enable the sparse attention mode;

[0012] Input texts with different context lengths to the large language model, record the execution time of each attention head in the layer of the large language model and the type of attention mode it adopts under different context lengths, and form a mapping relationship between the obtained execution time and the corresponding attention mode type, and then establish a statistical database;

[0013] Step 2, for the actual text , solve the multi-head attention load balancing allocation strategy of the large language model under, and the specific process is as follows:

[0014] For GPU devices, allocate each attention head in the layer under to different GPU devices for execution, as the multi-head attention load allocation scheme of the layer ; ;

[0015] Adopt the nearest neighbor matching method to find the text with the context length closest to in the statistical database, and use the execution time of each attention head in its layer as the execution time of the corresponding attention head in the layer under, and then calculate the multi-head attention execution time of the layer under the scheme layer under ; ;

[0016] Taking as the load balancing objective function, through a search algorithm, obtain the optimal multi-head attention load allocation scheme of the layer , and then obtain the multi-head attention load balancing allocation strategy of the large language model ;

[0017] Step 3: Split the query weight matrix, key weight matrix, and value weight matrix of the -th layer self-attention layer into sub-matrices concatenated column-wise respectively, to obtain the query weight sub-matrix, key weight sub-matrix, and value weight sub-matrix corresponding to each attention head;

[0018] According to the strategy , determine the attention heads to be executed on the same GPU device, concatenate their corresponding query weight sub-matrix, key weight sub-matrix, and value weight sub-matrix together to form a recombined matrix, and a total of recombined matrices are obtained, which are stored in the CPU memory in blocks; and based on the divided recombined matrices, a weight index table is established;

[0019] Step 4: Before processing the self-attention calculation (MHA calculation) of the -th layer, retrieve the query weight sub-matrix, key weight sub-matrix, and value weight sub-matrix corresponding to each attention head according to the scheme and the weight index table, and load them into the corresponding GPU device;

[0020] Step 5: The calculation of the -th layer includes MHA calculation and multi-layer perceptron module calculation (MLP calculation). By asynchronously preloading the MHA calculation weights and MLP calculation weights of adjacent layers, and cooperating with the Key-Value (KV) cache management, the load balancing of the inference process is achieved.

[0021] Furthermore, the sparse attention pattern described in Step 1 includes multiple of local attention pattern, block attention pattern, skip attention pattern, and retrieval attention pattern.

[0022] Furthermore, the formula for calculating in Step 2 is:

[0023] ;

[0024] In the formula, represents the GPU device number; represents the set of attention head numbers allocated on the -th GPU device; represents the execution time of the -th attention head in the -th layer.

[0025] Further, the search algorithm described in step 2 is a genetic algorithm, a dynamic programming algorithm, or a neural network algorithm.

[0026] Further, when the search algorithm is a genetic algorithm, the specific process of solving the multi-head attention load balancing allocation strategy is as follows:

[0027] Taking the scheme as an individual, initialize the population for the th layer, calculate the individual fitness according to the load balancing objective function ; through crossover operation and mutation operation, iteratively generate candidate solutions; and update the population through survival competition to obtain the scheme corresponding to the highest individual fitness of the th layer as the scheme; repeat the iterative process in the th layer to obtain the multi-head attention load balancing allocation strategy .

[0028] Further, the specific process of step 5 is as follows:

[0029] When calculating the th layer of the Transformer layer, first perform the MHA calculation of the th layer. At the same time, use the tensor parallel technology to asynchronously preload the MLP calculation weights of the th layer into the corresponding GPU device, and unload the generated KV cache in the CPU memory; after the MHA calculation of the th layer is completed, immediately perform the MLP calculation of the th layer, and asynchronously preload the MHA calculation weights of the th layer into the corresponding GPU device according to the process of step 4; after the MLP calculation of the th layer is completed, immediately perform the MHA calculation of the th layer to achieve pipelined parallel calculation for inference operation.

[0030] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0031] 1. The present invention proposes an inference operation acceleration method for processing extremely long texts in large language models, which can effectively reduce the performance bottleneck caused by uneven attention calculation load in the extremely long sequence inference scenario. Specifically, by splitting the model weights, each attention head can be independently scheduled and executed, realizing heterogeneous load balancing and parallel scheduling of the multi-head attention module, thereby effectively overcoming the problem of uneven calculation load in long text inference, greatly improving the throughput and stability of large language models in extremely long text inference tasks, and without modifying the model structure. This is the first time to systematically solve the problem of calculation imbalance in the sparse attention mechanism in the long text processing scenario.

[0032] 2. The present invention fully combines the existing sparse attention mechanism and model parallel technology, which can be used alone or in coordination with existing distributed training and inference frameworks, and has remarkable generality and scalability.

[0033] 3. Compared with the prior art, the multi-head attention load balancing, weight splitting and parallel scheduling scheme proposed by the present invention has the advantages of rapid deployment, automatic decision-making and orthogonality with other parallel mechanisms, and can be widely applied to fields such as programming assistance, legal text processing, and academic literature understanding that require processing of extremely long texts. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.

[0035] Figure 1 is a schematic flowchart of the inference operation acceleration method for processing extremely long texts in large language models proposed in Embodiment 1;

[0036] Figure 2 is a schematic diagram of weight splitting and storage in Embodiment 1;

[0037] Figure 3 is a schematic diagram of the weight index table in Embodiment 1;

[0038] Figure 4 is a schematic diagram of parallel scheduling and weight preloading in Embodiment 1. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0039] In order to further understand the present invention, the following describes the preferred implementation of the present invention in combination with embodiments. However, it should be understood that these descriptions are only for further illustrating the features and advantages of the present invention, rather than limiting the claims of the invention.

[0040] Example 1:

[0041] This example proposes an inference operation acceleration method for processing extremely long texts by large language models. The process is as Figure 1 shown and includes the following steps:

[0042] Step 1: Select a large language model, including layers of Transformer layers. Each self-attention layer of each layer includes attention heads, and enable the sparse attention mode, including multiple of local attention mode, block attention mode, skip attention mode, and retrieval attention mode; the large language model can be GPT or LLaMA;

[0043] Input texts with different context lengths to the large language model, record the execution time of each attention head in the layer of the large language model and the type of attention mode it adopts (such as local attention mode, block attention mode, skip attention mode, or retrieval attention mode) under different context lengths. The obtained execution time forms a mapping relationship with the corresponding attention mode type, and then a statistical database is established;

[0044] Assume that a total of texts with different context lengths are input. Denote the th text as , then the recorded execution time of the th attention head in the layer is , the corresponding attention mode type is , and the formed mapping relationship is . .

[0045] Step 2: For the actual text , solve the multi-head attention load balancing allocation strategy of the large language model under . The core goal is to "minimize the computing time of the slowest device", thereby improving the overall throughput of the long text inference process. The specific process is as follows:

[0046] Step 2.1: Assume that the total number of GPU devices is , and allocate the th attention head in the layer to the th GPU device for execution, denoted as , as the multi-head attention load allocation scheme of the layer; ;

[0047] Use the nearest neighbor matching method to find the context length closest to The text of ,but Next Layer Execution time of an attention head Approximately equal to The execution time of ,Right now:

[0048] ;

[0049] make For the The set of attention head numbers allocated on GPU devices, then the solution Next Multi-head attention execution time of the layer for:

[0050] ;

[0051] That is, in parallel computing, the execution time of multi-head attention is determined by the slowest GPU device;

[0052] Step 2.2, in In the Transformer layer, the self-attention layer, the multi-layer perception module layer and the RMS Norm normalization layer are combined together. If only the execution time of the multi-head attention is optimized, the solution is obtained. Next Total delay of the layer :

[0053] ;

[0054] In the formula, is the execution time of the MLP layer; is the execution time of the RMSNorm normalization layer;

[0055] Then get the solution Total latency for downloading large language models :

[0056] ;

[0057] because can be regarded as a constant term, so determining the multi-head attention load balancing allocation strategy only requires optimizing ;

[0058] Step 2.3: is the load balancing objective function. Through the search algorithm, we can obtain Optimal multi-head attention load allocation scheme for layers , and then obtain The multi-head attention load balancing allocation strategy for large language models ;

[0059] Specifically, when the search algorithm is a genetic algorithm, the specific process of solving the multi-head attention load balancing allocation strategy is as follows:

[0060] Using the scheme as an individual, initialize the population for the th layer. According to the load balancing objective function , calculate the individual fitness; through crossover operations and mutation operations, iteratively generate candidate solutions; and update the population through survival competition to obtain the scheme corresponding to the highest individual fitness of the th layer, as the scheme ; repeat the iterative process in the th layer to obtain the multi-head attention load balancing allocation strategy .

[0061] Step 3: To achieve heterogeneous parallelism, perform weight slicing and storage on the self-attention layer of the th layer. The principle is as shown in Figure 2 , Figure 2 where "head" is used to represent the attention head;

[0062] The specific process of the weight slicing and storage is as follows:

[0063] The linear transformation matrix of the self-attention layer of the th layer includes a query weight matrix , a key weight matrix , and a value weight matrix , , where is the hidden layer dimension of the large language model, is the dimension of a single attention head. Split , , and into sub-matrices concatenated column-wise respectively, that is:

[0064] W l,Q =[ W l,Q (0) , W l,Q (1) ,…, W l,Q (I-1) ] ;

[0065] W l,K =[ W l,K (0) , W l,K (1) ,…, W l,K (I-1) ] ;

[0066] W l,V =[ W l,V (0) , W l,V (1) ,…, W l,V (I-1) ] ;

[0067] In the formula, is the query weight sub - matrix of the th attention head; is the key weight sub - matrix of the th attention head; is the value weight sub - matrix of the th attention head;

[0068] According to the multi - head attention load - balancing distribution strategy , determine the attention heads to be executed on the same GPU device, and concatenate their corresponding query weight sub - matrices, key weight sub - matrices, and value weight sub - matrices to form a recombined matrix, obtaining a total of recombined matrices to reduce fragmented access and frequent scheduling during inference; The recombined matrices are stored in blocks in the CPU memory to ensure that the weights of the corresponding attention heads can be quickly found during the loading phase, avoiding frequent indexing or slicing operations on the original large matrix;

[0069] Based on the partitioned recombined matrices, establish a weight index table, as shown in Figure 3 , taking , as an example, each layer corresponds to 4 GPU devices, and each GPU device corresponds to a two - dimensional list of weights [ l,p ] , the elements it contains are the query weight sub - matrices, key weight sub - matrices, and value weight sub - matrices of the attention heads executed on this GPU device, and it has the function of indexing the weights to the corresponding GPU device when running in different layers.

[0070] Step 4. Before calculating the self - attention of the th layer, according to the scheme Retrieve the query weight sub - matrix, key weight sub - matrix, and value weight sub - matrix corresponding to each attention head from the weight index table and load them into the corresponding GPU device. Each GPU device only needs to read its own weight combination, reducing redundant loading time.

[0071] Step 5: In the long - text scenario, the computational overhead exceeds the I / O overhead, and the zig - zag pre - filling method has no practical advantage. Therefore, this embodiment proposes a parallel scheduling and weight pre - loading method, that is:

[0072] The layer's calculation includes MHA calculation and multi - layer perceptron module calculation. By asynchronously pre - loading the MHA calculation weights and MLP calculation weights of adjacent layers, and cooperating with the KV cache management, pipeline parallel computing for inference operation is achieved to significantly reduce the GPU memory pressure. The specific process is as follows:

[0073] As Figure 4 shown, when calculating the layer of the Transformer layer, first perform the MHA calculation of the layer. At the same time, using tensor parallel technology, asynchronously pre - load the MLP calculation weights of the layer into the corresponding GPU device, and unload the generated KV cache to the CPU memory. Although the KV cache management spans two stages of MHA calculation and MLP calculation, the full - duplex communication architecture ensures that data transfer between storage and loading operations does not cause blockage; after the MHA calculation of the layer is completed, immediately execute the MLP calculation of the layer, and asynchronously pre - load the MHA calculation weights of the layer into the corresponding GPU device according to the process of step 4; after the MLP calculation of the layer is completed, immediately execute the MHA calculation of the layer, and so on to achieve parallelization of the inference and loading processes, so that the calculation process is not interrupted by loading, thereby improving the inference efficiency.

[0074] Example: In the torch.bfloat16 precision, for the LLaMA3 - 8B model applying the inference operation acceleration method for processing extremely long texts in large language models described in this embodiment, the weight storage requirement can be reduced from 13 GB = MHA 2.5 GB+MLP 10.5 GB to 0.33 GB for each operation module, a reduction of 40 times. Through this parallel scheduling and weight pre - loading method, the I / O operation and calculation are completely overlapped, effectively converting the main cost factor from memory limitation to pure computational overhead, enabling larger models to be executed with limited hardware resources.

[0075] To comprehensively evaluate the performance of the inference operation acceleration method for processing extremely long texts by large language models, this embodiment conducts a large number of experiments on a variety of advanced large language models and long text inference algorithms. This experiment mainly conducts application tests on two major mainstream long text sparse algorithms: MInference (Minference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention) and DuoAttention (DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads Chunk), and comprehensively verifies the performance under highly sparse attention. In the MInference algorithm environment, the selected large language models are LLaMA3-8B, Yi-9B, and ChatGLM4-9B. In the DuoAttention algorithm environment, the selected large language models are LLaMA2-7B and LLaMA3-8B. Under different context lengths (768K and 160K), the inference efficiency of the inference operation acceleration method for processing extremely long texts by large language models proposed in this embodiment and three existing benchmark methods is compared, and the inference throughput (unit: token / second or its equivalent metric) data results shown in Table 1 are obtained. The larger the inference throughput value, the higher the inference efficiency. Among them, the three existing benchmark methods are: (1) Random allocation: For the multi-head attention mechanism of each layer, randomly allocate the attention heads to each GPU device; (2) Uniform partitioning (Megatron): Similar to the method of Megatron-LM (Training Multi-Billion Parameter Language Models Using Model Parallelism), evenly distribute the attention heads to each GPU device in order; (3) Random uniform allocation: Maintain an equal number of heads on each GPU device through restricted randomization.

[0076] Table 1:

[0077]

[0078] As can be seen from Table 1, under the long context of 768 K, when using MInference, the inference throughput of LLaMA3-8B can be significantly increased from 1896.93 to 2447.07 under the inference operation acceleration method proposed in this embodiment, significantly alleviating the inference bottleneck brought by long text input, and the same is applicable to other large language models. For a length of 160 K, the inference throughput of ChatGLM4-9B under the Random strategy is only 3885.62, while this embodiment can reach 4358.08, greatly improving the downstream inference efficiency. Similarly, in the DuoAttention environment, adopting the inference operation acceleration method proposed in this embodiment can increase the throughput of LLaMA3-8B from 991.13 to 1292.08, highlighting the effectiveness of "multi-head parallel load balancing" in the context of extremely long texts.

[0079] In summary, the inference operation acceleration method for processing extremely long texts in large language models proposed in this embodiment can effectively address the problem of inference load imbalance caused by extremely long sequences in the applications of multiple advanced large language models and long text algorithms (including MInference, DuoAttention, etc.), significantly reducing the end-to-end inference latency and increasing the throughput. Compared with the comparison methods such as uniform partitioning, random allocation, and random uniform allocation, this embodiment relies on core ideas such as offline data collection, attention head balance strategy solution, weight splitting, and parallel scheduling, and has both performance and generality. It can be widely applied to various large language model inference tasks that need to process ultra-long texts, such as programming assistance, legal text processing, and academic literature understanding, and has significant practical and application value.

[0080] The above embodiments are for better understanding of the present invention, and are not limited to the best implementation manner. They do not limit the content and protection scope of the present invention. Any product that is the same as or similar to the present invention obtained by anyone under the inspiration of the present invention or by combining the features of the present invention with other existing technologies is within the protection scope of the present invention.

Claims

1. A method for accelerating the inference operation of a large language model for processing extremely long texts, characterized in that: The following steps are involved: Step 1: Select a large language model, including Transformer layers, each self-attention layer includes attention heads and enable sparse attention mode; Input texts with different context lengths into the large language model and record the first The execution time of each attention head of the layer and the type of attention mode it adopts are mapped to the corresponding attention mode type, and then a statistical database is established; Step 2: For the actual text , solve The multi-head attention load balancing allocation strategy for the large language model is as follows: right GPU devices, Next Each attention head of the layer is assigned to different GPU devices for execution as the first Multi-head attention load distribution scheme for layers ; Use the nearest neighbor matching method to find the context length closest to , and The execution time of each attention head in the layer is Next The corresponding attention head execution time of the layer is calculated to obtain the solution Next Multi-head attention execution time of the layer ; by is the load balancing objective function. Through the search algorithm, we can obtain Optimal multi-head attention load distribution scheme for layers , and then obtain Multi-head attention load balancing allocation strategy for large language models ; Step 3: The query weight matrix, key weight matrix, and value weight matrix of the self-attention layer are split into The sub-matrices are spliced ​​in the column direction to obtain the query weight sub-matrix, key weight sub-matrix and value weight sub-matrix corresponding to each attention head; According to the strategy , determine the attention heads executed on the same GPU device, concatenate the corresponding query weight submatrix, key weight submatrix, and value weight submatrix together to form a reorganized matrix, and get The reorganized matrix is ​​stored in blocks in the CPU memory; and a weight index table is established based on the divided reorganized matrix; Step 4: According to the plan and weight index table, retrieve the query weight sub-matrix, key weight sub-matrix and value weight sub-matrix corresponding to each attention head, and load them into the corresponding GPU device; Step 5: The calculation of the layer includes MHA calculation and MLP calculation. The load balancing of the reasoning process is achieved by asynchronously preloading the MHA calculation weights and MLP calculation weights of adjacent layers and cooperating with KV cache management.

2. The method for accelerating the inference operation of a large language model for processing extremely long text according to claim 1, characterized in that: The sparse attention mode described in step 1 includes multiple of a local attention mode, a block attention mode, a jumping attention mode and a retrieval attention mode.

3. The method for accelerating the inference operation of a large language model for processing extremely long text according to claim 1, characterized in that: Calculated in step 2 The formula is: ; In the formula, Indicates the GPU device number; Indicates The set of attention head numbers assigned to the GPU devices; express Next Layer The execution time of an attention head.

4. The method for accelerating the inference operation of a large language model for processing extremely long text according to claim 1, characterized in that: The search algorithm in step 2 is a genetic algorithm, a dynamic programming algorithm or a neural network algorithm.

5. The method for accelerating the inference operation of a large language model for processing extremely long text according to claim 4, characterized in that: When the search algorithm is a genetic algorithm, solve the multi-head attention load balancing allocation strategy The specific process is: By plan For individuals, The layer initializes the population according to the load balancing objective function , calculate individual fitness; iteratively generate candidate solutions through crossover and mutation operations; And update the population through survival competition, and obtain the first The solution corresponding to the highest individual fitness of the layer , as a solution ;exist Repeat the iterative process in the layer to obtain a multi-head attention load balancing distribution strategy .

6. The method for accelerating the inference operation of a large language model for processing extremely long text according to claim 1, characterized in that: The specific process of step 5 is: In the When calculating the Transformer layer, first perform the The MHA calculation of the layer is performed by using tensor parallel technology to asynchronously preload the first The MLP calculation weights of the layer are transferred to the corresponding GPU device, and the generated KV cache is unloaded to the CPU memory; After the MHA calculation of the layer is completed, the The MLP calculation of the layer is performed, and the first The MHA calculation weights of the layer are transferred to the corresponding GPU device; After the MLP calculation of the layer is completed, the first The MHA calculation of the layer is used to realize the pipeline parallel computing of the reasoning operation.

Citation Information

Patent Citations

  • Complex task processing method of large language model based on plannable workflow

    CN117453915A

  • Dataset generation using large language models

    US20240185001A1