Large language model reasoning method based on unloading assembly line

Through fine-grained inference pipeline design and optimized data transmission and computing kernels, the memory and computing resource problems of large-scale data models when deploying on consumer-grade devices are solved, and high concurrent inference and efficient resource utilization are achieved.

CN120146191APending Publication Date: 2025-06-13NANJING UNIV
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510231932.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

When deploying large-scale data models on consumer-grade devices, it faces high requirements for computing resources and memory capacity. The existing offload framework has problems such as large-scale data loading overhead, high system memory footprint, and failure to fully utilize the offload potential of solid-state drives in local inference of large models.

Method used

Through fine-grained inference pipeline design and optimized data transmission and computing kernels, the optimal offload inference strategy is automatically calculated, and the multi-level storage scheduling model weights and key-value cache are used to optimize hardware resource usage and inference performance.

Benefits of technology

It realizes high concurrent inference and deployment of large models under limited hardware resources, improves GPU utilization and inference throughput, and significantly improves the local inference efficiency of large language models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146191A_ABST
    Figure CN120146191A_ABST
Patent Text Reader

Abstract

The invention discloses a large language model reasoning method based on an unloading pipeline, and the method comprises the steps: obtaining an optimal unloading reasoning strategy through calculation according to the obtained model structure information and model configuration information of a large language model, and the hardware specification information and system operation load information of reasoning equipment; and then tasks in the optimal unloading reasoning strategy are scheduled through a fine-grained pipeline to output an unloading reasoning result of the large language model. According to the method, reasoning tasks are automatically configured according to input large language model information and a system hardware environment, and hardware resource use and reasoning performance are optimized. For deployment of a large language model on local equipment, transmission optimization is carried out for unloading using a solid state disk, so that the data transmission speed is increased. Meanwhile, fine-grained assembly line task scheduling is carried out for unloading reasoning, and the reasoning concurrency and the model throughput can be remarkably improved through unloading reasoning of an assembly line model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a large language model inference method based on an offloading pipeline, belonging to the technical field of model data processing in artificial intelligence technology. Background Art

[0002] In recent years, large-scale data models have become an important breakthrough in the field of artificial intelligence due to their excellent performance in text generation, dialogue interaction, code generation, etc. With the increasing demand for privacy protection and the rising cost of cloud deployment, the need to deploy large-scale data models on local devices has become increasingly urgent. However, due to the large scale of large models (such as 30B and 400B parameter models), extremely high requirements are imposed on both computing resources and memory capacity. Taking a typical consumer-grade GPU (such as RTX3060 with 6GB video memory) as an example, its memory capacity is difficult to accommodate model weights (such as an 8B model requires approximately 17GB of storage) and dynamically generated key-value caches (which grow linearly with the sequence length and batch size), resulting in significant challenges for direct deployment.

[0003] Model offloading is a technology for optimizing the training and use of large models. This technology can "offload" part of the model parameters or calculation process from the main memory (usually the video memory) to other storage devices to save main memory space and improve training efficiency, especially suitable for large models whose video memory requirements exceed the physical video memory capacity. The offloading technology schedules model weights and key-value caches through multi-level storage (GPU video memory, system memory, hard disk), becoming the core means for deploying large models under limited hardware resources. When using the offloading technology, the GPU does not need to store all the data of the model, but dynamically loads the required weights and key-value caches unloaded to the system memory and hard disk onto the GPU for calculation during inference.

[0004] However, although the existing frameworks supporting offloading technology have achieved certain results in large model training. But there are still great problems in local inference of large models, especially in local deployment on consumer-grade devices. First, the data loading overhead brought by the offloading technology is large. Although some researchers have processed data loading and calculation concurrently, the task scheduling granularity is coarse and the degree of concurrency is low. Second, although consumer-grade devices are generally equipped with high-performance solid-state drives, the existing offloading frameworks still rely too much on the system memory, resulting in the model inference occupying up to dozens to hundreds of GB of system memory. At the same time, the offloading potential of solid-state drives is ignored and there is a lack of targeted research and optimization. Therefore, how to more efficiently schedule offloading inference tasks on local devices and make full use of hardware resources to achieve efficient model inference has become an important issue. Therefore, there is an urgent need to provide an artificial intelligence large model inference method based on an offloading pipeline. Summary of the Invention

[0005] The content of this application is partially used to introduce concepts in a brief form, which will be described in detail in the following detailed implementation section. The content of this application is not intended to identify the key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0006] Aiming at the problems and deficiencies in the prior art, the purpose of the present invention is to provide a large language model inference method based on an offloading pipeline. Through a fine-grained inference pipeline design, combined with optimized data transmission and computing kernels, it can achieve high-concurrency inference and deployment of large models whose scale exceeds the GPU video memory capacity. Compared with other inference frameworks that support offloading, it has better hardware support for consumer devices, higher GPU utilization rate and inference throughput. It is used to solve the problems raised in the above background technology.

[0007] To achieve the above purpose, the present invention provides the following technical solutions:

[0008] The present invention discloses a large language model inference method based on an offloading pipeline, including the following steps:

[0009] Step 1, obtain the model structure information and model configuration information of the large language model;

[0010] Step 2, obtain the hardware specification information and system operation load information of the inference device;

[0011] Step 3, automatically calculate the optimal offloading inference strategy based on the information obtained in Step 1 and Step 2;

[0012] Step 4, execute and schedule the tasks in the optimal offloading inference strategy through a fine-grained pipeline;

[0013] Step 5, finally output the offloading inference result of the large language model.

[0014] Preferably, in Step 2, the specification information and system operation load information of the inference device include the available amount of GPU video memory and the available amount of system memory, and also include the actual PCIe transmission speed of the GPU and the actual PCIe transmission speed of the solid-state drive.

[0015] Preferably, in Step 3, automatically calculating the optimal offloading inference strategy based on the information obtained in Step 1 and Step 2 further includes the following steps:

[0016] Step 3.1, calculate the total model weight W, the total key-value cache C, and the peak memory usage M during the inference stage according to the model structure information and model configuration information;

[0017] Step 3.2, select the target memory level for data offloading in combination with the hardware specification information and system operation load information;

[0018] Step 3.3, configure the framework inference pipeline policy based on the information obtained in Step 3.1 and Step 3.2;

[0019] Step 3.4, perform data transfer optimization and underlying computing kernel optimization for the framework inference pipeline policy to obtain the optimal offloading inference policy.

[0020] Preferably, in Step 3.2, the rule for selecting the target memory level for data offloading is that

[0021] if the sum of the total model weights and the peak memory usage during the inference phase is less than the available GPU video memory, the weights are offloaded to the GPU video memory;

[0022] if the sum of the total model weights and the total key-value cache is less than the available system memory, and the actual PCIe transfer speed of the solid-state drive is less than the actual PCIe transfer speed of the GPU, the weights are offloaded to the system memory;

[0023] if the sum of the total model weights and the total key-value cache is less than the available system memory, but the actual PCIe transfer speed of the solid-state drive is greater than the actual PCIe transfer speed of the GPU, the weights are offloaded to the solid-state drive.

[0024] Preferably, in Step 3.3, the configuration rule for the framework inference pipeline policy is that

[0025] if the peak memory usage during the inference phase is less than the available GPU video memory, select the high-performance pipeline policy;

[0026] if the peak memory usage during the inference phase is greater than the available GPU video memory, select the low-memory-overhead pipeline policy.

[0027] Preferably, in Step 4, the tasks in the optimal offloading inference policy include loading of offloaded weights, loading of system key-value caches, saving of system key-value caches, and computing optimization, and a high-performance pipeline scheduling policy is used to perform fine-grained scheduling.

[0028] Preferably, the calculation formula for the peak memory usage M during the inference phase is

[0029] M = max(Mmha, Mmlp, Membed);

[0030] where Wembed represents the weights of the embedding layer, Wmha represents the weights of the multi-head attention layer, and Wmlp represents the weights of the multi-layer perceptron layer.

[0031] Preferably, the calculation formula for the total model weights W is

[0032] W = 2Wembed + l·(Wmha + Wmlp);

[0033] where l represents the number of hidden layers.

[0034] Preferably, the formula for calculating the total key-value cache C is

[0035]

[0036] where d represents the input dimension, p represents the model precision, V represents the vocabulary size, b represents the batch size, and s represents the sum of the input sequence length and the output sequence length.

[0037] Compared with the prior art, the beneficial effects of the present invention are:

[0038] The present invention provides a large language model inference method based on an offloading pipeline. First, the model structure information and model configuration information of the large language model, as well as the hardware specifications information and system operating load information of the inference device, are obtained, and the optimal offloading inference strategy is automatically calculated through the above information. Then, multiple tasks in the optimal offloading inference strategy are scheduled through a fine-grained execution pipeline, and finally, the offloading inference result of the large language model is output. The method of the present invention automatically configures inference tasks according to the input large language model information and system hardware environment, optimizes the use of hardware resources and inference performance. For the deployment of large language models on local devices, transmission optimization is carried out for offloading using solid-state drives to improve data transmission speed. At the same time, the present invention also performs fine-grained pipeline task scheduling for offloading inference. By utilizing the offloading inference of the pipelined model, the inference concurrency can be significantly improved. Comparing the present invention with other inference frameworks that support offloading, the present invention has better hardware support for consumer devices, higher GPU utilization, and higher inference throughput. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] The accompanying drawings that form a part of this application are used to provide a further understanding of this application, making other features, objectives, and advantages of this application more obvious. The schematic embodiments and descriptions thereof of this application are used to explain this application and do not constitute an improper limitation of this application.

[0040] In the drawings:

[0041] Figure 1 is the connection block diagram of the main steps in the embodiment of the present invention;

[0042] Figure 2 is the flowchart of the main steps in the embodiment of the present invention;

[0043] Figure 3 is the schematic diagram of data transmission optimization in the embodiment of the present invention;

[0044] Figure 4 Schematic diagram of fine-grained pipeline execution scheduling in an embodiment of the present invention;

[0045] Figure 5 Schematic diagram of the system framework structure in an embodiment of the present invention. Detailed implementation manners

[0046] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0047] In addition, it should be noted that for the sake of description, only parts related to the relevant invention are shown in the drawings. Without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other.

[0048] The present invention discloses a large language model inference method based on an offloading pipeline. The present disclosure will be described in detail below with reference to the drawings and in combination with embodiments.

[0049] Referring to Figures 1 to 2 as shown, it mainly includes the following steps:

[0050] Step 1, obtain the model structure information and model configuration information of the large language model;

[0051] Step 2, obtain the hardware specification information and system operation load information of the inference device;

[0052] Step 3, automatically calculate the optimal offloading inference strategy based on the information obtained in Step 1 and Step 2;

[0053] Step 4, execute the tasks in the optimal offloading inference strategy through fine-grained pipeline scheduling;

[0054] Step 5, finally output the offloading inference result of the large language model.

[0055] This embodiment uses an electronic computer with a GPU graphics processor of RTX3060, configured with 6GB video memory, 32GB system memory, and 1TB solid-state drive. Among them, the measured PCIe transfer speed of the GPU graphics processor is 23GB / s, and the measured transfer speed of the solid-state drive is 2GB / s. Next, an efficient inference framework for artificial intelligence large models based on the offloading pipeline will be applied to deploy the inference large language model Llama3.1-8B. The Llama3.1-8B model is a large language model based on the Transformer architecture, which performs excellently in various language tasks, especially showing significant improvements in instruction following and function calling.

[0056] According to Step 1, first obtain the model structure information and model configuration information of the large language model. The model structure information of this large language model specifically includes the number of model layers, the number of hidden layers l, the input dimension d, etc. The model configuration information includes the model name, model type, model precision p, vocabulary size V, and batch size b, etc. In addition, the model configuration information may also include the input sequence length and output sequence length. Here, the sum of the input sequence length and output sequence length is set as s. Among them, most of the model structure information of the large language model is provided by the model itself, while most of the model configuration information is set by the user himself.

[0057] According to Step 2, obtain the hardware specification information and system operation load information of the inference device. Among them, the hardware specification information includes the available GPU video memory MGPU and the available system memory MCPU of the local inference device obtained by reading. The system operation load information includes the actual PCIe transfer speed BGPU of the GPU of the local inference device and the actual PCIe transfer speed BSSD of the solid-state drive obtained by testing. The above-mentioned available GPU video memory MGPU and available system memory MCPU are directly read, while the actual PCIe transfer speed BGPU of the GPU and the actual PCIe transfer speed BSSD of the solid-state drive are actually measured. At the same time, the data transfer method in the test uses a rewritten and optimized data transfer component.

[0058] According to Step 3, based on the information obtained in the above Steps 1 and 2, the optimal offloading inference strategy can be obtained through automatic calculation, which specifically includes the following steps:

[0059] Step 3.1, calculate the total model weight W, the total key-value cache C, and the peak memory usage M during the inference phase according to the model structure information and model configuration information;

[0060] Step 3.2, combined with the hardware specification information and system operation load information, select the target memory level for data offloading;

[0061] Step 3.3, configure the framework inference pipeline policy based on the information obtained in Steps 3.1 and 3.2;

[0062] Step 3.4, perform data transfer optimization and underlying computing kernel optimization for the framework inference pipeline policy to obtain the optimal offloading inference policy.

[0063] Specifically, for the Llama3.1-8B model, let the weight representation of its embedding layer be Wembed, the weight of the multi-head attention layer be Wmha, and the weight of the multi-layer perceptron layer be Wmlp. Then, the calculation formula for the total model weight W is expressed as

[0064] W = 2Wembed + l·(Wmha + Wmlp);

[0065] Among them, when calculating the embedding layer weight Wembed, the multi-head attention layer weight Wmha, and the multi-layer perceptron layer weight Wmlp, the parameters obtained in Step 1 are used, which are the number of hidden layers l, the input dimension d, the model accuracy p, the vocabulary size V, the batch size b, and the sum s of the input sequence length and the output sequence length. Specifically, the calculation formula for the embedding layer weight Wembed is Wembed = pdV, and the calculation formula for the multi-head attention layer weight Wmha is The calculation formula for the multi-layer perceptron layer weight Wmlp is Wmlp = pd·(3·dh + 1). In addition, the calculation formula for the total key-value cache is And the calculation formula for the peak memory usage M during the inference stage is M = max(Mmha, Mmlp, Membed), that is, the maximum value among the embedding layer weight Wembed, the multi-head attention layer weight Wmha, and the multi-layer perceptron layer weight Wmlp. Recalculate, the calculation formula for the multi-head attention layer weight Wmha is The calculation formula for the multi-layer perceptron layer weight Mmlp is The calculation formula for the embedding layer weight Wembed is Membed =

[0066] pbs·(V + d) + 2Wembed.

[0067] Combined with the hardware specification information and the system operation load information, select the target memory level for data offloading. The rule for selecting the target memory level for data offloading is: if the sum of the total model weight W and the peak memory usage M during the inference stage is less than the available GPU video memory M GPU , then the weights are offloaded to the GPU video memory. If the sum of the total model weight W and the total key-value cache C is less than the available system memory M CPU , and the actual PCIe transfer speed B of the solid-state drive SSD is less than the actual PCIe transfer speed B of the GPU GPU, the weights are offloaded to the system memory. If the sum of the total model weights W and the total key-value cache C is less than the available amount M of the system memory CPU , but the actual PCIe transfer speed B of the solid-state drive SSD is greater than the actual PCIe transfer speed B of the GPU GPU , then the weights are offloaded to the solid-state drive. In this embodiment, the sum of the total model weights W and the peak memory usage M during the inference phase is greater than the available amount M of the GPU video memory GPU , and at the same time, the sum of the total model weights W and the total key-value cache C is greater than the available amount M of the system memory CPU , so the weights of the Llama3.1-8B model are offloaded to the solid-state drive.

[0068] Next, we configure the framework inference pipeline strategy. Specifically, its configuration rule is: If the peak memory usage M during the inference phase is less than the available amount M of the GPU video memory GPU , i.e., M < M GPU , then the high-performance pipeline strategy is selected. If the peak memory usage M during the inference phase is greater than the available amount M of the GPU video memory GPU , i.e., M > M GPU , then the low-memory-overhead pipeline strategy is selected. In this embodiment, the peak memory usage M during the model inference phase of the Llama3.1-8B model is less than the GPU video memory size, so the high-performance pipeline strategy is selected.

[0069] Furthermore, data transfer optimization and underlying computing kernel optimization are carried out for the framework inference pipeline strategy. Specifically, data transfer optimization includes block transfer of data from the hard disk to the GPU video memory, multi-threaded parallel transfer within the data block, and combined transfer of weights within the same layer of the model. To help understand the specific implementation details of data transfer optimization, Figure 3 shows the data transfer schematic diagram of step 3.4. In this embodiment, data transfer optimization is applied to the loading of weights offloaded to the solid-state drive, the loading and saving of key-value caches in the system memory. Underlying computing kernel optimization refers to realizing matrix-vector multiplication without dequantization by customizing and modifying the CUDA underlying computing kernel. When the model uses quantized weights and the batch size is small, underlying computing kernel optimization will be applied to accelerate the matrix-vector multiplication of the multi-head attention layer. The customization modification refers to determining the parallel block size of the CUDA computing kernel according to the model specifications of the GPU to maximize the utilization of the computing performance of the GPU. In this embodiment, the Llama3.1-8B model uses quantized weights and the batch size is small, so computing optimization is applied.

[0070] Step 4, execute and schedule the tasks in the optimal offloading inference strategy through fine-grained pipelining. Further, referring to Figure 4, we divide the inference pipeline strategy into four tasks, namely the loading of offloaded weights, the loading of system key-value caches, the saving of system key-value caches, and the computation optimization task. During inference, these tasks are placed in a task queue. The worker threads in the thread pool retrieve tasks from the task queue and then, through fine-grained pipeline scheduling, improve the concurrency of task execution while avoiding data contention. Among them, the thread pool is used to schedule and manage the threads responsible for data transfer tasks, which include weight loading tasks, key-value cache loading tasks, and key-value cache saving tasks. Therefore, the size of the thread pool is set to 3. In addition to maintaining the thread pool, the main thread is also responsible for all computation tasks during inference, reducing the idle time of the main thread while waiting for data transfer. The fine-grained pipeline scheduling is divided into a high-performance pipeline scheduling strategy and a low-memory-overhead pipeline scheduling strategy. The granularity of all its synchronization operations is at the task level rather than the device level to ensure fine-grained pipeline control. In this embodiment, both the prefill and decoding processes of inference adopt the high-performance pipeline scheduling strategy.

[0071] The high-performance pipeline strategy means that when inferring a certain layer of a large model, the weight loading task of the next layer is first initialized. If the current layer is a multi-head attention layer, the key-value cache saving task of the next multi-head attention layer (i.e., the layer two levels below) is synchronized first, and then the key-value cache loading task of the next multi-head attention layer is initialized to ensure that the key-value cache can be correctly loaded and data contention is excluded. At the same time, the input data for the computation task of the current layer is prepared, and this process is executed in parallel with the above data transfer tasks. After the input data is prepared, all data transfer tasks of the current layer are synchronized to ensure that all the weights and key-value caches of the current layer are loaded onto the GPU before the computation starts. The main thread performs the computation, while other threads are responsible for the data transfer tasks of subsequent layers. When the computation is completed, if the current layer is a multi-head attention layer, the key-value cache saving task of the current layer is first initialized. Finally, the output data of the computation task of the current layer is saved and the next layer is entered.

[0072] The low-memory-overhead pipeline scheduling strategy means reducing the concurrency of framework inference in exchange for lower memory overhead during the inference process, which usually occurs during the prefill process of inference. The loading task of key-value cache pairs is postponed to the next layer to reduce the peak storage overhead during the inference of the multi-head attention layer. The weight loading is postponed to the next layer to reduce the peak storage overhead during the inference of all layers. The synchronization of the key-value cache saving task is advanced before the key-value cache loading task of the next multi-head attention layer to ensure that only one newly generated key-value cache pair exists on the GPU waiting to be saved at any time.

[0073] See Figure 5 , after the fine-grained pipeline execution scheduling, the offloaded inference results of the large language model are finally output.

[0074] Application of the Embodiment

[0075] The inference method of the large language model based on the offloading pipeline of the present invention has been experimented on multiple large language models. During the experiment, the high-performance inference framework of the artificial intelligence large language model based on the offloading pipeline was used to infer Llama 3.1 models and Opt models with different numbers of parameters, and the model throughput under the optimal configuration obtained by solving was statistically analyzed under different batch size settings. At the same time, the GPU utilization rate of the current hardware system during the inference of the large language model was statistically analyzed.

[0076]

[0077]

[0078] Table 1

[0079]

[0080] Table 2

[0081] The above Table 1 shows the throughput comparison of the present invention method for multiple large language models, and the above Table 2 shows the GPU utilization rate comparison after the existing benchmark framework and the framework of the present invention are executed. According to the above experimental results, it shows that compared with the existing best method, the method of the present invention can increase the inference throughput by up to 3.1 times at most, and can increase the GPU utilization rate from less than 40% to more than 90% at most.

[0082] Computer program code for performing the operations of some embodiments of the present invention can be written in one or more programming languages or combinations thereof. The above programming languages include object-oriented programming languages - such as Java, Smalltalk, C++; and also include conventional procedural programming languages: such as the "C" language or similar programming languages. The program code can be executed completely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or completely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any kind of network - including a local area network (LAN) or a wide area network (WAN) - or can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0083] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that, in some alternative implementations, the functions noted in the blocks may occur in a different order than that noted in the accompanying drawings.

[0084] The functions described above can be performed, at least in part, by one or more hardware logic components. By way of example, and without limitation, the types of hardware logic components that may be used include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0085] The above description is only some preferred embodiments of the present disclosure and an explanation of the technical principles applied. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the embodiments of the present disclosure.

Claims

1. A large language model reasoning method based on an offload pipeline, characterized in that: The steps include: Step 1, obtaining the model structure information and model configuration information of the large language model; Step 2: Obtain the hardware specification information and system operation load information of the inference device; Step 3, automatically calculating the optimal offloading reasoning strategy based on the information obtained in step 1 and step 2; Step 4, scheduling the tasks in the optimal offloading reasoning strategy through a fine-grained pipeline execution; Step 5: Finally, output the offloaded inference result of the large language model.

2. The large language model reasoning method based on offloading pipeline according to claim 1 is characterized in that: In step 2, the specification information and system operation load information of the inference device include the available amount of GPU video memory and the available amount of system memory, and also include the actual PCIe transmission speed of the GPU and the actual PCIe transmission speed of the solid state drive.

3. The large language model reasoning method based on offloading pipeline according to claim 2 is characterized in that: In step 3, the optimal offloading reasoning strategy is automatically calculated based on the information obtained in step 1 and step 2, and the following steps are also included: Step 3.1, calculating the total model weight W, the total key value cache C, and the peak memory usage M in the inference phase according to the model structure information and the model configuration information; Step 3.2, selecting a target memory level for data offloading based on the hardware specification information and the system operation load information; Step 3.3, configuring the framework reasoning pipeline strategy based on the information obtained in step 3.1 and step 3.2; Step 3.4, perform data transmission optimization and underlying computing kernel optimization for the framework reasoning pipeline strategy to obtain the optimal offload reasoning strategy.

4. The large language model reasoning method based on offloading pipeline according to claim 3 is characterized in that: In step 3.2, the rule for selecting the target memory level for data unloading is: If the sum of the total weight of the model and the peak memory usage during the inference phase is less than the available GPU memory, the weight is unloaded to the GPU memory; If the sum of the total model weight and the total key value cache is less than the available amount of system memory, and the actual PCIe transmission speed of the solid-state drive is less than the actual PCIe transmission speed of the GPU, the weight is unloaded to the system memory; If the sum of the total model weight and the total key value cache is less than the available system memory, but the actual PCIe transmission speed of the solid-state drive is greater than the actual PCIe transmission speed of the GPU, the weight is unloaded to the solid-state drive.

5. The large language model reasoning method based on offloading pipeline according to claim 4 is characterized in that: In step 3.3, the configuration rules of the configuration framework inference pipeline strategy are: If the peak memory usage during the inference phase is less than the available GPU memory, the high-performance pipeline strategy is selected; If the peak memory usage during the inference phase is greater than the available GPU memory, the low memory overhead pipeline strategy is selected.

6. The large language model reasoning method based on offloading pipeline according to claim 5 is characterized in that: In step 4, the tasks in the optimal offloading reasoning strategy include loading of offloading weights, loading of system key-value cache, preservation of system key-value cache and calculation optimization, and fine-grained scheduling is performed using a high-performance pipeline scheduling strategy.

7. The large language model reasoning method based on offloading pipeline according to claim 3 is characterized in that: The calculation formula for the peak memory usage M in the inference phase is: M = max(Mmha, Mmlp, Membed); Among them, Wembed represents the weight of the embedding layer, Wmha represents the weight of the multi-head attention layer, and Wmlp represents the weight of the multi-layer perceptron layer.

8. The large language model reasoning method based on offloading pipeline according to claim 7 is characterized in that: The calculation formula of the total weight W of the model is: W = 2Wembed + 1·(Wmha + Wmlp); Among them, l represents the number of hidden layers.

9. The large language model reasoning method based on offloading pipeline according to claim 8, characterized in that: The calculation formula of the total key value cache C is: Among them, d represents the input dimension, p represents the model accuracy, V represents the vocabulary size, b represents the batch size, and s represents the sum of the input sequence length and the output sequence length.

Citation Information

Cited By

  • Inference acceleration optimization method and system applied to intelligent dialogue large model

    CN120725158A

  • Self-adaptive voice vehicle control method and device, vehicle and storage medium

    CN121768390A

  • An adaptive voice control vehicle method and device, vehicle and storage medium

    CN121768390B

  • Ensemble communication unloading method, system, equipment and medium

    CN121979690A