Large language model reasoning optimization method, system and equipment and storage medium
By using KV header selective unloading and dynamic approximation caching technology, the problems of excessive video memory usage and large CPU-GPU data reading overhead in LLM inference are solved, achieving video memory optimization and performance improvement.
Patent Information
- Application Number
- CN202511511942.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-22
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-10-22
AI Technical Summary
During Large Language Model (LLM) inference, the excessive GPU memory usage of key-value (KV) cache data leads to insufficient GPU memory capacity, and the high overhead of reading KV data offloaded to CPU memory affects inference efficiency and hardware resource utilization.
It employs KV header selective offloading and dynamic approximation caching technology, where some KV data is cached in GPU memory and some is offloaded to CPU memory. It also reduces the amount of data read from CPU to GPU through top-k attention and manages KV data by combining dynamic approximation caching algorithm, prioritizing the reuse of cached data to reduce read overhead.
It effectively reduces the amount of GPU memory occupied by KV data, reduces the data reading overhead from CPU to GPU, ensures inference efficiency and hardware resource utilization, and improves the overall performance of LLM.
Smart Images

Figure CN120996208A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer GPU (Graphics Processing Unit) memory and computing efficiency optimization, and particularly relates to a large language model inference optimization method, system, device and storage medium. BACKGROUND
[0002] Large language model (LLM) is the focus of the field of artificial intelligence in recent years. LLM is a generative artificial intelligence model, and the process of generating new text according to the text sequence input by the user is called inference, which is an important link of LLM application. The text sequence input by LLM is composed of tokens, which will be mapped into word embedding vectors to participate in the calculation of the attention module. Specifically, the word embedding vectors are transformed by three different matrices to generate query (Query), key (Key) and value (Value) data (Query is referred to as Q data, and the latter two are collectively referred to as KV data), and then new tokens are generated through the steps of nonlinear softmax function (normalization exponential function) calculation and activation function. The aforementioned QKV data will generate h groups in total, each of which uses a different mapping matrix, and these groups are called attention heads (Heads). According to different data types, the attention heads are divided into two types: Q heads and KV heads. In early LLM, Q heads and KV heads are one-to-one corresponding and equal in number; more advanced LLM usually uses grouped-query attention technology (Grouped-Query Attention, GQA), which reduces the number of KV heads so that m Q heads share one KV head. In addition, the structure of LLM is multi-layered, each layer has the same module structure and needs to complete the above-mentioned attention calculation process, the input of the first layer is the word embedding vector mentioned above, and the input of other layers is the output of the previous layer. These input data can be collectively referred to as hidden state vectors (referred to as hidden vectors).
[0003] The inference process of the large language model is autoregressive, and each new token generated needs to process all existing tokens according to the above calculation process, which repeatedly calculates the KV data. To solve this problem, KV cache technology is proposed, which caches the calculated KV data in the GPU (Graphics Processing Unit) memory, eliminates repeated calculation, improves inference efficiency, and has become a standard technology for LLM inference. After using the KV cache technology, the inference process of LLM is divided into two stages: the pre-filling stage and the decoding stage. In the pre-filling stage, LLM calculates the Q data and KV data of all tokens in parallel, caches the KV data, and generates the first new token. The decoding stage only calculates the Q data and KV data of the latest token, adds the KV data to the KV cache, and generates the next new token. The decoding stage will be repeated until the end of generation. As can be seen, the size of the KV cache data increases linearly with the number of tokens processed in the inference process.
[0004] In recent years, the input size of LLMs has been growing. At the task level, there are more and more tasks that need to process long text sequences. At the inference efficiency level, batch processing technology is widely used, and multiple input sequences are packaged into a batch for joint inference to improve inference throughput. The growth of sequence length and batch size both leads to the need to process more tokens during inference, and the size of KV data also grows rapidly. Since KV data is cached in GPU memory, the memory it occupies will increase significantly, which may exceed the memory occupied by model parameters, and even exceed the memory capacity of the GPU. For example, when using the Llama3.1-8B model to infer a text sequence with a length of 256KB (kilobyte), the memory occupied by the KV cache data is 31.25GB (gigabyte), and the memory occupied by the model parameters is about 16GB. The sum of the two exceeds the memory capacity of most mid-range GPU devices. In summary, the high memory occupation of KV cache data has become a major bottleneck for the expansion of inference task size.
[0005] A mainstream method to solve this problem is to offload KV cache data from GPU memory to the host memory of the computer central processing unit (CPU), and use top-k attention technology to reduce the read amount of KV data based on the sparsity of attention calculation in which KV data participates. Specifically, top-k attention generates retrieval metadata based on KV data, which has low memory occupation and can be stored in GPU memory. During inference, Q data (or its variants) and retrieval metadata are used to query key KV data, and only these key data are read from CPU memory to participate in calculation. However, the high-speed serial computer expansion bus (PCIe) bandwidth of CPU-GPU memory interconnection is limited, and it will take a lot of time for GPU to read KV cache data from CPU memory, which will significantly exceed the GPU calculation time, resulting in a significant reduction in inference speed and a decrease in GPU hardware resource utilization.
[0006] Therefore, how to reduce the CPU-to-GPU KV cache data read overhead in the scenario of offloading KV cache data, optimize GPU memory occupation while ensuring inference efficiency and hardware resource utilization, is a key technical challenge in large-scale task inference of LLMs, and is one of the problems that need to be solved at the present stage.
[0007] In view of this, the present application is proposed. SUMMARY
[0008] The application aims to provide a large language model inference optimization method, system, device and storage medium, which can significantly reduce the GPU memory occupied by KV data, minimize the KV data reading overhead from CPU to GPU, and guarantee the inference efficiency and hardware resource utilization.
[0009] The application aims to achieve the above-mentioned purposes through the following technical solutions. The application aims to achieve the above-mentioned purposes through the following technical solutions. The application aims to achieve the above-mentioned purposes through the following technical solutions. The application aims to achieve the above-mentioned purposes through the following technical solutions.
[0010] The application aims to achieve the above-mentioned purposes through the following technical solutions. The application aims to achieve the above-mentioned purposes through the following technical solutions. A decoding unit based on dynamic approximate caching and data prefetching is used for autoregressive decoding combined with the latest word piece, which is the first new word piece generated in the prefilling stage or the new word piece generated in the last decoding step; the working process of the current layer of the large language model in a single decoding step is as follows: Q data and KV data are generated by using the hidden vector of the current layer, and the KV data is added to the GPU memory or CPU memory; the hidden vector of the current layer is the word embedding vector of the first new word piece or the calculation result of the last layer, and Q is a query; based on the hidden vector of the current layer, it is judged whether the next layer needs to prefetch KV data by taking the KV head as a unit; if so, the corresponding KV data is searched in the CPU memory, and after the KV data of the current layer is prefetched, the prefetching of the KV data of the next layer is started; the KV data is searched in the GPU memory by using the Q data generated by the current layer, and the calculation of the current layer is completed combined with the prefetched KV data; the process is repeated until the last layer, and the calculation result of the last layer is used to obtain a new word piece.
[0011] A processing device, comprising: one or more processors; a memory for storing one or more programs; Wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the foregoing method.
[0012] A readable storage medium, storing a computer program, when the computer program is executed by a processor, implementing the foregoing method.
[0013] As can be seen from the technical solutions provided by the above-mentioned application, in the inference process of the large language model, most of the KV data is offloaded to the CPU memory; for reading of the KV data from the CPU memory to the GPU memory, top-k attention is used to reduce the reading amount; and the KV data that has been read to the GPU memory is cached, and the approximate caching algorithm is used to manage the KV data by taking the KV head as a basic unit; when the KV data needs to be read in the inference process, it is first judged whether the data in the cache can be reused, if the cache hits, the data is directly reused, the data reading overhead from the CPU memory to the GPU memory is reduced, and if the cache misses and the data cannot be reused, data prefetching is performed. Thanks to the above improvements, the application can effectively reduce the memory occupied by the KV data, and minimize the KV data reading overhead from the CPU to the GPU, so that the inference performance reaches an ideal level. BRIEF DESCRIPTION OF DRAWINGS
[0014] In order to more clearly illustrate the technical solutions of the embodiments of the application, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the application, and those skilled in the art can also obtain other drawings according to these drawings without creating any inventive labor.
[0015] Figure 1 A flowchart of a large language model inference optimization method provided for an embodiment of the present application.
[0016] Figure 2 A flowchart of a prefill stage provided for an embodiment of the present application.
[0017] Figure 3 A flowchart of a decoding stage provided for an embodiment of the present application.
[0018] Figure 4 A schematic diagram of selective offloading of a KV head provided for an embodiment of the present application.
[0019] Figure 5 A schematic diagram of a dynamic approximate cache algorithm provided for an embodiment of the present application.
[0020] Figure 6 A schematic diagram of a GPU-centric data transfer synchronization method provided for an embodiment of the present application.
[0021] Figure 7 A schematic diagram of the overall architecture of a large language model inference optimization method provided for an embodiment of the present application.
[0022] Figure 8 A schematic diagram of a large language model inference optimization system provided for an embodiment of the present application.
[0023] Figure 9 A schematic diagram of a processing device provided for an embodiment of the present application. DETAILED DESCRIPTION
[0024] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the protection scope of the present application.
[0025] First, the terms that can be used in this text are explained as follows: The term “and / or” means either of the two or both at the same time, for example, X and / or Y means three cases including “X” or “Y” or “X and Y”.
[0026] The terms "comprising," "including," "containing," "having," or other similar semantic descriptions should be interpreted as non-exclusive inclusion. For example, including a technical feature element (such as raw material, component, ingredient, carrier, dosage form, material, size, part, component, mechanism, device, step, process, method, reaction conditions, processing conditions, parameter, algorithm, signal, data, product or article of manufacture, etc.) should be interpreted as including not only the expressly listed technical feature element, but also other technical feature elements that are not expressly listed and are well-known in the art.
[0027] The term "composed of" excludes any technical features not expressly listed. When used in a claim, it closes the claim to exclude all technical features other than those expressly listed, except for associated conventional impurities. If the term appears only in a clause of a claim, it limits the claim to the elements expressly listed in that clause; elements recited in other clauses are not excluded from the overall claim.
[0028] The following provides a detailed description of a large language model inference optimization method, system, device, and storage medium provided by this invention. Contents not described in detail in the embodiments of this invention are prior art known to those skilled in the art. Where specific conditions are not specified in the embodiments of this invention, they are performed according to conventional conditions in the art or conditions recommended by the manufacturer. Where the manufacturers of the instruments used in the embodiments of this invention are not specified, they are all conventional products that can be purchased commercially.
[0029] Example 1 This invention provides a method for optimizing reasoning in large language models, such as... Figure 1 As shown, it mainly includes the following steps: Step 1: Pre-filling stage based on selective unloading of KV header.
[0030] In this embodiment of the invention, during the pre-filling stage, considering the difficulty of reusing the KV header, a portion of the KV data of the KV header is cached in the GPU memory, while the other portion of the KV data of the KV header is unloaded to the CPU memory.
[0031] Specifically, in this step: the Q data and the KV data corresponding to the input token sequence are calculated, and the first new token is generated. There are multiple attention heads in each layer of the large language model, which are divided into Q heads and KV heads. Each KV head corresponds to a set of KV data, and each Q head corresponds to a set of Q data. The reuse difficulty of each KV head is quantified, and the number of KV heads that can be prefetched in each layer is calculated. In combination with the number of KV heads that can be prefetched in each layer and the reuse difficulty of the corresponding KV head, the KV heads residing in the GPU memory are selected, and the corresponding KV data is stored in the GPU memory. The remaining KV data corresponding to the KV head is unloaded to the CPU memory. The areas storing KV data in the GPU memory and the CPU memory are collectively referred to as KV cache.
[0032] The main process involved in this step is shown in Figure 2 as shown, mainly including: Step 11, generating Q data and KV data.
[0033] In the embodiment of the application, the Q data and the KV data corresponding to the input text sequence are generated, and the principle and generation process have been introduced in the foregoing background technology, which will not be repeated here.
[0034] Step 12, generating retrieval metadata based on KV data.
[0035] In the embodiment of the application, the meaning of the retrieval metadata has been introduced in the foregoing background technology, and its generation method is provided by the specific top-k attention calculation. The optimization method of the embodiment of the application focuses on the management and reading method optimization of the KV data, and the top-k attention technology is not within the optimization range, which can be realized using the effective algorithm proposed in the art.
[0036] Step 13, selectively unloading KV data to CPU memory.
[0037] In the embodiment of the application, the unloading of the KV data is selective. The contribution degree of the KV data of different KV heads to the inference generation quality and the data distribution characteristics are different. By analyzing these indicators, and accordingly taking customized unloading and data reading strategies for different KV heads, specifically: using the average cosine similarity of offline analysis and the preset reuse threshold of the KV head, the reuse difficulty of the KV head is calculated; by pre-running the large language model inference, the number of KV heads that can be prefetched in each layer is calculated; for each layer, the number of KV heads with a reuse difficulty greater than 0 is counted, and the number of KV heads that can be prefetched is subtracted to determine the number of KV heads residing in the GPU memory. Then, the KV head with the highest reuse difficulty is selected to reside in the GPU memory. If the GPU memory is not sufficient to store, then within the range that the memory can accommodate, the batch of KV heads with the highest reuse difficulty is selected to reside in the GPU memory.
[0038] Step 14, the remaining calculation steps of the pre-filling stage are completed.
[0039] In the embodiment of the present application, the remaining steps of the pre-filling stage involve attention calculation, residual link, feedforward network calculation, etc., which are consistent with the existing process, so no further description is given.
[0040] Step 2, decoding stage based on dynamic approximate cache and data prefetching.
[0041] In the embodiment of the present application, in the decoding stage, autoregressive decoding is performed in combination with the latest token, which is the first new token generated in the pre-filling stage or the new token generated in the previous decoding step; the working process of the current layer of the large language model in a single decoding step is as follows: Q data and KV data are generated using the hidden vector of the current layer, and the KV data is added to the GPU memory or CPU memory; the hidden vector of the current layer is the word embedding vector of the first new token or the calculation result of the previous layer, and Q is the query; based on the hidden vector of the current layer, it is judged whether the next layer needs to prefetch KV data with KV head as a unit, if yes, the corresponding KV data is searched in the CPU memory, and after the KV data of the current layer is prefetched, the prefetching of the KV data of the next layer is started; the KV data is searched in the GPU memory using the Q data generated by the current layer, and the calculation of the current layer is completed in combination with the prefetched KV data; the process is repeated until the last layer, and the calculation result of the last layer is used to obtain a new token.
[0042] For the convenience of description, the current layer is denoted as the layer, the previous layer is denoted as the layer, and the next layer is denoted as the layer, and the decoding stage can be described as follows: in this step: autoregressive decoding is performed using the latest generated token and the KV cache, and the next new token is continuously generated until the generation is completed; if it is the first decoding step, the latest generated token is the first new token, which is generated by the pre-filling stage; if it is not the first decoding step, the latest generated token is the new token generated in the previous decoding step; autoregressive decoding includes the cooperation between different layers of the large language model, and for the layer, the working process is as follows: first, the corresponding Q data and KV data are generated using the input hidden vector of the layer, and the KV data is added to the KV cache; when , the hidden vector is the word embedding vector corresponding to the first new token, and when , the hidden vector is the output of the layer; then, the KV data of the layer is prefetched: based on the similarity of the hidden vectors between layers, the hidden vector of the layer is used to calculate the Approximate Q data of the layer, for KV data offloaded to CPU memory, in units of KV header, combined with the first Approximate Q data of the layer, to determine whether the first Layer needs to prefetch KV data, if so, through top-k attention technology to retrieve the corresponding KV data as the first Layer KV data that needs to be prefetched, when the first Layer KV data prefetching is completed, start the second Layer KV data prefetching, wherein, through top-k attention technology to retrieve the corresponding KV data refers to retrieve the k KV data with the highest attention score; then, in the second Layer KV header resident in GPU memory for KV data retrieval, when , using the generated Q data and the KV data retrieved from the KV header resident in the GPU memory to complete the first Layer calculation, when , using the generated Q data, the KV data retrieved from the KV header resident in the GPU memory, and the second Layer KV data prefetched to complete the first Layer calculation; repeatedly until the last layer, using the output of the last layer to obtain the next new word.
[0043] As shown in Figure 3 , the main workflow of the decoding stage is shown, which involves the cooperation between different layers of the large language model. Here, only the decoding stage workflow of the current layer (referred to as the first Layer, same below) is introduced. Similarly, this process also involves the previous layer (referred to as the first Layer, same below) and the next layer (referred to as the first Layer, same below). The workflow of other layers is the same, so it is not described here. The main workflow is as follows: Step 21, generate the Q data and KV data of the first Layer new word.
[0044] In the embodiment of the application, the Q data and KV data generation process of the decoding stage is consistent with the existing process, which has been introduced in the foregoing background art. Here, it is not described again. The newly generated KV data is also added to the KV cache. Specifically, for the unloaded KV header, the newly generated KV data is added to its data buffer in the CPU memory; for the unloaded KV header, the newly generated KV data is added to its data buffer in the GPU memory.
[0045] Step 22, generate the approximate Q data of the first Layer.
[0046] In the embodiment of the application, the decoding stage generates the approximate Q data of the first Layer computation, the first Layer KV cache data reading to reduce the exposure of KV data reading overhead on the execution flow critical path, which is called KV data prefetching. Since the top-k attention algorithm is used to retrieve the key KV in the embodiment of the application, the first Layer Q data should be calculated in advance for top-k retrieval and prefetching. For this purpose, the embodiment of the application uses the similarity of the inter-layer hidden vectors of the large language model to use the first Layer hidden vector and the Q data mapping matrix of the first Layer to calculate the approximate Q data of the first Layer, and to perform top-k retrieval in advance to support prefetching.
[0047] Step 23, using the approximate Q data to perform cache query of the first Layer.
[0048] In the embodiment of the application, for the KV data that has been offloaded to the CPU memory, in addition to the KV prefetching technology, a dynamic approximate cache algorithm is also used, which saves the top-k KV data in the GPU memory in units of KV headers, further reducing the data reading overhead. The dynamic cache algorithm needs to use Q data to query whether the cached KV header in the GPU memory (specifically, the approximate cache data buffer of the first Layer, which will be described later) hits, if the KV header hits, the KV data of the KV header can be directly reused, if it does not hit, it cannot be reused, and the first Layer KV data needs to be prefetched. To combine with prefetching, the Q data used for querying here is for the first Layer, and the approximate Q data calculated in step 22 is used, specifically: in units of KV headers, combined with the approximate Q data of the first Layer, query in the approximate cache data buffer of the first Layer, to determine whether the cosine similarity of the Q data corresponding to the KV header and the approximate Q data of the first Layer is greater than a pre-set reuse threshold, if yes, it means that the KV header hits, otherwise, it means that it does not hit; the reuse threshold is calculated using the KV header importance algorithm, and the specific calculation method will be described later.
[0049] Step 24, using the approximate Q data to retrieve the top-k KV data of the first Layer that does not hit.
[0050] In the embodiment of the application, only when the KV data is not cached or does not hit (corresponding to cache initialization and cache update in turn, which will be described later) does top-k retrieval and prefetching need to be performed. The retrieval here is for the first The index of the KV data required to be read will be generated; this part of retrieval occurs in the GPU based on the retrieval metadata stored in the GPU memory, which has been described in the foregoing background and will not be described here.
[0051] Step 25, waiting for the first layer to complete the KV data prefetching.
[0052] In the embodiment of the present application, in addition to the KV data of the first layer, the KV data of the other layers may need to be prefetched. As described in the foregoing step 23, if the KV data is hit, it can be directly reused without the need for prefetching; if it is not hit, the KV data needs to be prefetched. Specifically, the KV data of the second layer is prefetched by the first layer, and the KV data of the third layer is prefetched by the second layer. To avoid the conflict and competition of PCIe bandwidth of CPU-GPU interconnection caused by different layers of prefetching, the KV data prefetching of the second layer needs to start after the KV data prefetching of the first layer is completed, and therefore, the second layer needs to use the CPU-GPU data transmission synchronization operation to wait for the completion of data prefetching.
[0053] Step 26, triggering the data prefetching of the second layer.
[0054] In the embodiment of the present application, when the KV data prefetching of the first layer is completed, the KV data prefetching of the second layer will be started immediately using the index of the KV data obtained in step 24. The data will be prefetched into the approximate cache data buffer area (located in the GPU memory) of the second layer.
[0055] Step 27, performing top-k retrieval in the KV header data of the second layer that is not offloaded.
[0056] In the embodiment of the present application, part of the KV header data will not be offloaded, which has been described in step 13 of the foregoing prefilling stage. Since these data are located in the GPU memory, there is no need for prefetching. Therefore, the embodiment of the present application uses the real Q data (i.e., the Q data generated in step 21) of the second layer to retrieve the index of the top-k KV data of these KV headers.
[0057] Step 28, completing the remaining calculation steps of the second layer.
[0058] In the embodiment of the application, the remaining steps of the decoding stage involve attention calculation, residual link, feedforward network calculation, etc., all of which are consistent with the original process before optimization except for attention calculation. Attention calculation needs to use data from two cache areas, which are: the data buffer area of the KV head unloaded by the first layer and the approximate cache data buffer area of the second layer. For the former, only the top-k KV data needs to be read, and their positions in the buffer area are identified by the index of the KV data given in step 27; the latter contains the top-k KV data required for calculation, without further screening.
[0059] The above-mentioned scheme provided by the embodiment of the application offloads most of the KV data to the CPU memory during the large language model inference process. For the reading of KV data from the CPU memory to the GPU, the system reduces the reading amount by using top-k attention. And the KV data that has been read to the GPU memory is cached, and a dynamic approximate cache algorithm is used to manage it according to the KV head as the basic unit. When the KV data needs to be read during the inference process, it is judged whether the data in the cache can be reused first. If the cache hits, the data in the cache is reused directly, reducing the data reading overhead from the CPU memory to the GPU memory. If the cache misses and the data cannot be reused, data prefetching is performed. Through the three optimization methods of selective offloading, dynamic approximate caching and data prefetching, the application can effectively reduce the memory occupied by KV data and minimize the KV data reading overhead from the CPU to the GPU, so that the inference performance reaches an ideal level.
[0060] In order to more clearly show the technical solutions provided by the application and the technical effects produced, the following mainly introduces the selective offloading of KV head, the dynamic approximate caching algorithm, and the prefetching and synchronization of KV data in detail. The processes not introduced in detail can be referred to the conventional technology.
[0061] I. Selective offloading of KV head
[0062] In the embodiment of the present application, the hit rate of the dynamic approximate caching algorithm is mainly determined by the importance of the KV head and the Q data distribution characteristics corresponding to the KV head (which determines the similarity level of the Q head in the continuous decoding step). The decoding step refers to the generation process of a word element, covering the working process of all layers of the model, and defines that for the same layer, the working process of the same layer is continuous when the adjacent word element is obtained by decoding; for example, assuming that the model has 3 layers, i.e. the 0th layer, the 1st layer and the 2nd layer, in the process of decoding to obtain two word elements a and b, six processes a0, a1, a2, b0, b1 and b2 are required, wherein a0, a1 and a2 refer to the working processes of the 0th layer, the 1st layer and the 2nd layer when the word element a is decoded, which is a decoding step, and b0, b1 and b2 refer to the working processes of the 0th layer, the 1st layer and the 2nd layer when the word element b is decoded, which is also a decoding step. In this example, a0 and b0 are defined as continuous processes, and a1 and b1, a2 and b2 are also defined as continuous processes.
[0063] Considering that there may be KV heads with high importance and low similarity level in the large language model, the cache hit rate of these KV heads is difficult to reach a high level (referred to as difficult-to-reuse KV heads), and the KV cache data of these KV heads still needs to be read from the CPU memory. Even if the prefetching technology is adopted, the reading delay of the KV data of these KV heads may still be difficult to be completely masked by calculation, which will slow down the inference speed. The embodiment of the present application further uses a selective offloading method to cope with this problem, and the difficult-to-reuse KV heads are resident in the GPU video memory, and only other KV heads are offloaded, and the overall process is as shown in Figure 4 The attention calculation unit is responsible for attention calculation in the decoding stage.
[0064] In the embodiment of the present application, the reuse difficulty of each KV head is quantified in advance: for each KV head of each layer, the corresponding average cosine similarity is obtained through offline analysis, wherein in the offline analysis, a plurality of sequences are randomly sampled from the selected evaluation data set and input into the large language model, and for each KV head of each layer, the average cosine similarity is calculated using all the outputs; for a single KV head i, the average cosine similarity is denoted as , the reuse threshold and the error tolerance difference are set, and the reuse difficulty is: ; wherein, is the reuse difficulty of a single KV head i. The physical meaning of the above formula is that the KV head with an actual cosine similarity lower than the reuse threshold is considered to be difficult to reuse, and the greater the numerical difference between the two, the higher the difficulty of cache reuse. The reuse threshold can be calculated by using the KV head importance algorithm, which will be introduced later.
[0065] For example, the offline analysis can use 30 sequences randomly sampled from the LongBench benchmark dataset (which is a bilingual and multi-task dataset with various sequences of different lengths, distributions, patterns, languages, and domains for comprehensive evaluation of long context understanding ability), The value of can be 0.05.
[0066] For each layer, after quantifying the multiplexing difficulty of each KV head, the KV heads are ranked in descending order according to the multiplexing difficulty; the number of KV heads with a multiplexing difficulty greater than 0 (i.e. ) is counted .
[0067] In addition, the number of KV heads that can be prefetched (the read latency can be completely masked by calculation) needs to be analyzed in advance. Specifically, for the user-set running configuration, the large language model is pre-run once for inference, and the decoding time of each layer is counted . Then, combined with the size of the KV data corresponding to each KV head in each layer and the PCle bandwidth, the number of KV heads that can be prefetched is calculated and represented as: ; wherein, is the number of KV heads that can be prefetched in each layer, is the size of the KV data corresponding to a single KV head, is the PCle bandwidth, and PCle is a high-speed serial computer expansion bus standard.
[0068] After calculating the number and , the selective offloading decision can be made. For a certain layer of the large language model, the number of KV heads that can be prefetched in the corresponding layer is determined by subtracting the number of KV heads that cannot be prefetched from the total number of KV heads in the layer, and the number of KV heads that reside in the GPU memory is represented as: ; wherein, the max function outputs the maximum value in the parentheses.
[0069] After determining the number , the corresponding number of KV heads from the front end of the sorted results of the KV heads are selected, and the selected KV heads are used as the KV heads that reside in the GPU memory, and the KV data corresponding to the selected KV heads are stored in the GPU memory. If the GPU memory is insufficient and cannot accommodate the KV data of KV heads, then within the range that can be accommodated by the GPU memory, the KV data of the batch of KV heads with the highest multiplexing difficulty are preferentially selected from the sorted results of the KV heads to reside in the GPU memory to maximize the performance gain. The KV data of the remaining KV heads are all offloaded to the CPU memory, and the selective offloading of the KV heads is completed.
[0070] II. Dynamic approximate caching algorithm.
[0071] The basis of the dynamic approximate caching strategy is that for the same Q head of the same layer of a large language model, the Q data in consecutive decoding steps has high directional similarity (quantified using cosine similarity, which is a floating-point number between -1 and 1, the closer the cosine similarity of two vectors to 1, the more similar their directions are). For example, on the Llama3-8B-1048K model, the cosine similarity of Q vectors in consecutive inference steps is higher than 0.85 for 94% of Q heads. The directional similarity of Q data makes the KV data index generated by top-k retrieval using Q data and K data also have high similarity. Specifically, top-k retrieval uses Q data and K data to calculate the score of each KV data by the following formula: where q is the Q data of the latest token in the decoding stage, which generally represents the K data of any token in the KV cache, , corresponding to q, the modulus of q, and the cosine value of the angle between q and . Obviously, given a set of KV data of tokens, the relative size order of their scores is only affected by and because they share the same in their calculation process. Then, if two Q data are directionally similar, their angle size with the same K data will also be relatively close, that is, by calculating with their respective results will be highly similar. Therefore, the relative size order of the final calculated scores will also be highly similar, making the results of top-k retrieval close. From the above analysis, an important conclusion can be drawn: in the inference process of a large language model, affected by the directional similarity of Q data, the top-k KV data required in consecutive decoding steps has high overlap. In this case, the top-k KV data read to the GPU memory by the historical decoding steps can be reused by the later decoding steps without re-reading.
[0072] Based on the above conclusion, the embodiment of the present application proposes a dynamic approximate caching algorithm based on Q similarity, as shown in FIG. 1, which shows the principle of the dynamic approximate caching algorithm. It mainly includes the following steps: cache initialization; cache query; cache update. Figure 5
[0073] 1. Cache initialization.
[0074] At the start of the inference process, the cache is empty. In this embodiment of the invention, in the first decoding step, each layer except layer 0 is initialized with cache, that is: the top-k KV data is retrieved and read from the KV header unloaded to the CPU memory using the generated Q data, and stored in the approximate cache data buffer of the corresponding layer to complete the cache initialization; wherein, the first decoding step refers to the first decoding step that occurs after the pre-filling stage ends, which uses the first new lexical generated in the pre-filling stage to generate the next new lexical.
[0075] 2. Cache query.
[0076] After the first decoding step, for a new decoding step, every key-value header in every layer of the large language model except for layer 0 needs to be cached and queried. For the current layer (e.g., layer 0), the cache lookup is performed. (Layer), using KV headers as the unit, using the calculated next layer (e.g., the first layer). The cosine similarity is calculated between the approximate Q data of the first layer and the Q data corresponding to the KV data cached in the approximate cache data buffer of the next layer. If the cosine similarity is higher than the preset reuse threshold, it is considered a cache hit, indicating that the KV data of the corresponding KV header can be directly reused; otherwise, it is considered a cache miss and the cache update step is initiated.
[0077] Those skilled in the art will understand that the KV data cached in the approximate cached data buffer is partitioned by KV header, and each KV header has a Q data as its label (saved during initialization or update). In a single decoding step, one Q header corresponds to only one Q data. During cache query, for each KV header, the cosine similarity between the approximate Q data and the Q data label in the cache is calculated to determine whether a match has occurred. For the GQA model, one KV header corresponds to multiple Q headers, so it has multiple Q data labels. The cosine similarity calculation will produce multiple results, which need to be aggregated. This will be introduced in detail later in the section on supporting grouped query attention.
[0078] 3. Cache update.
[0079] For KV headers that are determined to be missed in the cache lookup, the next layer to be read is retrieved using approximate Q data (e.g., the first...). The index of the KV data of the layer is obtained (see the explanation of step 24 above). Then, the corresponding KV data is read from the CPU memory through KV data prefetching (see the explanation of step 26 above), and overwritten to the approximate cache data buffer of the next layer (see the explanation of step 26 above) to replace the KV data of the missing KV header, thus completing the cache update.
[0080] Figure 5 In the input is referred to as the Q data (specifically, approximate Q data) input by the algorithm, the Q data in the cache is referred to as the K data (key data), is referred to as the V data (value data), and the two are the KV data in the approximate cache data buffer, is referred to as the Q data corresponding to the KV header, and is used and The cosine similarity and the size of the multiplexing threshold are used to determine whether a hit occurs. If so, the corresponding KV data can be multiplexed. Otherwise, the corresponding KV data needs to be retrieved from the KV data of the CPU (K and V in the bottom block refer to the retrieved KV data), and the approximate cache data buffer is updated.
[0081] The above three steps are collectively referred to as cache management. Compared with traditional LRU (least recently used algorithm) / LFU (least frequently used algorithm) cache algorithms, the dynamic approximate cache algorithm provided by the present application greatly simplifies the query and update process of the cache. The query process only needs to calculate the cosine similarity based on the Q header as the basic unit, which is lower in cost and more suitable for the hardware characteristics of the GPU than the table lookup (such as a hash table) of the traditional cache algorithm. In the cache update process, the KV data is directly overwritten to the data buffer of the cache, which reduces a data copy within the GPU memory compared with the traditional algorithm. At the same time, the cache metadata update only needs to overwrite the Q data, while the traditional algorithm needs to perform a low concurrency and time-consuming linked list update operation by the CPU. The cache metadata here, i.e., the data that assists the cache in querying and updating, is the Q data possessed by each KV header as mentioned above. In cache querying, the cosine similarity is calculated between the next layer of approximate Q data and the Q data in the approximate cache data buffer. The former is new data newly input, and the latter is old data existing in the cache. When a miss occurs, not only the KV data needs to be re-read and updated, but also the Q data in the cache is considered to be invalid, and therefore needs to be updated together. In summary, the dynamic approximate cache algorithm provided by the present application significantly reduces the overhead of cache management. In the single-layer single-step reasoning of a large language model, the total time consumption is only 5-10 microseconds, which accounts for less than 1% of the total time consumption and can be ignored. Therefore, compared with traditional algorithms, the dynamic approximate cache algorithm is more suitable for large language model reasoning with strict time consumption requirements.
[0082] Preferably, the application further provides a calculation scheme of multiplexing threshold for the aforementioned cache query step. Specifically, a multiplexing threshold is set for each KV head in advance to determine whether the cache hits in the cache query process and is applied to the selective offloading of the KV head. The embodiment of the application adopts an algorithm (referred to as a KV head importance algorithm) capable of quantitatively analyzing the contribution of the KV head to the inference accuracy to assist the multiplexing threshold setting. Specifically, before the inference starts, the embodiment of the application analyzes the importance of all KV heads in a given large language model using the KV head importance algorithm. For a single KV head i, the importance is denoted as , , the closer to 1, the more important the KV head is, and then is used as the upper limit of the multiplexing threshold (exemplarily, the value of may be set to 0.5), and the multiplexing threshold of each KV head is set using the following formula : ; ; ; wherein the parameter , p is an attenuation speed reinforcement index, exemplarily, p is 2 or 3, which helps to highlight the difference between the heads; and are intermediate variables generated in the calculation process.
[0083] The principle of the above formula is to convert the multiplexing threshold into an angle, and then attenuate it according to the importance of the KV head. The more important the KV head is, the closer its final multiplexing threshold is to the upper limit ; the less important it is, the closer it is to the lower limit -1.
[0084] Exemplarily, the KV head importance algorithm can select the Duo-Attention algorithm, which is an algorithm capable of quantitatively analyzing the contribution of the KV head to the inference accuracy.
[0085] Preferably, the dynamic approximate cache algorithm provided by the application also supports group query attention (which has been introduced in the background section). The cosine similarity is calculated based on the Q head, but the multiplexing threshold is set based on the KV head, and the Q head and the KV head are not one-to-one corresponding in the model using the GQA technology. Therefore, it is necessary to aggregate the calculated cosine similarity within the Q head group. Based on the KV head importance algorithm, the embodiment of the application also analyzes the importance of all Q heads. In the cache query step, the importance is preferably used as the weight, and the harmonic mean method is used to aggregate the Q heads within the group, and the specific formula is: ; wherein Importance of Q head j, Cosine similarity of Q data corresponding to Q head j and approximate Q data, m is the number of Q heads corresponding to one KV head, s is the aggregated cosine similarity, and by comparing with a multiplexing threshold to determine whether the m Q data corresponding to the KV head hit. The harmonic mean helps to highlight the influence of the Q head with the lowest similarity in the group on the aggregation result, ensuring the accuracy of the cache hit judgment.
[0086] III. KV data prefetching and synchronization.
[0087] In the embodiment of the application, the KV data prefetching technology is used to read the KV head that fails to cache hit from the CPU memory. The algorithm principle of this technology is derived from InfiniGen (it is a KV cache offload reasoning system that uses prefetching and top-k attention technology to improve reasoning performance). Compared with InfiniGen, the embodiment of the application has two innovations, one is to use a zero-copy data transmission method for data prefetching, and the other is to use a GPU-centered data transmission synchronization method.
[0088] In the embodiment of the application, the zero-copy data transmission method refers to the use of a zero-copy method based on GPU Direct (direct communication technology) to perform KV data prefetching between CPU and GPU. Zero-copy means that during data copying, data is directly copied from the source location to the target location without any temporary buffer for transfer. In the embodiment of the application, the GdrCopy library based on GPU Direct (it is an open source library mainly used to realize high-speed data transmission between GPU memory and CPU memory) is used to directly copy the required KV data from the CPU memory to the specified GPU memory buffer without first collecting the discrete distributed KV data into a continuous data block in the CPU memory for copying. This zero-copy data transmission method bypasses the high-cost CPU operation and provides a CPU-to-GPU transmission bandwidth of up to 21 GB / sec (gigabyte per second).
[0089] In the embodiment of the application, the GPU-centered data transmission synchronization method is also used to optimize the KV data transmission process, such as Figure 6The data transmission between the CPU and the GPU needs to be synchronized, and the progress of the data transmission is informed to the GPU, so that the GPU is in a continuous waiting state before the data transmission is completed. The traditional data transmission synchronization method takes the CPU as the core, blocks the start of the GPU kernel function responsible by the CPU, and blocks the operation of the subsequent kernel function of the GPU, so as to make the GPU fall into a continuous waiting state. However, this will cause the GPU to be idle for a period of time after the synchronization is completed: the GPU kernel function can start to execute only after the CPU starts, and the CPU cannot start the subsequent kernel function until the synchronization is completed. After the synchronization is completed, this start process will be exposed to the critical path of the execution flow, causing the GPU to be idle. The embodiment of the application changes the data transmission synchronization to be centered on the GPU, solving this problem. Specifically, using the NVIDIA (NVIDIA) Unified Virtual Addressing (UVA) technology, a shared memory that can be directly accessed by the CPU and the GPU is established in the CPU memory, and a signal variable is set in the shared memory to describe the data transmission state; for a KV header, when the transmission is not completed, the value of the signal variable is 0; when the transmission is completed, the value of the signal variable is 1; the GPU polls the signal variable through a kernel function, if the value of the signal variable is 0, the GPU will be in a busy waiting state, otherwise, the GPU exits the polling and executes the operation of the subsequent kernel function. This synchronization strategy does not block the start of the kernel function after the synchronization of the CPU, completely eliminating the GPU idle caused by the data transmission synchronization.
[0090] Based on the above introduction, the application provides an overall scheme architecture as shown in the figure. Figure 7 Figure 7 Among them, the KV cache manager is mainly responsible for managing the storage, reading, caching of all KV data, including selective offloading, configuration and execution of dynamic approximate caching algorithm, and retrieval of top-k KV cache data; the data transmission engine is mainly responsible for pre-fetching KV cache data from the CPU memory to the GPU memory, and performing data transmission in a zero-copy manner; it contains a plurality of data transmission threads; the data transmission synchronization controller is mainly responsible for synchronizing the CPU-GPU data transmission, and adopts a synchronization method centered on the GPU, without affecting the start of the GPU kernel function after the synchronization; the kernel function start thread is located on the CPU and is responsible by an independent thread, without being disturbed by the data transmission synchronization; the inference execution module (the weights of the large language model are included in the inference execution unit) is mainly responsible for executing the inference operation of the large language model.
[0091] Through the above description of the embodiments, those skilled in the art can clearly understand that the above embodiments can be implemented by software, or by using software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solutions of the above embodiments can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, mobile hard drive, etc.), including several instructions to cause a computer device (such as a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0092] Example 2 This invention also provides a large language model inference optimization system, which is mainly used to implement the methods provided in the foregoing embodiments, such as... Figure 8 As shown, the system mainly includes: The pre-filling unit based on selective offloading of KV headers (applied to the pre-filling stage) is used to combine the difficulty of KV header reuse, cache part of the KV data of the KV header in GPU memory, and offload the KV data of the other part of the KV header to CPU memory; where KV header is a type of attention header in large language models, KV stands for key and value, GPU stands for graphics processor, and CPU stands for central processing unit. The decoding unit based on dynamic approximation caching and data prefetching (applied to the decoding stage) is used to perform autoregressive decoding by combining the latest lexical units. The latest lexical unit is the first new lexical unit generated in the pre-filling stage or the new lexical unit generated in the previous decoding step. The working process of the current layer of the large language model in a single decoding step is as follows: Q data and KV data are generated using the latent vector of the current layer. The KV data is added to the GPU memory or CPU memory. The latent vector of the current layer is the word embedding vector of the first new lexical unit or the calculation result of the previous layer. Q is the query. Based on the latent vector of the current layer, the KV head is used as the unit to determine whether the next layer needs to prefetch KV data. If so, the corresponding KV data is retrieved in the CPU memory. After the KV data prefetching of the current layer is completed, the prefetching of the KV data of the next layer is started. The Q data generated by the current layer is used to retrieve KV data in the GPU memory. The calculation of the current layer is completed by combining the prefetched KV data. This process is repeated until the last layer. The calculation result of the last layer is used to obtain the new lexical unit.
[0093] Since the main technical details of the above system have been described in detail in the previous embodiments, they will not be repeated here.
[0094] Those skilled in the art will understand that, for the sake of convenience and brevity, the above-described division of functional modules is used as an example. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the system can be divided into different functional modules to complete all or part of the functions described above.
[0095] Embodiment three The application further provides a processing device, as shown in the accompanying drawings, which mainly comprises one or more processors, a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method provided by the foregoing embodiments. Figure 9
[0096] Further, the processing device further comprises at least one input device and at least one output device; in the processing device, the processor, the memory, the input device and the output device are connected through a bus.
[0097] In the embodiments of the application, the specific types of the memory, the input device and the output device are not limited; for example: The input device can be a touch screen, an image acquisition device, a physical button or a mouse, etc. The output device can be a display terminal. The memory can be a random access memory (RAM) or a non-volatile memory such as a disk memory.
[0098] Embodiment four The application further provides a readable storage medium storing a computer program, when the computer program is executed by a processor, the method provided by the foregoing embodiments is implemented.
[0099] In the embodiments of the application, the readable storage medium as the computer readable storage medium can be arranged in the foregoing processing device, for example, as the memory in the processing device. In addition, the readable storage medium can also be a U disk, a mobile hard disk, a read-only memory (ROM), a magnetic disk or an optical disk and various storage program codes.
[0100] The above description is only the preferred embodiment of the application, but the protection scope of the application is not limited to this, any changes or replacements within the technical range disclosed by the application can be easily thought by those skilled in the art, which should be covered in the protection scope of the application. Therefore, the protection scope of the application should be subject to the protection scope of the claims. The information disclosed in the background section of this document is only intended to deepen the understanding of the general background of the application, and should not be regarded as acknowledging or implying in any form that the information constitutes the prior art known by those skilled in the art.
Claims
1. A large language model inference optimization method, characterized in that, The method comprises: a pre-filling stage, in combination with the multiplexing difficulty of a KV head, a part of KV data of the KV head is cached in GPU memory, and another part of KV data of the KV head is unloaded to CPU memory; wherein the KV head is a type of attention head in a large language model, KV is key and value, GPU is a graphics processing unit, and CPU is a central processing unit; a decoding stage, in combination with the latest word element, autoregressive decoding is performed, and the latest word element is the first new word element generated in the pre-filling stage or a new word element generated in the last decoding step; the working process of the current layer of the large language model in a single decoding step is as follows: Q data and KV data are generated by using the hidden vector of the current layer, the KV data is added to the GPU memory or the CPU memory, the hidden vector of the current layer is the word embedding vector of the first new word element or the calculation result of the last layer, and Q is a query; based on the hidden vector of the current layer, whether the next layer needs to prefetch KV data is determined by taking the KV head as a unit, if yes, the corresponding KV data is searched in the CPU memory, after the KV data of the current layer is prefetched, the prefetching of the KV data of the next layer is started; the KV data is searched in the GPU memory by using the Q data generated by the current layer, and the calculation of the current layer is completed in combination with the prefetched KV data; the process is repeatedly performed until the last layer, and the calculation result of the last layer is used to obtain a new word element.
2. The large language model inference optimization method of claim 1, wherein, The method comprises: the multiplexing difficulty of each KV head is quantified, and the number of KV heads that can be prefetched in each layer is calculated; in combination with the number of KV heads that can be prefetched in each layer and the multiplexing difficulty of the corresponding KV head, the KV heads that reside in the GPU memory are selected, and the corresponding KV data of the KV heads is cached in the GPU memory, and the KV data of the remaining KV heads is unloaded to the CPU memory; wherein each layer of the large language model has multiple attention heads, which are divided into two types of Q heads and KV heads; The method comprises:
3. The large language model inference optimization method of claim 2, wherein, the multiplexing difficulty of each KV head is quantified, and the number of KV heads that can be prefetched in each layer is calculated; in combination with the number of KV heads that can be prefetched in each layer and the multiplexing difficulty of the corresponding KV head, the KV heads that reside in the GPU memory are selected, and the corresponding KV data of the KV heads is cached in the GPU memory, and the KV data of the remaining KV heads is unloaded to the CPU memory; wherein each layer of the large language model has multiple attention heads, which are divided into two types of Q heads and KV heads; The method comprises: For a single KV head i, let its average cosine similarity be denoted as The corresponding multiplexing threshold is calculated in advance using the KV head importance algorithm and an error tolerance difference is set, then the multiplexing difficulty is: ; wherein, multiplexing difficulty for a single KV head i; The calculating the number of KV heads that each layer can prefetch comprises: performing one inference pre-run on the large language model, and counting the decoding time of each layer In combination with the size of the KV data corresponding to a single KV head in each layer and the PCIe bandwidth, the number of KV heads that can be prefetched is calculated and represented as: ; wherein, is the number of KV heads that can be prefetched for each layer, is the size of KV data corresponding to a single KV head, is the PCle bandwidth, PCle is a high-speed serial computer expansion bus standard.
4. The large language model inference optimization method of claim 1, wherein, The method comprises: The method comprises: Let the current layer be denoted as the th. The next layer is denoted as the first layer. Layer, based on the similarity of latent vectors between layers, utilizes the 1st layer... The latent vectors of the layer are calculated to obtain the first layer. The approximate Q data of the layer is used with a dynamic approximation caching algorithm, taking the KV header as the unit, to determine the first... The system checks if the key-value (KV) data in the KV header of the approximate cached data buffer of the layer is matched. If the KV header is matched, it means that the corresponding KV data in the KV header can be reused; if it is not matched, it cannot be reused. This is combined with the first... The approximate Q-data of the layer, using the key-value header as the unit, is retrieved from CPU memory using top-k attention technology, and used as the first layer to be prefetched. Layer KV data, wherein retrieving the corresponding KV data through top-k attention technology means retrieving the k KV data with the highest attention scores, the approximate cache data buffer is located in GPU video memory and is used to cache prefetched KV data, wherein the KV data is partitioned according to KV header, and the approximate cache data buffer is initialized in the first decoding step; The algorithm uses a dynamic approximation caching mechanism, with the key-value header as the unit, to determine the first... Whether the KV data in the KV header of the approximate cached data buffer of the layer is hit includes: taking the KV header as the unit, combined with the first Approximate Q-data of the layer, in the first layer The query is performed in the approximate cached data buffer of the layer to determine whether the Q data corresponding to the KV header is the same as that of the first layer. If the cosine similarity of the approximate Q data of the layer is greater than a preset reuse threshold, it indicates that the KV header has been hit; otherwise, it indicates that it has not been hit. The reuse threshold is calculated using the KV header importance algorithm.
5. The large language model inference optimization method of claim 4, wherein, The dynamic approximate cache algorithm comprises cache initialization, cache query and cache update. The cache initialization is that, in the first decoding step, each layer except the 0th layer is respectively subjected to cache initialization, that is, the top-k KV data is retrieved and read from the KV header unloaded to the CPU memory by using the generated Q data, and is saved in the approximate cache data buffer of the corresponding layer, so that the cache initialization is completed; wherein the first decoding step refers to the first decoding step occurring after the pre-filling stage is completed. The cache query is that, for a new decoding step, each KV header of each layer except the 0th layer in the large language model is subjected to cache query; the decoding step refers to the generation process of a word element, and covers the working process of all layers of the model; for the current layer, the cosine similarity is calculated by using the approximate Q data of the next layer calculated and the Q data corresponding to the KV header in the approximate cache data buffer of the next layer, so as to determine whether the KV data is hit; if not, the cache update step is entered; The cache update is that, for the KV header determined as not hit in the cache query, the index of the next layer KV data to be prefetched is retrieved, the corresponding KV data is read from the CPU memory by KV data prefetching, and is overwritten to the approximate cache data buffer of the next layer, so that the cache update is completed.
6. The large language model inference optimization method according to any one of claims 3-5, characterized in that, The calculation method of the multiplexing threshold comprises: The importance of all KV heads in the large language model is analyzed using the KV head importance algorithm. For a single KV head i, the importance is recorded as , , and then, taking as the upper limit of the reuse threshold, the reuse threshold of the KV head is set as follows : ; ; ; where the parameters p is an attenuation rate enhancement exponent, and are intermediate variables generated during the calculation process.
7. The large language model inference optimization method of claim 1, wherein, In the KV data prefetching process, a zero-copy method based on GPU Direct is adopted to perform the prefetching of KV data between CPU and GPU, and a GPU-centered data transmission synchronization method is used; The zero-copy method based on GPU Direct directly copies the required KV data from the CPU memory to the specified GPU memory buffer; the GPU Direct is a GPU direct communication technology; The GPU-centered data transmission synchronization method comprises: setting the data transmission synchronization as GPU-centered, using the NVIDIA unified virtual addressing technology to establish a shared memory in the CPU memory which can be directly accessed by the CPU and the GPU, and setting a signal variable in the shared memory to describe the data transmission state; for a KV header, when the transmission is not completed, the signal variable value is 0; when the transmission is completed, the signal variable value is 1; the GPU polls the signal variable by a kernel function, if the signal variable value is 0, the GPU will be in busy waiting state, otherwise, the GPU exits the polling.
8. A large language model inference optimization system, characterized in that, The method for implementing any one of claims 1-7 comprises: The pre-filling unit based on selective unloading of KV headers is used to buffer the KV data of a part of KV headers in the GPU memory and unload the KV data of another part of KV headers to the CPU memory in combination with the multiplexing difficulty of the KV headers; wherein the KV header is a type of attention header in the large language model, KV is key and value, GPU is a graphics processing unit, and CPU is a central processing unit. A decoding unit based on dynamic approximate caching and data prefetching is used for autoregressive decoding combined with the latest token, which is the first new token generated in the prefilling stage or the new token generated in the last decoding step; the working process of the current layer of the large language model in a single decoding step is as follows: Q data and KV data are generated by using the hidden vector of the current layer, and the KV data is added to the GPU memory or CPU memory; the hidden vector of the current layer is the word embedding vector of the first new token or the calculation result of the last layer, and Q is a query; based on the hidden vector of the current layer, it is judged whether the next layer needs to prefetch KV data by taking the KV head as a unit; if yes, the corresponding KV data is searched in the CPU memory; after the KV data of the current layer is prefetched, the prefetching of the KV data of the next layer is started; the KV data is searched in the GPU memory by using the Q data generated by the current layer, and the calculation of the current layer is completed combined with the prefetched KV data; the process is repeated until the last layer, and the calculation result of the last layer is used to obtain a new token.
9. A processing device, characterized by Comprise: One or more processors; Memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method of any one of claims 1-7.
10. A readable storage medium, storing a computer program, characterized in that, When the computer program is executed by the processor, the method of any one of claims 1-7 is implemented.
Citation Information
Patent Citations
Language model reasoning optimization method, electronic equipment, storage medium and program product
CN118396128A
Model reasoning method and device and electronic equipment
CN119129746A
Method, medium, device and program product for calculating attention score based on CPU
CN120297328A
Text inference acceleration method applied to large language model and related device
WO2025194554A1