Model reasoning method, computer program product and chip
By storing the hidden state of each word element during the model inference process and restoring the key-value cache based on this state, the problem of computing time and bandwidth limitation in the key-value cache recovery process in the prior art is solved, and more efficient model inference is achieved.
Patent Information
- Application Number
- CN202510323473.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-06-24
AI Technical Summary
In the key-value cache recovery process in the model inference stage, the calculation time overhead or bandwidth limitation leads to the first word delay problem, resulting in low model inference efficiency.
During the model inference process, the hidden state of each word element at each layer is stored, and the key-value cache of each word element at each layer of the model is restored based on the hidden state and key-value projection weight matrix, using the chip's computing resources and data transmission bandwidth resources.
Through this method, the recovery efficiency of key-value cache can be greatly improved, the calculation time and data transmission delay can be reduced, and the inference efficiency of the model can be improved.
Smart Images

Figure CN120197702A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of large models, and in particular, to a model inference method, a computer program product, and a chip. Background Art
[0002] In related technologies, during the model inference stage, there are mainly two methods for restoring the key-value cache (i.e., KV cache). The first method is to recalculate the key-value cache of each token based on the historical text sequence. This method requires repeating all previous calculation operations during the restoration process, resulting in a high computational time overhead. The second method is to offload the key-value cache, that is, save the key-value cache to an external storage unit of the model inference chip (such as an SSD, etc.). When the key-value cache is needed later, the key-value cache is transferred back from the external storage unit to the model inference chip memory. This method will cause a first-word delay problem due to limited bandwidth during the transmission process. Therefore, it is necessary to provide a more efficient key-value cache restoration scheme to improve the efficiency of the model inference process. Summary of the Invention
[0003] This application provides a model inference method, a computer program product, and a chip.
[0004] According to the first aspect of the embodiments of this application, a model inference method is provided. The model is a model based on the Transformer architecture. The method is executed by a target chip, and the method includes:
[0005] During the model inference stage, for the current network layer of the model, if the calculation of the current network layer requires the key-value cache of the tokens in the historical text sequence, obtain the pre-stored hidden state of the tokens in the current network layer from an external storage unit of the target chip;
[0006] Based on the hidden state and the key-value projection weight matrix of the current network layer, determine the key-value cache of the tokens in the current network layer, and complete the calculation of the current network layer based on the key-value cache; where the key-value projection weight matrix is learned during the model training stage.
[0007] According to the second aspect of the embodiments of this application, a computer program product is provided. The computer program product includes a computer program, and when the computer program is executed, it implements the method mentioned in the first aspect above.
[0008] According to the third aspect of the embodiments of this application, a chip is provided. The chip includes a processor, a memory, and computer instructions stored in the memory and executable by the processor. When the processor executes the computer instructions, it can implement the method mentioned in the first aspect above.
[0009] According to a fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided, on which computer instructions are stored, and when the computer instructions are executed, the method mentioned in the first aspect above is implemented.
[0010] In the embodiments of the present application, during the model inference process, the hidden state of each token at each layer (i.e., the input of each token at each layer) can be stored. When the key-value cache of these tokens is needed, the key-value cache of each token at each layer of the model can be restored based on the hidden state and the key-value projection weight matrix (i.e., the model parameters learned during the model training process). The key-value cache restoration scheme provided by the embodiments of the present application can utilize the computing resources and data transmission bandwidth resources of the chip at the same time with lower overhead, rather than using only one type of resource to implement the restoration of the key-value cache, which can greatly improve the restoration efficiency of the key-value cache, and further improve the inference efficiency of the model.
[0011] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The drawings herein are incorporated into the specification and form a part of this specification. These drawings illustrate embodiments consistent with the present application and, together with the specification, are used to explain the technical solutions of the present application.
[0013] Figure 1 It is a schematic diagram of an application scenario of the embodiments of the present application.
[0014] Figure 2 It is a flowchart of a model inference method according to the embodiments of the present application.
[0015] Figure 3 It is a schematic diagram of a two-stage storage strategy for hidden states according to the embodiments of the present application.
[0016] Figure 4 It is a schematic diagram showing that the computing speed in the embodiments of the present application is greater than the I / O speed, resulting in an idle period of computing resources.
[0017] Figure 5 It is a schematic diagram showing that the computing speed in the embodiments of the present application is less than the I / O speed, resulting in an idle period of I / O resources.
[0018] Figure 6 It is a schematic diagram of the combined use of the strategy of restoring the key-value cache based on the hidden state and the strategy of recalculating the key-value cache based on the historical text sequence in the embodiments of the present application.
[0019] Figure 7 It is a schematic diagram of the combined use of the strategy of restoring the key-value cache based on the hidden state and the strategy of reading the key-value cache from an external storage unit in the embodiments of the present application.
[0020] Figure 8 It is a schematic diagram of the logical structure of a chip according to an embodiment of the present application. Detailed implementation manners
[0021] Here, exemplary embodiments will be described in detail, and examples thereof are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0022] The terms used in the present application are for the purpose of describing specific embodiments only and are not intended to limit the present application. The singular forms "a", "the", and "said" used in the present application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and includes any or all possible combinations of one or more of the associated listed items. Additionally, the term "at least one" as used herein represents any one of a plurality or any combination of at least two of a plurality.
[0023] It should be understood that although the terms first, second, third, etc. may be used in the present application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0024] In order to enable those skilled in the art of the present technology to better understand the technical solutions in the embodiments of the present application and to make the above-mentioned objects, features, and advantages of the embodiments of the present application more obvious and understandable, the technical solutions in the embodiments of the present application will be further described in detail below with reference to the drawings.
[0025] Models adopting the Transformer architecture such as large language models usually include N repeated network layers (i.e., Transformer layers). Each Transformer layer contains two main modules: an attention module (AttentionModule) and a feed-forward network module (Feed-Forward Network, FFN). The attention module is used to establish an information interaction relationship between tokens, and the feed-forward network module is used to perform non-linear transformation and enhancement on the feature information of the tokens.
[0026] During the forward computation of the model, each token is represented as a high-dimensional vector at each layer of the model, called Hidden States. Among them, the hidden states of the first layer come from the input embedding table, that is, the embedding vectors of each token, while the hidden states of subsequent layers come from the output of the previous layer. In the attention module, the hidden state of each token is mapped into three tensors: Q (Query, query vector), K (Key, key vector), and V (Value, value vector), and information interaction is completed through the attention mechanism.
[0027] For example, assume that the hidden state of the i-th token in the text sequence generated by the model at the L-th layer of the model is (i.e., the output of the L-1-th layer of the model). Then, in the computation of the L-th layer, the query projection weight matrix key projection weight matrix value projection weight matrix of the model at the L-th layer and the hidden state of this token at the L-th layer are used to determine the query vector key vector and value vector of this token at the L-th layer respectively. Specifically as follows:
[0028]
[0029] Among them, the query projection weight matrix key projection weight matrix value projection weight matrix are model parameters learned during the model training phase.
[0030] After obtaining the Q, K, and V of each token, each token calculates the inner product of its own Q and the K of all tokens in the historical text sequence to generate attention scores. The attention scores are subjected to softmax normalization processing, and the V of other tokens is weighted and calculated according to the normalized scores, and the final attention result is generated through the output projection matrix. The feed-forward network module generally consists of two linear transformations (linear projections) and an activation function, and its output is used as the input of the next layer.
[0031] Although the Transformer architecture has efficient feature modeling capabilities, during the model inference phase, since the attention calculation of each token needs to depend on the K values and V values of all tokens before this token in the historical text sequence, for the inference of long text sequences, this calculation method leads to significant time and space overheads, limiting the inference efficiency of the model.
[0032] To solve the above problems, related technologies have proposed an optimization method called Key-Value Cache (KV Cache), that is, during the inference process, the K values and V values of each token are cached on the model inference chip. For subsequent generated tokens, in the calculation of the attention mechanism, the existing K values and V values in the cache can be directly accessed without repeated calculation, thus significantly improving the inference efficiency. Therefore, the key-value cache has become an intermediate state that needs to be maintained during the model inference process.
[0033] For example, assume that a large language model based on the Transformer architecture is used for text generation, and the task is to generate the sentence "I like to eat apples". The model outputs the words "I", "like", "eat", and "apples" one by one. During the generation process, each time the model generates a word, it needs to calculate the attention scores between the query vector (Q) of the current word and the key vectors (K) and value vectors (V) of all historical words. Without the key-value cache, each time a new word is generated, the model needs to recalculate the K and V of all historical words, which will result in a huge computational overhead.
[0034] Taking the key-value cache of a certain layer as an example:
[0035] When generating "I", it is necessary to calculate Q_I, K_I, and V_I.
[0036] When generating "like", without caching, it is necessary to recalculate K_I, V_I, and then calculate the attention scores between Q_like and K_I, K_like, V_I, and V_like.
[0037] The core idea of the key-value cache is to cache the K and V of historical tokens to avoid repeated calculation. The specific process is as follows:
[0038] When generating "I", calculate Q_I, K_I, and V_I, and store K_I and V_I in the key-value cache.
[0039] When generating "like", only calculate Q_like, obtain K_I and V_I from the cache, calculate the attention score between Q_like and K_I, calculate K_like and V_like, and update the cache.
[0040] The generation process of subsequent tokens is similar. By continuously caching the K values and V values of historical tokens, the computational efficiency can be improved.
[0041] However, as the length of the text sequence increases, the K values and V values that need to be cached also increase. However, the storage space of the chip (such as GPU) used to execute the model inference method is limited, and the key-value cache cannot be fully stored in the chip's memory. Therefore, during the inference process, it is necessary to restore the key-value cache.
[0042] Currently, there are mainly two methods for restoring the key-value cache. The first method is to recalculate the key-value cache for each token based on the historical text sequence. Although this method is simple and direct, during the restoration process, it is necessary to repeat all previous calculation operations, including the calculations in the attention module and the feed-forward network, which will bring a relatively high computational time overhead, especially when dealing with contexts of longer input lengths. The second method is to offload the key-value cache. This method is to save the key-value cache in an external storage unit outside the model inference chip (such as an SSD, etc.). When the key-value cache is needed subsequently, the key-value cache is transferred back to the memory of the model inference chip from the external storage unit. This method improves the flexibility and scalability of data access by reducing the storage amount in the model inference chip's memory, but there are also certain overheads, especially the problem of the first-word latency caused by limited bandwidth during the transmission process.
[0043] Considering the two current methods for restoring the key-value cache, the first method only utilizes the computing power of the model inference chip. Due to the limited computing power of the model inference chip, and recalculating the key-value cache for each token based on the historical text sequence requires repeating all previous calculations, with a relatively large amount of computation, resulting in low efficiency. The second method only utilizes the data transmission capacity of the model inference chip, and due to the limited data transmission bandwidth of the model inference chip, there is a delay in data transmission. The core defect of the above two methods is that when restoring the key-value cache, only one of the computing power or data transmission capacity of the model inference chip is inefficiently utilized.
[0044] To improve the restoration efficiency of the key-value cache during the model inference process, the applicant thought that the restoration of the key-value cache could be achieved by simultaneously leveraging the computing power and data transmission capacity of the chip. Considering that the key-value cache of each token at each layer of the model can be calculated based on the hidden state of the token at that layer, for each token, the data volume of its hidden state is half of the data volume of the key-value cache. If the hidden state is stored, compared with directly storing the key-value cache, the storage space and data transmission bandwidth occupied will be reduced by half. In addition, calculating the key-value cache based on the hidden state only requires KV projection calculation and does not go through the FFN calculation. Therefore, compared with directly recalculating the key-value cache based on the historical text sequence, the computational amount of this restoration method will be greatly reduced, at most 1 / 6 of the computational amount of recalculating the key-value cache based on the historical text sequence.
[0045] Based on this, the embodiments of the present application provide a model inference method. During the model inference process, the hidden state of each token at each layer (i.e., the input of each token at each layer) can be stored. When the key-value cache of these tokens is needed, the key-value cache of each token at each layer of the model can be restored based on this hidden state and the key-value projection weight matrix (i.e., the model parameters learned during the model training process). The key-value cache restoration solution provided by the embodiments of the present application can utilize the computing resources and data transmission bandwidth resources of the chip with low overhead simultaneously, rather than using only one type of resource to implement the restoration of the key-value cache, which can greatly improve the restoration efficiency of the key-value cache, and further improve the inference efficiency of the model.
[0046] The model in the embodiments of the present application can be various models based on the Transformer architecture. For example, it can be a large language model, which can be used to process various natural language processing tasks, such as text generation tasks, text translation tasks, multi-turn dialogue tasks, etc., and the embodiments of the present application do not make limitations.
[0047] The model inference method of the embodiments of the present application can be executed by a target chip, which can be various chips capable of performing calculations during the model inference process, such as GPU, CPU, etc., and the present application does not make limitations. Among them, this method can be executed by a single chip or by multiple chips in cooperation, and the present application does not make limitations.
[0048] For example, as Figure 1 shown, it is a schematic diagram of an application scenario of the embodiments of the present application. The model can be deployed in a target device, which can include a CPU, one or more GPUs, and one or more storage units (such as SSD). This method can be executed by the GPU, or this method can be jointly implemented by the CPU and the GPU, that is, some steps of the method are executed by the CPU and some steps are executed by the GPU, and it can be flexibly set based on actual needs.
[0049] The following Figure 2 introduces the model inference method of the embodiments of the present application. It should be noted that the following takes the inference process of a certain layer of the model (i.e., the current network layer) as an example for explanation. It is not difficult to understand that the method introduced below is applicable to any network layer of the model.
[0050] As Figure 2 shown, the model inference method can include the following steps:
[0051] S202. In the model inference stage, for the current network layer of the model, if the calculation of the current network layer requires the key-value cache of the tokens in the historical text sequence, obtain the pre-stored hidden state of the tokens in the current network layer from the storage unit outside the target chip.
[0052] Considering that the memory of the target chip is limited, in the scenario of multiple historical text sequences, it is impossible to store the key-value caches of all tokens in all historical text sequences in the memory of the target chip. Therefore, during the inference process, for the key-value caches of tokens in some sequences that cannot be stored in the memory of the target chip, the hidden states of these tokens can be stored in a storage unit located outside the target chip, hereinafter collectively referred to as the external storage unit. For example, in the application scenario shown below, this inference method can be executed by a GPU, and this external storage unit can be the memory of a CPU or other memories such as an SSD. Figure 1 Taking the application scenario shown as an example, this inference method can be executed by a GPU, and this external storage unit can be the memory of a CPU or other memories such as an SSD.
[0053] In step S202, in the model inference stage, for the current network layer of the model, if the calculation of this network layer requires the key-value cache of a token in a certain historical text sequence, the pre-stored hidden state of the token in this sequence corresponding to the current network layer can be obtained from the external storage unit. Among them, if the current network layer is the first layer of the model, the hidden state of each token corresponding to the current network layer can be the embedding vector of the token. If the current network layer is not the first layer of the model, the hidden state of each token corresponding to the current network layer is the output of the previous network layer of the current network layer, that is, the input of the current network layer.
[0054] S204. Determine the key-value cache of the token in the current network layer based on the hidden state and the key-value projection weight matrix of the current network layer, and complete the calculation of the current network layer based on the key-value cache; wherein, the key-value projection weight matrix is learned in the model training stage.
[0055] After obtaining the hidden states of each token in the current network layer, the key-value cache of each token in the current network layer can be determined based on the key-value projection weight matrix of the current network layer. Among them, the key-value projection matrix is a weight matrix learned in the model training stage. The key-value projection matrix includes a key projection weight matrix and a value projection weight matrix. The key projection weight matrix is used to generate the key vector of the token based on the hidden state of the token, and the value projection weight matrix is used to generate the value vector of the token based on the hidden state of the token.
[0056] For example, assume is the hidden state of the i-th token in the L-th layer of the model, and respectively represent the key-value cache of this token in this layer, and Respectively representing the key projection weight matrix and the value projection weight matrix of the model at this layer, the key-value cache of the token at this layer can be determined by the following formula.
[0057]
[0058] After determining the key-value cache of each token in the current network layer, the calculation of the current network layer can be completed based on this key-value cache. For example, a context vector characterizing the association relationship between the new token and each token in the historical text sequence can be calculated as the output of the current network layer.
[0059] In some embodiments, after obtaining the key-value cache of each token based on the hidden state of each token in the current network layer, the position encoding information of each token can be used to correct the key-value cache, and subsequent calculations can be completed based on the corrected key-value cache. Among them, the position encoding information of each token is used to represent the position of each token in the historical text sequence. By using the position encoding information to correct the key-value cache, the position relationship of the token in the context can be introduced, thereby enhancing the model's perception ability of the token position and improving the accuracy and efficiency of reasoning.
[0060] In some embodiments, when storing the hidden states of the tokens in the historical text sequence in each network layer of the model in an external storage unit, an appropriate storage strategy can be selected based on the task type of the model to improve the processing efficiency. For example, taking Figure 1 the application scenario shown as an example, assuming that the calculation of the key-value cache is executed by the GPU and the hidden states are stored in the CPU memory or an external memory (such as an SSD), if the task processed by the model is a multi-turn dialogue task, then after the key-value cache of the token is calculated by the GPU, the key-value cache of the token can be directly stored in the CPU memory or an external memory through direct memory access technology. If the task processed by the model is a RAG (Retrieval-Augmented Generation) task, then through an offline preprocessing method, the hidden states of each token in the relevant context can be pre-generated and stored in the CPU memory or an external memory.
[0061] In some embodiments, to improve the inference efficiency of the model, the operations of reading the hidden state from the external storage unit and calculating the key-value cache based on the hidden state can be performed in parallel. For example, in the process of determining the key-value cache of each token in the current network layer based on the hidden state and the key-value projection weight matrix of the current network layer, the hidden state of each token corresponding to the next network layer in the current network layer can be obtained from the external storage unit in parallel. That is, while restoring the hidden state of the current network layer, the hidden state corresponding to the next network layer is read from the external storage unit in advance, so that when calculating the next network layer, the required data is already prepared, which can improve the processing efficiency.
[0062] In some embodiments, for the scenario where there are multiple external storage units, to improve the efficiency of reading the hidden state from the external storage unit. When storing the hidden state of each token in the same network layer of the model, the hidden state of the same network layer can be divided into multiple data blocks, where each data block can include the hidden states of one or more tokens, and then the multiple data blocks can be assigned to each external storage unit, so that each external storage unit stores one or more of the multiple data blocks. By adopting the above storage method, when reading the hidden state of each token corresponding to the current network layer, the data blocks stored in each of the multiple external storage units can be read in parallel to obtain the hidden state of each token corresponding to the current network layer. By storing the hidden state corresponding to the same network layer in multiple external storage units in a block-by-block manner, IO bandwidth aggregation can be achieved when reading the hidden state, thereby improving the data reading efficiency.
[0063] In some embodiments, considering that the hidden state of each token in the same network layer can be stored in multiple external storage units in a block-by-block manner, to quickly read the hidden state of each token in each network layer from the external storage unit, a metadata index table can be constructed. When reading the hidden state of each token corresponding to the current network layer, the metadata index table can be queried first, so as to determine the identification information of each external storage unit used to store the hidden state of each token in the current network layer and the storage location of the hidden state in each external storage unit based on the metadata index table. Among them, the metadata index table can record the storage address information of the hidden state of the tokens in the historical text sequence in each network layer of the model in each external storage unit through a quadruple, where the quadruple can be expressed as follows:
[0064] (N layer , Addr start, N data blocks, M storage units)
[0065] Among them, N layerRepresents the number of the model network layer. Addr_start represents the starting storage address of the data block in this external storage unit. N_data_block represents the number of data blocks stored in this external storage unit. M_storage_unit represents the identification information of this external storage unit.
[0066] For example, taking Figure 1 the application scenario shown as an example, assuming the target chip is a GPU, then this metadata index table can be stored in the memory of the CPU. When performing the calculation of the current network layer of the model, the CPU can determine the storage address information of each token's hidden state in the current network layer based on this metadata index table, and obtain this hidden state based on this storage address information, and then send it to the GPU.
[0067] In some embodiments, the target chip may include a CPU and a GPU. Among them, the calculation of the hidden state and the key-value cache can be completed by the GPU. When storing the hidden states of each token in each network layer, considering that if the GPU stores the hidden state of each token it calculates into the external storage unit, the transmission bandwidth resources cannot be maximally utilized. In order to maximize the utilization of transmission bandwidth resources, the embodiments of this application propose a two-stage storage strategy, as Figure 3 shown. Among them, the storage in the first stage means that after calculating the hidden state of each token in the current network layer using the GPU, the hidden state of each token in the current network layer can be stored in the memory of the CPU through the Direct Memory Access (DMA) technology. For example, a temporary cache can be established in the memory of the CPU to temporarily store the hidden states of each token calculated by the GPU (such as Figure 3 ① in it). The storage in the second stage means that the CPU can determine whether the number of hidden states stored in the current memory meets a preset condition. If so, the hidden states stored in the current memory are stored in this external storage unit. Among them, the preset condition can be that the number of hidden states stored in the current memory reaches a certain number, or the hidden states stored in the current memory include the hidden states of all tokens in the same layer, which can be specifically set based on actual needs. For example, the CPU can determine whether the hidden states stored in the current memory already include the hidden states of all tokens in a single network layer. If so, the hidden states can be block-processed (for example, the hidden states of 64 tokens are used as a data block), and then these multiple data blocks are stored in multiple SSDs (such as Figure 3 ② in it).
[0068] In some embodiments, the multiple data blocks can be stored in the multiple SSDs through a polling algorithm, so that the multiple data blocks can be evenly distributed to the multiple SSDs.
[0069] Through the above two-stage storage strategy, in the first stage, a temporary cache for storing hidden states is stored in the CPU memory. The hidden states at the token level generated during the forward pass are first written into the temporary cache in the CPU memory. In the second stage, a block-based persistent storage strategy is adopted. The hidden states of all tokens within a single network layer are divided into data blocks of a fixed size (e.g., 64 tokens / block). Through a polling algorithm, the data blocks of hidden states in the memory are stored on multiple different SSDs, so that when reading the hidden states, they can be read in parallel from multiple SSDs, improving the reading efficiency.
[0070] Considering that the time required for the target chip to read the hidden states of each layer from the external storage unit may not be consistent with the time required for the target chip to restore the key-value cache of that layer based on the hidden states of that layer, that is, the computing speed and I / O speed of the target chip do not match, resulting in idle periods in the computing resources or I / O transmission resources. For example, as Figure 4 shown, if the time to read the hidden states is greater than the time to restore the key-value cache based on the hidden states (i.e., the computing speed is greater than the I / O speed), then after completing the restoration of the key-value cache of the current layer, it may be necessary to wait for the reading of the hidden states of the next layer to be completed before the key-value cache of the next layer can be restored, and there will be a certain idle period in the computing resources. As Figure 5 shown, if the time to read the hidden states is less than the time to restore the key-value cache based on the hidden states (i.e., the computing speed is less than the I / O speed), then after reading the hidden states of the current layer, it is necessary to wait for the calculation of the key-value cache of the previous layer to be completed before the calculation of the key-value cache of the current layer can be performed, that is, there will be a certain idle period in the I / O transmission resources. That is, the reading of the hidden states and the calculation of the key-value cache cannot be carried out side by side without idle periods, and the idle periods will waste the hardware resources of the device to a certain extent and cannot achieve the maximum utilization of resources.
[0071] Considering the existing strategy of directly recalculating the key-value cache based on the historical text sequence, although it has a large amount of calculation, it does not require data transmission. And the existing strategy of reading the pre-stored key-value cache from the external storage unit, although it takes more time to complete the data transmission, can directly obtain the key-value cache without additional calculation. In order to minimize the idle periods of the hardware resources and achieve the maximum utilization of the hardware resources, the present application comes up with the idea of combining the strategy of restoring the key-value cache based on the hidden states with the existing two strategies, and can dynamically adjust the combination strategy according to the computing power and I / O capabilities of the target chip. For example, when implementing the restoration of the key-value cache, the key-value cache restoration strategy can be configured at the granularity of each network layer of the model, that is, different key-value cache restoration strategies can be configured for different network layers.
[0072] For example, assume that the model has a total of L layers. At least two of the above key-value cache recovery policies can be configured for these L layers. For example, the policy of recovering the key-value cache based on the hidden state and the policy of recalculating the key-value cache based on the historical text sequence can be combined, or the policy of recovering the key-value cache based on the hidden state and the policy of reading the pre-stored key-value cache from an external storage unit can be combined, or the above three key-value cache recovery policies can be combined. Among them, the specific combination method and which recovery policy to configure for each network layer in the model can be flexibly set based on the characteristics of the target chip, and the embodiments of this application do not make any restrictions.
[0073] In some embodiments, for the scenario of configuring different key-value cache recovery policies for different network layers, before obtaining the pre-stored hidden state of each token in the current network layer, the key-value cache recovery policy corresponding to the current network layer can be determined first. Among them, the key-value cache recovery policy can include one or more of the following: the policy of recovering the key-value cache based on the hidden state, the policy of recalculating the key-value cache based on the historical text sequence, and the policy of reading the pre-stored key-value cache from an external storage unit. If it is determined that the key-value cache recovery policy corresponding to the current network layer is the policy of recovering the key-value cache based on the hidden state, then the pre-stored hidden state of each token in the current network layer is obtained.
[0074] Of course, if the key-value cache recovery policy corresponding to the current network layer is the policy of recalculating the key-value cache based on the historical text sequence, then the key-value cache of each token in the current network layer can be recalculated based on the historical text sequence. If the key-value cache recovery policy corresponding to the current network layer is the policy of reading the pre-stored key-value cache from an external storage unit, then the pre-stored key-value cache of each token corresponding to the current network layer can be read without calculation. Among them, if the pre-configured key-value cache recovery policy for some network layers of the model is the policy of reading the pre-stored key-value cache from an external storage unit, then during the model inference process, when calculating the key-value cache of tokens in these network layers, it can be pre-stored in the external storage unit so that it can be obtained from the external storage unit when needed.
[0075] Considering each layer of the model, the strategy of restoring key-value cache based on hidden states proposed in the embodiments of this application requires less computing time than the existing strategy of directly recalculating key-value cache based on the historical text sequence, but has a longer data transmission time. Compared with the strategy of reading the pre-stored key-value cache from an external storage unit, it requires more computing time but has a shorter data transmission time. In order to keep the overall data reading time and computing time as consistent as possible during the model inference process, minimize the idle period of hardware resources during the inference process, and maximize the utilization of hardware resources, the key-value cache restoration strategy corresponding to each network layer can be configured comprehensively based on the key-value cache calculation speed and hidden state I / O speed of the target chip.
[0076] For example, in some embodiments, for the scenario where the key-value cache calculation speed of the target chip is greater than the hidden state I / O speed, the strategy of restoring key-value cache based on hidden states and the strategy of recalculating key-value cache based on the historical text sequence can be combined. For example, the M-th layer of the model can be used as the demarcation layer. For the network layers before and including the M-th layer, the strategy of recalculating key-value cache based on the historical text sequence can be adopted, and for the network layers after the M-th layer, the strategy of restoring key-value cache based on hidden states can be used, where M is an integer. Considering that the key-value cache calculation speed of the target chip is greater than the hidden state I / O speed, in order to avoid waiting for reading the hidden state from the external storage unit before calculating the key-value cache during the inference process, the key-value cache restoration strategy for the first M layers of the model can be configured as the strategy of recalculating key-value cache based on the historical text sequence, and the remaining layers can be configured as the strategy of restoring key-value cache based on hidden states. Moreover, when calculating the key-value cache of the first M layers using the strategy of recalculating key-value cache based on the historical text sequence, the hidden states of each token in the network layers after the M-th layer of the model can be obtained in advance from the external storage unit, so that when performing the calculation of the network layers after the M-th layer, the required hidden states are already prepared, avoiding long waiting times and resource waste. As Figure 6 shown, by combining the above two strategies, it is possible to achieve a non-idle parallel execution of the calculation process and the transmission process.
[0077] Similarly, for the scenario where the key-value cache calculation speed of the target chip is less than the hidden state I / O speed, the strategy of restoring the key-value cache based on the hidden state and the strategy of reading the pre-stored key-value cache from the external storage unit can be combined. For example, the Nth layer of the model can be used as the demarcation layer. For the network layers at the Nth layer and before the Nth layer, the strategy of restoring the key-value cache based on the hidden state can be adopted, and for the network layers after the Nth layer, the strategy of reading the pre-stored key-value cache from the external storage unit can be used, where N is an integer. Considering that the key-value cache calculation speed of the target chip is less than the hidden state I / O speed, in order to avoid consuming too much time waiting for the calculation of the key-value cache, the key-value cache restoration strategy for the first N layers of the model can be configured as the strategy of restoring the key-value cache based on the hidden state, and the remaining layers can be configured as the strategy of reading the pre-stored key-value cache from the external storage unit. Moreover, after reading the hidden state of the token in the first N layers of the model from the external storage unit, the key-value cache of the token in the network layers after the Nth layer can be read from the external storage unit. Thus, when performing the calculation of the network layers after the Nth layer, the key-value cache it needs is already prepared, avoiding long waiting times and resource waste. As Figure 7 shown, by combining the above two strategies, the calculation process and the transmission process can be carried out side by side without idle periods.
[0078] After adopting the above combined strategy, in some embodiments, when determining the type of the key-value cache restoration strategy corresponding to the current network layer, for the scenario where the time for the target chip to read the hidden state from the external storage unit is greater than the time for restoring the key-value cache based on the hidden state (i.e., the key-value cache calculation speed of the target chip is greater than the hidden state I / O speed), if the current network layer is the Mth layer or a network layer before the Mth layer of the model, it is determined that the key-value cache restoration strategy corresponding to the current network layer is the strategy of recalculating the key-value cache based on the historical text sequence; if the current network layer is a network layer after the Mth layer of the model, it is determined that the key-value cache restoration strategy corresponding to the current network layer is the strategy of restoring the key-value cache based on the hidden state.
[0079] In some embodiments, the M minimizes the difference between the total reading time of the hidden states required to complete the calculations of all network layers of the model and the total restoration time of the key-value caches required to complete the calculations of all network layers of the model. That is, when determining from which layer of the model the key-value cache restoration strategy should be divided, M can be determined such that the difference between the data reading time and the calculation time during the key-value cache restoration process is minimized, that is, the time when the hardware resources are in the idle state is minimized, so as to maximize the utilization of the hardware resources.
[0080] Among them, M can be determined based on the following method:
[0081] Assume that the transmission time of the hidden state is IO H, the computing time from the token to the key-value cache is C Token , the computing time from the hidden state to the key-value cache is C H , the model has a total of L layers, and an optimization function can be constructed as follows:
[0082] argmin M max(IO H M, C Token M + C H (L - M))
[0083] Among them, by minimizing the difference between the total reading duration of the hidden states required to complete the calculations of all network layers of the model and the total restoration duration of the key-value caches required to complete the calculations of all network layers of the model, that is, by minimizing the maximum of the total data reading duration or the computing duration, an optimal M can be determined by constructing a min-max optimization function and solving the above optimization function.
[0084] In some embodiments, when determining the key-value cache restoration policy corresponding to the current network layer, for a target chip, if the duration of reading the hidden state from the external storage unit is less than the duration of restoring the key-value cache based on the hidden state (i.e., the key-value cache computing speed of the target chip is less than the hidden state I / O speed), if the current network layer is the Nth layer or a network layer before the Nth layer of the model, then the key-value cache restoration policy is the policy of restoring the key-value cache based on the hidden state; if the current network layer is a network layer after the Nth layer of the model, then the key-value cache restoration policy is the policy of reading the pre-stored key-value cache from the external storage unit.
[0085] In some embodiments, in order to minimize the idle duration of the hardware resources as much as possible and reduce the resource waste during the inference process, this N minimizes the difference between the total reading duration of the hidden states and the key-value caches required to complete the calculations of all network layers of the model and the total restoration duration of the key-value caches required to complete the calculations of all network layers of the model.
[0086] Among them, N can be determined based on the following method:
[0087] Assume that the transmission time of the hidden state is IO H , the computing time from the hidden state to the key-value cache is C H , the transmission time of the key-value cache is IO KV , the model has a total of L layers, and an optimization function can be constructed as follows:
[0088] argmin N max(C H N, IO H N + IO KV (L - N))
[0089] Among them, the difference between the total reading duration of the hidden states and the key-value cache required to complete the calculations of all network layers of the model and the total restoration duration of the key-value cache required to complete the calculations of all network layers of the model is minimized, that is, the maximum of the total data reading duration or the calculation duration is minimized. By constructing a min-max optimization function and solving the above optimization function, the optimal N can be determined.
[0090] It is not difficult to understand that the solutions described in the above embodiments can be freely combined to obtain new solutions in the case of no conflict. Due to space limitations, they are not listed one by one in the embodiments of this application.
[0091] Correspondingly, the embodiments of this application also provide a computer program product, which includes a computer program that implements the method mentioned in any of the above embodiments when executed.
[0092] Furthermore, the embodiments of this application also provide a chip, as Figure 8 shown. The device includes a processor 81, a memory 82, and computer instructions stored in the memory 82 and executable by the processor 81. When the processor 81 executes the computer instructions, the method described in any one of the above embodiments is implemented. Among them, the chip can be a CPU, a GPU, etc., and the embodiments of this application do not make limitations.
[0093] The embodiments of this application also provide a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the method described in any of the foregoing embodiments is implemented.
[0094] Computer-readable media includes both permanent and non-permanent, removable and non-removable media and can be implemented by any method or technology for information storage. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile discs (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0095] As can be seen from the description of the above embodiments, those skilled in the art can clearly understand that the embodiments of the present application can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments of the present application.
[0096] The systems, devices, modules or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer, and the specific form of the computer can be a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email transceiver device, a game console, a tablet computer, a wearable device, or a combination of any several of these devices.
[0097] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiments. The device embodiments described above are only illustrative. The modules described as separate components may or may not be physically separated. When implementing the solutions of the embodiments of the present application, the functions of the various modules can be implemented in one or more software and / or hardware. It is also possible to select some or all of the modules according to actual needs to achieve the purpose of the solutions of this embodiment. Those of ordinary skill in the art can understand and implement without creative efforts.
[0098] The above are only the specific implementation manners of the embodiments of the present application. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the embodiments of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the embodiments of the present application.
Claims
1. A model reasoning method, characterized in that: The model is a model based on the Transformer architecture, and the method is executed by a target chip, and the method includes: In the model inference stage, for the current network layer of the model, if the calculation of the current network layer requires the use of the key value cache of the word in the historical text sequence, the hidden state of the word in the current network layer stored in advance is obtained from the storage unit outside the target chip; Based on the hidden state and the key-value projection weight matrix of the current network layer, the key-value cache of the word element in the current network layer is determined, and the calculation of the current network layer is completed based on the key-value cache; wherein the key-value projection weight matrix is learned during the model training phase.
2. The method according to claim 1, characterized in that The completing the calculation of the current network layer based on the key-value cache includes: The key-value cache is modified using the position encoding information of the word element, wherein the position encoding information is used to represent the position of the word element in the historical text sequence; The calculation of the current network layer is completed based on the modified key-value cache.
3. The method according to claim 1, characterized in that: In the process of determining the key value cache of the word element in the current network layer based on the hidden state and the key value projection weight matrix of the current network layer, the hidden state corresponding to the word element in the next network layer of the current network layer is obtained from the external storage unit in parallel.
4. The method according to claim 1, characterized in that: The hidden state is stored in a plurality of external storage units located outside the target chip, the hidden state is divided into a plurality of data blocks, each data block includes the hidden state of a plurality of word units, and each external storage unit stores one or more of the plurality of data blocks; The obtaining of the pre-stored hidden state of the word element corresponding to the current network layer includes: The data blocks stored in each of the plurality of external storage units are read in parallel to obtain the hidden state of the word unit in the current network layer.
5. The method according to claim 4, characterized in that The identification information of the multiple external storage units and the storage addresses of the hidden states in the multiple external storage units are determined by querying a pre-built metadata index table, wherein the metadata index table records the storage address information of the hidden states of the words in the historical text sequence at each network layer of the model in each external storage unit through a quadruple, wherein the quadruple is as follows: (N layer , Addr start, N data blocks, M storage units) Among them, N layer Indicates the number of the model network layer, Addr start indicates the starting storage address of the data block in the external storage unit, N data block indicates the number of data blocks stored in the external storage unit, and M storage unit indicates the identification information of the external storage unit.
6. The method according to claim 1, characterized in that The target chip includes a CPU and a GPU, and the hidden state is stored in the external storage unit in the following manner: After the GPU calculates the hidden state of each word in the current network layer, the hidden state of the word in the current network layer is stored in the memory of the CPU through direct memory access technology; The CPU determines whether the number of hidden states currently stored in the memory meets a preset condition, and if so, stores the hidden states currently stored in the memory in the external storage unit.
7. The method according to claim 1, characterized in that Before obtaining the pre-stored hidden state corresponding to the word element in the current network layer, the method further includes: Determine a key-value cache recovery strategy corresponding to the current network layer, wherein the key-value cache recovery strategy includes one or more of the following: a strategy for recovering the key-value cache based on a hidden state, a strategy for recalculating the key-value cache based on a historical text sequence, and a strategy for reading a pre-stored key-value cache from the external storage unit; If the key-value cache recovery strategy is a strategy for recovering the key-value cache based on a hidden state, then the pre-stored hidden state corresponding to the word element in the current network layer is obtained.
8. The method according to claim 7, characterized in that Determining the key-value cache recovery strategy corresponding to the current network layer includes: In the case where the duration of the target chip reading the hidden state from the external storage unit is longer than the duration of the target chip restoring the key-value cache based on the hidden state, if the current network layer is the Mth layer of the model or a network layer before the Mth layer, the key-value cache recovery strategy is a strategy of recalculating the key-value cache based on a historical text sequence; if the current network layer is a network layer after the Mth layer of the model, the key-value cache recovery strategy is a strategy of restoring the key-value cache based on the hidden state, where M is an integer; In which, in the process of restoring the hidden state corresponding to the word element in the network layer of the Mth layer or before the Mth layer of the model by recalculating the key value cache strategy based on the historical text sequence, the hidden state corresponding to the network layer of the word element after the Mth layer of the model is read in advance from the external storage unit.
9. The method according to claim 8, characterized in that The M minimizes the difference between the total time required to read the hidden state required to complete the calculation of all network layers of the model and the total time required to restore the key-value cache required to complete the calculation of all network layers of the model.
10. The method according to claim 7, characterized in that Determining the key-value cache recovery strategy corresponding to the current network layer includes: In the case that the duration of the target chip reading the hidden state from the external storage unit is less than the duration of the target chip restoring the key-value cache based on the hidden state, if the current network layer is the Nth layer of the model or a network layer before the Nth layer, the key-value cache recovery strategy is a strategy of restoring the key-value cache based on the hidden state; if the current network layer is a network layer after the Nth layer of the model, the key-value cache recovery strategy is a strategy of reading a pre-stored key-value cache from an external storage unit, where N is an integer; In which, in response to reading the hidden state of the word element corresponding to the Nth layer and the network layer before the Nth layer of the model from the external storage unit, the key-value cache corresponding to the network layer after the Nth layer of the model of the word element is read from the external storage unit.
11. The method according to claim 10, characterized in that The method further comprises: The N minimizes the difference between the total time required to read the hidden state and key-value cache required to complete the calculation of all network layers of the model and the total time required to restore the key-value cache required to complete the calculation of all network layers of the model.
12. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed, the method according to any one of claims 1 to 11 is implemented.
13. A chip, characterized in that: The chip includes a processor, a memory, and computer instructions stored in the memory and executable by the processor. When the processor executes the computer instructions, the method according to any one of claims 1 to 11 is implemented.
14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 11 is implemented.
Citation Information
Cited By
Fault tolerance method for key value cache in model, product, equipment and medium
CN120508433A
KV cache compression and eviction lexical element recovery method and system for large-scale language model reasoning
CN120975245A
Large model data processing method and device, equipment and medium
CN121029652A
Memory data processing method, electronic equipment and computer readable storage medium
CN121166376A
Memory data processing method, electronic device, and computer-readable storage medium
CN121166376B