KV cache compression and eviction lexical element recovery method and system for large-scale language model reasoning

By using a preset threshold to control the eviction and restoration of the KV cache during the reasoning process of a large language model, and combining this with a pre-computed matrix to reduce the restoration latency, the problem of difficulty in coordinating the optimization of memory usage and computational efficiency in existing technologies is solved, and efficient reasoning under limited memory conditions is achieved.

CN120975245AActive Publication Date: 2025-11-18HANGZHOU INTERNATIONAL INNOVATION INSTITUTE OF BEIHANG UNIVERSITY

Patent Information

Application Number
CN202511475772.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-16
Publication Date
2025-11-18
Estimated Expiration
2045-10-16

AI Technical Summary

Technical Problem

Existing technologies struggle to achieve coordinated optimization of memory usage and computational efficiency in key-value cache management for large language models while maintaining model inference quality. In particular, recovery techniques suffer from computational latency and accuracy loss.

Method used

By using a preset threshold to control the eviction and restoration of the KV cache during the inference process of a large language model, and combining a pre-computed matrix to reduce the restoration latency, the KV cache memory usage is dynamically managed. The value vector is regenerated and restored using a pre-computed fusion weight matrix.

Benefits of technology

It achieves the goal of maintaining model inference quality and improving computational efficiency under limited video memory conditions, dynamically manages KV cache video memory usage, reduces video memory usage and computational latency, and improves inference efficiency in long context scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120975245A_ABST
    Figure CN120975245A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence and natural language processing, in particular to a KV cache compression and eviction lexical element recovery method and system for large language model reasoning, and the method comprises the steps: for a query vector of an ith lexical element calculated by a current Transform layer, calculating attention scores of the query vector and all key vectors; executing a V cache dynamic updating operation based on the attention score; and for the ith lexical element, after the V cache dynamic updating operation is completely completed, performing attention calculation by using the updated value vector set in the V cache storage pool and the pre-calculated attention score partial product P. According to the technical scheme, the technical problem that in the prior art, collaborative optimization of video memory occupancy and calculation efficiency is difficult is solved, and the method has the advantages that KV cache video memory occupancy is dynamically managed, the model reasoning quality is kept, and the calculation efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the fields of artificial intelligence and natural language processing technology, specifically to a method and system for KV cache compression and evicted lexical recovery for large-scale language model inference. Background Technology

[0002] With the widespread deployment of large-scale language models in various application scenarios, the memory bottleneck in their inference process is becoming increasingly prominent. As a key data structure for storing historical key-value pairs in the autoregressive generation process, the memory usage of the key-value cache increases linearly with the increase of context length, which seriously restricts the deployment capability of the model in resource-constrained environments.

[0003] Current mainstream key-value (KV) cache compression technologies primarily employ three approaches: static truncation methods based on sliding windows, such as StreamingLLM, which simply discard historical tokens exceeding the window, leading to the permanent loss of long-range dependency information; dynamic eviction strategies based on attention scores, such as H2O, which, while filtering based on token importance, lack an effective recovery mechanism; and paged-attention schemes based on paging management, which borrow from the operating system's paging mechanism to manage the KV cache in pages, evictioning entire pages of KV pairs according to a strategy, rather than individual tokens. Its eviction strategy can be compatible with two types: direct eviction (page deletion); and swapping the evicted page pairs into CPU memory to free up space for another task, but this has certain limitations.

[0004] In terms of recovery techniques, existing solutions generally face two core drawbacks: traditional recalculation methods require reloading the entire projection matrix. and This results in significant computational latency; while CPU-based memory offload solutions are limited by finite memory bandwidth. Of particular note is that while the matrix transformation theory proposed by Slim-Attention has compression potential, its matrix inversion operation introduces unacceptable precision loss in practical quantization deployment scenarios. These technical limitations collectively lead to a key contradiction in current KV cache management: achieving coordinated optimization of memory usage and computational efficiency while maintaining model inference quality. Summary of the Invention

[0005] To address the problems in the related technologies, this disclosure provides a method and system for KV cache compression and evicted lexical recovery for large-scale language model inference.

[0006] In a first aspect, embodiments of this disclosure provide a method for KV cache compression and evicted lexical recovery for large-scale language model inference, including: During LLM inference, the following operations are performed for at least one Transformer layer: For the query vector of the i-th word computed in the current Transformer layer Calculate its key vectors and all key vectors calculated in the current Transformer layer. Attention score : Where j = 1, 2, ..., n; n is the total number of tokens calculated in the current Transformer layer; Perform a dynamic update operation on the V cache based on the attention score, including: when At that time, perform an eviction operation, removing the value vector of the j-th term from the V cache storage pool. ; when And the value vector of the j-th word element If expelled, perform a recovery operation, including: via calculation. Regenerate the value vector of the j-th word. And add it to the V cache storage pool; among which, For the pre-computed fusion weight matrix, and These are the key projection matrix and value projection matrix of the current Transformer layer, respectively; T is the preset threshold.

[0007] In one embodiment of this disclosure, the preset threshold T is set as follows: ,in, For configurable hyperparameters, satisfying 0 < <1.

[0008] In one embodiment of this disclosure, the method further includes: Configure a fixed window to be retained; The size of the fixed retention window is a preset constant. Tokens located within the fixed retention window do not participate in the eviction and restoration operations, but their attention scores are used to calculate the preset threshold.

[0009] In one embodiment of this disclosure, the pre-computed fusion weight matrix The model is calculated once during the initialization phase and remains resident in GPU memory; during the recovery operation, it is only loaded. Instead of loading separately and .

[0010] In one embodiment of this disclosure, the maximum capacity of the V cache storage pool is configured with a preset parameter. m Used to store the most m A vector of values; Accordingly, the recovery operation is performed before the eviction operation, and when the number of value vectors in the V cache storage pool exceeds m after the recovery operation is performed, a forced eviction operation is triggered to evict value vectors with low attention scores, so that the number of value vectors retained in the end does not exceed m.

[0011] In one embodiment of this disclosure, the method performs an evicting and recovery operation during the pre-filling phase to compress the long context key-value cache and improve response efficiency to new problems.

[0012] In one embodiment of this disclosure, the method performs an evicting and recovery operation after each wp inference iteration during the decoding phase, where wp is a preset window length.

[0013] In one embodiment of this disclosure, the method further includes: For the i-th word element, based on all After completing the V-cache dynamic update operation, the attention value is calculated using the updated set of value vectors in the V-cache storage pool and the pre-computed attention score partial product P. : V new ,in, V new For the updated set of value vectors, .

[0014] Secondly, this disclosure provides a large-scale language model inference system, including: The attention score calculation module is configured to calculate the query vector of the i-th word in the current Transformer layer. Calculate its key vectors and all key vectors calculated in the current Transformer layer. Attention score : Where j = 1, 2, ..., n; n is the total number of tokens calculated in the current Transformer layer; The V-cache dynamic management module is configured to perform V-cache dynamic update operations based on the attention score, including: when At that time, perform an eviction operation, removing the value vector of the j-th term from the V cache storage pool. ; when And the value vector of the j-th word element If expelled, perform a recovery operation, including: via calculation. Regenerate the value vector of the j-th word. And add it to the V cache storage pool; among which, For the pre-computed fusion weight matrix, and These are the key projection matrix and value projection matrix of the current Transformer layer, respectively; T is the preset threshold.

[0015] Thirdly, embodiments of this disclosure provide an electronic device including a memory and a processor, wherein the memory is used to store one or more computer instructions, wherein the one or more computer instructions are executed by the processor to implement the method as described in any of the first aspects.

[0016] Fourthly, this disclosure provides a computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the method as described in any of the first aspects.

[0017] The technical effects provided by the embodiments of this disclosure may include the following beneficial effects: According to the technical solution provided in the embodiments of this disclosure, the KV cache compression and eviction token recovery method for large-scale language model inference controls the eviction and recovery of the KV cache by setting a preset threshold and reducing the recovery latency by pre-computing a matrix. This solves the technical problem that it is difficult to coordinate the optimization of memory usage and computational efficiency in the prior art. It has the advantages of dynamically managing KV cache memory usage, maintaining model inference quality, and improving computational efficiency.

[0018] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0019] Other features, objects, and advantages of this disclosure will become more apparent from the following detailed description of non-limiting embodiments, taken in conjunction with the accompanying drawings. The following is a description of the accompanying drawings.

[0020] Figure 1 A flowchart illustrating a method for KV cache compression and evicted lexical recovery for large language model inference according to an embodiment of the present disclosure is shown.

[0021] Figure 2 The present disclosure illustrates the dynamic management process of the V cache during the Transformer decoding stage according to an embodiment of the present disclosure.

[0022] Figure 3 The diagram illustrates the attention scores and value vector evicting retention when decoding to different tokens.

[0023] Figure 4 A structural block diagram of a large-scale language model inference system according to an embodiment of the present disclosure is shown.

[0024] Figure 5 A structural block diagram of an electronic device according to an embodiment of the present disclosure is shown.

[0025] Figure 6 A schematic diagram of the structure of a computer system suitable for implementing the method according to embodiments of the present disclosure is shown. Detailed Implementation

[0026] In the following, exemplary embodiments of the present disclosure will be described in detail with reference to the accompanying drawings to enable those skilled in the art to readily implement them. Furthermore, for clarity, portions unrelated to the description of exemplary embodiments have been omitted from the drawings.

[0027] In this disclosure, it should be understood that terms such as “comprising” or “having” are intended to indicate the presence of features, figures, steps, behaviors, components, parts or combinations thereof disclosed in this specification, and are not intended to exclude the possibility of the presence or addition of one or more other features, figures, steps, behaviors, components, parts or combinations thereof.

[0028] It should also be noted that, unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other. This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.

[0029] Currently, mainstream key-value (KV) cache compression technologies primarily employ three approaches: static truncation methods based on sliding windows, such as StreamingLLM, which simply discard historical tokens exceeding the window size, leading to the permanent loss of long-range dependency information; dynamic eviction strategies based on attention scores, such as H2O, which, while filtering based on token importance, lack an effective recovery mechanism; and Paged-Attention schemes based on paging management, which borrow from the operating system's paging mechanism to manage the KV cache by evicting entire pages of KV pairs, rather than individual tokens. These schemes can employ two eviction strategies: direct eviction (page deletion) and swapping evicted page pairs into CPU memory to free up space for other tasks, but with certain limitations. Regarding recovery techniques, existing solutions generally face two core drawbacks: traditional recomputation methods require reloading the entire projection matrix. and This results in significant computational latency; while CPU-memory-based offload schemes are limited by finite memory bandwidth. Of particular note is that while the matrix transformation theory proposed by Slim-Attention has compression potential, its matrix inversion operation introduces unacceptable precision loss in practical quantization deployment scenarios. These technical limitations collectively lead to a key contradiction in current KV cache management: achieving coordinated optimization of memory usage and computational efficiency while maintaining model inference quality.

[0030] In view of the above-mentioned shortcomings, the KV cache compression and eviction token recovery method for large-scale language model inference provided in this disclosure controls the eviction and recovery of KV cache by setting a preset threshold and reducing the recovery latency by pre-computing a matrix. It solves the technical problem that it is difficult to coordinate the optimization of memory usage and computational efficiency in the prior art, and has the advantages of dynamically managing KV cache memory usage, maintaining model inference quality, and improving computational efficiency.

[0031] Figure 1 A flowchart illustrating a method for KV cache compression and evicted lexical recovery for large language model inference according to an embodiment of the present disclosure is shown.

[0032] like Figure 1 As shown, the KV cache compression and evicted lexical recovery method for large-scale language model inference specifically includes the following steps S110-S130: In step S110, the query vector of the i-th token is calculated for the current layer (e.g., the current Transformer layer). Calculate its relationship with all key vectors calculated in the current layer. Attention score; In step S120, a V-cache dynamic update operation is performed based on the attention score; In step S130, for the i-th word, based on all After completing the V-cache dynamic update operation, attention is calculated using the updated set of value vectors in the V-cache storage pool and the pre-computed partial product P of the attention score.

[0033] In step S110, the attention score is calculated using the following formula. : ,in, Let j be the dimension of the key vector, j = 1, 2, ..., n; n is the total number of tokens calculated in the current layer. In step S120, when At that time, perform an eviction operation, removing the value vector of the j-th term from the V cache storage pool. ; when And the value vector of the j-th word element If expelled, perform a recovery operation, including: via calculation. Regenerate the value vector of the j-th word. And add it to the V cache storage pool; among which, For the pre-computed fusion weight matrix, and These are the key projection matrix and value projection matrix of the current layer, respectively; T is a preset threshold. In step S130, the attention calculation formula is as follows: V new ,in, V new For the updated set of value vectors, Let i be the set of attention scores for the i-th word. .

[0034] The inventors discovered that existing technologies cannot effectively balance memory compression and model performance. By analyzing the distribution of attention scores, they found that some values ​​in the value vectors of evicted tokens still have potential importance. Further research revealed that directly recalculating the value vectors requires simultaneously loading the original weight matrix, leading to increased memory bandwidth pressure. Therefore, they proposed eliminating the weight loading bottleneck through matrix pre-computation and combining it with a preset threshold mechanism to achieve selective recovery, thereby maintaining model inference quality under limited memory conditions.

[0035] Therefore, this application proposes a scheme in which the above-described method steps are performed for at least one Transformer layer during LLM inference, and preferably for each Transformer layer. In this scheme, firstly, an attention score is calculated for each Token; then, the value vectors of low-scoring tokens are evicted according to a preset threshold; the value vectors of historical tokens that were evicted but scored above the preset threshold are regenerated using a pre-computed fusion weight matrix and restored; finally, attention is calculated based on the updated set of value vectors.

[0036] In this disclosure, the attention score calculation refers to a numerical evaluation metric generated through a multi-head attention mechanism, which can be implemented using scaled dot product attention to quantify the importance of each token to the current inference task.

[0037] The preset threshold is a dynamically adjusted critical value. Specifically, it can be set proportionally based on the maximum value of the current layer's attention score. For example, setting the preset threshold to 30% of the maximum value can be used to filter the set of tokens that need to be managed.

[0038] Pre-computed fusion weight matrix This refers to the weight matrix The inverse matrix and The composite matrix obtained by multiplication can be calculated and stored offline during the model initialization phase, and then directly loaded during the recovery process to avoid repeatedly loading the original weight matrix and reduce computational latency.

[0039] Specifically, during model inference, each Transformer layer performs standard attention calculations on the current sequence and extracts the attention score set of all tokens. The system dynamically generates a preset threshold based on the maximum score calculated in real time; for example, if the maximum score is 0.85, the preset threshold is set to 0.255. Tokens with scores higher than the preset threshold are added to the recovery queue. During the recovery operation, the pre-stored fusion weight matrix is ​​used. Perform matrix multiplication with the retained key vector to regenerate the value vector, write it back to the display memory, and complete the reconstruction of key information.

[0040] Compared to existing technologies, dynamic eviction schemes such as H2O only perform one-way eviction operations, while this scheme achieves reversible token state management through a preset threshold mechanism. Compared to the weight-based distribution scheme of KVPR, this method eliminates the weight loading stage through pre-computation, thus reducing the time consumption of the recovery operation.

[0041] Through the above technical solutions, this application reduces KV cache usage while maintaining model inference accuracy. Fine-grained token management is achieved through a dynamic threshold mechanism, and the application of pre-calculated fusion weight matrices reduces recovery operation time, significantly improving inference efficiency in long-context scenarios.

[0042] In one embodiment of this disclosure, the preset threshold T is set as follows: ,in, For configurable hyperparameters, satisfying 0 < <1.

[0043] In this disclosure method, Defined as an expulsion ratio coefficient, this coefficient is used to dynamically adjust the expulsion threshold based on the current attention score. It can be implemented using a preset fixed value or an adaptive adjustment algorithm. This coefficient forms a dynamic threshold benchmark by proportionally scaling down the highest attention score. Hyperparameter The numerical range is limited to the interval (0, 1) to ensure the validity of the threshold. This refers to the maximum attention score of the current layer, which is used to establish a baseline reference value for threshold calculation. Specifically, it can be achieved by iterating through the attention scores of all tokens in the current layer and taking the maximum value.

[0044] Traditional methods like H2O employ fixed thresholds or static scaling factors based on cumulative attention, which cannot adapt to the differences in attention distribution across different layers. For example, StreamingLLM uses a fixed window length for eviction, and its threshold is implicitly set in the window boundary, lacking dynamic awareness of the semantic importance of the current context. In contrast, this solution calculates the current layer's attention peak in real time and dynamically generates a scaling threshold, enabling the compression strategy to accurately match the semantic features of the current layer, reducing the false eviction of important tokens while maintaining computational efficiency.

[0045] Through the above technical solution, this application achieves the technical effect of dynamically adjusting the compression intensity based on the current context, solving the problem of model performance fluctuation caused by static thresholds. By correlating the threshold with the peak value of the attention score calculated in real time, the system can automatically adapt to the length changes and semantic density differences of different input sequences, maintaining model inference accuracy while compressing GPU memory usage.

[0046] In one embodiment of this disclosure, the method further includes: Configure a fixed window to be retained; The size of the fixed retention window is a preset constant. Tokens located within the fixed retention window do not participate in the eviction and restoration operations, but their attention scores are used to calculate the preset threshold.

[0047] In this disclosed method, the fixed reserved window refers to a protected area consisting of several recently generated consecutive tokens. Specifically, it can be implemented using a sliding window mechanism, where the starting position of the window is dynamically updated as the sequence is generated, while the window length remains fixed. This window is used to protect tokens that have a significant impact on subsequent reasoning, preventing the loss of critical information due to misjudgment.

[0048] The preset constant refers to a fixed value that the window size is pre-set based on the hardware configuration before model inference. This value can be determined empirically or through offline experiments, for example, set to 2, 4, or other values. This setting avoids the additional computational overhead of dynamically adjusting the window while ensuring controllable video memory usage.

[0049] Specifically, during model inference, tokens within a fixed retention window remain protected; their corresponding value vectors are neither evicted nor involved in the recovery operation. Tokens outside the window are managed according to a dynamically calculated preset threshold, the calculation of which includes the attention score data of tokens within the window. When the sequence length exceeds the window size, only tokens exceeding the window range are subject to evicting and recovery operations. The window size remains constant and does not adjust with changes in sequence length, effectively reducing the number of dynamic memory allocations.

[0050] Through the above technical solutions, this application effectively protects important tokens that have a continuous impact on model inference while achieving dynamic compression of KV cache, avoiding semantic coherence disruption caused by excessive eviction. The fixed window mechanism reduces the complexity of dynamic memory management, and the preset constant window size ensures predictable GPU memory usage, providing stable resource consumption guarantees for edge deployment. Attention data of tokens within the window participates in threshold calculation, making eviction and recovery decisions more aligned with the current context state, thus improving the model's accuracy in long sequence generation tasks.

[0051] In one embodiment of this disclosure, the pre-computed fusion weight matrix The model is calculated once during the initialization phase and remains resident in GPU memory; during the recovery operation, it is only loaded. Instead of loading separately and .

[0052] In this disclosure, the model initialization phase refers to the preparation phase after the model is loaded onto the computing device and before processing the input sequence begins. Specifically, it can be generated by calling a one-time matrix operation function. This matrix ensures that subsequent recovery operations do not require recalculation. Resident memory refers to... The matrix is ​​continuously stored in the graphics processor's video memory space. Specifically, the video memory allocation interface can be used to lock the matrix data in a specific area of ​​video memory, thereby ensuring that the matrix data can be quickly accessed during the recovery operation.

[0053] In one embodiment of this disclosure, the maximum capacity of the V cache storage pool is configured with a preset parameter. m Used to store the most m A vector of values; Accordingly, the recovery operation is performed before the eviction operation, and when the number of value vectors in the V cache storage pool exceeds m after the recovery operation is performed, a forced eviction operation is triggered to evict value vectors with low attention scores, so that the number of value vectors retained in the end does not exceed m.

[0054] In this disclosure method, preset parameters are used. m This refers to a pre-defined limit on video memory capacity, which can be determined through system configuration parameters or hardware resource constraints. It controls the total amount of video memory used by the key-value cache. This feature prevents inference interruptions caused by resource exhaustion by forcibly constraining video memory usage boundaries.

[0055] The "recover first, then evict" operation order means that the value vector of the token to be recovered is processed first, and then evictment is performed based on the updated attention score. This can be implemented through a state tag queue. This feature ensures that eviction decisions are based on the latest attention distribution by prioritizing the recovery of tokens that are likely to be of higher importance, thus avoiding the erroneous eviction of important tokens.

[0056] Specifically, when the video memory capacity reaches a fixed upper limit k, the value vectors of the tokens to be recovered are first filtered according to a preset threshold, and then a pre-computed matrix is ​​used. Regenerate the value vectors. After recovery, retain the m value vectors with the highest scores and evict the rest. For example, when m is set to 4000, the recovery phase may add 300 value vectors. In this case, the evictment phase eliminates the 300 value vectors with the lowest scores, ultimately retaining 4000 value vectors. This order ensures that the attention scores of the recovered tokens participate in the evictment decision, avoiding memory overflow caused by the recovery operation.

[0057] Through the above technical solutions, this application achieves better KV cache utilization under fixed GPU memory capacity. By optimizing the execution order of recovery and eviction operations, it prioritizes the retention of high-value tokens, avoiding model performance degradation due to GPU memory limitations. Simultaneously, this solution can dynamically adapt to changes in attention distribution at different stages, maintaining a balance between inference quality and efficiency in resource-constrained scenarios.

[0058] In one embodiment of this disclosure, the method performs an evicting and recovery operation during the pre-filling phase to compress the long context key-value cache and improve response efficiency to new problems.

[0059] In this disclosure, the pre-filling stage refers to the stage where the model receives user input and generates the initial sequence embedding, at which time the KV cache of all input tokens is fully computed and stored.

[0060] Specifically, after calculating the attention score of the initial sequence during the pre-filling phase, a preset threshold is dynamically determined based on the attention scores of all tokens in the current layer. Among tokens outside the fixed retention window, V-cache tokens with attention scores below the eviction threshold are marked as releaseable regions, while evictiond value vectors with attention scores above the recovery threshold are processed through a pre-calculated matrix. A new key-value (KV) cache is generated. This operation is performed only once, pre-compressing the memory usage of non-critical tokens in long contexts while preserving the core tokens within the window for full-precision storage. Therefore, subsequent decoding stages can directly perform inference based on the compressed KV cache, reducing memory resource contention.

[0061] In one embodiment of this disclosure, the method performs an evicting and recovery operation after each wp inference iteration during the decoding phase, where wp is a preset window length.

[0062] In this disclosed method, the window length refers to the number of inference operations required between two eviction and recovery operations. It can be implemented using a fixed value or a dynamic adjustment strategy. This parameter controls the operation execution frequency to balance memory usage and computational overhead. The eviction and recovery operations refer to the process of dynamically managing the key-value cache in video memory based on attention scores. Periodic execution of these operations avoids performance degradation caused by frequent triggering.

[0063] Specifically, during the decoding phase, when generating new tokens, a memory management process is triggered after each preset number of inference iterations. When the window length is reached, the system statistically calculates the attention scores of all current tokens and filters the key-value caches that need to be retained or restored based on dynamic thresholds. This periodic operation mechanism allows the model to effectively control memory resource consumption while maintaining inference coherence.

[0064] Through the above technical solution, this application effectively reduces the memory management overhead in long sequence generation scenarios and balances computational efficiency and resource utilization by reasonably setting the operation trigger frequency. This mechanism is particularly suitable for application scenarios that require the continuous generation of a large number of tokens, such as question-answering scenarios, reducing inference latency caused by frequent memory operations while maintaining generation quality.

[0065] Figure 2 The present disclosure illustrates the dynamic management process of the V cache during the Transformer decoding stage according to an embodiment of the present disclosure.

[0066] like Figure 2 As shown, during the decoding phase of the current Transformer layer, after a window period (wp), the V buffer is dynamically managed during the multi-head attention (MHA) computation process of the self-attention mechanism. This process includes the following steps: Input sequence: The user inputs a question, and a sequence embedding is generated based on the question pre-filling process as the current token input sequence.

[0067] Calculate QKV and Calculate the Query, Key, and Value vectors and use a pre-computed fusion weight matrix. This is to prepare for the possible subsequent value vector recovery.

[0068] Update KV cache: Update the Key and Value vectors calculated in the current step to the KV cache. Here, the K cache is fully preserved, while the V cache is dynamically managed.

[0069] Calculate attention score: Calculate the attention score between the current query vector and all key vectors.

[0070] Evicting from the V cache: Based on the attention score and a set threshold, evict Value vectors with scores below the threshold T (remove them from the V cache storage pool).

[0071] Restore the V cache: For Value vectors whose attention scores have reached the threshold T but have been evicted, use the formula... Recalculate and restore to the V cache storage pool.

[0072] After the recovery operation, if the number of Value vectors in the V cache storage pool exceeds the preset capacity m, the Value vector with the lowest score is evicted to ensure that the capacity does not exceed m.

[0073] Calculate MHA: Use the updated V cache (V new And a complete K-cachrystal for multi-head attention computation.

[0074] Feedforward Network (FFN): The result of attention calculation is input into the feedforward network for further processing to obtain the output token of the current layer, which is then added to the Sequence Embedding.

[0075] Figure 3 This diagram illustrates the attention scores and value vector retention / evicting when decoding to different tokens. The vertical axis represents Query Embedding, and the horizontal axis represents Key Embedding. Green squares represent windows of size 2; yellow represents value vectors selected for retention; blue represents value vectors selected for recovery; white represents value vectors selected for eviction; and green represents retained value vectors, where only the corresponding V buffer is evicted. Black represents the mask area.

[0076] Threshold hyperparameter When the value is 0.6, taking the fourth token as an example, the current value is... The value is 0.4. At this time, the preset threshold T = 0.6 × 0.4 = 0.24. According to the preset threshold, the first value vector (0.03) is extracted, and the second value vector (0.32) is retained.

[0077] Taking the fifth token as an example again, currently... The value is 0.3. At this time, the preset threshold T = 0.6 × 0.3 = 0.18. According to the threshold, the second value vector (0.06) is eliminated, the third value vector (0.24) is retained, and the first value vector (0.18) is restored.

[0078] Figure 4 A structural block diagram of a large-scale language model inference system according to an embodiment of the present disclosure is shown. This device can be implemented as part or all of an electronic device through software, hardware, or a combination of both.

[0079] like Figure 4 As shown, the large-scale language model inference system 400 includes: The attention score calculation module 410 is configured to calculate the query vector of the i-th word in the current Transformer layer. Calculate its key vectors and all key vectors calculated in the current Transformer layer. Attention score : ,in, Let j be the dimension of the key vector, j = 1, 2, ..., n; n is the total number of tokens computed in the current Transformer layer. The V-cache dynamic management module 420 is configured to perform a V-cache dynamic update operation based on the attention score, including: when At that time, perform an eviction operation, removing the value vector of the j-th term from the V cache storage pool. ; when And the value vector of the j-th word element If expelled, perform a recovery operation, including: via calculation. Regenerate the value vector of the j-th word. And add it to the V cache storage pool; among which, For the pre-computed fusion weight matrix, and These are the key projection matrix and value projection matrix of the current Transformer layer, respectively; T is a preset threshold. Attention calculation module 430 is configured to perform attention calculation on the i-th word based on all After completing the V-cache dynamic update operation, the attention value is calculated using the updated set of value vectors in the V-cache storage pool and the pre-computed attention score partial product P. : V new ,in, V new For the updated set of value vectors, .

[0080] According to the technical solution provided in the embodiments of this disclosure, by controlling the eviction and restoration of the KV cache by preset threshold and reducing the restoration latency by pre-computing the matrix, the technical problem of difficulty in coordinating the optimization of video memory usage and computing efficiency in the prior art is solved. It has the advantages of dynamically managing the video memory usage of the KV cache, maintaining the model inference quality, and improving computing efficiency.

[0081] In one embodiment of this disclosure, the preset threshold T is set as follows: ,in, For configurable hyperparameters, satisfying 0 < <1.

[0082] In one embodiment of this disclosure, the system further includes: The configuration module is configured to keep the window permanently open. The size of the fixed retention window is a preset constant. Tokens located within the fixed retention window do not participate in the eviction and restoration operations, but their attention scores are used to calculate the preset threshold.

[0083] In one embodiment of this disclosure, the pre-computed fusion weight matrix The model is calculated once during the initialization phase and remains resident in GPU memory; during the recovery operation, it is only loaded. Instead of loading separately and .

[0084] In one embodiment of this disclosure, the maximum capacity of the V cache storage pool is configured with a preset parameter. m Used to store the most m A vector of values; Accordingly, the recovery operation is performed before the eviction operation, and when the number of value vectors in the V cache storage pool exceeds m after the recovery operation is performed, a forced eviction operation is triggered to evict value vectors with low attention scores, so that the number of value vectors retained in the end does not exceed m.

[0085] In one embodiment of this disclosure, the system performs an evicting and recovery operation during the pre-filling phase to compress the long context key-value cache and improve response efficiency to new problems.

[0086] In one embodiment of this disclosure, the system performs an evicting and recovery operation after each wp inference iteration during the decoding phase, where wp is a preset window length.

[0087] This disclosure also discloses an electronic device. Figure 5A structural block diagram of an electronic device according to an embodiment of the present disclosure is shown.

[0088] like Figure 5 As shown, the electronic device includes a memory and a processor, wherein the memory is used to store one or more computer instructions, wherein the one or more computer instructions are executed by the processor to implement the method according to embodiments of the present disclosure.

[0089] The KV cache compression and evicted lexical recovery method for large-scale language model inference includes: During LLM inference, the following operations are performed for at least one Transformer layer: For the query vector of the i-th word computed in the current Transformer layer Calculate its key vectors and all key vectors calculated in the current Transformer layer. Attention score : ,in, Let j be the dimension of the key vector, j = 1, 2, ..., n; n is the total number of tokens computed in the current Transformer layer. Perform a dynamic update operation on the V cache based on the attention score, including: when At that time, perform an eviction operation, removing the value vector of the j-th term from the V cache storage pool. ; when And the value vector of the j-th word element If expelled, perform a recovery operation, including: via calculation. Regenerate the value vector of the j-th word. And add it to the V cache storage pool; among which, For the pre-computed fusion weight matrix, and These are the key projection matrix and value projection matrix of the current Transformer layer, respectively; T is a preset threshold. For the i-th word element, based on all After completing the V-cache dynamic update operation, the attention value is calculated using the updated set of value vectors in the V-cache storage pool and the pre-computed attention score partial product P. : V new ,in, V new For the updated set of value vectors, .

[0090] In one embodiment of this disclosure, the preset threshold T is set as follows: ,in, For configurable hyperparameters, satisfying 0 < <1.

[0091] In one embodiment of this disclosure, the method further includes: Configure a fixed window to be retained; The size of the fixed retention window is a preset constant. Tokens located within the fixed retention window do not participate in the eviction and restoration operations, but their attention scores are used to calculate the preset threshold.

[0092] In one embodiment of this disclosure, the pre-computed fusion weight matrix The model is calculated once during the initialization phase and remains resident in GPU memory; during the recovery operation, it is only loaded. Instead of loading separately and .

[0093] In one embodiment of this disclosure, the maximum capacity of the V cache storage pool is configured with a preset parameter. m Used to store the most m A vector of values; Accordingly, the recovery operation is performed before the eviction operation, and when the number of value vectors in the V cache storage pool exceeds m after the recovery operation is performed, a forced eviction operation is triggered to evict value vectors with low attention scores, so that the number of value vectors retained in the end does not exceed m.

[0094] In one embodiment of this disclosure, the method performs an evicting and recovery operation during the pre-filling phase to compress the long context key-value cache and improve response efficiency to new problems.

[0095] In one embodiment of this disclosure, the method performs an evicting and recovery operation after each wp inference iteration during the decoding phase, where wp is a preset window length.

[0096] Figure 6 A schematic diagram of the structure of a computer system suitable for implementing the method according to embodiments of the present disclosure is shown.

[0097] like Figure 6As shown, the computer system includes a processing unit that can execute various methods described above based on a program stored in a read-only memory (ROM) or a program loaded from a storage portion into a random access memory (RAM). The RAM also stores various programs and data required for the operation of the computer system. The processing unit, ROM, and RAM are interconnected via a bus. Input / output (I / O) interfaces are also connected to the bus.

[0098] The following components are connected to the I / O interface: input sections including keyboards, mice, etc.; output sections including cathode ray tubes (CRTs), liquid crystal displays (LCDs), and speakers; storage sections including hard disks; and communication sections including network interface cards such as LAN cards and modems. The communication section performs communication processes via a network such as the Internet. Drives are also connected to the I / O interface as needed. Removable media, such as disks, optical disks, magneto-optical disks, semiconductor memories, etc., are installed on the drive as needed so that computer programs read from them can be installed into the storage section as required. The processing unit can be implemented as a CPU, GPU, TPU, FPGA, NPU, etc.

[0099] In particular, according to embodiments of this disclosure, the methods described above can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program tangibly embodied on a machine-readable medium, the computer program containing program code for performing the methods described above. In such embodiments, the computer program can be downloaded and installed from a network via a communication component, and / or installed from a removable medium.

[0100] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0101] The units or modules described in the embodiments of this disclosure can be implemented in software or programmable hardware. The described units or modules can also be located in a processor, and the names of these units or modules do not necessarily constitute a limitation on the unit or module itself.

[0102] In another aspect, this disclosure also provides a computer-readable storage medium, which may be a computer-readable storage medium included in the electronic device or computer system described above; or it may be a standalone computer-readable storage medium not assembled into a device. The computer-readable storage medium stores one or more programs, which are used by one or more processors to perform the methods described in this disclosure.

[0103] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

Claims

1. A method for KV cache compression and eviction lexical recovery for large-scale language model inference, characterized in that, include: During LLM inference, the following operations are performed for at least one Transformer layer: For the query vector of the i-th word computed in the current Transformer layer Calculate its key vectors and all key vectors calculated in the current Transformer layer. Attention score : Where j = 1, 2, ..., n; n is the total number of tokens calculated in the current Transformer layer; Perform a dynamic update operation on the V cache based on the attention score, including: when At that time, perform an eviction operation, removing the value vector of the j-th term from the V cache storage pool. ; when And the value vector of the j-th word element If expelled, perform a recovery operation, including: via calculation. Regenerate the value vector of the j-th word. And add it to the V cache storage pool; among which, For the pre-computed fusion weight matrix, and These are the key projection matrix and value projection matrix of the current Transformer layer, respectively; T is the preset threshold.

2. The method according to claim 1, characterized in that, The preset threshold T is set as follows: ,in, For configurable hyperparameters, satisfying 0 < <1.

3. The method according to claim 1 or 2, characterized in that, The method further includes: Configure a fixed window to be retained; The size of the fixed retention window is a preset constant. Tokens located within the fixed retention window do not participate in the eviction and restoration operations, but their attention scores are used to calculate the preset threshold.

4. The method according to claim 1, characterized in that, The pre-calculated fusion weight matrix The model is calculated once during the initialization phase and remains resident in GPU memory; during the recovery operation, it is only loaded. Instead of loading separately and .

5. The method according to claim 1, characterized in that, The maximum capacity of the V cache storage pool is configured with a preset parameter. m Used to store the most m A vector of values; Accordingly, the recovery operation is performed before the eviction operation, and when the number of value vectors in the V cache storage pool exceeds m after the recovery operation is performed, a forced eviction operation is triggered to evict value vectors with low attention scores, so that the number of value vectors retained in the end does not exceed m.

6. The method according to claim 1, characterized in that, The method performs an eviction and recovery operation during the pre-filling phase to compress the long context KV cache and improve the response efficiency to new problems.

7. The method according to claim 1, characterized in that, The method performs an evicting and recovery operation after each wp inference iteration during the decoding phase, where wp is a preset window length.

8. The method according to claim 1, characterized in that, The method further includes: For the i-th word element, based on all After completing the V-cache dynamic update operation, the attention value is calculated using the updated set of value vectors in the V-cache storage pool and the pre-computed attention score partial product P. : V new ,in, V new For the updated set of value vectors, .

9. A large-scale language model reasoning system, characterized in that, include: The attention score calculation module is configured to calculate the query vector of the i-th word in the current Transformer layer. Calculate its key vectors and all key vectors calculated in the current Transformer layer. Attention score : Where j = 1, 2, ..., n; n is the total number of tokens calculated in the current Transformer layer; The V-cache dynamic management module is configured to perform V-cache dynamic update operations based on the attention score, including: when At that time, perform an eviction operation, removing the value vector of the j-th term from the V cache storage pool. ; when And the value vector of the j-th word element If expelled, perform a recovery operation, including: via calculation. Regenerate the value vector of the j-th word. And add it to the V cache storage pool; among which, For the pre-computed fusion weight matrix, and These are the key projection matrix and value projection matrix of the current Transformer layer, respectively; T is the preset threshold.

10. A computer-readable storage medium storing computer instructions thereon, characterized in that, When executed by a processor, the computer instructions implement the method described in any one of claims 1-7.

Citation Information

Patent Citations

  • Large-scale language model KV Cache optimization method based on recent query attention information

    CN119396995A

  • Key value caching method and device, equipment, storage medium and product

    CN119941879A

  • Model reasoning method, computer program product and chip

    CN120197702A

  • Question and answer reasoning method and device based on key value cache compression, equipment and medium

    CN120598057A

  • Dynamic quantization and memory management of key-value cache for serving large language models

    US20250061316A1

Cited By

  • Multi-layer attention model reasoning acceleration method and device based on anchor point compression mechanism

    CN121351903A

  • Model reasoning optimization method and electronic equipment

    CN122133816A

  • Model inference optimization method and electronic device

    CN122133816B