Large language model text answering method integrated with draft answering and KV cache expelling
By incorporating draft answers and KV cache eviction into the text-based answering method of large language models, the problem of decreased answer quality in the KV cache eviction method is solved, more accurate answers are achieved under low cache conditions, and the memory usage of the model when generating answers is reduced.
Patent Information
- Application Number
- CN202511011151.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-22
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-07-22
AI Technical Summary
Existing key-value cache eviction methods, in long-context scenarios, fail to fully reflect the overall contextual text information due to the limited key-value caches retained, leading to a decline in the quality of model responses.
This paper proposes a text-based answering method for large language models that incorporates draft answers and key-value caching. By segmenting and encoding long text sequences, it retains the query vectors at the end of the query vector set. Combined with attention score calculation, it retains important key vectors and value vectors, performs autoregressive operations to generate draft answers, and finally generates accurate final answers.
At the same level of answer accuracy, the use of key-value cache is reduced, the accuracy of the answer and the comprehensiveness of information consideration are improved, and the memory usage of the model when generating the answer is reduced.
Smart Images

Figure CN120849565A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of large language model text response generation. Background Technology
[0002] In recent years, with the widespread application of large language models for input text problems, generative tasks (such as multi-turn question answering and content creation) have placed higher demands on model response speed and contextual understanding capabilities. To improve inference efficiency, mainstream Transformer architectures generally introduce a key-value cache (KV) mechanism. This mechanism can avoid repeatedly calculating historical context information, significantly accelerating the model's inference speed. However, in long context scenarios, there is a lot of redundant information in the KV cache that can be evicted, and a longer KV cache will increase model decoding latency and machine memory usage.
[0003] Existing key-value (KV) cache eviction methods, such as StreamingLLM, typically maintain only the first and last KV caches of long text sequences. SnapKV, on the other hand, calculates attention scores for the preceding text using the last observation window in the long text sequence, retaining the KV caches corresponding to the high-scoring tokens. The drawback of these existing methods is that the retained KV caches are incomplete and fail to reflect the overall contextual text information, which is inconsistent with the KV caches the model focuses on during the response. This leads to a significant decrease in response quality after evicting the KV caches. These problems urgently need to be addressed. Summary of the Invention
[0004] The purpose of this invention is to address the problem of low response quality in large language model response methods based on existing key-value (KV) cache eviction techniques. This invention provides a text response method for large language models that incorporates draft responses and KV cache eviction.
[0005] A text-based response method for large language models that incorporates draft responses and key-value cache eviction, comprising:
[0006] S1. After segmenting and encoding the long sequence text sentence X, multiple question words are obtained and fed into the large language model. After the forward operation of the model, the key vector set K, value vector set V and query vector set Q corresponding to the long sequence text sentence X are obtained.
[0007] X={x 1 ,…,x t}, K = {K 1 ,…,K t}, V={V 1 ,…,V t}, Q={Q 1 ,…,Q t}, x i For the i-th question term, Ki 、V i and Q i x i The key vector, value vector, and query vector, i = 1, 2, ..., t; t is the total number of question terms;
[0008] S2. Evict the key vector set K and the value vector set V from the cache, resulting in the evicted key vector set K1 and value vector set V1. And retain the query vectors at the tail of the query vector set Q, to obtain the query vector set Q1 = {Q} t-8 ,…,Q t};
[0009] S3. Perform autoregression on the key vector set K1 and value vector set V1 to obtain a draft answer consisting of the first n answer tokens generated during the autoregression process. Let j be the word element of the j-th answer, where j = 1, 2, ..., n;
[0010] Perform a forward operation on the n answer tokens to obtain a query vector set consisting of the query vectors of the n answer tokens. Let j be the query vector for the j-th answer word element;
[0011] S4. Based on Q1 and Q2, construct the query vector set Q3 = Q2 or Q3 = (Q1∪Q2);
[0012] Using the query vector set Q3, calculate the attention score A for each question word in the long sequence text sentence X. i And retain the key vectors and value vectors corresponding to the B question words with high attention scores, forming a key vector set composed of the key vectors and value vectors of the B question words respectively. Sum value vector set
[0013] A i For x i Attention score Let be the key vector of the b-th question word in K2. Let b be the value vector of the b-th question word in V2, where b = 1, 2, ..., B;
[0014] S5. Perform autoregression on the key vector set K2 and the value vector set V2 to obtain the final answer composed of all answer words. This final answer is the answer corresponding to the long sequence text sentence X output by the large language model.
[0015] Preferably, in step S1, Q i =x i W Q K i =xi W K V i =x i W V ;
[0016] Among them, W K 、W V and W Q These are the parameter matrices for the key vector, value vector, and query vector of the large language model, respectively.
[0017] Preferably, in step S2, the principle of cache eviction is as follows:
[0018] Retaining the first 4 and last M-4 vectors in the sorting direction from the key vector set K and value vector set V, we obtain the expelled key vector set. Sum value vector set
[0019] Let be the key vector of the m-th question word in the key vector set K1. Let m be the value vector of the m-th question word in the value vector set V1, where m = 1, 2, ..., M, M is the total number of vectors in set K1 or V1, and M ≤ B.
[0020] Preferably, in step S3,
[0021] W Q This is the parameter matrix of the query vector for the large language model.
[0022] Preferably, in step S4,
[0023]
[0024] h is the index of the query vector in the query vector set Q3, h = 1, 2, ..., |Q3|, where |Q3| is the length of the query vector set Q3.
[0025] A large language model text response system incorporating draft responses and key-value (KV) cache eviction includes a storage device, a processor, and a computer program stored in the storage device and executable on the processor. The processor executes the computer program to implement the large language model text response method incorporating draft responses and KV cache eviction.
[0026] A computer-readable storage device stores a computer program that, when executed, implements the large language model text response method incorporating draft responses and KV cache eviction.
[0027] A computer program product includes a computer program that, when executed by a processor, implements the large language model text response method incorporating draft responses and KV cache eviction.
[0028] The beneficial effects of this invention are:
[0029] The large language model text answering method described in this invention, which incorporates draft answers and KV cache eviction, uses information from draft answers, making the remaining small portion of the KV cache (K2 and V2) more important. Furthermore, attention scores are introduced during the process of obtaining the remaining small portion of the KV cache (K2 and V2), resulting in a more comprehensive consideration of information. Compared with existing methods that use KV cache eviction techniques to generate answers from the model, this method can obtain more accurate answers.
[0030] The text-based answering method of this invention, compared with the method of generating answers using existing key-value (KV) cache eviction techniques, requires less KV cache while maintaining the same answer accuracy. Attached Figure Description
[0031] Figure 1 This is a flowchart of the large language model text response method incorporating draft responses and KV cache eviction as described in this invention;
[0032] Figure 2 This is a schematic diagram illustrating the principle of generating the key vector set K, the value vector set V, and the query vector set Q;
[0033] Figure 3 This is a schematic diagram illustrating the principle of generating draft answers using K1 and V1;
[0034] Figure 4 This is a schematic diagram illustrating the principle of generating the final answer using K2 and V2. Detailed Implementation
[0035] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0036] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments may be combined with each other.
[0037] The present invention will be further described below with reference to the accompanying drawings and specific embodiments, but this is not intended to limit the scope of the invention.
[0038] Specific Implementation Method 1: Combination Figure 1As shown in this embodiment, the large language model text answering method incorporating draft answers and KV cache eviction includes:
[0039] S1. After segmenting and encoding the long sequence text sentence X, multiple question words are obtained and fed into the large language model. After the forward operation of the model, the key vector set K, value vector set V and query vector set Q corresponding to the long sequence text sentence X are obtained.
[0040] X={x 1 ,…,x t}, K = {K 1 ,…,K r}, V={V 1 ,…,V t}, Q={Q 1 ,…,Q t}, x i For the i-th question term, K i 、V i and Q i x i The key vector, value vector, and query vector, i = 1, 2, ..., t; t is the total number of question terms; see also Figure 2 ;
[0041] S2. Evict the key vector set K and the value vector set V from the cache, resulting in the evicted key vector set K1 and value vector set V1. And retain the query vectors at the tail of the query vector set Q, to obtain the query vector set Q1 = {Q} t-8 ,…,Q t See also Figure 3 ;
[0042] S3. Perform autoregression on the key vector set K1 and value vector set V1 to obtain a draft answer consisting of the first n answer tokens generated during the autoregression process. Let j be the word element of the j-th answer, where j = 1, 2, ..., n;
[0043] Perform a forward operation on the n answer tokens to obtain a query vector set consisting of the query vectors of the n answer tokens. Let j be the query vector for the j-th answer word element;
[0044] S4. Based on Q1 and Q2, construct the query vector set Q3 = Q2 or Q3 = (Q1∪Q2);
[0045] Using the query vector set Q3, calculate the attention score A for each question word in the long sequence text sentence X. iAnd retain the key vectors and value vectors corresponding to the B question words with high attention scores, forming a key vector set composed of the key vectors and value vectors of the B question words respectively. Sum value vector set
[0046] A i For x i Attention score Let be the key vector of the b-th question word in K2. Let b be the value vector of the b-th question word in V2, where b = 1, 2, ..., B;
[0047] S5. Perform autoregression on the key vector set K2 and value vector set V2 to obtain the final answer, which is composed of all answer words. This final answer serves as the response to the long sequence text sentence X output by the large language model. See [link / reference]. Figure 4 .
[0048] This preferred embodiment first performs cache eviction on the sets K and V corresponding to the obtained long sequence text sentence X, retaining a small portion of the KV cache. Then, it uses the retained sets K1 and V1 for autoregression to generate a draft answer. Based on the draft answer, it calculates the attention score, thereby obtaining the key vector set K2 and value vector set V2 of the B most important question words in the original set. This provides a more comprehensive consideration of information. Furthermore, by using K2 and V2 for autoregression, a more accurate answer can be obtained.
[0049] Furthermore, in step S1, Q i =x i W Q K i =x i W K V i =x i W V ;
[0050] Among them, W K 、W V and W Q These are the parameter matrices for the key vector, value vector, and query vector of the large language model, respectively.
[0051] Furthermore, in step S2, the principle of cache eviction is as follows:
[0052] Retaining the first 4 and last M-4 vectors in the sorting direction from the key vector set K and value vector set V, we obtain the expelled key vector set. Sum value vector set
[0053] Let be the key vector of the m-th question word in the key vector set K1. Let m be the value vector of the m-th question word in the value vector set V1, where m = 1, 2, ..., M, M is the total number of vectors in set K1 or V1, and M ≤ B.
[0054] Furthermore, in step S3,
[0055] W Q This is the parameter matrix of the query vector for the large language model.
[0056] Furthermore, in step S4,
[0057]
[0058] h is the index of the query vector in the query vector set Q3, h = 1, 2, ..., |Q3|, where |Q3| is the length of the query vector set Q3.
[0059] This preferred embodiment calculates attention scores based on draft answers, resulting in a more comprehensive consideration of information. This leads to the acquisition of the key vector set K2 and value vector set V2 of the B most important question words in the original set, providing a more comprehensive understanding of information. Furthermore, by utilizing K2 and V2 autoregression, a more accurate answer can be obtained.
[0060] Specific Implementation Method Two: The large language model text response system incorporating draft responses and KV cache eviction described in this implementation method includes a storage device, a processor, and a computer program stored in the storage device and executable on the processor. The processor executes the computer program to implement the large language model text response method incorporating draft responses and KV cache eviction as described in Specific Implementation Method One.
[0061] Specific Implementation Method 3: A computer-readable storage device storing a computer program, wherein when the computer program is executed, it implements the large language model text response method incorporating draft responses and KV cache eviction as described in Specific Implementation Method 1.
[0062] Specific implementation method four: A computer program product, including a computer program that, when executed by a processor, implements the large language model text response method incorporating draft responses and KV cache eviction as described in specific implementation method one.
[0063] Experimental results:
[0064] The test results for the "needle in a haystack" task with a context length of 32K are shown in Table 1 below.
[0065] The test models include Mistral-7B-v0.2-Instruct, Llama3.1-8B-Instruct, and Qwen2.5-7B-Instruct;
[0066] The KV budget size settings for the test include three options: 64, 96, and 128.
[0067] The comparison methods include H2O, SnapKV, PyramidKV, and the method of this invention.
[0068] Table 1 Test Results
[0069]
[0070]
[0071] Referring to Table 1, the text response method of the present invention, compared with the method of generating responses using existing KV cache eviction techniques, occupies less KV cache while maintaining the same response accuracy; and in the case of low-budget KV cache, it can obtain more accurate responses than the method of generating responses using existing KV cache eviction techniques, greatly saving the GPU memory occupied by the model when generating responses.
[0072] Although the present invention is described herein with reference to specific embodiments, it should be understood that these embodiments are merely illustrative of the principles and applications of the invention. It should be understood that many modifications may be made to the illustrative embodiments, and that other arrangements may be devised, without departing from the spirit and scope of the invention as defined by the appended claims. It should be understood that the various dependent claims and features described herein may be combined in ways other than those described in the original claims. It should also be understood that features described in conjunction with individual embodiments may be employed in conjunction with other described embodiments.
Claims
1. A text-based response method for large language models that incorporates draft responses and key-value (KV) cache eviction, characterized by: The method includes: S1. After segmenting and encoding the long sequence text sentence X, multiple question words are obtained and fed into the large language model. After the forward operation of the model, the key vector set K, value vector set V and query vector set Q corresponding to the long sequence text sentence X are obtained. X={x 1 ,…,x t }, K = {K 1 ,…,K t }, V={V 1 ,…,V t }, Q={Q 1 ,…,Q t }, x i For the i-th question term, K i 、V i and Q i x i The key vector, value vector, and query vector, i = 1, 2, ..., t; t is the total number of question terms; S2. Evict the key vector set K and the value vector set V from the cache, resulting in the evicted key vector set K1 and value vector set V1. And retain the query vectors at the tail of the query vector set Q, to obtain the query vector set Q1 = {Q} t-8 ,…,Q t }; S3. Perform autoregression on the key vector set K1 and value vector set V1 to obtain a draft answer consisting of the first n answer tokens generated during the autoregression process. Let j be the word element of the j-th answer, where j = 1, 2, ..., n; Perform a forward operation on the n answer tokens to obtain a query vector set consisting of the query vectors of the n answer tokens. Let j be the query vector for the j-th answer word element; S4. Based on Q1 and Q2, construct a query vector set Q3 = Q2 or Q3 = (Q1∪Q2); Using the query vector set Q3, calculate the attention score A for each question word in the long sequence text sentence X. i And retain the key vectors and value vectors corresponding to the V question words with high attention scores, forming a key vector set composed of the key vectors and value vectors of the B question words respectively. Sum value vector set A i For x i Attention score Let be the key vector of the b-th question word in K2. Let b be the value vector of the b-th question word in V2, where b = 1, 2, ..., B; S5. Perform autoregression on the key vector set K2 and value vector set V2 to obtain the final answer composed of all answer words. This final answer is the answer corresponding to the long sequence text sentence X output by the large language model.
2. The text response method for large language models incorporating draft responses and KV cache eviction as described in claim 1, characterized in that, In step S1, Q i =x i W Q K i =x i W K V i =x i W V ; Among them, W K W V and W Q These are the parameter matrices for the key vector, value vector, and query vector of the large language model, respectively.
3. The text response method for large language models incorporating draft responses and KV cache eviction as described in claim 1, characterized in that, In step S2, the principle of cache eviction is as follows: Retaining the first 4 and last M-4 vectors in the sorting direction from the key vector set K and value vector set V, we obtain the expelled key vector set. Sum value vector set Let be the key vector of the m-th question word in the key vector set K1. Let m be the value vector of the m-th question word in the value vector set V1, where m = 1, 2, ..., M, M is the total number of vectors in set K1 or V1, and M ≤ B.
4. The text response method for large language models incorporating draft responses and KV cache eviction as described in claim 1, characterized in that, In step S3, W Q This is the parameter matrix of the query vector for the large language model.
5. The text response method for large language models incorporating draft responses and KV cache eviction as described in claim 1, characterized in that, In step S4, h is the index of the query vector in the query vector set Q3, h = 1, 2, ..., |Q3|, where |Q3| is the length of the query vector set Q3.
6. A large language model text response system incorporating draft responses and key-value (KV) cache eviction, comprising a storage device, a processor, and a computer program stored in the storage device and executable on the processor, characterized in that, The processor executes a computer program to implement the large language model text response method as described in any one of claims 1 to 6, which incorporates draft responses and KV cache eviction.
7. A computer-readable storage device storing a computer program, characterized in that, When the computer program is executed, it implements the large language model text response method as described in any one of claims 1 to 6, which incorporates draft responses and KV cache eviction.
8. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the large language model text response method as described in claims 1 to 6, which incorporates draft responses and KV cache eviction.
Citation Information
Patent Citations
Method, device and equipment for realizing knowledge questions and answers
CN119621900A
Caching method for large language model questions and answers
CN119719336A
Document question and answer processing method and system, electronic equipment, storage medium and computer program product
CN119917628A
Personalized question answering using semantic caching
US20240070489A1