Method and system for improving reasoning length and performance of large model and medium

By adding gating signals to the Attention layer and filtering key_states and value_states, the problems of token diversity and retention in the existing technology are solved, the memory usage of KV cache is reduced, and the accuracy and inference performance of generated content are improved.

CN120197699APending Publication Date: 2025-06-24KYLIN CORP
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510299384.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The existing technology cannot meet the diversity of tokens when eliminating tokens, resulting in unfair retention of tokens, unable to effectively reduce the memory usage of KV cache, and at the same time affect the accuracy of generated content.

Method used

Add gating signals to the Attention layer, filter key_states and value_states through the gating signals, generate the filtered key state tensor c and value state tensor d, and assign them to key_states and value_states to perform attention mechanism calculation.

Benefits of technology

It reduces the complexity of attention mechanism calculation, reduces the memory usage of KV cache, and improves the accuracy of generated content and improves inference performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197699A_ABST
    Figure CN120197699A_ABST
Patent Text Reader

Abstract

The invention discloses a method and system for improving the reasoning length and performance of a large model and a medium, and the method comprises the steps: adding a gating signal in an Attention layer, enabling the gating signal to screen data in a key state tensor and a value state tensor, and carrying out the attention mechanism calculation through employing the key and the value after screening and assignment. Due to the fact that the lengths of the screened keystates and values are fixed, the complexity of attention mechanism calculation is reduced, the KV cache occupies a video memory, the reasoning performance is improved, the token corresponding to the key matrix reserved in the keystates is key information, and the accuracy of the generated content is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of large model video memory optimization, and particularly relates to a method, system and medium for improving the inference length and performance of large models. Background Art

[0002] Currently, large model technology has developed rapidly. As a result, the model size has been continuously increasing, the input and output lengths of the model have been continuously lengthening, and the consumption of inference resources has been rising rapidly. To accelerate inference, the main method is to use KV cache. The widespread application of KV cache has greatly improved the inference speed of the model, but it has also brought huge challenges to video memory resources. Suppose the length of the model input sequence is s, the length of the output sequence is n, the model depth is l, and the dimension is h. Saving the KV cache in FP16, then the peak video memory occupancy size of the KV cache is b(s + n)h×l×2×2 = 4blh(s + n). Taking GPT3 (175B) as an example, compare the video memory occupancy of the KV cache with that of the model parameters. The video memory occupancy of the GPT3 model weight is 350GB (FP16), the number of layers l is 96, and the dimension h is 12888. As the batch size, input sequence length, and output sequence length increase, the video memory overhead occupied by the KV cache increases rapidly and may even exceed the model itself.

[0003] The prior art often uses the method of removing tokens to reduce the video memory overhead occupied by the KV cache. For example, "EFFICIENT STREAMING LANGUAGE MODELS WITH ATTENTION SINKS" proposes a StreamingLLM that can remove tokens. However, since StreamingLLM removes fixed tokens, it cannot meet the diversity of tokens, nor does it conform to human subjective thinking. The token correlation in the sentence is not consistent, resulting in poor accuracy of the generated content; "Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models" proposes a method of removing tokens by accumulating the values of K for different Qs, but because the probabilities behind are sparse, fairness for each token cannot be achieved. Summary of the Invention

[0004] Technical problems to be solved by the present invention: In view of the above problems of the prior art, a method, system and medium for improving the inference length and performance of a large model are provided. The present invention aims to solve the problems that the existing method of eliminating tokens cannot meet the diversity of tokens and the tokens are unfairly retained, and can reduce the video memory size occupied by the KV cache while improving the accuracy of the generated content.

[0005] To solve the above technical problems, the technical solution adopted by the present invention is as follows: A method for improving the inference length and performance of a large model, which adds a gating signal in the Attention layer, and the execution process of the gating signal includes the following steps: Step S1, restore the dimensions of the key state tensor key_states and the value state tensor value_states; Step S2, perform a linear transformation on the key_states through a linear layer to generate a gating score tensor key_states_zip corresponding to the token positions of the input sequence, and its dimension is [batchsize, num_attention_heads, seqlen, 1], where batchsize is the batch size, num_attention_heads is the number of attention heads, and seqlen is the number of tokens in the input sequence; Step S3, obtain the indices of the top M positions with the highest gating scores in each batch and each attention head of the key_states_zip in the seqlen dimension, where M is the screening threshold; Step S4, based on the indices, extract the corresponding key matrices and value matrices from the key_states and value_states to generate a filtered key state tensor c and a value state tensor d, assign the c to the key_states, and assign the d to the value_states; Step S5, based on the assigned key_states and value_states, perform the attention mechanism calculation.

[0006] Optionally, in step S3, if seqlen is greater than M, obtain the indices of the top M positions with the highest gating scores in each batch and each attention head of the key_states_zip in the seqlen dimension; if seqlen is less than or equal to M, obtain the indices of each position in each batch of the key_states_zip in the seqlen dimension.

[0007] Optionally, after the assignment is recorded in the cache, the gating scores of each key matrix in the key_states and the storage location information of the key matrix in the video memory are stored in the key matrix selection queue. When new Query information enters the calculation, the gating score of the latest key matrix is calculated and compared with the gating scores of each key matrix in the cache queue. If the gating score of the latest key matrix is less than the gating scores of all key matrices in the key matrix selection queue, the gating score of the latest key matrix is discarded; otherwise, the smallest one of the gating scores of all key matrices in the key matrix selection queue and its corresponding storage location information of the key matrix in the video memory are discarded, and the gating score of the latest key matrix and its storage location information in the video memory are saved to the last position of the key matrix selection queue.

[0008] Optionally, when the gating score of the key matrix is discarded in the key matrix selection queue, according to the storage location information of the key matrix stored in the key matrix selection queue in the video memory, the storage location of the key matrix corresponding to the discarded gating score in the video memory is found, and the latest key matrix is used to replace the discarded key matrix and stored at this location.

[0009] Optionally, when the latest key matrix replaces the original key matrix in the video memory, the latest value matrix is replaced according to the corresponding replacement index with the original value matrix corresponding to the original key matrix.

[0010] Optionally, pre-allocate the video memory according to the screening threshold M and the lengths of each key matrix and value matrix.

[0011] Optionally, divide the pre-allocated video memory into multiple video memory blocks, and dynamically allocate the video memory blocks according to the actual number of key matrices and value matrices.

[0012] In addition, the present invention also provides a system for improving the inference length and performance of a large model, including a microprocessor and a memory connected to each other, and the microprocessor is programmed or configured to execute the method for improving the inference length and performance of the large model.

[0013] In addition, the present invention also provides a computer-readable storage medium, in which a computer program or instruction is stored, and the computer program or instruction is programmed or configured to execute the method for improving the inference length and performance of the large model through a processor.

[0014] In addition, the present invention also provides a computer program product, including a computer program or instruction, and the computer program or instruction is programmed or configured to execute the method for improving the inference length and performance of the large model through a processor.

[0015] Compared with the prior art, the present invention mainly has the following advantages: In the present invention, a gating signal is added to the Attention layer. The lengths of the key_states and value_states after being screened by the gating signal are fixed and will not increase with the increase of the input sequence length, reducing the complexity of the attention mechanism calculation, improving the inference performance, reducing the VRAM occupied by the KV cache. At the same time, the gating scores of the key matrices retained in the screened key_states are higher than those of the eliminated key matrices, which means that the tokens corresponding to the retained key matrices are key information, facilitating the improvement of the accuracy of the generated content. Description of the Drawings

[0016] Figure 1 It is a partial flowchart of the attention mechanism calculation in the method of the embodiment of the present invention.

[0017] Figure 2 It is a schematic diagram of the training effect of the Attention of the original large model.

[0018] Figure 3 It is a schematic diagram of the training effect of the Attention of the large model adopting the method of the embodiment of the present invention. Detailed Embodiment

[0019] Next, the technical solution of the present invention will be further described in detail in conjunction with the drawings.

[0020] As Figure 1 shown, the method for improving the inference length and performance of the large model in this embodiment adds a gating signal to the Attention layer. The execution process of the gating signal includes the following steps: Step S1, restore the dimensions of the key state tensor key_states and the value state tensor value_states; Step S2, perform a linear transformation on key_states through a linear layer to generate a gating score tensor key_states_zip corresponding to the token positions of the input sequence, whose dimension is [batchsize, num_attention_heads, seqlen, 1], where batchsize is the batch size, num_attention_heads is the number of attention heads, and seqlen is the number of tokens in the input sequence; Step S3, obtain the indexes of the top M positions with the highest gating scores in each batch and each attention head of key_states_zip in the seqlen dimension, where M is the screening threshold; Step S4: Based on the index, extract the key matrix and value matrix at the corresponding positions from key_states and value_states, generate the filtered key state tensor c and value state tensor d, assign c to key_states, and assign d to value_states; Step S5: Based on the assigned key_states and value_states, perform the attention mechanism calculation.

[0021] Figure 1 In it, Key is the key matrix, Value is the value matrix, and Query is the query matrix.

[0022] This method adds a gating signal in the Attention layer. The lengths of key_states and value_states after being filtered by the gating signal are fixed and will not increase with the increase of the input sequence length, reducing the complexity of the attention mechanism calculation, improving the inference performance, reducing the video memory occupied by the KV cache. At the same time, the gating scores of the key matrices retained in the filtered key_states are higher than those of the eliminated key matrices, which means that the tokens corresponding to the retained key matrices are key information, conducive to improving the accuracy of the generated content; in addition, since the length of the KV cache is fixed, no additional padding information needs to be added during the processing of the Batch, improving the calculation efficiency.

[0023] Corresponding to the above embodiment, the gating signal is a trainable linear layer gating signal, and its training pseudo-code is as follows: # In the forward function of the Qwen2Attention class … Restore the dimensions of key_states and value_states; Create a linear layer from key_states to key_states_zip for outputting the number N of gating score values, with the dimension of [batchsize, num_attention_heads, seqlen, 1]; Set the threshold of the number of gating score values to M; IF N>M Obtain the indices k_indices of the largest M gating score values in each batch and each attention head in seqlen; Create temporary tensors c and d with dimensions equal to key_states, and then replace the size of the third dimension with M; FOR i In Range(batchsize): FOR j In Range(num_attention_heads): FOR k In Range(M): Get the index k_indices[i, j, k], recorded as idx; Assign a value to c, c[i, j, k] is equal to key_states[i, j, idx, :]; Assign a value to d, d[i, j, k] is equal to value_states[i, j, idx, :]; Replace key_states with c; Replace value_states with d; #At this point, the dimension problem is solved, and subsequent calculations can be performed normally.

[0024] As an optional embodiment, in this embodiment, in step S3, if seqlen is greater than M, the indexes of the top M positions with the highest gating scores in each attention head of each batch of key_states_zip under the seqlen dimension are obtained; if seqlen is less than or equal to M, the index of each position of each batch of key_states_zip under the seqlen dimension is obtained.

[0025] As an optional embodiment, the gating scores of each key matrix in the key_states after the assignment and the storage location information of the key matrix in the video memory are recorded in the cache and stored in the key matrix selection queue. When new Query information enters the calculation, the gating score of the latest key matrix is ​​calculated and compared with the gating scores of each key matrix in the cache queue. If the gating score of the latest key matrix is ​​less than the gating scores of all key matrices in the key matrix selection queue, the gating score of the latest key matrix is ​​discarded; otherwise, the smallest one of the gating scores of all key matrices in the key matrix selection queue and the storage location information of the corresponding key matrix in the video memory are discarded, and the gating score of the latest key matrix and the storage location information of the key matrix in the video memory are saved to the last bit of the key matrix selection queue. By recording the gating scores of each key matrix in the key_states after the assignment and the storage location information of the key matrix in the video memory in the cache, it is only necessary to calculate the gating score of the latest key matrix and compare it with the gating score in the cache. According to the comparison result, it is selected to be retained or discarded, which can reduce the calculation amount of the gating signal and improve the reasoning speed.

[0026] As an alternative embodiment, when the key matrix gating score is discarded in the key matrix selection queue, based on the storage location information of the stored key matrix in the video memory in the key matrix selection queue, the storage location of the key matrix corresponding to the discarded key matrix gating score in the video memory is found, and the latest key matrix is used to replace the discarded key matrix and stored at this location. There is no need to reorder the original video memory content, reducing the video memory access time. It should be noted that the key matrices in the video memory are stored in the form of a queue. Since the discarded key matrices are not stored in the video memory, the space of the video memory can be reduced, allowing the system to increase the selection queue of the key matrices in the video memory to maintain more requests.

[0027] As an alternative embodiment, when the latest key matrix replaces the original key matrix in the video memory, the latest value matrix is replaced with the original value matrix corresponding to the original key matrix according to the corresponding replacement index. There is no need to store the value matrix and its access location data in the video memory in the cache queue, saving cache space.

[0028] Figure 2 Schematic diagram of the training effect of the original large model Attention Figure 3 Schematic diagram of the training effect of the large model Attention using the method of this embodiment. From Figure 3 it can be seen that the large model using this method can achieve good convergence, and there is no obvious gap in the Attention training effect compared with the Attention training effect of the original large model.

[0029] As an alternative embodiment, pre-allocate the video memory according to the screening threshold M and the lengths of each key matrix and value matrix. The number of video memory allocations can be reduced.

[0030] As an alternative embodiment, divide the pre-allocated video memory into multiple video memory blocks, and dynamically allocate the video memory blocks according to the actual quantities of the key matrix and value matrix, which can reduce the generation of video memory fragmentation and improve the video memory utilization rate.

[0031] In addition, this embodiment also provides a system for improving the inference length and performance of a large model, including a microprocessor and a memory connected to each other. The microprocessor is programmed or configured to execute the method for improving the inference length and performance of the large model.

[0032] In addition, this embodiment also provides a computer-readable storage medium, in which a computer program or instruction is stored, and the computer program or instruction is programmed or configured to execute the method for improving the inference length and performance of the large model through a processor.

[0033] In addition, this embodiment also provides a computer program product, including a computer program or instruction, and the computer program or instruction is programmed or configured to execute the method for improving the inference length and performance of the large model through a processor.

[0034] Those skilled in the art should understand that the technical solutions provided by the embodiments of the present application can be in the form of a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code. The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks. These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks. These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Therefore, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0035] The above is only the preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions within the idea of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art in this technical field, several improvements and refinements made without departing from the principle of the present invention should also be regarded as within the protection scope of the present invention.

Claims

1. A method for improving the inference length and performance of large models, characterized in that: A gating signal is added to the Attention layer. The execution process of the gating signal includes the following steps: Step S1, restore the dimensions of the key state tensor key_states and the value state tensor value_states; Step S2, perform a linear transformation on the key_states through a linear layer to generate a gated fraction tensor key_states_zip corresponding to the token position of the input sequence, whose dimension is [batchsize, num_attention_heads, seqlen, 1], where batchsize is the batch size, num_attention_heads is the number of attention heads, and seqlen is the number of tokens in the input sequence; Step S3, obtaining the indexes of the top M positions with the highest gating scores in each attention head of each batch of the key_states_zip under the seqlen dimension, where M is the screening threshold; Step S4, based on the index, extract the key matrix and value matrix of the corresponding position from key_states and value_states, generate the filtered key state tensor c and value state tensor d, and assign the c to the key_states, and assign the d to the value_states; Step S5, performing attention mechanism calculation based on the assigned key_states and value_states.

2. The method for improving the inference length and performance of large models according to claim 1, characterized in that: In step S3, if seqlen is greater than M, the indexes of the top M positions with the highest gating scores in each attention head of each batch of key_states_zip under the seqlen dimension are obtained; If seqlen is less than or equal to M, then get the index of each position of each batch of key_states_zip under the seqlen dimension.

3. The method for improving the inference length and performance of large models according to claim 1, characterized in that: The gating scores of each key matrix in the key_states after assignment and the storage location information of the key matrix in the video memory are recorded in the cache and stored in the key matrix selection queue. When new Query information enters the calculation, the gating score of the latest key matrix is ​​calculated and compared with the gating scores of each key matrix in the cache queue. If the gating score of the latest key matrix is ​​less than the gating scores of all key matrices in the key matrix selection queue, the gating score of the latest key matrix is ​​discarded; otherwise, the smallest one among all the key matrix gating scores in the key matrix selection queue and the storage location information of the corresponding key matrix in the video memory are discarded, and the gating score of the latest key matrix and the storage location information of the key matrix in the video memory are saved to the last position of the key matrix selection queue.

4. The method for improving the inference length and performance of large models according to claim 3, characterized in that: When the key matrix gating score is discarded in the key matrix selection queue, the storage position of the key matrix corresponding to the discarded key matrix gating score in the video memory is found according to the storage position information of the key matrix stored in the key matrix selection queue in the video memory, and the latest key matrix replaces the discarded key matrix and is stored in the position.

5. The method for improving the inference length and performance of large models according to claim 4, characterized in that: When the latest key matrix appears in the video memory to replace the original key matrix, the latest value matrix replaces the original value matrix corresponding to the original key matrix according to the corresponding replacement index.

6. The method for improving the inference length and performance of large models according to claim 1, characterized in that: The video memory is pre-allocated according to the screening threshold M and the length of each key matrix and value matrix.

7. The method for improving the inference length and performance of large models according to claim 1, characterized in that: The pre-allocated video memory is divided into a plurality of video memory blocks, and the video memory blocks are dynamically allocated according to the actual number of key matrices and value matrices.

8. A system for improving the inference length and performance of large models, comprising a microprocessor and a memory connected to each other, characterized in that: The microprocessor is programmed or configured to execute the method for improving the reasoning length and performance of large models as described in any one of claims 1 to 7.

9. A computer-readable storage medium having a computer program or instruction stored therein, characterized in that: The computer program or instruction is programmed or configured to execute the method for improving the reasoning length and performance of large models as described in any one of claims 1 to 7 through a processor.

10. A computer program product comprising a computer program or instructions, characterized in that The computer program or instruction is programmed or configured to execute the method for improving the reasoning length and performance of large models as described in any one of claims 1 to 7 through a processor.

Citation Information

Cited By

  • Large model reasoning length and performance optimization method

    CN121882287A

  • A large model inference length and performance optimization method

    CN121882287B