An attention mechanism calculation method, calculation system and storage medium

By optimizing the attention mechanism calculation process in the large language model, especially in the forward and backpropagation stages, the strategies of operator fusion and computing process rearrangement are adopted to solve the problem of inefficient attention mechanism calculation on the existing technology in high computing memory access ratio GPU, and a significant performance improvement is achieved.

CN118333167BActive Publication Date: 2025-06-27SHANGHAI JIAOTONG UNIV

Patent Information

Application Number
CN202410439736.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-12
Publication Date
2025-06-27
Estimated Expiration
2044-04-12

AI Technical Summary

Technical Problem

The attention mechanism acceleration optimization technology in the existing large language models is mainly suitable for GPUs with relatively low computing memory fetching. It fails to effectively optimize GPUs with relatively high computing memory fetching, resulting in inefficient computing efficiency of attention mechanisms on these GPUs.

Method used

By optimizing the forward propagation and backpropagation stages of the attention mechanism, the strategies of operator fusion and calculation process rearrangement are adopted to reduce the number of memory accesses and inventory accesses, and the calculation order and memory access order are adjusted to improve efficiency.

Benefits of technology

The efficiency of attention computing process in large language models is improved, and the training and inference performance of the model on the GPU on high computing memory access ratios is improved. Experiments show that forward propagation is accelerated by about 30% and backpropagation is accelerated by about 10%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118333167B_ABST
    Figure CN118333167B_ABST
Patent Text Reader

Abstract

The present invention discloses a calculation method, a calculation system and a storage medium for an attention mechanism. The calculation method includes a forward propagation stage and a backward propagation stage; in the forward propagation stage, operators for QKV mapping, updating KV cache and rotary position encoding are fused, and the internal calculation process of the fused operator is adjusted to reduce the time overhead of memory access; in the backward propagation stage, the calculation order and the memory access order are adjusted to improve the efficiency of the backward propagation process. By optimizing the forward propagation and backward propagation processes of attention calculation, and respectively adopting the strategies of operator fusion and calculation process rearrangement, this method reduces the number of memory accesses and the amount of memory access in the calculation, and improves the training and inference performance of the large language model by accelerating the attention calculation process in the large language model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and particularly to a method and system for calculating an attention mechanism and a storage medium. Background Art

[0002] The influence of large language models (LLMs) is increasing day by day. They have very good text processing capabilities and can be applied to various scenarios by training or fine-tuning the model parameters, including machine translation, text generation, intelligent assistants, chatbots, etc. However, the number of parameters of a large language model is much larger than that of traditional neural networks, usually reaching the level of billions, resulting in very long training and inference time. How to accelerate the calculation process of large language models is the key. The calculation processes of existing large language models such as Llama2, OPT, and GPT are mainly divided into two stages: prefill and decode. During the prefill stage or the decode stage of long text input, the proportion of attention calculation is very high, often reaching 1 / 3 or more of the overall calculation process. Taking Llama2-7B as an example, the entire network structure consists of 32 Transformer Blocks + normalization layers + linear layers. The Transformer Block mainly includes attention calculation and FFN (Feed-Forward Neural Network) calculation.

[0003] The existing attention calculation acceleration and optimization technologies are mainly flash attention. Through a series of optimization schemes, such as using segmented softmax so as not to use all input data, and not storing the intermediate attention matrix during backpropagation, the time for reading and writing the attention matrix from HBM (High Bandwidth Memory) is fully shortened. By accelerating a single operator in the attention mechanism, the overall process can also be accelerated, such as GEMM (General Matrix Multiplication), GEMV (General Matrix Vector Multiplication), etc.

[0004] Multiplication), GEMV (General Matrix Vector

[0005] Multiplication), etc.

[0006] The following problems still exist in the acceleration and optimization of the attention mechanism in existing large language models:

[0007] Most of the existing attention mechanism optimization techniques are deployed on GPUs with relatively low computational memory access (e.g., GPUs from NVIDIA), and there is no optimization for GPUs with relatively high computational memory access. For example: The NVIDIA A100 GPU has a fp32 peak computing power of 19.5 TFLOPS (one trillion (=10^12) floating-point operations per second), a bandwidth of 1935 GB / s, and a computational memory access ratio of 10.32. While the AMD MI210 has a peak computing power of 22.6 TFLOPS, but its bandwidth is only 1638 GB / s, and the computational memory access ratio is as high as 14.13. If the optimization techniques applicable to GPUs with relatively low computational memory access are applied to AMD GPUs or other manufacturers' GPUs with relatively high computational memory access, there will be a situation where the acceleration effect is very small or even negative optimization occurs. Data shows that the utilization rate of flash attention calculation on the NVIDIA A100 GPU is about 50%, but the utilization rate on the AMD MI210 GPU is only 30%. After analysis, the memory access volume is relatively large during the forward and backward propagation processes of the attention mechanism, and the number of memory accesses is large. A large storage space is required to import various data into registers and shared memory for calculation. The existing technologies do not consider the problem of increased memory access overhead caused by multiple memory accesses for GPUs with relatively slow memory access speeds.

[0008] The main optimization process of existing technologies such as flash attention is the softmax(Q*K) process after Q (query), K (key), and V (value) mappings. However, most of the entire attention calculation has not been optimized. Especially after the new LLaMa model was launched, in addition to QKV mappings, the updates of RoPE (Rotary Position Encoding) and KV cache (KV cache) cannot be ignored in the attention calculation. Summary of the Invention

[0009] In view of the above-mentioned defects of the existing technologies, the present invention provides an attention mechanism calculation method, a calculation system, and a storage medium. By optimizing the forward propagation stage and the backward propagation stage of the attention mechanism, the utilization rate of the computing unit is improved. To achieve the above objectives, the technical solutions of the present invention include:

[0010] An attention mechanism calculation method, which includes a forward propagation stage and a backward propagation stage;

[0011] In the forward propagation stage, the operators of QKV mapping, updating KV cache, and rotary position encoding are fused, and the internal calculation process of the fused operator is adjusted to reduce the time overhead of memory access;

[0012] In the backward propagation stage, the calculation order and memory access order are adjusted to improve the efficiency of the backward propagation process.

[0013] A further improvement of the present invention lies in that: the internal calculation process of the fused operator includes three stages:

[0014] The first stage: perform K mapping and V mapping on the input matrix X to obtain matrix K and matrix V;

[0015] The second stage: perform Q mapping on the input matrix X to obtain matrix Q; and update the KV cache using matrix K and matrix V;

[0016] The third stage: perform rotational position encoding on matrix Q and matrix K.

[0017] A further improvement of the present invention lies in that: in the first stage and the second stage, the input matrix X is temporarily stored on the chip; in the second stage, matrix K and matrix V are temporarily stored on the chip; in the third stage, matrix Q and matrix K are temporarily stored on the chip.

[0018] A further improvement of the present invention lies in that:

[0019] During the first stage, the mapping matrix W for K mapping K and the mapping matrix W for V mapping V ;

[0020] During the second stage, write and update the KV cache using matrix K and matrix V, and read in the mapping matrix W for Q mapping Q .

[0021] A further improvement of the present invention lies in that: the on-chip storage space includes shared memory.

[0022] A further improvement of the present invention lies in that: the output in the forward propagation stage is expressed as:

[0023]

[0024] where V, Q, and K are matrix V, matrix Q after rotational position encoding, and matrix K respectively; d k is the normalization constant.

[0025] A further improvement of the present invention lies in that: the calculations in the backpropagation process include:

[0026] S1: Import lse data for softmax calculation and calculate S = Q * K T ; where lse is the denominator for softmax calculation;

[0027] S2: Perform softmax calculation using S and lse data;

[0028] S3: Calculate dP = dY × V T ;

[0029] S4: Calculate dV = P T × dY, and import Y dY;

[0030] S5: Calculate dS = P × (dP - YdY);

[0031] S6: Calculate dQ = dS × K;

[0032] S7: Calculate dK = dS T × Q;

[0033] S8: Calculate atomic addition of dQ; Atomic addition of dQ means accumulating the partial results obtained from dQ.

[0034] The present invention also provides a computing system, which includes:

[0035] A memory for storing computer-executable instructions; and

[0036] A processor for executing the computer-executable instructions to implement the above-mentioned attention mechanism calculation method.

[0037] A further improvement of the present invention lies in that: the processor includes a graphics processing unit (GPU), a neural network processing unit (NPU), and a field-programmable gate array (FPGA) chip.

[0038] The present invention also provides a computer-readable storage medium storing executable instructions, and when a computer executes the executable instructions, it can implement the above-mentioned attention mechanism calculation method.

[0039] The technical solution provided by the present invention has the following technical effects: By optimizing the forward and backward propagation processes of attention calculation, and respectively adopting the strategies of operator fusion and calculation process rearrangement, the number of memory accesses and the amount of memory access in the calculation are reduced, and by accelerating the attention calculation process in the large language model, the training and inference performance of the large language model are improved.

[0040] The following will further illustrate the concept, specific structure and technical effects generated by the present invention with reference to the accompanying drawings, so as to fully understand the purpose, features and effects of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 Shows a schematic diagram of the forward propagation calculation process of the attention mechanism in the prior art;

[0042] Figure 2 Shows a schematic diagram of the forward propagation calculation process of the optimized attention mechanism in the present invention;

[0043] Figure 3 The figure shows a schematic diagram of the internal calculation process of the fusion operator in the present invention;

[0044] Figure 4 The figure shows a schematic diagram of the calculation order of the backpropagation process in the prior art and the present invention. Detailed implementation manners

[0045] The following uses specific examples to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0046] It should be noted that the diagrams provided in the following embodiments only schematically illustrate the basic concept of the present invention. Therefore, only the components related to the present invention are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The types, numbers, and proportions of the components in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.

[0047] For the purpose of illustration, some exemplary embodiments of the present invention are described. It should be understood that the present invention can be implemented in other ways not specifically shown in the drawings.

[0048] The embodiment of the present invention provides an attention mechanism calculation method, which improves the utilization rate on computing units with a high computational memory access ratio by optimizing the forward propagation and backward propagation processes of the attention mechanism. In this embodiment, taking the attention mechanism in the LLaMa2 model as an example for optimization, its calculation formula:

[0049]

[0050] Among them, the matrices Q, K, and V are all mapped from the input matrix X, that is, Q = W Q X, K = W K X, V = W V X. Actually, three linear transformations are performed to map X from the k1 dimension to the k2 dimension, where W Q , W K , are all mapping matrices, d k is a normalization constant related to the dimension of the matrix K. On this basis, LLaMa2 also adds the RoPE (Rotary Position Encoding) and the process of updating the KV cache, such asFigure 1 As shown in the figure. Among them, RoPE is a calculation that performs element-wise multiplication of fixed parameters on matrices Q and K, that is, the input and output sizes are the same. Updating the KV cache actually stacks the mapped matrices K and V with the previously generated matrices K and V (placed in the KV cache). Figure 1 The figure shows a schematic diagram of the calculation method of the unoptimized attention mechanism in the prior art.

[0051] As Figure 2 As shown in the figure, in this embodiment, in the forward propagation stage, the operators of QKV mapping, updating the KV cache, and rotary position encoding are fused to reduce the number of memory accesses and the memory access size; and the internal calculation process of the fused operator is adjusted to reduce the time overhead of memory access.

[0052] As Figure 1 As shown in the figure, in the existing forward propagation stage, matrix K undergoes three calculations, while matrices Q and V both undergo two calculations. This is mainly because obtaining matrix K needs to participate in the update of the KV cache and also perform RoPE calculation. If the three steps of mapping, updating the KV cache, and RoPE are executed sequentially, there is no calculation step in the stage of updating the KV cache, but instead a large number of memory access operations, resulting in waste of computing resources. Therefore, the internal calculation process is optimized within the operator as Figure 3 shown in the figure.

[0053] The internal calculation process of the fused operator includes three stages:

[0054] The first stage: Perform K mapping and V mapping on the input matrix X to obtain matrix K and matrix V;

[0055] The second stage: Perform Q mapping on the input matrix X to obtain matrix Q; and update the KV cache using matrices K and V;

[0056] The third stage: Perform rotary position encoding on matrices Q and K.

[0057] Operator fusion reduces the number of memory accesses and the memory access size: The operations of mapping matrix X to matrices Q, K, and V are matrix multiplications of the same size (GEMM). The two RoPE operations and the two operations of updating the KV cache have the same computational process. Therefore, they can be fused into one operator through operator fusion, thus avoiding the overhead of multiple memory accesses. If calculated separately, the total memory access volume is 3×(k1×k2 + n×k1 + n×k2) + 2×(n×k2×2) + 2×(n×k2×2) = (3k1k2 + 3nk1 + 11nk2), while if it is a fused operator, the total memory access volume is 3×(k1×k2 + n×k2) + n×k1 + 2×(n×k2) + 2×(n×k2) = (3k1k2 + nk1 + 7nk2). In addition, it also reduces the overhead of multiple operator calls and effectively hides part of the computational latency. The fused operator is as Figure 2 shown.

[0058] By optimizing the computational process within the operator, it is possible to achieve load balancing between computation and memory access. Through the above two optimizations, it is possible to mask the time overhead brought by memory access and reduce the impact of the slow memory access speed of GPUs with relatively small computational memory access (such as some AMD GPU chips) on the forward propagation calculation speed.

[0059] In the backpropagation stage, the computational order and the memory access order are adjusted to improve the efficiency of the backpropagation process. The computations in the backpropagation process include:

[0060] S1: Import lse data for softmax and calculate S = Q * K T ; where lse is the denominator used for softmax calculation, obtained by applying the log-sum-exp function to S, and the lse data is stored for reuse;

[0061] S2: Use S and lse data for softmax calculation, where P = softmax(QK T ) = exp(S - lse);

[0062] S3: Calculate dP = dY × V T , where Y = PV, Y is the output of the attention mechanism and can also be regarded as the input in the backpropagation stage;

[0063] S4: Calculate dV = P T × dY, and import YdY;

[0064] S5: Calculate dS = P × (dP - YdY);

[0065] S6: Calculate dQ = dS × K;

[0066] S7: Calculate dK = dST ×Q;

[0067] S8: Calculate the atomic addition of dQ; the atomic addition of dQ means accumulating the partial results obtained from dQ.

[0068] The backpropagation process is optimized in the present invention, and the backpropagation process based on flash attention2 is improved. By analyzing the dependency relationships between various steps in detail, after calculating dP = V T Calculating dS = P × (dP - YdY) immediately after calculating × dY will cause repeated calls to dP, and performing a high-synchronization-overhead operation such as the atomic addition of dQ immediately after calculating dQ = dS × K will greatly delay the subsequent calculation of dK.

[0069] An embodiment of the present invention further provides a computing system, which includes:

[0070] A memory for storing computer-executable instructions; and

[0071] A processor for executing the computer-executable instructions to implement the above-mentioned attention mechanism calculation method.

[0072] A further improvement of the present invention lies in that: the processor includes a GPU, an NPU, and an FPGA chip.

[0073] An embodiment of the present invention further provides a computer-readable storage medium storing executable instructions, and when a computer executes the executable instructions, it can implement the above-mentioned attention mechanism calculation method.

[0074] During the optimization process, the following two key strategies are adopted in this embodiment:

[0075] 1. Remove the dependency relationship: According to the dependencies between the various steps of backpropagation, the calculation order is adjusted to reduce the possibility of memory access conflicts. Such optimization can effectively improve the calculation efficiency, avoid performance bottlenecks caused by continuously using the same block of space, and at the same time alleviate the high synchronization overhead caused by the mutual waiting between threads when multiple threads execute in parallel.

[0076] 2. Alternate parallel calculation and memory access: A staggering mechanism is introduced in the backpropagation calculation process, that is, try to ensure that only the corresponding data is imported in advance before the calculation. Such an operation can avoid the frequent swapping-out problem caused by the simultaneous storage in registers and shared memory, and at the same time can cleverly hide the latency between calculation and memory access. By staggering the calculation and memory access, the parallel computing power of the GPU is more effectively utilized, and the performance of the overall backpropagation process is improved.

[0077] Experiments show that this solution is an improvement based on flash attention, accelerating the forward propagation, i.e., the inference phase, by approximately 30%, and also accelerating the backward propagation by 10%.

[0078] The above embodiments are merely illustrative of the principles and effects of the present invention and are not intended to limit the present invention. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or changes made by those with ordinary knowledge in the technical field without departing from the spirit and technical ideas disclosed by the present invention should still be covered by the claims of the present invention.

Claims

1. A method for calculating an attention mechanism, characterized in that: Including forward propagation stage and back propagation stage; In the forward propagation phase, the operators of QKV mapping, updating KV cache, and rotating position encoding are fused, and the internal calculation process of the fused operators is adjusted to reduce the time overhead of memory access; In the back-propagation phase, adjust the calculation order and memory access order to improve the efficiency of the back-propagation process; The internal calculation process of the fused operator includes three stages: The first stage: perform K mapping and V mapping on the input matrix X to obtain the matrix K and matrix V; The second stage: perform Q mapping on the input matrix X to obtain the matrix Q; And use the matrix K and matrix V to update the KV cache; The third stage: perform rotation position encoding on the matrix Q and the matrix K; The calculations during backpropagation include: S1: Import lse data for softmax calculation and calculate ; where lse is the denominator used for softmax calculation; S2: Use S and lse data to perform softmax calculation; S3: Compute ;in: , is the output of the attention mechanism; S4: Calculation , and import ; S5: Calculation ; S6: Calculation ; S7: Calculation ; S8: Calculation Atomic Add; Atomic addition means adding The partial results obtained are accumulated; In the first and second stages, the input matrix X is temporarily stored in the on-chip storage space; in the second stage, the matrix K and the matrix V are temporarily stored in the on-chip storage space; in the third stage, the matrix Q and the matrix K are temporarily stored in the on-chip storage space; In the first stage, the mapping matrix for K mapping is read in and the mapping matrix for V mapping ; In the second phase, the KV cache is updated using the matrix K and the matrix V, and the mapping matrix for Q mapping is read in. .

2. The attention mechanism calculation method according to claim 1, characterized in that: On-chip storage space includes shared memory.

3. The attention mechanism calculation method according to claim 1 or 2, characterized in that: The output of the forward propagation phase is expressed as: ; in, , , The matrices are , the matrix Q and matrix K after rotation position encoding; is the normalization constant.

4. A computing system, characterized in that: The computing system comprises: a memory for storing computer-executable instructions; and A processor, wherein the processor is used to execute the computer executable instructions to implement the attention mechanism calculation method according to any one of claims 1-3.

5. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores executable instructions, and when a computer executes the executable instructions, it can implement the attention mechanism calculation method according to any one of claims 1-3.

Citation Information

Patent Citations

  • Multi-head attention mechanism fusion calculation distribution method based on acceleration processor

    CN116431562A

Cited By

  • Method and system for reasoning and calculating graph attention mechanism

    CN122334470A