Transform accelerator-oriented multi-stage dynamic sparse optimization method
Through the multi-stage dynamic sparse optimization method, a mask matrix is generated for each computing stage of the Transformer model, which solves the problem of high energy consumption in the inference stage of the Transformer accelerator, and realizes the reduction of energy consumption and the improvement of computing efficiency.
Patent Information
- Application Number
- CN202510217775.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing Transformer accelerators have high energy consumption in the inference stage, especially in scenarios where resource constraints or real-time requirements are high, it is difficult to meet the needs of computing resources and energy.
The multi-stage dynamic sparse optimization method is adopted to independently generate mask matrix for the Q, K, V generation stage, Attention calculation stage and FFN stage of the Transformer model, and guide matrix operations to reduce unnecessary calculations. The specific methods include approximate calculation based on fixed quantization values and sampling-based threshold generation method.
By reducing the computational cost required for mask generation and optimizing matrix operations, the energy consumption in the inference stage of the Transformer model is significantly reduced and the computing efficiency is improved. It is suitable for scenarios with limited resource or high real-time requirements.
Smart Images

Figure CN120146105A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a multi-stage dynamic sparse optimization method, and particularly to a multi-stage dynamic sparse optimization method for a Transformer accelerator, belonging to the technical field of accelerator optimization of deep neural networks. Background Art
[0002] In recent years, various deep learning large models based on the Transformer architecture, such as the GPT series of OpenAI and the BERT model of Google, have shone brightly in various applications. Not only in various traditional machine learning tasks such as image classification and speech recognition, by using pre-training and fine-tuning, they have repeatedly broken records in various benchmark tests, promoting the frontier development of various fields, but also with their powerful learning and understanding capabilities, tasks such as autonomous driving, artificial intelligence (AI) dialogue, and AI creation have more development prospects. It can be predicted that the Transformer model will reshape the digital processes of all walks of life, thus having a profound impact on all aspects of society; however, while achieving high performance, various models based on Transformer have also brought an exponential increase in the number of parameters and a doubling of computational complexity. Taking the ChatGPT model as an example, the stable operation of the ChatGPT model is supported by more than 30,000 top-level AI computing power chips, and its energy consumption is relatively high, which is sufficient to illustrate the demand of the ChatGPT model for computing resources and energy. To meet the needs of model inference, the hardware platform needs to have conditions such as a large computing throughput, rich storage resources, and high data bandwidth. However, in some cases where resources are limited (such as embedded devices and mobile terminals) or the real-time requirements of tasks are relatively high, these conditions cannot be met. Therefore, it is urgent to optimize the Transformer accelerator to reduce the energy consumption required in its inference stage.
[0003] To solve the energy consumption problem in the inference stage of the Transformer accelerator, the field has begun to turn its attention to approximate computing technology to alleviate the demand of the Transformer for computing resources. Approximate computing is a technology that, on the premise that the accuracy of the application can be accepted, actively reduces a part of the computing accuracy in exchange for lower system energy consumption or higher operating efficiency; approximate computing technology can be applied to each module of the Transformer model. On the premise that the accuracy degradation of the model is within an acceptable range, it can significantly reduce the computing resources occupied by the system or reduce the system energy consumption. At the same time, the core idea of some approximate computing technologies is to reduce some calculations that have relatively little impact on the results, and it can also directly reduce the time required for system operation.
[0004] Multi-Head Self-Attention (MHSA) is the main bottleneck restricting the further optimization of Transformer accelerators. Its optimization methods can be mainly divided into sparse attention and linear attention. Linear attention converts the calculation of attention with complexity O(n 2 ) into a linear calculation through a kernel function, but it also introduces other calculations and has limited optimization effect when the length n of the input data is small. Sparse attention achieves the optimization effect by predicting the positions of the calculations that can be sparsified in the Attention calculation process and skipping the calculations of those positions. According to whether its prediction method is static or dynamic, it is divided into static sparse attention and dynamic sparse attention. The static sparse attention method is very simple and efficient, but it has a greater impact on accuracy due to its excessive dependence on data distribution, while dynamic sparse attention can balance accuracy and optimization effect by accurately identifying the positions of unimportant calculations.
[0005] In addition to optimizing the attention in the multi-head attention model, more and more researchers have realized that the FFN (feed-forward neural network) part also occupies a large part of the total computational amount. And compared with only optimizing a certain module, better optimization effect can be achieved by optimizing the three main computational parts in the multi-head attention model: the generation stage of Q, K, V (query, key, and value), the calculation stage of Attention, and the FFN stage. However, the existing research uses a multi-stage dynamic sparse optimization method based on derivation, and the mask matrices in the generation stage of QKV and the FFN stage are derived from the mask in the Attention calculation part, thus failing to fully utilize the redundancy in the linear layer and limiting the overall optimization effect.
[0006] In summary, a multi-stage dynamic sparse optimization method for Transformer accelerators is needed. Summary of the Invention
[0007] A brief overview of the present invention is given below to provide a basic understanding of certain aspects of the present invention. It should be understood that this overview is not an exhaustive overview of the present invention. It is not intended to identify the key or important parts of the present invention, nor is it intended to limit the scope of the present invention. Its purpose is merely to present certain concepts in a simplified form as a prelude to the more detailed description to follow.
[0008] In view of this, to solve the problems of low efficiency and poor effect in the traditional acceleration optimization method in the inference stage of Transformer accelerators in the prior art, the present invention provides a multi-stage dynamic sparse optimization method for Transformer accelerators.
[0009] The technical solution is as follows: A multi-stage dynamic sparsity optimization method for Transformer accelerators, including the following steps:
[0010] S1. Determine the input matrix, and perform independent dynamic sparsity optimization methods on the generation stages of Q, K, and V, the calculation stage of Attention, and the FFN stage in the inference stage of the Transformer model, that is, generate a mask matrix for guiding matrix multiplication to guide the matrix operations in each stage;
[0011] Specifically:
[0012] S11. Use an approximate calculation method based on fixed quantization values to obtain the quantized input matrix, and then obtain an approximate value of the matrix operation result, that is, the approximate result;
[0013] S12. According to the approximate result, use a sampling-based threshold generation method to obtain a mask matrix for guiding matrix multiplication;
[0014] S2. According to the generation process of the mask matrix for guiding matrix multiplication, adjust the data flow in the Transformer model through an early mask generation method, generate the mask matrices required for each stage in advance, and accelerate the matrix operation through the pre-generated mask matrices to obtain the matrix operation result.
[0015] Further, in S11, the independent dynamic sparsity optimization method is expressed as: after determining the input matrix, obtain an approximate value of the matrix operation result of the input matrix, that is, the approximate result, through approximate calculation technology, and mark the mask matrix for guiding matrix multiplication by determining the positions of the set elements in the approximate result to obtain the mask matrix for guiding matrix multiplication;
[0016] In the matrix operation, for the attention score QK in the calculation stage of attention Attention T , when the mask matrix mask ij for guiding matrix multiplication is 1, use exact multiplication operation for calculation, and when the mask matrix mask ij for guiding matrix multiplication is 0, skip the calculation of the corresponding position values;
[0017] For the matrix operations in the generation stages of Q, K, and V and the FFN stage, when the mask matrix mask ij for guiding matrix multiplication is 1, multiply the input matrix by the weight retaining the highest two significant digits, and when the mask matrix mask ij for guiding matrix multiplication is 0, multiply the input matrix by the weight retaining the highest one significant digit.
[0018] Further, in S11, during the process of calculating the approximate result, the first threshold Thre1 is set to 3×θ, the second threshold Thre2 is set to 2×θ, where θ is the average value of the input matrix. The input matrix is quantized, that is, the input with an absolute value greater than the first threshold Thre1 is quantized to 4, the input greater than the second threshold Thre2 but less than the first threshold Thre1 is quantized to 2, and the input less than the second threshold Thre2 is quantized to 1. Then, while keeping its positive and negative signs unchanged, the quantized input matrix is obtained.
[0019] According to the quantized input matrix, matrix multiplication operations are performed through logical operations to obtain the approximate result.
[0020] Further, in S12, according to the quantized input matrix, the first input matrix A is sampled at intervals of 8 to obtain a vector set {A 0 , A 8 , A 16 …}. The dot product is taken between the vector set and the elements B j of the second input matrix B to obtain the minimum value sam_min and the average value sam_mean of the dot product results corresponding to the elements B j of the second input matrix B. And by calculating the value of (sam_mean - sam_min)×k + sam_min, the third threshold Thre3 is obtained, where k is the approximate degree evaluation value.
[0021] All elements of the first input matrix A and the second input matrix B are traversed for multiplication calculation to obtain the multiplication calculation result S ij , S ij = A i × B j , where A i is the element of the first input matrix A. The multiplication calculation result S ij is compared with the third threshold Thre3. If the multiplication calculation result S ij is greater than the third threshold Thre3, the mask matrix Mask ij corresponding to its position for guiding matrix multiplication is marked as 1, otherwise it is marked as 0.
[0022] Further, in S2, for the generation stage of Q, K, and V, the mask matrices of the generated query Q and value V, that is, the mask matrices for guiding matrix multiplication in the generation stage of Q, K, and V, are used to guide the matrix operations to generate the mask matrices mask Q and mask KMeanwhile, the matrix operation results of the input matrices of query Q and key K obtained by the approximate calculation method, i.e., the approximate results, are compared with the set new threshold, and the quantized input matrices of query Q and key K are generated in advance. When generating query Q and key K through the mask matrices of query Q and key K, the quantized input matrices of query Q and key K generated in advance are used to generate the guiding attention score QK in advance T The mask matrix mask for operation qk ;
[0023] For the Attention part in the calculation stage of Attention, according to the mask matrices of query Q and key K and the guiding attention score QK T The mask matrix for operation directly performs the attention score QK operation T For the linear projection part in the calculation stage of Attention, according to the process of generating the mask matrix for the guiding attention score QK operation in the generation stage of Q, K, and V, the quantized input matrix of the residual connection in the FFN stage is generated in advance; T
[0024] For the FFN stage, while performing the residual connection and Norm() function calculations, the quantized input matrix of the residual connection obtained in the linear projection part is added to the corresponding quantized matrix of the input matrix of the residual connection to generate the quantized input matrix for the subsequent Fc1 operation in advance, and the mask matrix mask for guiding the Fc1 operation is generated in advance according to the quantized input matrix in the Fc1 operation fc1 During the process of generating the mask matrix mask for guiding the Fc1 operation fc1 The quantized input matrix required for the Gelu() function calculation is generated in advance, and during the Gelu() function calculation, all negative numbers in the quantized input matrix required for the Gelu() function calculation are converted to -1, positive numbers remain unchanged, and the converted quantized input matrix Output′ required for the Gelu() function calculation fc1 is multiplied by the weight matrix weight in the quantized Fc2 operation to generate the mask matrix mask for guiding the Fc2 operation in advance fc2 and input into the Fc2 operation to complete the matrix operation acceleration and obtain the matrix operation result
[0025] The beneficial effects of the present invention are as follows: By generating respective mask matrices independently in three main calculation stages of the Transformer model to guide its matrix operations, the present invention can fully explore the approximable space in each stage of the Transformer model; the present invention reduces the computational cost required for mask generation. On the one hand, by converting the multiplication operations required for generating approximate results into logical operations of selecting from a finite result set, the computational amount required for generating approximate results is greatly reduced. On the other hand, through the sampling-based threshold generation method and by comparing the generated mask matrices, the hardware-unfriendly TopK() operation is avoided, and the mask matrix can be obtained by traversing only once; while generating the mask matrix of the matrix operation result, the present invention generates the quantization matrix of the matrix operation result, and then generates the mask matrix for the next matrix operation participated by the result matrix in advance, so that the next matrix operation does not need to wait for the generation of its mask matrix, but directly uses the pre-generated mask matrix for operation, accelerating the inference stage of the Transformer model. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:
[0027] Figure 1 is a schematic flow chart of a multi-stage dynamic sparsity optimization method for a Transformer accelerator;
[0028] Figure 2 is a schematic flow chart of an independent dynamic sparsity optimization method;
[0029] Figure 3 is a schematic flow chart of a sampling-based threshold generation method;
[0030] Figure 4 is a schematic flow chart of the pre-generation of the mask matrix in step S2. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0031] In order to make the technical solutions and advantages in the embodiments of the present invention clearer and more understandable, the following further details the exemplary embodiments of the present invention with reference to the drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than an exhaustive list of all embodiments. It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0032] Refer to Figures 1-4Detailed description of this embodiment, a multi-stage dynamic sparse optimization method for Transformer accelerators, specifically including the following steps:
[0033] S1. Determine the input matrix, and perform independent dynamic sparse optimization methods on the generation stages of Q, K, and V, the calculation stage of Attention, and the FFN stage in the inference stage of the Transformer model, that is, generate a mask matrix for guiding matrix multiplication to guide the matrix operations in each stage;
[0034] Specifically:
[0035] S11. Use an approximate calculation method based on fixed quantization values to obtain the quantized input matrix, and then obtain an approximate value of the matrix operation result, that is, the approximate result;
[0036] S12. According to the approximate result, use a sampling-based threshold generation method to obtain a mask matrix for guiding matrix multiplication;
[0037] S2. According to the generation process of the mask matrix for guiding matrix multiplication, adjust the data stream in the Transformer model through an early mask generation method, generate the mask matrices required for each stage in advance, and accelerate the matrix operation through the pre-generated mask matrices to obtain the matrix operation result.
[0038] Furthermore, in S11, the independent dynamic sparse optimization method is expressed as: after determining the input matrix, obtain an approximate value of the matrix operation result of the input matrix, that is, the approximate result, through approximate calculation techniques, and mark the mask matrix for guiding matrix multiplication by determining the positions of the set (smaller value) elements in the approximate result, and mark the corresponding positions in the all-1 matrix of the same scale as 0 to obtain the mask matrix for guiding matrix multiplication;
[0039] In the matrix operation, for the attention score QK in the calculation stage of attention Attention T , when the mask matrix mask ij for guiding matrix multiplication is 1, perform the calculation using exact multiplication operations. When the mask matrix mask ij for guiding matrix multiplication is 0, skip the calculation of the corresponding position values; for the matrix operations in the generation stages of Q, K, and V and the FFN stage, when the mask matrix mask ij for guiding matrix multiplication is 1, multiply the input matrix by the weight retaining the highest two significant digits. When the mask matrix mask ij for guiding matrix multiplication is 0, multiply the input matrix by the weight retaining the highest one significant digit.
[0040] Specifically, in this embodiment, an independent dynamic sparse optimization method is implemented for the three calculation processes in the inference stage of the Transformer model. After generating the mask matrix, it can be used to guide matrix operations. For the calculation of the values at the positions marked as 1 in the mask matrix, a more accurate operation method is adopted, while for the calculation of the values at the positions marked as 0 in the mask matrix, a more approximate operation method is adopted, or the operation result is directly skipped and set to 0.
[0041] Further, in S11, during the process of calculating the approximate result, the first threshold Thre1 is set to 3×θ, the second threshold Thre2 is set to 2×θ, where θ is the average value of the input matrix. The input matrix is quantized by comparison, that is, the input with an absolute value greater than the first threshold Thre1 is quantized to 4, the input greater than the second threshold Thre2 but less than the first threshold Thre1 is quantized to 2, and the input less than the second threshold Thre2 is quantized to 1, and then its positive and negative signs are kept unchanged to obtain the quantized input matrix.
[0042] According to the quantized input matrix, matrix multiplication operations are performed through logical operations to obtain an approximate result.
[0043] Specifically, the first threshold Thre1 and the second threshold Thre2 are obtained through statistics, that is, before optimizing the Transformer model, according to the average value θ of each input matrix obtained from the test data set, the first threshold Thre1 and the second threshold Thre2 are set.
[0044] Since the input matrix is quantized into three fixed numbers, the types of possible values of its multiplication result are limited, which are 16, 8, 4, 2, and 1 respectively. After obtaining the quantized input matrix through comparison, the logical operation of selecting the corresponding result according to the input matrix is used to replace the traditional multiplication operation.
[0045] The process of the logical operation is expressed as:
[0046] y[3] = x1[2] ^ x2[2]
[0047] y[2] = x1[1] & (x1[0] | x2[0]) & x2[2]
[0048] y[1] = (x1[1] & (∼x2[1]) & x2[0]) | ((∼x1[1]) & x1[1] & x2[1]) | (x1[1] & (∼x1[0]) & x2[1] & (∼x2[0]))
[0049] y[0] = (x1[0] & x2[0]) | (x1[1] & (∼x1[0]) & x2[1] & (∼x2[0]))
[0050] Among them, y represents the output, the range of the output is 4 bits, x1 represents the first input, x2 represents the second input, and the range of the input is 3 bits.
[0051] Further, in S12, according to the quantized input matrix, the first input matrix A is sampled at an interval of 8 to obtain a vector set {A 0 , A 8 , A 16 …}. The dot product is performed between the vector set and the element B j of the second input matrix B to obtain the minimum value sam_min of the dot product results corresponding to the element B j of the second input matrix B and the average value sam_mean of the dot product results. The value of (sam_mean - sam_min) × k + sam_min is calculated to obtain the third threshold Thre3. Among them, k is an approximation degree evaluation value, and the value of k is related to the approximation degree of the Transformer model. The larger the value of k, the higher the approximation degree of the Transformer model, and the smaller the value of k, the lower the approximation degree of the model;
[0052] All elements of the first input matrix A and the second input matrix B are traversed for multiplication calculation to obtain the multiplication calculation result S ij , S ij = A i × B j , where A i is the element of the first input matrix A. The multiplication calculation result S ij is compared with the third threshold Thre3. If the multiplication calculation result S ij is greater than the third threshold Thre3, the mask matrix Mask ij used to guide matrix multiplication at its corresponding position is marked as 1, otherwise it is marked as 0.
[0053] Further, in S2, for the generation stages of Q, K, and V, the mask matrices of the generated query Q and value V, that is, the mask matrices used to guide matrix multiplication in the generation stages of Q, K, and V, are used to guide matrix operations. While generating the mask matrices mask Q and mask K of the query Q and key K, the matrix operation results of the input matrices of the query Q and key K obtained by the approximate calculation method, that is, the approximate results, are compared with a set new threshold to generate the quantized input matrices of the query Q and key K in advance. While generating the query Q and key K through the mask matrices of the query Q and key K, the mask matrix mask T used to guide the attention score QK qk operation is generated in advance;
[0054] For the attention part in the calculation stage of attention, according to the mask matrices of the query Q and the key K and the guiding attention score QK T Perform the attention score QK operation directly on the mask matrix of the operation T For the linear projection part in the calculation stage of attention, according to the process of generating the mask matrix of the guiding attention score QK in the generation stages of Q, K, and V T Generate the quantized input matrix of the residual connection in the FFN stage in advance;
[0055] For the FFN stage, while performing the residual connection and the Norm() function calculation, add the quantized input matrix of the residual connection obtained in the linear projection part to the corresponding quantized matrix of the input matrix of the residual connection to generate the quantized input matrix for the subsequent Fc1 operation in advance, and generate the mask matrix mask for guiding the Fc1 operation in advance according to the quantized input matrix in the Fc1 operation fc1 During the process of generating the mask matrix mask for guiding the Fc1 operation fc1 Generate the quantized input matrix required for the Gelu() function calculation in advance, and during the Gelu() function calculation, convert all negative numbers in the quantized input matrix required for the Gelu() function calculation to -1, keep the positive numbers unchanged, and obtain the converted quantized input matrix Output′ for the Gelu() function calculation fc1 Multiply it by the weight matrix weight in the quantized Fc2 operation to generate the mask matrix mask for guiding the Fc2 operation in advance fc2 And input it into the Fc2 operation to complete the matrix operation acceleration and obtain the matrix operation result
[0056] Specifically, step S2 solves the time overhead introduced by the inability to parallelize the mask generation with the matrix operations before and after due to data dependencies;
[0057] Due to the mask matrix of the attention score QK T operation being generated in advance, therefore, in the calculation stage of attention, there is no need to perform the mask matrix operation anymore;
[0058] In the generation stages of Q, K, and V, the new thresholds set are obtained through statistics in advance. After obtaining the average value θ of the approximate result through sampling, multiply it by α and β respectively, and set the first threshold Thre1 and the second threshold Thre2 for comparing quantization to αθ and βθ respectively. When taking statistics, take the value of α as when the data greater than αθ accounts for 10% of the total data, and take the value of β as when the data greater than βθ accounts for 20% of the total data;
[0059] Reference Figure 4 ,θ 1 represents a limit value for taking the value of α, θ 2 represents a limit value for taking the value of β, Q′ ij represents the element in the i-th row and j-th column of the matrix operation result Q′ of the input matrix of query Q, X i ′W′ Qj represents the input matrix X in the generation stage of the quantized Q, K, V ′ where the i-th row of Q is multiplied by the quantized input matrix W′ of the query Q generated in advance
[0060] In the pre-generation process of the Attention part, X represents the input matrix in the generation stage of Q, K, V, X′ represents the input matrix in the generation stage of the quantized Q, K, V, W Q and W K respectively represent the input matrices for generating the query Q and the value V, W′ Q and W′ K respectively represent the quantized input matrices of the query Q and the key K generated in advance, Q′ and K′ respectively represent the matrix operation results of the input matrices of the query Q and the key K, QK represents the masked matrix mask T for the operation of the guided attention score QK qk the matrix operation results of the guided query Q and key K;
[0061] In the pre-generation process in the linear layer, proj represents the linear projection part, Output′ proj represents the output matrix of the linear projection part, that is, the quantized input matrix in the residual connection in the FFN stage, Y represents the input matrix of the residual connection, Y′ represents the quantized matrix corresponding to the input matrix of the residual connection, Fc1 represents the first fully connected layer, and Fc2 represents the second fully connected layer.
[0062] Although the present invention has been described with reference to a limited number of embodiments, those skilled in the art in this technical field will understand that, within the scope of the present invention thus described, other embodiments can be envisaged. In addition, it should be noted that the language used in this specification is mainly selected for readability and teaching purposes, rather than for the purpose of explaining or limiting the subject matter of the present invention. Therefore, many modifications and changes are obvious to those of ordinary skill in the art in this technical field without departing from the scope and spirit of the appended claims. For the scope of the present invention, the disclosure of the present invention is illustrative rather than restrictive, and the scope of the present invention is defined by the appended claims.
Claims
1. A multi-stage dynamic sparse optimization method for Transformer accelerator, characterized in that: The following steps are involved: S1. Determine the input matrix and perform independent dynamic sparse optimization methods on the generation phase of Q, K, V, the calculation phase of Attention, and the FFN phase of the Transformer model inference phase, that is, generate a mask matrix for guiding matrix multiplication to guide the matrix operations at each stage; Specific: S11. Using an approximate calculation method based on a fixed quantized value, a quantized input matrix is obtained, and then an approximate value of the matrix operation result, that is, an approximate result, is obtained; S12. According to the approximation result, a sampling-based threshold generation method is used to obtain a mask matrix for guiding matrix multiplication; S2. According to the generation process of the mask matrix used to guide matrix multiplication, the data flow in the Transformer model is adjusted through the advance mask generation method, the mask matrix required for each stage is generated in advance, and the matrix operation is accelerated by the mask matrix generated in advance to obtain the matrix operation result.
2. A multi-stage dynamic sparse optimization method for Transformer accelerator according to claim 1, characterized in that: In the S11, the independent dynamic sparse optimization method is expressed as follows: after determining the input matrix, an approximate value of the matrix operation result of the input matrix, i.e., an approximate result, is obtained by using an approximate calculation technique, and a mask matrix for guiding matrix multiplication is marked by determining the position of a set element in the approximate result, to obtain a mask matrix for guiding matrix multiplication; In the matrix operation, the attention score QK in the calculation phase of attention is T , when used as the mask matrix mask to guide matrix multiplication ij When it is 1, the exact multiplication operation is used for calculation. When the mask matrix mask used to guide the matrix multiplication ij When it is 0, skip the calculation of the value at the corresponding position; For the matrix operations in the generation phase of Q, K, V and the FFN phase, the mask matrix mask used to guide the matrix multiplication ij When it is 1, the input matrix is multiplied by the weights that retain the two most significant digits, and the mask matrix mask used to guide the matrix multiplication ij When 0, multiply the input matrix by the weight that retains the most significant digit.
3. A multi-stage dynamic sparse optimization method for Transformer accelerator according to claim 2, characterized in that: In the S11, in the process of calculating the approximate result, the first threshold Thre1 is set to 3×θ, the second threshold Thre2 is set to 2×θ, θ is the average value of the input matrix, and the input matrix is quantized, that is, the input whose absolute value is greater than the first threshold Thre1 is quantized to 4, the input whose absolute value is greater than the second threshold Thre2 but less than the first threshold Thre1 is quantized to 2, and the input whose absolute value is less than the second threshold Thre2 is quantized to 1, and then the positive and negative signs are kept unchanged to obtain the quantized input matrix; According to the quantized input matrix, matrix multiplication operation is performed through logical operation to obtain an approximate result.
4. A multi-stage dynamic sparse optimization method for Transformer accelerator according to claim 3, characterized in that: In S12, the first input matrix A is sampled at intervals of 8 according to the quantized input matrix to obtain a vector set {A0, A8, A 16 ...}, which adds the vector set to the elements of the second input matrix B j Perform dot product to get element B of the second input matrix B j The minimum value sam_in of the corresponding dot product result and the average value sam_mean of the dot product result are calculated, and the third threshold Thre3 is obtained by calculating the value of (sam_mean-sam_min)×k+sam_min, where k is the approximation evaluation value; Traverse all elements of the first input matrix A and the second input matrix B to perform multiplication calculations and obtain the multiplication result S ij , S ij =A i ×B j , where A i The elements of the first input matrix A are multiplied by the result S ij Compared with the third threshold Thre3, if the multiplication result S ij If it is greater than the third threshold Thre3, the mask matrix Mask used to guide matrix multiplication at its corresponding position ij Marked as 1, otherwise marked as 0.
5. A multi-stage dynamic sparse optimization method for Transformer accelerator according to claim 4, characterized in that: In the S2, for the generation stage of Q, K, V, the generated query Q, value V mask matrix, i.e., the mask matrix used to guide matrix multiplication in the generation stage of Q, K, V, is used to guide matrix operations to generate the query Q, key K mask matrix mask Q and mask K At the same time, the matrix operation result of the input matrix of the query Q and key K obtained by the approximate calculation method, that is, the approximate result, is compared with the set new threshold, and the quantized input matrix of the query Q and key K is generated in advance. While the query Q and key K are generated through the mask matrix of the query Q and key K, the guided attention score QK is generated in advance through the quantized input matrix of the query Q and key K generated in advance. T The mask matrix mask of the operation qk ; For the Attention part in the Attention calculation phase, according to the query Q, the mask matrix of the key K and the guided attention score QK T The mask matrix of the operation is directly used for the attention score QK T Operation, for the linear projection part in the Attention calculation phase, follow the guidance attention score QK in the generation phase of Q, K, V T The mask matrix of the operation is generated in advance, and the quantized input matrix of the residual connection in the FFN stage is generated in advance; For the FFN stage, while performing residual connection and Norm() function calculations, the quantized input matrix of the residual connection obtained in the linear projection part is added to the quantized matrix corresponding to the input matrix of the residual connection, and the quantized input matrix in the subsequent Fc1 operation is generated in advance, and the mask matrix mask used to guide the Fc1 operation is generated in advance according to the quantized input matrix in the Fc1 operation. fc1 , in generating the mask matrix mask used to guide the Fc1 operation fc1 In the process of , the quantized input matrix required for the Gelu() function calculation is generated in advance, and in the process of the Gelu() function calculation, all negative numbers in the quantized input matrix required for the Gelu() function calculation are converted to -1, and the positive numbers remain unchanged, and the converted quantized input matrix required for the Gelu() function calculation is Output' fc1 Multiply it with the weight matrix weight in the quantized Fc2 operation to generate the mask matrix mask used to guide the Fc2 operation in advance fc2 , and input it into the Fc2 operation to complete the matrix operation acceleration and obtain the matrix operation result.
Citation Information
Cited By
Text analysis method and device, computer equipment and storage medium
CN120409496A
Data processing method and device, electronic equipment, computer readable storage medium and computer program product
CN120449950A
Matrix multiplier, chip, device, data processing method, medium and product
CN120994946A
Transform model accelerator for ocean drifting buoy
CN121683911A
A transformer model accelerator for ocean drift buoys
CN121683911B