Method for sparse attention operation

Through dynamic pruning strategies and attention probability-based token selection criteria, the problem of high computational cost of transformer models in large data processing and real-time computing is solved, achieving more efficient computing and better model performance.

CN120179972APending Publication Date: 2025-06-20IND TECH RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311759045.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-20
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The transformer model requires a large amount of computing resources when processing large data and real-time computing, resulting in high computational costs and difficult to maintain model performance.

Method used

Through the dynamic pruning strategy, the pruning ratio is dynamically adjusted according to the characteristics of the input text, the amount of data involved in the calculation is reduced, and a token selection criteria based on attention probability is proposed to better retain the semantic information of the input text.

Benefits of technology

Maintain model performance while reducing computational costs and outperforms traditional static pruning methods in most situations, achieving greater flexibility and adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179972A_ABST
    Figure CN120179972A_ABST
Patent Text Reader

Abstract

The invention relates to a method for sparse attention operation. Comprising the following steps: calculating an attention probability matrix according to a query matrix and a key matrix; calculating a pruning ratio according to the attention probability matrix; pruning the attention probability matrix according to the pruning ratio to obtain a pruning attention probability matrix; pruning the value matrix according to the pruning ratio to obtain a pruning value matrix; and multiplying the pruning attention probability matrix by the pruning value matrix to obtain an attention matrix.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for sparse attention operation. Background Art

[0002] With the increasing applications in natural language processing and other fields, the models using transformers are also gradually increasing. These models require a large amount of computing resources and need high performance when processing large-scale data or performing real-time calculations. Therefore, an algorithm method that can reduce the computing cost while maintaining the model performance has broad application value. Summary of the Invention

[0003] The present invention provides a method for sparse attention operation to reduce the model calculation amount and maintain the model accuracy at the same time.

[0004] According to some embodiments of the present invention, a method for sparse attention operation is provided, including the following steps: calculating an attention probability matrix according to a query matrix and a key matrix; calculating a pruning ratio according to the attention probability matrix; pruning the attention probability matrix according to the pruning ratio to obtain a pruned attention probability matrix; pruning a value matrix according to the pruning ratio to obtain a pruned value matrix; multiplying the pruned attention probability matrix by the pruned value matrix to obtain an attention matrix.

[0005] In some embodiments of the present invention, the query matrix Q, the key matrix K, and the attention probability matrix O have the following relationship:

[0006]

[0007] where d k is the number of rows of the key matrix.

[0008] In some embodiments of the present invention, in the step of calculating the pruning ratio according to the attention probability matrix, it includes: weighting each element of the attention probability matrix to obtain an attention probability weighted matrix; calculating the sum of elements of each column of the attention probability weighted matrix to obtain the column weighted sum of each column of the attention probability weighted matrix; calculating the sum of elements of each row of the attention probability weighted matrix to obtain the row weighted sum of each row of the attention probability weighted matrix; summing up the column weighted sums of each column of the attention probability weighted matrix to obtain the attention information of the attention probability weighted matrix; determining the pruning ratio according to the attention information.

[0009] In some embodiments of the present invention, in the step of weighting each element of the attention probability matrix to obtain the attention probability weighted matrix, the maximum value of the elements of the attention probability matrix is Vmax, the minimum value of the elements of the attention probability matrix is Vmin, the maximum value of the elements after weighting is Vmax', and the minimum value of the elements after weighting is Vmin'. Then, the following relationship holds:

[0010] In some embodiments of the present invention, in the step of weighting each element of the attention probability matrix, it includes weighting each element by a non-linear method.

[0011] In some embodiments of the present invention, in the step of weighting each element of the attention probability matrix, it includes calculating the square value of each element.

[0012] In some embodiments of the present invention, in the step of determining the pruning ratio according to the attention information, it includes normalizing the attention information in an interval to obtain the pruning ratio.

[0013] In some embodiments of the present invention, the normalization includes normalizing by a linear method or by a non-linear method.

[0014] In some embodiments of the present invention, the pruning ratio can be expressed by the following formula:

[0015]

[0016] Wherein, R P is the pruning ratio, α is a given parameter, S I is the attention parameter, and n is the number of tokens in the attention probability matrix.

[0017] In some embodiments of the present invention, in the step of pruning the attention probability matrix according to the pruning ratio, it includes: calculating the row pruning quantity according to the pruning ratio and the number of rows of the attention probability matrix, arranging the row weighted sums of each row of the attention probability matrix in ascending order, and pruning the rows corresponding to the row weighted sums in sequence until the number of pruned rows is equal to the row pruning quantity. Pruning a row means setting the element values of all elements in the row to 0.

[0018] According to some other embodiments of the present invention, a method for sparse attention operation is provided, including the following steps: pruning the query matrix according to a first pruning ratio to obtain a pruned query matrix, where the first pruning ratio is the pruning ratio of the previous pruning layer; pruning the key matrix according to the first pruning ratio to obtain a pruned key matrix; calculating an attention probability matrix according to the pruned query matrix and the pruned key matrix; calculating a second pruning ratio according to the attention probability matrix; pruning the attention probability matrix according to the second pruning ratio to obtain a pruned attention probability matrix; pruning the value matrix according to the second pruning ratio to obtain a pruned value matrix; multiplying the pruned attention probability matrix and the pruned value matrix to obtain an attention matrix.

[0019] In some embodiments of the present invention, the pruned query matrix Q’, the pruned key matrix K’ and the attention probability matrix O have the following relationship:

[0020]

[0021] where d k is the number of rows of the key matrix.

[0022] In some embodiments of the present invention, in the step of calculating the pruning ratio according to the attention probability matrix, it includes: weighting each element of the attention probability matrix to obtain an attention probability weighted matrix; calculating the sum of elements of each column of the attention probability weighted matrix to obtain the column weighted sum of each column of the attention probability weighted matrix; calculating the sum of elements of each row of the attention probability weighted matrix to obtain the row weighted sum of each row of the attention probability weighted matrix; summing up the column weighted sums of each column of the attention probability weighted matrix to obtain the attention information of the attention probability weighted matrix; determining the second pruning ratio according to the attention information.

[0023] In some embodiments of the present invention, in the step of weighting each element of the attention probability matrix to obtain the attention probability weighted matrix, the maximum value of the elements of the attention probability matrix is Vmax, the minimum value of the elements of the attention probability matrix is Vmin, the maximum value of the elements after weighting is Vmax’, and the minimum value of the elements after weighting is Vmin’, then the following relationship exists:

[0024] In some embodiments of the present invention, in the step of weighting each element of the attention probability matrix, it includes weighting each element in a non-linear method.

[0025] In some embodiments of the present invention, in the step of weighting each element of the attention probability matrix, it includes calculating the square value of each element.

[0026] In some embodiments of the present invention, in the step of determining the pruning ratio according to the attention information, it includes normalizing the attention information in an interval to obtain the pruning ratio.

[0027] In some embodiments of the present invention, the normalization includes normalizing in a linear method or in a non-linear method.

[0028] In some embodiments of the present invention, the pruning ratio can be expressed by the following formula:

[0029]

[0030] where, R P is the pruning ratio, α is a given parameter, S I is the attention parameter, and n is the number of tokens in the attention probability matrix.

[0031] In some embodiments of the present invention, in the step of pruning the attention probability matrix according to the pruning ratio, it includes: calculating the row pruning quantity according to the pruning ratio and the number of rows of the attention probability matrix, arranging the row weighted sums of each row of the attention probability matrix in ascending order, pruning the rows corresponding to the row weighted sums in sequence until the number of pruned rows is equal to the row pruning quantity, where pruning the row means setting the element values of all elements of the row to 0.

[0032] Based on the above, the present invention aims to meet the requirements of real-time computing, improve energy efficiency, adapt to the trend of mobile device computing, and promote the development of research. The present invention provides a dynamic pruning strategy that dynamically adjusts the pruning ratio according to the characteristics of the input text to reduce the amount of data participating in the calculation. Compared with the traditional static token pruning method, this method shows higher flexibility and adaptability. Experiments prove that the present invention can maintain the model performance while reducing the computing cost, and is superior to the traditional static pruning method in most scenarios. In addition, the present invention also proposes a token selection criterion based on the element-wise square of the attention probability to better retain the semantic information of the input text. In summary, the present invention provides an efficient dynamic pruning strategy for accelerating and optimizing Transformer-based models, with broad application prospects. Description of the Drawings

[0033] Figure 1 is a flowchart of a method according to some embodiments of the present invention.

[0034] Figure 2 is a flowchart of a method according to some embodiments of the present invention.

[0035] Figure 3A is the attention probability matrix of a method according to the prior art.

[0036] Figure 3B is the attention probability weighted matrix of a method according to some embodiments of the present invention.

[0037] Figure 4A is Figure 3A the result after pruning the attention probability matrix of.

[0038] Figure 4B is Figure 3B the result after pruning the attention probability weighted matrix of.

[0039] Figure 5 is a comparison between a method of the prior art and a method according to some embodiments of the present invention. Detailed implementation manners

[0040] To make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to specific embodiments and the accompanying drawings.

[0041] Figure 1 is a flowchart of a method according to some embodiments of the present invention. Please refer to Figure 1 . The flowchart 100 is a method for sparse attention operation. The method has three input matrices, namely the query matrix 102, also referred to as the Q matrix 102; the key matrix 104 or referred to as the K matrix 104; and the value matrix 120, or referred to as the V matrix 120.

[0042] The query matrix 102, the key matrix 104, and the value matrix 120 are all MxN matrices with M columns and N rows. Among them, the number of columns M is the number n of tokens in the input text, and the number of rows N is the dimension of the model, denoted as d model .

[0043] The method includes the following steps.

[0044] Step 1: Calculate the attention probability matrix 114 according to the query matrix 102 and the key matrix 104.

[0045] Specifically, first multiply the query matrix 102 by the transpose of the key matrix 104 to obtain matrix 112. Matrix 112 can be expressed as QK T , which is an MxM matrix.

[0046] Next, perform a softmax operation on each element of matrix 112 to obtain an attention probability matrix 114. The Softmax function is also known as the normalized exponential function, which is used to make the element values of each element in each column of the attention probability matrix 114 between 0 and 1, and the sum of the element values of each element in each column is 1. The higher the element value, the higher the probability (attention).

[0047] Therefore, the query matrix 102, the key matrix 104, and the attention probability matrix 114 have the following relationship:

[0048]

[0049] where d k is the number of rows of the key matrix.

[0050] Step 2: Input the attention probability matrix 114 into operation 116 to calculate the pruning ratio R P . The specific steps of operation 116 are as follows.

[0051] Weight each element of the attention probability matrix 114 to obtain an MxM attention weighted matrix.

[0052] Specifically, in the step of weighting each element of the attention probability matrix 114, after weighting, the difference between the maximum value and the minimum value of all elements in the attention probability matrix 114 can be amplified to assist in screening out elements with larger weights. Therefore, in this step, if the maximum value of the elements in the attention probability matrix 114 is Vmax, the minimum value of the elements in the attention probability matrix 114 is Vmin, the maximum value Vmax of the elements after weighting is Vmax’, and the minimum value Vmin of the elements after weighting is Vmin’, then the following relationship holds: where the equal sign appears when the element values of all elements are the same. In this case, the weights of all elements are the same and pruning cannot be performed.

[0053] Therefore, in the weighting step, a function for weighting the element values of the attention probability matrix 114 can be selected according to actual requirements. In some embodiments, it includes weighting each element in a non-linear manner. In some embodiments, the non-linear method includes calculating the square value of each element, or other suitable methods.

[0054] The following is illustrated by an example. Table 1 is an embodiment of the attention probability matrix 114. The attention probability matrix 114 is a 6x6 matrix.

[0055] Table 1

[0056]

[0057] As shown in Table 1, for the matrix in Table 1, the sum of the elements in each column is 1, indicating that the sum of the probabilities in each column is 1.

[0058] On the other hand, in the first column of Table 1, all element values are 1 / 6, indicating that the probabilities or attentions of each item are equal, that is, each element has the same importance. In the second column, the first five elements are all 1 / 12, and the sixth element is 7 / 12, that is, the probability or attention of the sixth element is greater than that of the first five elements. In the third column, the first five elements are all 1 / 24, and the sixth element is 19 / 24, that is, the probability or attention of the sixth element is greater than that of the first five elements, and compared with the second column, the proportion of the sixth element is higher, indicating that the sixth element is more important.

[0059] Next, for all the elements in Table 1, calculate the square value of each element to weight each element. The result is shown in Table 2. Table 2 is also called the attention weighted matrix or the attention square matrix.

[0060] Table 2

[0061] 1 / 36 1 / 36 1 / 36 1 / 36 1 / 36 1 / 36 1 / 144 1 / 144 1 / 144 1 / 144 1 / 144 49 / 144 1 / 576 1 / 576 1 / 576 1 / 576 1 / 576 381 / 576 1 / 576 1 / 576 1 / 576 1 / 576 1 / 576 381 / 576 1 / 576 1 / 576 1 / 576 1 / 576 1 / 576 381 / 576 1 / 576 1 / 576 1 / 576 1 / 576 1 / 576 381 / 576

[0062] As shown in Table 2, in the first column of Table 2, all element values are 1 / 36. That is to say, before weighting (Table 1), when the probabilities of each item are equal, after weighting, the relative magnitudes of each item still remain unchanged, that is, each element has the same importance.

[0063] In the second column, the first five elements are all 1 / 144, while the sixth element is 49 / 144, that is, the probability or attention of the sixth element is greater than that of the first five elements. In particular, before weighting, the ratio of the maximum value 7 / 12 to the minimum value 1 / 12 of the elements in the second column is 7, and after weighting, the ratio of the maximum value 49 / 144 to the minimum value 1 / 144 of the elements in the second column is 49, which is greater than the ratio before weighting. Therefore, it can be seen that after weighting, the elements with greater attention (the maximum value of the elements) can be highlighted.

[0064] In the third column, the first five elements are all 1 / 576, while the sixth element is 361 / 576, that is, the probability or attention of the sixth element is greater than that of the first five elements. In particular, before weighting, the ratio of the maximum value 19 / 24 to the minimum value 1 / 24 of the elements in the second column is 19, and after weighting, the ratio of the maximum value 361 / 576 to the minimum value 1 / 576 of the elements in the second column is 361, which is much greater than the ratio before weighting. Therefore, it can be seen that after weighting, the elements with greater attention (the maximum value of the elements) can be highlighted.

[0065] Next, for each column of the attention probability weighted matrix (shown in Table 2), calculate the sum of its elements to obtain the row weighted sum of each column of the attention probability weighted matrix. In addition, for each row of the attention probability weighted matrix, calculate the sum of its elements to obtain the column weighted sum of each row of the attention probability weighted matrix, as shown in Table 3.

[0066] Table 3

[0067]

[0068] Please refer to Table 3. When calculating the sum of the elements of each column of the attention probability weighted matrix (shown in Table 2) to obtain the row weighted sum of each column of the attention probability weighted matrix, the row weighted sum of each column can be used to represent the probability distribution of each column. A smaller row weighted sum indicates that the probability distribution of the column is more uniform. On the other hand, a larger row weighted sum indicates that the probability distribution of the column is more uneven. Therefore, taking the first column, the second column, and the third column in Table 3 as examples, the row weighted sums of the first three columns are 6 / 36, 54 / 144, and 366 / 576 respectively, and the relative order of their magnitudes is 6 / 36 < 54 / 144 < 366 / 576. Therefore, it indicates that the distribution of the third column is more uneven than that of the first column and the second column.

[0069] On the other hand, when calculating the sum of elements in each row of the attention probability weighted matrix (as shown in Table 2) to obtain the row weighted sum of each row of the attention probability weighted matrix, the row weighted sum of each row is called the importance score of each row, which can be used to represent the importance of each row. The higher the row weighted sum, the higher the importance of the corresponding row; the lower the row weighted sum, the lower the importance of the corresponding row. The lower the importance score, the higher the chance that the corresponding row will be pruned during pruning.

[0070] Please refer to Table 3 again. Sum up the column weighted sums of each column of the attention probability weighted matrix (such as Table 3) to obtain the attention information of the attention probability weighted matrix. As shown in Table 3, summing up the column weighted sums of each column gives 1776 / 576 = 3.083. In some embodiments, the attention information is also called sparsity, which is used to measure the degree of dispersion in the attention probability weighted matrix.

[0071] Finally, determine the pruning ratio R according to the attention information P . Specifically, normalize the attention information within an interval to obtain the pruning ratio. The interval is the range of the upper and lower limits of the attention information.

[0072] For example, when the number of tokens is n, in an extreme case, if the probabilities of all tokens are evenly distributed, the probability of each token is 1 / n. If each probability is weighted by the square method and the attention information (i.e., sparsity) is calculated, the attention information can be obtained as In another extreme case, if the probability of one token is 1 and the probabilities of other tokens are all 0, then the attention information can be obtained as S = 1 × n = n.

[0073] Therefore, according to the above calculations, it can be known that when using the square method for weighting, if the number of tokens is n, the upper and lower limits of the sparsity S can be obtained as 1 ≤ S ≤ n.

[0074] Normalizing the attention information includes normalizing it in a linear method or a non - linear method. For example, as mentioned above, if weighted by the square method, the upper and lower limits of the sparsity S can be obtained as 1 ≤ S ≤ n, that is, the sparsity S is between 1 and n. Therefore, the pruning ratio can be expressed by the following formula:

[0075]

[0076] where RP is the pruning ratio, α is a given parameter, SI is the attention parameter, and n is the number of tokens in the attention probability matrix.

[0077] Step 3, according to the pruning ratio R- P prune the attention probability matrix 114 to obtain a pruned attention probability matrix 118. Specifically, the steps include according to the pruning ratio R- P calculate the number of rows to be pruned based on the number of rows of the attention probability matrix 114. Among them, the number of rows to be pruned is the largest integer less than or equal to the product of the pruning ratio multiplied by the number of rows of the attention probability matrix 114. This number is the number of rows that need to be pruned in the attention probability matrix 114.

[0078] Subsequently, arrange the row weighted sums of each row of the attention probability matrix 114 in ascending order, and prune the rows corresponding to the row weighted sums in ascending order until the number of pruned rows is equal to the number of rows to be pruned, so as to obtain the pruned attention probability matrix 118. Among them, pruning a row means setting the element values of all elements of the row to 0.

[0079] Therefore, after pruning, the size of the pruned attention probability matrix 118 is the same as that of the original matrix, that is, the attention probability matrix 114, but the elements of the pruned rows are all 0, so the amount of computation during matrix operations can be greatly reduced.

[0080] Step 4, in operation 122, prune the value matrix 120 according to the pruning ratio to obtain a pruned value matrix 124. Specifically, step 4 prunes the value matrix 120 in the same way as operation 116 in step 2. The difference is that the value matrix 120 only needs to calculate the corresponding row weighted sum, and uses the pruning ratio R- P , and prune the value matrix 120 in a method similar to operation 116 to obtain a pruned value matrix 124.

[0081] Step 5, multiply the pruned attention probability matrix 118 by the pruned value matrix 124 to obtain an attention matrix 126.

[0082] Therefore, according to the method as Figure 1 shown, it is possible to calculate and prune the tokens corresponding to the input text to greatly reduce the amount of computation for matrix operations.

[0083] Figure 1 The method shown is the case of only having one pruning layer. If there are multiple pruning layers, starting from the second pruning layer, the query matrix and key matrix can be pruned first using the pruning ratio obtained from the previous pruning layer to reduce the amount of computation.

[0084] Figure 2 is a flowchart of a method according to some embodiments of the present invention. Please refer to Figure 2 . Figure 2 The method shown is the same asFigure 1 are similar, so some steps are the same as those in Figure 1 which is similar.

[0085] is similar to Figure 1 which is similar. Figure 2 The method of has three input matrices, namely the query matrix 202, also known as the Q matrix 202; the key matrix 204 or the K matrix 104, and the value matrix 220, or the V matrix 220. The query matrix 202, the key matrix 204, and the value matrix 220 are all MxN matrices with M columns and N rows. Among them, the number of columns M is the number of tokens n in the input text, and the number of rows N is the dimension of the model, denoted as d model .

[0086] Step 1: Prune the query matrix 202 according to the first pruning ratio in the operation 206 to obtain the pruned query matrix 208, where the first pruning ratio is the pruning ratio of the previous pruning layer. Prune the key matrix 204 according to the first pruning ratio to obtain the pruned key matrix 210.

[0087] Specifically, the first pruning ratio here is the pruning ratio calculated in the previous pruning layer. Generally, in different pruning layers, the query matrix 202 and the key matrix 204 are similar to the query matrix and the key matrix of the previous layer. Therefore, the pruning ratio calculated in the previous pruning layer can be used to prune the query matrix 202 and the key matrix 204 first to obtain the pruned query matrix 208 and the pruned key matrix 210. The pruning method is similar to that of Figure 1 the operation 216 in , which will not be elaborated here.

[0088] Step 2: Calculate the attention probability matrix 214 according to the pruned query matrix 208 and the pruned key matrix 210. Compared with Figure 1 , Figure 1 calculates the attention probability matrix using the unpruned query matrix 102 and the key matrix 104, while Figure 2 calculates the attention probability matrix 214 using the pruned query matrix 208 and the pruned key matrix 210. Therefore, Figure 2 in the calculation amount of calculating the attention probability matrix 214 will be less than Figure 1 in calculating the attention probability matrix 114 in .

[0089] Step 3: Input the attention probability matrix 214 into the operation 216 to calculate the second pruning ratio R- P . The specific steps of the operation 216 are similar to those of Figure 1 the operation 116 in , which will not be elaborated here.

[0090] Step 4: According to the second pruning ratio R- PPrune the attention probability matrix 214 to obtain the pruned attention probability matrix 218. The specific steps for calculating the pruned attention probability matrix 218 are similar to those of the pruned attention probability matrix 118 in Figure 1 and will not be elaborated here.

[0091] Step 5, in operation 222, prune the value matrix 220 according to the second pruning ratio R- P to obtain the pruned value matrix 224. The specific steps for calculating the pruned value matrix 224 are similar to those of operation 122 in Figure 1 and will not be elaborated here. When pruning the value matrix 220 here, the pruning ratio used is the second pruning ratio R- P obtained from operation 216 in this pruning layer.

[0092] Step 6, multiply the pruned attention probability matrix 218 by the pruned value matrix 224 to obtain the attention matrix 126.

[0093] Therefore, according to the method as Figure 2 shown, the tokens corresponding to the input text can be calculated and pruned based on the input text. In particular, the input query matrix 202 and key matrix 204 can be pruned in advance according to the pruning ratio calculated in the previous pruning layer, so that the computational amount of calculating the attention probability matrix 214 can be significantly reduced, and the computational amount of subsequent matrix operations can be reduced.

[0094] The following is illustrated with an example and compared with the prior art.

[0095] Figure 3A is the attention probability matrix of a method according to the prior art. Figure 3B is the attention probability weighted matrix of a method according to some embodiments of the present invention. Figure 4A is Figure 3A the result after pruning the attention probability matrix. Figure 4B is Figure 3B the result after pruning the attention probability weighted matrix.

[0096] Please also refer to Figure 3A and Figure 3B . In this example, both the prior art and the method provided by the present invention operate on the same text "A gorgeous, witty, seductive movie." This text is divided into 9 tokens, namely "A", "gorgeous", ",", "witty", ",", "sed", "uctive", "movie", ".".

[0097] Figure 3Ais the attention probability matrix calculated according to a method of the prior art. Figure 3B is the attention probability weighted matrix calculated according to the method shown in the present invention, that is, Figure 3A all elements in are squared and weighted. Figure 3A and Figure 3B both calculate the column weighted sum (rSum) and row weighted sum (cSum) of the matrix.

[0098] According to Figure 3A 's method, the calculated column pruning number is 2, that is, the two columns with the smallest column weighted sum cSum need to be removed, corresponding to the tokens "movie" and "." respectively. The pruning result is as Figure 4A shown.

[0099] According to Figure 3B 's method, that is, the method provided by the present invention, the dispersion S I = 5.628, and the pruning ratio R- P is 0.578. Therefore, the column pruning number is 9 * 0.578 = 5, that is, the five columns with the smallest column weighted sum cSum are removed, corresponding to the tokens ",", "witty", ",", "uctive", and "." respectively. The pruning result is as Figure 4B shown.

[0100] Therefore, from Figure 3A , Figure 3B , Figure 4A , Figure 4B 's results, it can be seen that the method provided by the present invention can effectively prune the text to greatly reduce the matrix operation amount.

[0101] Figure 5 is a comparison between a method of the prior art and a method according to some embodiments of the present invention.

[0102] Please refer to Figure 5 . Figure 5 For Figure 3A , Figure 3B , Figure 4A , Figure 4B 's results, that is, a comparison between a method of the prior art and the method of the present invention and the case without pruning.

[0103] The comparison method is to evaluate with BLEU (BiLingual Evaluation Understudy). The BLEU score ranges from 0 to 1. The higher the Blue score, the higher the efficiency.

[0104] Please refer to Figure 5. In the case of no pruning, performing operations on "A gorgeous, witty, seductive movie." is equivalent to, in Figure 1 , omitting the cases of operation 116 and operation 222. The BLEU score at this time is 0.702672.

[0105] When using the prior art as shown in Figure 3A and Figure 4A to perform calculations on the same text, the average pruning ratio is 0.2. At this time, the Bleu score decreases by 3.7% compared to the non-pruned case. The number of operations saved is 2.26% compared to the non-pruned case.

[0106] When using the method shown in the present invention as shown in Figure 3B and Figure 4B to perform calculations on the same text, the average pruning ratio is 0.5. At this time, the Bleu score decreases by 0.1% compared to the non-pruned case. The number of operations saved is 43.7% compared to the non-pruned case.

[0107] Therefore, it can be seen that the method shown in the present invention can obtain a better pruning ratio, that is, remove more unnecessary tokens, but the impact on the BLEU score is only one thirty-seventh of the prior art, so it means that it is almost the same as the result obtained without pruning. And the number of operations saved exceeds 40% of that without pruning. Therefore, the method shown in the present invention can significantly improve the pruning ratio, save the number of operations, and only have a minimal impact on the text.

[0108] Through the method provided by the present invention, it can be applied in many fields.

[0109] In the requirements of real-time operations, many modern applications, such as machine translation, voice assistants or real-time communication tools, need to provide feedback in a very short time. Therefore, accelerating the calculation speed through pruning can significantly improve the speed of real-time operations.

[0110] In terms of energy efficiency considerations, current large models require a large amount of computing resources during training and inference. This not only increases the operating cost, but also puts pressure on the environment. Using the method provided by the present invention can significantly reduce the energy consumption during the operation of these models.

[0111] In the requirements of mobile device operations, with the progress of technology, more and more operations are starting to shift to edge devices, such as smartphones, drones or other embedded systems. These devices often have limited computing power and battery life. Through the method provided by the present invention, the operation requirements can be significantly reduced, and these applications can be more smooth and persistent on these devices.

[0112] In the need to promote the development of research, the research on deep learning and natural language processing is advancing rapidly. Researchers need to continuously adjust and optimize models, and the instruction cycle is often their bottleneck. The method provided by the present invention can conduct experiments and iterations more quickly, thus accelerating the research process.

[0113] In summary, the present invention provides a dynamic pruning strategy that dynamically adjusts the pruning ratio according to the characteristics of the input text to reduce the amount of data involved in the calculation. Compared with the traditional static token pruning method, this method shows higher flexibility and adaptability. Experiments prove that, in the case of substantial pruning, the present invention can maintain the model accuracy while reducing the computing cost, and is superior to the traditional static pruning method in most scenarios.

[0114] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for sparse attention operation, characterized in that, It includes the following steps: Calculate the attention probability matrix according to the query matrix and the key matrix; Calculate the pruning ratio according to the attention probability matrix; Prune the attention probability matrix according to the pruning ratio to obtain a pruned attention probability matrix; Prune the value matrix according to the pruning ratio to obtain a pruned value matrix; Multiply the pruned attention probability matrix by the pruned value matrix to obtain an attention matrix.

2. The method according to claim 1, characterized in that, The query matrix Q, the key matrix K, and the attention probability matrix O have the following relationship: where d k is the number of rows of the key matrix.

3. The method according to claim 1, characterized in that, In the step of calculating the pruning ratio according to the attention probability matrix, it includes: Weight each element of the attention probability matrix to obtain a weighted attention probability matrix; For each column of the weighted attention probability matrix, calculate the sum of its elements to obtain the column weighted sum of each column of the weighted attention probability matrix; For each row of the weighted attention probability matrix, calculate the sum of its elements to obtain the row weighted sum of each row of the weighted attention probability matrix; Sum up the column weighted sums of each column of the weighted attention probability matrix to obtain the attention information of the weighted attention probability matrix; Determine the pruning ratio according to the attention information.

4. The method according to claim 2, characterized in that, In the step of weighting each element of the attention probability matrix to obtain the weighted attention probability matrix, the maximum value of the elements of the attention probability matrix is Vmax, the minimum value of the elements of the attention probability matrix is Vmin, the maximum value of the elements after weighting is Vmax’, and the minimum value of the elements after weighting is Vmin’, then they have the following relationship:

5. The method according to claim 2, characterized in that, In the step of weighting each element of the attention probability matrix, it includes weighting each element by a non-linear method.

6. The method according to claim 2, characterized in that, In the step of weighting each element of the attention probability matrix, it includes calculating the square value of each element.

7. The method according to claim 2, characterized in that, In the step of determining the pruning ratio according to the attention information, it includes normalizing the attention information in an interval to obtain the pruning ratio.

8. The method according to claim 7, characterized in that, The normalization includes normalizing by a linear method or by a non-linear method.

9. The method according to claim 2, characterized in that, The pruning ratio can be expressed by the following formula: Among them, R P is the pruning ratio, α is a given parameter, S I is the attention parameter, and n is the number of tokens in the attention probability matrix.

10. The method according to claim 1, characterized in that, In the step of pruning the attention probability matrix according to the pruning ratio, it includes: Calculate the row pruning quantity according to the pruning ratio and the number of rows of the attention probability matrix, Arrange the row weighted sums of each row of the attention probability matrix in ascending order, and sequentially prune the rows corresponding to the row weighted sums until the number of pruned rows is equal to the row pruning quantity, where pruning the row means setting the element values of all elements of the row to 0.

11. A method for sparse attention operation, characterized in that, It includes the following steps: Prune the query matrix according to the first pruning ratio to obtain a pruned query matrix, where the first pruning ratio is the pruning ratio of the previous pruning layer; Prune the key matrix according to the first pruning ratio to obtain a pruned key matrix; Calculate the attention probability matrix according to the pruned query matrix and the pruned key matrix; Calculate a second pruning ratio according to the attention probability matrix; Prune the attention probability matrix according to the second pruning ratio to obtain a pruned attention probability matrix; Prune the value matrix according to the second pruning ratio to obtain a pruned value matrix; Multiply the pruned attention probability matrix by the pruned value matrix to obtain an attention matrix.

12. The method according to claim 11, wherein, The pruned query matrix Q’, the pruned key matrix K’ and the attention probability matrix O have the following relationship: where d k is the number of rows of the key matrix.

13. The method according to claim 11, wherein, In the step of calculating the pruning ratio according to the attention probability matrix, it includes: Weight each element of the attention probability matrix to obtain an attention probability weighted matrix; For each column of the attention probability weighted matrix, calculate the sum of its elements to obtain the column weighted sum of each column of the attention probability weighted matrix; For each row of the attention probability weighted matrix, calculate the sum of its elements to obtain the row weighted sum of each row of the attention probability weighted matrix; Sum up the column weighted sums of each column of the attention probability weighted matrix to obtain the attention information of the attention probability weighted matrix; Determine the second pruning ratio according to the attention information.

14. The method according to claim 12, wherein, In the step of weighting each element of the attention probability matrix to obtain the attention probability weighted matrix, the maximum value of the elements of the attention probability matrix is Vmax, the minimum value of the elements of the attention probability matrix is Vmin, the weighted maximum value of the elements is Vmax’, and the weighted minimum value of the elements is Vmin’. Then, they have the following relationship:

15. The method according to claim 12, wherein, In the step of weighting each element of the attention probability matrix, it includes weighting each element by a non-linear method.

16. The method according to claim 12, wherein, In the step of weighting each element of the attention probability matrix, it includes calculating the square value of each element.

17. The method according to claim 12, wherein, In the step of determining the pruning ratio according to the attention information, it includes normalizing the attention information in an interval to obtain the pruning ratio.

18. The method according to claim 17, wherein, The normalization includes normalization by a linear method or normalization by a non-linear method.

19. The method according to claim 12, wherein, The pruning ratio can be expressed by the following formula: where R P is the pruning ratio, α is a given parameter, and S I is the attention parameter, and n is the number of tokens in the attention probability matrix.

20. The method according to claim 11, wherein, In the step of pruning the attention probability matrix according to the pruning ratio, it includes: Calculate the row pruning quantity according to the pruning ratio and the number of rows of the attention probability matrix, Arrange the row weighted sums of each row of the attention probability matrix in ascending order, and sequentially prune the rows corresponding to the row weighted sums until the number of pruned rows is equal to the row pruning quantity. Pruning a row means setting the element values of all elements of the row to 0.