A method, device, storage medium and program product for caching resource reuse
By reusing cache resources in attention mechanism calculation, the problem of insufficient cache resources is solved, and more efficient cache utilization and computing performance improvement is achieved.
Patent Information
- Application Number
- CN202411336688.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-24
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2044-09-24
AI Technical Summary
During the calculation of attention mechanism, the cache resources of the on-chip cache are limited, resulting in insufficient resources and ineffective utilization.
By performing matrix multiplication calculations on query representation and key representation input matrix multiplication, the calculation results are obtained and saved in the on-chip cache. Then the result is input to the normalization operator for operation, and the cache area is multiplexed to save the normalization result, realizing cache resource multiplexing between different operators.
Improve the utilization rate of cache resources, avoid the shortage of cache resources, and ensure the accuracy and performance of attention mechanism calculation.
Smart Images

Figure CN118860963B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of artificial intelligence technology, and in particular, to a cache resource reuse method, device, storage medium, and program product. Background Art
[0002] At present, large language models (LLMs) are developing rapidly, and such models are increasingly applied to various scenarios requiring language processing, such as machine translation, intelligent query, code debugging, etc. The core computing part of large language models is the attention mechanism. The attention mechanism forms a deep structure through stacking and serves as the feature representation part of models such as text classification, text clustering, and relation extraction.
[0003] During the execution of the attention mechanism operator, in order to load data faster during data calculation, the data required by multiple operators included in the attention mechanism operator is loaded into the on-chip cache. To avoid data trampling, related technologies allocate different cache resources in the on-chip cache for the data required during the calculation process.
[0004] However, the cache resources of the on-chip cache are limited. Saving different data required during the calculation process on different cache resources will lead to the problem of insufficient cache resources. Summary of the Invention
[0005] Embodiments of the present application provide a cache resource reuse method, device, storage medium, and program product, which are used to improve the utilization rate of cache resources and avoid the problem of insufficient cache resources.
[0006] On the one hand, embodiments of the present application provide a cache resource reuse method, including:
[0007] Input the query representation and key representation into the first matrix multiplication operator for matrix multiplication calculation to obtain a first calculation result; save the first calculation result in the first buffer area allocated for the first matrix multiplication operator in the on-chip cache;
[0008] Read the first calculation result from the first buffer area, and input the first calculation result into the normalization operator for normalization operation to obtain a normalization result; reuse the first buffer area to save the normalization result, where the first matrix multiplication operator and the normalization operator are operators in a target fusion operator;
[0009] Perform matrix multiplication calculation on the normalization result and the value representation to obtain an output tensor.
[0010] On the one hand, embodiments of the present application provide a cache resource reuse device, including:
[0011] A computing module, configured to input a query representation and a key representation into a first matrix multiplication operator for matrix multiplication calculation to obtain a first calculation result; and save the first calculation result in a first buffer area allocated for the first matrix multiplication operator in the on-chip cache.
[0012] A normalization module, configured to read the first calculation result from the first buffer area, and input the first calculation result into a normalization operator for normalization operation to obtain a normalization result; and reuse the first buffer area to save the normalization result, where the first matrix multiplication operator and the normalization operator are operators in a target fusion operator.
[0013] The computing module is further configured to perform matrix multiplication calculation on the normalization result and a value representation to obtain an output tensor.
[0014] Optionally, the computing module is specifically configured to:
[0015] Read the normalization result from the first buffer area for data type conversion to obtain an intermediate conversion result, where the data type of the intermediate conversion result is the same as the data type of the value representation.
[0016] Save the intermediate conversion result in a second buffer area in the on-chip cache.
[0017] Read the intermediate conversion result from the second buffer area, and perform matrix multiplication calculation with the value representation to obtain an output tensor.
[0018] Optionally, the computing module is specifically configured to:
[0019] Read the intermediate conversion result from the second buffer area, and perform matrix multiplication calculation with the value representation to obtain a second calculation result; and reuse the first buffer area to save the second calculation result.
[0020] Read the second calculation result from the first buffer area for data type conversion to obtain an output tensor, where the data type of the output tensor is the same as the data type of the value representation; and reuse the second buffer area to save the output tensor.
[0021] Optionally, the query representation is one query block among multiple query blocks traversed by an outer loop, the key representation is one key block among multiple key blocks traversed by an inner loop corresponding to the one query block, the value representation is one value block among multiple value blocks traversed by the inner loop corresponding to the one query block, the output tensor is the calculation result of the round of the inner loop traversal where the one key block and the one value block are located, and the one query block corresponds to multiple rounds of inner loop traversal.
[0022] Optionally, the computing module is specifically configured to:
[0023] Read the normalization result from the first buffer and perform matrix multiplication calculation with the value representation to obtain a second calculation result; reuse the first buffer to save the second calculation result;
[0024] Read the second calculation result from the first buffer, and accumulate it with the calculation result obtained from the previous inner loop traversal to obtain the output tensor.
[0025] Optionally, the computing module is specifically configured to:
[0026] Read the normalization result from the first buffer and perform data type conversion to obtain an intermediate conversion result, where the data type of the intermediate conversion result is the same as the data type of the value representation;
[0027] Save the intermediate conversion result in a second buffer in the on-chip cache;
[0028] Read the intermediate conversion result from the second buffer, and perform matrix multiplication calculation with the value representation to obtain the second calculation result.
[0029] Optionally, the computing module is further configured to:
[0030] After reading the second calculation result from the first buffer, and accumulating it with the calculation result obtained from the previous inner loop traversal to obtain the output tensor, when the output tensor is the calculation result obtained from the last inner loop traversal in the multiple inner loop traversals, perform data type conversion on the output tensor to obtain the target calculation result for the outer loop traversal round where the one query block is located, where the data type of the target calculation result is the same as the data type of the value representation.
[0031] Optionally, the computing module is further configured to:
[0032] Reuse the second buffer to save the target calculation result.
[0033] Optionally, the computing module is specifically configured to:
[0034] Read the second calculation result from the first buffer; and obtain the calculation result obtained from the previous inner loop traversal from a third buffer in the on-chip cache;
[0035] Accumulate the second calculation result and the calculation result obtained from the previous inner loop traversal to obtain an output tensor;
[0036] Reuse the third buffer to save the output tensor.
[0037] Optionally, the computing module is further configured to:
[0038] Save the one query block in a fourth buffer area in the on-chip cache until the execution of multiple rounds of inner loop traversal corresponding to the one query block ends.
[0039] On the one hand, an embodiment of the present application provides a computer device, including a memory, an artificial intelligence chip, and a computer program stored on the memory and executable on the artificial intelligence chip. When the artificial intelligence chip executes the computer program, the steps of the above cache resource reuse method are implemented.
[0040] On the one hand, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program executable by a computer device. When the computer program runs on the computer device, the computer device is enabled to execute the steps of the above cache resource reuse method.
[0041] On the one hand, an embodiment of the present application provides a computer program product. The computer program product includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer device, the computer device is enabled to execute the steps of the above cache resource reuse method.
[0042] In the embodiment of the present application, the query representation and the key representation are input into a first matrix multiplication operator for matrix multiplication calculation to obtain a first calculation result; then the first calculation result is saved in a first buffer area allocated for the first matrix multiplication operator in the on-chip cache. The first calculation result is read from the first buffer area, and the first calculation result is input into a normalization operator for normalization operation to obtain a normalization result. Since the data types of the first calculation result and the normalization result are the same, and after obtaining the normalization result, the first calculation result is no longer required for subsequent attention mechanism calculation, and the first matrix multiplication operator and the normalization operator are operators in the same target fusion operator, therefore, the first buffer area originally storing the first calculation result can be reused to store the normalization result, which not only ensures the accuracy of the attention mechanism calculation, but also realizes the reuse of cache resources between different operators in a fusion operator, thereby improving the utilization rate of cache resources and effectively alleviating the situation of tight use of cache resources in the process of operator implementation. Description of the Drawings
[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0044] Figure 1 Schematic diagram of the structure of an artificial intelligence chip provided by an embodiment of the present application;
[0045] Figure 2 Schematic flowchart of a cache resource reuse method provided by an embodiment of the present application;
[0046] Figure 3 Schematic flowchart of an attention mechanism calculation provided by an embodiment of the present application;
[0047] Figure 4 Schematic flowchart of an inner loop traversal provided by an embodiment of the present application;
[0048] Figure 5 Schematic diagram of the usage situation of cache resources provided by an embodiment of the present application;
[0049] Figure 6 Schematic diagram of the structure of a cache resource reuse device provided by an embodiment of the present application;
[0050] Figure 7 Schematic diagram of the structure of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0051] In order to make the objectives, technical solutions and beneficial effects of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0052] For the convenience of understanding, the nouns involved in the embodiments of the present application are explained below.
[0053] FlashAttention: An accelerated attention calculation method. This method first divides the input into multiple chunks and loads each chunk from the video memory into the on-chip cache. During the attention mechanism calculation, some chunks need to be reused. At this time, the same chunk can be read from the on-chip cache multiple times for the attention mechanism calculation; compared with reading the same chunk from the video memory multiple times for the attention mechanism calculation, it reduces the access to the video memory, thereby improving the calculation speed and reducing the video memory access overhead.
[0054] A brief introduction to the system architecture diagram applicable to the technical solutions of the embodiments of the present application is given below. It should be noted that the system architecture diagram introduced below is only used to illustrate the embodiments of the present application and is not a limitation.
[0055] Refer to Figure 1, which is a structural diagram of an artificial intelligence chip applicable to an embodiment of the present application. The artificial intelligence chip 100 at least includes: a video memory 101, an on-chip cache 102, and multiple execution units 103.
[0056] The video memory 101 can be a High Bandwidth Memory (HBM) or other types of memories.
[0057] The cache resources in the on-chip cache 102 include: Group-Shared Memory (GSM) and Thread-Local Register (TLR).
[0058] Multiple execution units 103 share the GSM, and each execution unit 103 corresponds to an exclusive TLR. The number of TLRs exclusive to each execution unit 103 can be set according to actual situations. Compared with the GSM, the capacity of the TLR is smaller than that of the GSM, but the data exchange speed is faster than that of the GSM.
[0059] The execution unit 103 can be used to execute various operators, such as an attention mechanism operator, a convolution operator, etc. During the process of executing an operator, the execution unit 103 loads relevant data from the video memory 101 into the on-chip cache 102, and then reads the relevant data from the on-chip cache 102 for calculation to obtain a calculation result. In the actual calculation process, a certain data may need to be used for calculation multiple times. Therefore, first load the data from the video memory 101 into the on-chip cache 102, and then read the data from the on-chip cache 102 multiple times for calculation. This method can greatly improve the data reading speed and calculation speed compared with reading the data from the video memory 101 multiple times.
[0060] In addition, the final calculation result or the intermediate result generated during the calculation process can also be saved in the on-chip cache 102 and then written back from the on-chip cache 102 to the video memory 101.
[0061] In addition to the two types of cache resources, GSM and TLR, the on-chip cache 102 also includes other types of cache resources. Regarding this, the present application does not make specific limitations.
[0062] The artificial intelligence chip 100 can be: a Graphics Processing Unit (GPU), a General-purpose computing on graphics processing units (GPGPU), a Domain Specific Architecture (DSA), etc.
[0063] In practical applications, during the execution of the attention mechanism operator, in order to load data faster during data calculation, the data required by multiple operators included in the attention mechanism operator will be loaded into the on-chip cache. However, if two data required during the calculation process are loaded onto the same cache resource successively, then the later-loaded data will overwrite the earlier-loaded data. When the later operator needs to use the earlier-loaded data, it cannot read the data from the on-chip cache and must read the data from the video memory, that is, data trampling occurs.
[0064] To avoid data trampling, different cache resources in the on-chip cache need to be allocated for different data required during the calculation process. For example, the query representation is saved in 32 TLRs, and the matrix multiplication result of the query representation and the key representation is saved in another 64 TLRs.
[0065] However, the cache resources of the on-chip cache are limited. Saving different data required during the calculation process on different cache resources will lead to the problem of insufficient cache resources.
[0066] In view of this, based on Figure 1 the architecture diagram of the artificial intelligence chip shown, a process of a cache resource reuse method is provided, as Figure 2 shown. The process of this method is executed by the artificial intelligence chip and includes the following steps:
[0067] Step 201, input the query representation and the key representation into the first matrix multiplication operator for matrix multiplication calculation to obtain a first calculation result; save the first calculation result in the first buffer area allocated for the first matrix multiplication operator in the on-chip cache.
[0068] Specifically, the query representation refers to the specific content of the query (Q) in the attention mechanism. The query representation can specifically be a scalar, a vector, a multi-dimensional array, etc. Similarly, the key representation refers to the specific content of the key (K) in the attention mechanism. The key representation can specifically be a scalar, a vector, a multi-dimensional array, etc. The value representation refers to the specific content of the value (V) in the attention mechanism. The value representation can specifically be a scalar, a vector, a multi-dimensional array, etc.
[0069] The execution unit loads the query representation, key representation, and value representation from the video memory into the GSM respectively. Compared with the GSM, the capacity of the TLR is smaller than that of the GSM, but the data exchange speed is faster than that of the GSM. Therefore, one or more of the query representation, key representation, and value representation can be loaded from the GSM into the TLR.
[0070] For example, the query representation is loaded from the GSM into the first number of TLRs allocated for the query representation. In this way, during the calculation of the attention mechanism, the key representation is read from the GSM, the query representation is read from the first number of TLRs, and then the read query representation and key representation are subjected to matrix multiplication calculation to obtain the first calculation result. Compared with reading the query representation from the GSM, reading the query representation from the TLR is faster, so the data reading speed can be further improved.
[0071] It should be noted that the embodiments of the present application may also only use the GSM to save the query representation, key representation, and value representation. In this way, during the calculation of the attention mechanism, the key representation and query representation are read from the GSM, and then the read query representation and key representation are input into the first matrix multiplication operator for matrix multiplication calculation to obtain the first calculation result. The present application does not make specific limitations on this.
[0072] Next, the first calculation result is saved in the first buffer area allocated for the matrix multiplication operator in the on-chip cache, where the size of the first buffer area is associated with the data type of the first calculation result to be saved. That is to say, the size of the first buffer area is determined according to the data type of the first calculation result.
[0073] Specifically, based on the data type of the first calculation result in advance, the storage space size required by the first calculation result is determined. The data type of the first calculation result can be: FP32 (i.e., 32-bit floating point number), BF16 (i.e., 16-bit floating point number), etc.
[0074] Then, according to the storage space size required by the first calculation result, the size of the first buffer area for saving the first calculation result is determined, where the size of the first buffer area is greater than or equal to the storage space size required by the first calculation result. Finally, according to the size of the first buffer area, the first buffer area for saving the first calculation result is divided from the on-chip cache. The first buffer area can be a storage area in the GSM or the second number of TLRs.
[0075] Step 202, read the first calculation result from the first buffer area, and input the first calculation result into the normalization operator for normalization operation to obtain the normalization result, and reuse the first buffer area to save the normalization result.
[0076] Specifically, the first matrix multiplication operator and the normalization operator are operators in a target fusion operator, that is, at least the first matrix multiplication operator and the normalization operator are fused to obtain the target fusion operator. A total buffer is pre-allocated in the on-chip cache for the target fusion operator, and each operator within the target fusion operator can use the cache resources in this total buffer. For example, a first buffer is allocated for the matrix multiplication operator from this total buffer.
[0077] When reusing a buffer to store data, at least the following reuse conditions need to be met: Reuse condition one, the data type of the old data originally stored in the buffer is the same as the data type of the new data to be stored by reusing this buffer, that is, the storage space size required by the old data is the same as that required by the new data. Reuse condition two, the old data originally stored in this buffer is not required in the subsequent attention mechanism calculation process, which can avoid data trampling. Reuse condition three, the operator that originally used this buffer and the operator that reuses this buffer are operators in the same fusion operator.
[0078] Taking the reuse of the first buffer to store the normalization result as an example, the data type of the normalization result (i.e., the new data) is the same as the quantity type of the first calculation result (i.e., the old data), that is, reuse condition one is met. In addition, after obtaining the normalization result by performing the normalization operation on the first calculation result, the subsequent attention mechanism calculation no longer requires the first calculation result, that is, reuse condition two is met. The first matrix multiplication operator and the normalization operator are operators in the target fusion operator, that is, reuse condition three is met.
[0079] Therefore, the first buffer can be reused to store the normalization result, that is, the normalization result overwrites the first calculation result originally stored in the first buffer.
[0080] In some embodiments, during the attention mechanism calculation process, if a scaling operation is included before the normalization operation (that is, the scaling operator is executed before the normalization operator, and the scaling operator is also an operator in the target fusion operator), then the first calculation result is read from the first buffer, divided by a constant value, to obtain the scaling operation result, and then the first buffer is reused to store the scaling operation result. Then, the scaling operation result is read from the first buffer to perform the normalization operation to obtain the normalization result; and the first buffer is reused to store the normalization result.
[0081] In the attention mechanism calculation process, if a masking operation is included before the normalization operation (that is, the masking operator is executed before the normalization operator, and the masking operator is also an operator in the target fusion operator), then the first calculation result is read from the first buffer to perform the masking operation to obtain the masking operation result, and then the first buffer is reused to store the masking operation result. Then, the masking operation result is read from the first buffer to perform the normalization operation to obtain the normalization result; and the first buffer is reused to store the normalization result.
[0082] During the calculation process of the attention mechanism, if there are scaling and masking operations before the normalization operation (that is, the masking operator and the scaling operator are executed before the normalization operator, and the masking operator and the scaling operator are also operators in the target fusion operator), then the first calculation result is read from the first buffer, divided by a constant value to obtain the scaling operation result, and then the first buffer is reused to save the scaling operation result. Then, the scaling operation result is read from the first buffer for masking operation to obtain the masking operation result, and then the first buffer is reused to save the masking operation result. Finally, the masking operation result is read from the first buffer for normalization operation to obtain the normalization result; and the first buffer is reused to save the normalization result.
[0083] Step 203: Obtain the output tensor by performing matrix multiplication on the normalization result and the value representation.
[0084] Specifically, the value representation is read from the GSM, and the normalization result is read from the first buffer. Then, the read value representation and the normalization result are input into the second matrix multiplication operator for matrix multiplication calculation to obtain the output tensor. If the data type of the output tensor is the same as that of the normalization result, and the second matrix multiplication operator is also an operator in the target fusion operator, then the first buffer can be reused to save the output tensor. When the first buffer is the second number of TLRs, the output tensor will be written from the first buffer to the GSM subsequently. Of course, the output tensor can also be directly written to the GSM after obtaining the output tensor.
[0085] In the embodiment of the present application, the query representation and the key representation are input into the first matrix multiplication operator for matrix multiplication calculation to obtain the first calculation result; then the first calculation result is saved in the first buffer allocated for the first matrix multiplication operator in the on-chip cache. The first calculation result is read from the first buffer and input into the normalization operator for normalization operation to obtain the normalization result; since the data type of the first calculation result is the same as that of the normalization result, and after obtaining the normalization result, the first calculation result is no longer required for subsequent attention mechanism calculation, and the first matrix multiplication operator and the normalization operator are operators in a target fusion operator, therefore, the first buffer originally used to save the first calculation result can be reused to save the normalization result, which not only ensures the accuracy of the attention mechanism calculation, but also realizes the reuse of cache resources between different operators within a fusion operator, thereby improving the utilization rate of cache resources and effectively alleviating the situation of tight cache resource usage during the implementation of the operator.
[0086] In some artificial intelligence chips, the data types of query representations, key representations, and value representations are the same. However, the data type of the first calculation result obtained by performing matrix multiplication on the query representation and the key representation may be different from the data type of the value representation. Further, the data type of the normalization result obtained by performing a normalization operation on the first calculation result may also be different from the data type of the value representation. Since the data type of the normalization result is different from the data type of the value representation, some artificial intelligence chips cannot directly perform matrix multiplication on the normalization result and the value representation.
[0087] In view of this, before performing matrix multiplication on the normalization result and the value representation, the present application reads the normalization result from the first buffer for data type conversion to obtain an intermediate conversion result, and the data type of the intermediate conversion result is the same as the data type of the value representation. The intermediate conversion result is stored in a second buffer in the on-chip cache. The intermediate conversion result is read from the second buffer and matrix multiplication is performed with the value representation to obtain an output tensor.
[0088] Specifically, the normalization result is input into a first type conversion operator for data type conversion to obtain an intermediate conversion result. The data type of the intermediate conversion result is different from the data type of the normalization result; correspondingly, the storage space size required for the intermediate conversion result is also different from the storage space size required for the normalization result, that is, the first reuse condition is not satisfied. Therefore, the first buffer cannot be directly reused to store the intermediate conversion result, and a second buffer needs to be partitioned in the on-chip cache to store the intermediate conversion result. When the first type conversion operator is also an operator in the target fusion operator, the second buffer can be allocated from the total buffer allocated to the target fusion operator to the first type conversion operator.
[0089] The size of the second buffer is associated with the data type of the intermediate conversion result to be stored, that is, the size of the second buffer is determined according to the data type of the intermediate conversion result.
[0090] Specifically, based on the data type of the intermediate conversion result, the storage space size required for the intermediate conversion result is determined in advance; then, according to the storage space size required for the intermediate conversion result, the size of the second buffer for storing the intermediate conversion result is determined, where the size of the second buffer is greater than or equal to the storage space size required for the intermediate conversion result. Finally, according to the size of the second buffer, a second buffer for storing the intermediate conversion result is partitioned from the on-chip cache. The second buffer can also be a storage area in the GSM and can be the TLR of the third quantity.
[0091] In some embodiments, in addition to ensuring that the data types of the normalization result and the value representation are the same, it is also necessary to ensure that the data types of the input tensor and the output tensor calculated by the attention mechanism are the same. The input tensor calculated by the attention mechanism includes: query representation, key representation, and value representation, where the data types corresponding to the query representation, key representation, and value representation are the same. Therefore, it is necessary to ensure that the data type of the output tensor is the same as that of any one of the query representation, key representation, and value representation.
[0092] In view of this, in the process of calculating the attention mechanism in the present application, the intermediate conversion result is read from the second buffer, and matrix multiplication is performed with the value representation to obtain a second calculation result; then the first buffer is reused to save the second calculation result.
[0093] The second calculation result is read from the first buffer for data type conversion to obtain an output tensor, and the data type of the output tensor is the same as that of the value representation; then the second buffer is reused to save the output tensor.
[0094] Specifically, the intermediate conversion result and the value representation are input into a second matrix multiplication operator for matrix multiplication to obtain a second calculation result. The second matrix multiplication operator is an operator in the target fusion operator. The data type of the second calculation result is the same as that of the normalization result, that is, the first reuse condition is satisfied. After the normalization result is subjected to data type conversion to obtain the intermediate conversion result, the normalization result is no longer required for subsequent attention mechanism calculations, that is, the second reuse condition is satisfied. The second matrix multiplication operator and the normalization operator are both operators in the target fusion operator, that is, the third reuse condition is satisfied. Therefore, the first buffer can be reused to save the second calculation result. After the second calculation result is saved in the first buffer, the normalization result originally saved in the first buffer will be overwritten.
[0095] In addition, the data type of the second calculation result is not the same as that of the value representation. Therefore, the second calculation result is read from the first buffer, and the second calculation result is input into a second type conversion operator for data type conversion to obtain an output tensor, and the data type of the output tensor is the same as that of the value representation, so as to ensure that the data types of the input tensor and the output tensor calculated by the attention mechanism are the same. The first type conversion operator and the second type conversion operator are operators in the same fusion operator, and this fusion operator can be the target fusion operator described above or other fusion operators.
[0096] Since the data type of the output tensor is the same as that of the intermediate conversion result, that is, the first reuse condition is satisfied. In addition, after the intermediate conversion result is read from the second buffer and matrix multiplication is performed with the value representation to obtain the second calculation result, the intermediate conversion result is no longer required for subsequent attention mechanism calculations, that is, the second reuse condition is satisfied. The second type conversion operator and the first type conversion operator are operators in the same fusion operator, that is, the third reuse condition is satisfied.
[0097] Therefore, the second buffer can be reused to store the output tensor. After the output tensor is stored in the second buffer, the intermediate conversion result originally stored in the second buffer will be overwritten.
[0098] For example, referring to Figure 3 , it is assumed that the data types of the query representation, key representation, and value representation are all BF16 (i.e., 16-bit floating point numbers). The key representation and value representation are stored in the GSM, and the query representation is stored in a specified buffer (32 TLRs).
[0099] Perform matrix multiplication on the query representation and the key representation to obtain a first calculation result of FP32 (32-bit floating point number), and store the first calculation result of FP32 in the first buffer (64 TLRs).
[0100] Read the first calculation result of FP32 from the first buffer for normalization operation to obtain a normalized result of FP32, and reuse the first buffer to store the normalized result of FP32.
[0101] Read the normalized result of FP32 from the first buffer for data type conversion to obtain an intermediate conversion result of BF16, and store the intermediate conversion result of BF16 in the second buffer (32 TLRs). The above-mentioned specified buffer, first buffer, and second buffer are different buffers.
[0102] Read the intermediate conversion result of BF16 from the second buffer and perform matrix multiplication with the value representation of BF16 to obtain a second calculation result of FP32; reuse the first buffer to store the second calculation result of FP32.
[0103] Read the second calculation result of FP32 from the first buffer for data type conversion to obtain an output tensor of BF16, and reuse the second buffer to store the output tensor of BF16. Write the output tensor of BF16 from the second buffer to the GSM.
[0104] In the embodiments of the present application, both the first buffer and the second buffer are reused multiple times between different operators within the fusion operator. That is, the first buffer is respectively used to store the first calculation result output by the first matrix multiplication operator, the normalized result output by the normalization operator, and the second calculation result output by the second matrix multiplication operator. The second buffer is respectively used to store the intermediate conversion result output by the first type conversion operator and the output tensor output by the second type conversion operator. This greatly improves the utilization rate of cache resources, effectively alleviates the situation of tight cache resource usage during the implementation of operators, and avoids the problem of insufficient cache resources.
[0105] In some embodiments, during the Flash Attention calculation process, the entire query input is sliced into i query chunks, the entire key input is sliced into m key chunks, and the entire value input is sliced into m value chunks, where both i and m are greater than 1.
[0106] Perform i rounds of outer loop traversal for the i query chunks. During each round of outer loop traversal, it includes m rounds of inner loop traversal for the m key chunks and m value chunks.
[0107] Taking the outer loop traversal of the j-th query chunk among the i query chunks as an example (i.e., the round of outer loop traversal where the j-th query chunk is located is the j-th round of outer loop traversal), during the j-th round of outer loop traversal, it includes the following m rounds of inner loop traversal:
[0108] Perform attention mechanism calculation on the j-th query chunk, the first key chunk, and the first value chunk to obtain the calculation result of the first round of inner loop traversal, where the rounds of inner loop traversal where the first key chunk and the first value chunk are located are the first round of inner loop traversal.
[0109] Perform attention mechanism calculation on the j-th query chunk, the second key chunk, and the second value chunk, and accumulate the obtained calculation result with the calculation result of the first round of inner loop traversal to obtain the calculation result of the second round of inner loop traversal, where the rounds of inner loop traversal where the second key chunk and the second value chunk are located are the second round of inner loop traversal.
[0110] And so on, until performing attention mechanism calculation on the j-th query chunk, the m-th key chunk, and the m-th value chunk, and accumulating the obtained calculation result with the calculation result of the m - 1-th round of inner loop traversal to obtain the calculation result of the m-th round of inner loop traversal. The calculation result of the m-th round of inner loop traversal is the calculation result of the round of outer loop traversal where the j-th query chunk is located, where j is greater than 0 and less than or equal to i.
[0111] The cache resource reuse method in this application is also applicable to each round of inner loop traversal process of the above Flash Attention calculation. Taking one query chunk as an example, the query in the cache resource reuse method of this application is represented as one query chunk among the multiple query chunks of the outer loop traversal, the key is represented as one key chunk among the multiple key chunks of the inner loop traversal corresponding to this query chunk, the value is represented as one value chunk among the multiple value chunks of the inner loop traversal corresponding to this query chunk, the output tensor is the calculation result of the round of inner loop traversal where one key chunk and one value chunk are located, and one query chunk corresponds to multiple rounds of inner loop traversal.
[0112] Correspondingly, the attention mechanism calculation performed in each round of inner loop traversal includes the following steps:
[0113] Perform matrix multiplication on the query representation and the key representation to obtain a first calculation result. That is, input a query chunk and a key chunk into a first matrix multiplication operator for matrix multiplication calculation to obtain a first calculation result, and then save the first calculation result in a first buffer area allocated for the first matrix multiplication operator in the on-chip cache. Read the first calculation result from the first buffer area and input the first calculation result into a normalization operator for normalization operation to obtain a normalization result, where the first matrix multiplication operator and the normalization operator are operators in a target fused operator.
[0114] For the Flash Attention scenario, when reusing a buffer area to save data, at least the following reuse conditions need to be met: Reuse condition one: The data type of the old data originally saved in the buffer area is the same as the data type of the new data saved by reusing this buffer area, that is, the storage space size required by the old data is the same as the storage space size required by the new data. Reuse condition two: The subsequent calculation process in the current round of inner loop traversal does not need to use the old data originally saved in this buffer area, so as to avoid data trampling. Reuse condition three: The operator that originally used this buffer area and the operator that reuses this buffer area are operators in the same fused operator. Reuse condition four: The old data originally saved in the buffer area is not needed in the next round of inner loop traversal.
[0115] Taking the reuse of the first buffer area to save the normalization result as an example, the data type of the normalization result (new data) is the same as the quantity type of the first calculation result (old data), that is, it meets reuse condition one. After performing the normalization operation on the first calculation result to obtain the normalization result, the subsequent calculation in the current round of inner loop traversal no longer needs the first calculation result, that is, it meets reuse condition two. The first matrix multiplication operator and the normalization operator are both operators in the target fused operator, that is, it meets reuse condition three. The first calculation result obtained in the current round of inner loop traversal is not needed in the next round of inner loop traversal, that is, it meets reuse condition four. Therefore, reuse the first buffer area to save the normalization result.
[0116] Next, read the normalization result from the first buffer area and perform matrix multiplication with the value representation to obtain a second calculation result. That is, input the normalization result and a value chunk into a second matrix multiplication operator for matrix multiplication calculation to obtain a second calculation result, where the second matrix multiplication operator and the normalization operator are both operators in the target fused operator.
[0117] Since the data type of the second calculation result is the same as that of the normalization result, that is, the first reuse condition is satisfied. After obtaining the second calculation result by performing matrix multiplication on the normalization result and the value representation, the subsequent calculations in this round of inner loop traversal no longer require the normalization result, that is, the second reuse condition is satisfied. Both the second matrix multiplication operator and the normalization operator are operators in the target fusion operator, that is, the third reuse condition is satisfied. The next round of inner loop traversal does not need to use the normalization result obtained in this round of inner loop traversal, that is, the fourth reuse condition is satisfied. Therefore, the first buffer is reused to store the second calculation result.
[0118] Read the second calculation result from the first buffer and accumulate it with the calculation result obtained in the previous round of inner loop traversal to obtain the output tensor.
[0119] Specifically, read the second calculation result from the first buffer, and input the second calculation result and the calculation result obtained in the previous round of inner loop traversal into the accumulation operator for accumulation to obtain the output tensor.
[0120] In some embodiments, since the accumulation operator needs to be used to accumulate the calculation result obtained in the previous round of inner loop traversal in each round of inner loop traversal process, that is, the next round of inner loop traversal needs to use the calculation result obtained in this round of inner loop traversal, the fourth reuse condition can never be satisfied. That is to say, the buffer for storing the output tensor of the accumulation operator (that is, the calculation result obtained in each round of inner loop traversal) cannot be reused during multiple rounds of inner loop traversal.
[0121] In this case, a third buffer is separately allocated from the on-chip cache for the accumulation operator to store the calculation result obtained in each round of inner loop traversal; and during multiple rounds of inner loop traversal, the third buffer cannot be used to store the calculation results of other operators, that is, it cannot be reused by other operators to avoid data trampling. When the accumulation operator is also an operator in the target fusion operator, the third buffer can be allocated from the total buffer allocated to the target fusion operator to the accumulation operator.
[0122] In specific implementation, during each round of inner loop traversal, read the second calculation result from the first buffer; and obtain the calculation result obtained in the previous round of inner loop traversal from the third buffer in the on-chip cache; input the second calculation result and the calculation result obtained in the previous round of inner loop traversal into the accumulation operator for accumulation to obtain the output tensor (that is, the calculation result obtained in this round of inner loop traversal), and then reuse the third buffer to store the output tensor, overwriting the calculation result obtained in the previous round of inner loop traversal, that is, the third buffer always stores the calculation result obtained in the latest round of inner loop traversal.
[0123] In some embodiments, since the same query chunk needs to be used in multiple rounds of inner loop traversal corresponding to a query chunk, that is, the query chunk originally stored in the buffer also needs to be used in the next round of inner loop traversal, the fourth reuse condition can never be met. That is to say, the buffer for storing query chunks cannot be reused during multiple rounds of inner loop traversal.
[0124] In this case, a query chunk is stored in the fourth buffer area of the on-chip cache until the multiple rounds of inner loop traversal corresponding to the query chunk are completed.
[0125] Specifically, a fourth buffer area is separately allocated from the on-chip cache to store the query chunk, and during multiple rounds of inner loop traversal, this fourth buffer area cannot be used to store other data, that is, it cannot be reused, to avoid data trampling. When the multiple rounds of inner loop traversal corresponding to a query chunk are completed, the output tensor obtained from the last round of inner loop traversal is written out to the GSM. In practical applications, the fourth buffer area can be allocated from the total buffer area assigned to the target fusion operator to store the query chunk; or the fourth buffer area can be allocated from other caches in the on-chip cache except the above total buffer area to store the query chunk. Regarding this, the present application does not make specific limitations.
[0126] It should be noted that during the first round of inner loop traversal, the calculation result obtained from the previous round of inner loop traversal is preset initial data. In addition, the attention mechanism calculation performed in each round of inner loop traversal may further include: a scaling operation and / or a masking operation. The specific calculation process has been explained above and will not be elaborated here.
[0127] In some embodiments, during each round of inner loop traversal, since the data type of the normalization result is different from the data type of the value representation, some artificial intelligence chips cannot directly perform matrix multiplication calculation on the normalization result and the value representation. In view of this, before performing matrix multiplication calculation on the normalization result and the value representation in the present application, the normalization result is read from the first buffer area for data type conversion to obtain an intermediate conversion result, and the data type of the intermediate conversion result is the same as the data type of the value representation. The intermediate conversion result is stored in the second buffer area of the on-chip cache. The intermediate conversion result is read from the second buffer area and matrix multiplication calculation is performed with the value representation to obtain a second calculation result.
[0128] Specifically, the normalization result is input into a first type conversion operator for data type conversion to obtain an intermediate conversion result. As described above, the first buffer area cannot be directly reused to store the intermediate conversion result, and a second buffer area needs to be divided in the on-chip cache to store the intermediate conversion result. The size of the second buffer area is associated with the data type of the intermediate conversion result that needs to be stored.
[0129] The intermediate conversion result and the value representation are input into the second matrix multiplication operator for matrix multiplication calculation to obtain a second calculation result. The second matrix multiplication operator is an operator in the target fusion operator. The data type of the second calculation result is the same as that of the normalization result, that is, the first reuse condition is satisfied. After performing data type conversion on the normalization result to obtain the intermediate conversion result, the subsequent calculations in this round of inner loop traversal no longer require the normalization result, that is, the second reuse condition is satisfied. The second matrix multiplication operator and the normalization operator are both operators in the target fusion operator, that is, the third reuse condition is satisfied. The next round of inner loop traversal does not need to use the normalization result obtained in this round of inner loop traversal, that is, the fourth reuse condition is satisfied. Therefore, the first buffer can be reused to store the second calculation result, that is, the second calculation result is stored in the first buffer to overwrite the normalization result originally stored in the first buffer.
[0130] In some embodiments, when the output tensor is the calculation result obtained in the last round of inner loop traversal in multiple rounds of inner loop traversal, the data type of the output tensor is not the same as that of the value representation. If the output tensor is directly used as the target calculation result for the outer loop traversal round where the query block is located, it will cause the data types of the input and output of the attention mechanism calculation to be different.
[0131] In view of this, the present application performs data type conversion on the output tensor to obtain the target calculation result for the outer loop traversal round where the query block is located. The data type of the target calculation result is the same as that of the value representation, thereby ensuring that the data types of the input and output of the attention mechanism calculation are the same.
[0132] Specifically, the output tensor is input into the second type conversion operator for data type conversion to obtain the target calculation result for the outer loop traversal round where the query block is located. The first type conversion operator and the second type conversion operator are operators in the same fusion operator. The fusion operator can be the target fusion operator described above or other fusion operators.
[0133] Since the data type of the target calculation result is also the same as that of the intermediate conversion result, that is, the first reuse condition is satisfied. The subsequent calculations in this round of inner loop traversal no longer require the intermediate conversion result, that is, the second reuse condition is satisfied. The first type conversion operator and the second type conversion operator are operators in the same fusion operator, that is, the third reuse condition is satisfied. The next round of inner loop traversal does not need to use the intermediate conversion result obtained in this round of inner loop traversal, that is, the fourth reuse condition is satisfied. Therefore, when the execution of multiple rounds of inner loop traversal corresponding to a query block ends, the second buffer can be reused to store the target calculation result for the outer loop traversal round where the query block is located.
[0134] When the second buffer is TLR of the third quantity, subsequently, the target calculation result of the outer loop traversal round where the query block is located is written from the second buffer to the GSM; of course, after obtaining the target calculation result, the output tensor can also be directly written to the GSM.
[0135] For example, it is assumed that the Flash Attention operator is a fused operator obtained by fusing at least a first matrix multiplication operator, a normalization operator, a first type conversion operator, a second matrix multiplication operator, an accumulation operator, and a second type conversion operator. A total buffer is allocated for the Flash Attention operator in the on-chip cache.
[0136] During the execution of the Flash Attention operator, the entire query input is split into 10 query blocks, the entire key input is split into 8 key blocks, and the entire value input is split into 8 value blocks.
[0137] Ten rounds of outer loop traversals are performed for the 10 query blocks. During each round of outer loop traversal, eight rounds of inner loop traversals for the 8 key blocks and 8 value blocks are included.
[0138] Taking the outer loop traversal (i.e., the first round of outer loop traversal) for query block 1 (i.e., the first query block) as an example, during the first round of outer loop traversal, eight rounds of inner loop traversals are included. The usage of TLR resources during each round (taking the r-th round as an example) of inner loop traversal is as Figure 4 shown, and the usage of TLR resources during each round of inner loop traversal is as Figure 5 shown. Figure 5 The abscissa in Figure 5 represents the life cycle of the buffer, and Figure 4 and Figure 5 are specifically introduced below in combination with
[0139] The BF16 query block 1, the BF16 key block r (i.e., the r-th key block), and the BF16 value block r (i.e., the r-th value block) are loaded from the video memory to the GSM. The query block 1 is loaded from the GSM to buffer 1 (32 TLRs), where r is greater than or equal to 1 and less than or equal to 8. Since query block 1 is required in all eight rounds of inner loop traversals, during the eight rounds of inner loop traversals, buffer 1 is allocated for query block 1 in the total buffer. Buffer 1 is always used to store query block 1 until the end of the eight rounds of inner loop traversals. For details, see Figure 5 .
[0140] The first matrix multiplication operator is input with query block 1 and key block r for matrix multiplication calculation to obtain the first calculation result in FP32, and the first calculation result in FP32 is saved in buffer area 2 (64 TLRs) allocated for the first matrix multiplication operator in the total buffer.
[0141] The first calculation result in FP32 is read from buffer area 2, and the first calculation result is input to the normalization operator for normalization operation to obtain the normalization result in FP32, and buffer area 2 is reused to save the normalization result in FP32.
[0142] The normalization result in FP32 is read from buffer area 2, and the normalization result is input to the first type conversion operator for data type conversion to obtain the intermediate conversion result in BF16, and the intermediate conversion result in BF16 is saved in buffer area 3 (32 TLRs) allocated for the first type conversion operator in the total buffer.
[0143] The intermediate conversion result in BF16 is read from buffer area 3, and the intermediate conversion result and the value block r in BF16 in GSM are input to the second matrix multiplication operator for matrix multiplication calculation to obtain the second calculation result in FP32; and buffer area 2 is reused to save the second calculation result in FP32.
[0144] The second calculation result in FP32 is read from buffer area 2, and the second calculation result and the calculation result of the (r - 1)-th inner loop traversal saved in buffer area 4 (64 TLRs) are input to the accumulation operator for accumulation to obtain the calculation result of the r-th inner loop traversal. Since in the 8 inner loop traversal processes, the calculation result obtained in each inner loop traversal process needs to be accumulated with the calculation result obtained in the previous inner loop traversal process, therefore, buffer area 4 is allocated for the accumulation operator in the total buffer. During the 8 inner loop traversal processes, buffer area 4 is always used to save the calculation result obtained in each inner loop traversal until the end of the 8 inner loop traversal processes. For details, see Figure 5 。
[0145] After obtaining the calculation result in FP32 of the 8th inner loop traversal, the calculation result of the 8th inner loop traversal is input to the second type conversion operator for data type conversion to obtain the target calculation result (data type is BF16) of the outer loop traversal round (the 1st outer loop traversal) where query block 1 is located. Buffer area 3 is reused to save the target calculation result in BF16, and then the target calculation result in BF16 is written from buffer area 3 to GSM.
[0146] In the embodiments of the present application, when performing Flash Attention calculation, both the first buffer and the second buffer are reused multiple times among different operators within the Flash Attention operator (i.e., the fusion operator). That is, in each round of the outer loop traversal process, the first buffer is respectively used to store the first calculation result output by the first matrix multiplication operator, the normalization result output by the normalization operator, and the second calculation result output by the second matrix multiplication operator. The second buffer is respectively used to store the intermediate conversion result output by the first type conversion operator and the target calculation result output by the second type conversion operator. This greatly improves the utilization rate of buffer resources, effectively alleviates the situation of tight buffer resource usage during the implementation of operators, avoids the problem of insufficient buffer resources, and at the same time improves the performance of Flash Attention calculation.
[0147] Based on the same technical concept, the embodiments of the present application provide a schematic structural diagram of a buffer resource reuse device, as Figure 6 shown. The buffer resource reuse device 600 includes:
[0148] A calculation module 601, configured to input a query representation and a key representation into a first matrix multiplication operator for matrix multiplication calculation to obtain a first calculation result; store the first calculation result in a first buffer area allocated for the first matrix multiplication operator in the on-chip cache;
[0149] A normalization module 602, configured to read the first calculation result from the first buffer area and input the first calculation result into a normalization operator for normalization operation to obtain a normalization result; reuse the first buffer area to store the normalization result, where the first matrix multiplication operator and the normalization operator are operators in a target fusion operator;
[0150] The calculation module 601 is further configured to obtain an output tensor by performing matrix multiplication calculation on the normalization result and a value representation.
[0151] Optionally, the calculation module 601 is specifically configured to:
[0152] Read the normalization result from the first buffer area for data type conversion to obtain an intermediate conversion result, where the data type of the intermediate conversion result is the same as the data type of the value representation;
[0153] Store the intermediate conversion result in a second buffer area in the on-chip cache;
[0154] Read the intermediate conversion result from the second buffer area and perform matrix multiplication calculation with the value representation to obtain an output tensor.
[0155] Optionally, the calculation module 601 is specifically configured to:
[0156] Read the intermediate conversion result from the second buffer, perform matrix multiplication calculation with the value representation to obtain a second calculation result; reuse the first buffer to save the second calculation result;
[0157] Read the second calculation result from the first buffer for data type conversion to obtain an output tensor, where the data type of the output tensor is the same as the data type of the value representation; reuse the second buffer to save the output tensor.
[0158] Optionally, the query representation is one query block among multiple query blocks traversed by an outer loop, the key representation is one key block among multiple key blocks traversed by an inner loop corresponding to the one query block, the value representation is one value block among multiple value blocks traversed by an inner loop corresponding to the one query block, the output tensor is the calculation result of the round of inner loop traversal where the one key block and the one value block are located, and the one query block corresponds to multiple rounds of inner loop traversal.
[0159] Optionally, the calculation module 601 is specifically configured to:
[0160] Read the normalization result from the first buffer and perform matrix multiplication calculation with the value representation to obtain a second calculation result; reuse the first buffer to save the second calculation result;
[0161] Read the second calculation result from the first buffer and accumulate it with the calculation result obtained in the previous round of inner loop traversal to obtain the output tensor.
[0162] Optionally, the calculation module 601 is specifically configured to:
[0163] Read the normalization result from the first buffer for data type conversion to obtain an intermediate conversion result, where the data type of the intermediate conversion result is the same as the data type of the value representation;
[0164] Save the intermediate conversion result in the second buffer in the on-chip cache;
[0165] Read the intermediate conversion result from the second buffer and perform matrix multiplication calculation with the value representation to obtain the second calculation result.
[0166] Optionally, the calculation module 601 is further configured to:
[0167] After reading the second calculation result from the first buffer and accumulating it with the calculation result obtained from the previous round of inner loop traversal to obtain the output tensor, when the output tensor is the calculation result obtained from the last round of inner loop traversal in the multiple rounds of inner loop traversal, perform a data type conversion on the output tensor to obtain the target calculation result for the outer loop traversal round where the one query block is located, and the data type of the target calculation result is the same as the data type represented by the value.
[0168] Optionally, the calculation module 601 is further configured to:
[0169] Reuse the second buffer to save the target calculation result.
[0170] Optionally, the calculation module is specifically configured to:
[0171] Read the second calculation result from the first buffer; and obtain the calculation result obtained from the previous round of inner loop traversal from the third buffer in the on-chip cache;
[0172] Accumulate the second calculation result and the calculation result obtained from the previous round of inner loop traversal to obtain an output tensor;
[0173] Reuse the third buffer to save the output tensor.
[0174] Optionally, the calculation module 601 is further configured to:
[0175] Save the one query block in the fourth buffer in the on-chip cache until the execution of the multiple rounds of inner loop traversal corresponding to the one query block ends.
[0176] In the embodiments of the present application, the query representation and the key representation are input into the first matrix multiplication operator for matrix multiplication calculation to obtain the first calculation result; then the first calculation result is saved in the first buffer allocated for the first matrix multiplication operator in the on-chip cache. Read the first calculation result from the first buffer and input the first calculation result into the normalization operator for normalization operation to obtain the normalization result. Since the data type of the first calculation result is the same as that of the normalization result, and after obtaining the normalization result, the first calculation result is no longer required for subsequent attention mechanism calculations, and the first matrix multiplication operator and the normalization operator are operators in the same target fusion operator, therefore, the first buffer that originally saved the first calculation result can be reused to save the normalization result, which not only ensures the accuracy of the attention mechanism calculation, but also realizes the reuse of cache resources between different operators in a fusion operator, thereby improving the utilization rate of cache resources and effectively alleviating the situation of tight use of cache resources during the implementation of operators.
[0177] Based on the same inventive concept, an embodiment of the present application provides a computer device, such as Figure 7 shown, including at least one artificial intelligence chip 100 and a memory 701 connected to the at least one artificial intelligence chip 100. In the embodiments of the present application, the specific connection medium between the artificial intelligence chip 100 and the memory 701 is not limited. Figure 7 Taking the example that the artificial intelligence chip 100 and the memory 701 are connected through a bus. The bus can be divided into an address bus, a data bus, a control bus, etc.
[0178] In the embodiments of the present application, the memory 701 stores instructions executable by the at least one artificial intelligence chip 100. By executing the instructions stored in the memory 701, the at least one artificial intelligence chip 100 can perform the steps of the above-mentioned cache resource reuse method.
[0179] Among them, the artificial intelligence chip 100 is the control center of the computer device. It can use various interfaces and circuits to connect various parts of the computer device, and by running or executing the instructions stored in the memory 701 and calling the data stored in the memory 701, cache resource reuse can be achieved. Optionally, the artificial intelligence chip 100 may include one or more processing units. The artificial intelligence chip 100 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the artificial intelligence chip 100. In some embodiments, the artificial intelligence chip 100 and the memory 701 may be implemented on the same chip, and in some embodiments, they may also be separately implemented on independent chips.
[0180] The artificial intelligence chip 100 may be a general-purpose processor, such as a digital signal processor, an application specific integrated circuit (ASIC), a field programmable gate array, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, which can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor.
[0181] The memory 701 is a non-volatile computer-readable storage medium and can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. The memory 701 may include at least one type of storage medium. For example, it may include flash memory, hard disks, multimedia cards, card-type memories, random access memory (RAM), static random access memory (SRAM), programmable read-only memory (PROM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), magnetic memories, magnetic disks, optical disks, and so on. The memory 701 is any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer device, but is not limited thereto. The memory 701 in the embodiments of the present application may also be a circuit or any other device capable of implementing a storage function for storing program instructions and / or data.
[0182] Based on the same inventive concept, an embodiment of the present application provides a computer-readable storage medium storing a computer program executable by a computer device. When the computer program runs on the computer device, the computer device is caused to execute the steps of the above-mentioned cache resource reuse method.
[0183] Based on the same inventive concept, an embodiment of the present application provides a computer program product. The computer program product includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer device, the computer device is caused to execute the steps of the above-mentioned cache resource reuse method.
[0184] Those skilled in the art should understand that the embodiments of the present invention may be provided as a method, or a computer program product. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.
[0185] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general purpose computer, special purpose computer, embedded processor or other programmable data processing device to produce a machine, such that the instructions executed by the processor of the computer device or other programmable data processing device generate means for implementing the specified functions in one flow Figure 1 one flow or more flows and / or blocks Figure 1 or means for implementing the specified functions in one block or more blocks.
[0186] These computer program instructions can also be stored in a computer-readable memory that can direct a computer device or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the specified functions in one flow Figure 1 one flow or more flows and / or blocks Figure 1 or means for implementing the specified functions in one block or more blocks.
[0187] These computer program instructions can also be loaded onto a computer device or other programmable data processing device, such that a series of operation steps are executed on the computer device or other programmable device to produce a process implemented by the computer device, so that the instructions executed on the computer device or other programmable device provide steps for implementing the specified functions in one flow Figure 1 one flow or more flows and / or blocks Figure 1 or means for implementing the specified functions in one block or more blocks.
[0188] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they know the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present invention.
[0189] Obviously, those skilled in the art can make various changes and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. A cache resource reuse method, characterized in that, Including: Loading query representations, key representations, and value representations from video memory into the group shared memory in the on-chip cache; The query representation is one query chunk among multiple query chunks traversed by an outer loop, the key representation is one key chunk among multiple key chunks traversed by an inner loop corresponding to the one query chunk, and the value representation is one value chunk among multiple value chunks traversed by the inner loop corresponding to the one query chunk; Loading the query representation from the group shared memory into the first number of thread-local registers in the on-chip cache until the execution of multiple rounds of inner loop traversal corresponding to the one query representation ends; Inputting the query representation and the key representation into a first matrix multiplication operator for matrix multiplication calculation to obtain a first calculation result; storing the first calculation result in a first buffer area allocated for the first matrix multiplication operator in the total buffer area, the total buffer area being the buffer resources allocated in the on-chip cache for the target fusion operator, the first matrix multiplication operator being an operator in the target fusion operator, and the first buffer area being the second number of thread-local registers; Reading the first calculation result from the first buffer area and inputting the first calculation result into a normalization operator for normalization operation to obtain a normalization result; When it is determined that the data type of the normalization result is the same as the data type of the first calculation result, the normalization operator is an operator in the target fusion operator, and the first calculation result is no longer used, reusing the first buffer area to store the normalization result; Reading the value representation from the group shared memory and performing matrix multiplication calculation with the normalization result to obtain an output tensor; the output tensor is the calculation result of the round of inner loop traversal where the one key chunk and the one value chunk are located, and is stored in a third buffer area of the on-chip cache allocated separately.
2. The method according to claim 1, wherein The step of reading the value representation from the group shared memory and performing matrix multiplication calculation with the normalization result to obtain an output tensor includes: Reading the normalization result from the first buffer area and performing matrix multiplication calculation with the value representation to obtain a second calculation result; reusing the first buffer area to store the second calculation result; Reading the second calculation result from the first buffer area and accumulating it with the calculation result obtained in the previous round of inner loop traversal to obtain the output tensor.
3. The method according to claim 2, characterized in that, The step of reading the normalization result from the first buffer area and performing matrix multiplication calculation with the value representation to obtain a second calculation result includes: Reading the normalization result from the first buffer area for data type conversion to obtain an intermediate conversion result, the data type of the intermediate conversion result being the same as the data type of the value representation; Storing the intermediate conversion result in a second buffer area in the on-chip cache; Reading the intermediate conversion result from the second buffer area and performing matrix multiplication calculation with the value representation to obtain the second calculation result.
4. The method according to claim 3, wherein After the step of reading the second calculation result from the first buffer area and accumulating it with the calculation result obtained in the previous round of inner loop traversal to obtain the output tensor, further including: When the output tensor is the calculation result obtained in the last inner loop traversal of the multi-round inner loop traversal, perform a data type conversion on the output tensor to obtain the target calculation result of the outer loop traversal round where the one query block is located, and the data type of the target calculation result is the same as the data type represented by the value.
5. The method according to claim 4, wherein It further includes: Reusing the second buffer to store the target calculation result.
6. The method according to claim 3, characterized in that, The step of reading the second calculation result from the first buffer and accumulating it with the calculation result obtained in the previous inner loop traversal to obtain the output tensor includes: Reading the second calculation result from the first buffer; and obtaining the calculation result obtained in the previous inner loop traversal from the third buffer in the on-chip cache; Accumulating the second calculation result and the calculation result obtained in the previous inner loop traversal to obtain the output tensor; Reusing the third buffer to store the output tensor.
7. The method according to any one of claims 1 to 6, characterized in that, It further includes: Storing the one query block in the fourth buffer in the on-chip cache until the execution of the multi-round inner loop traversal corresponding to the one query block ends.
8. A computer device, comprising a memory, an artificial intelligence chip, and a computer program stored in the memory and executable on the artificial intelligence chip, characterized in that, When the artificial intelligence chip executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
9. A computer-readable storage medium, characterized in that, It stores a computer program executable by a computer device. When the computer program runs on the computer device, it causes the computer device to execute the steps of the method according to any one of claims 1 to 7.
10. A computer program product, characterized in that, The computer program product includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer device, they cause the computer device to execute the steps of the method according to any one of claims 1 - 7.
Citation Information
Patent Citations
Mixed precision floating point multiplication device and mixed precision floating point number processing method
CN116795324A
Attention operation processing method and device
CN118585249A