Attention mechanism fusion method and device based on data stream architecture accelerator

By integrating Attention calculation steps on the data flow architecture chip, using pre-incoming transposed data and block transmission technology, the problems of long instruction configuration time and large memory access overhead in large-scale model inference calculations are solved, and efficient Attention calculation is achieved.

CN119940434AActive Publication Date: 2025-05-06INST OF COMPUTING TECH CHINESE ACAD OF SCI

Patent Information

Application Number
CN202510009132.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-05-06
Estimated Expiration
2045-01-03

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently process large-scale inference calculations of different scales on data flow architecture chips, especially in Attention calculations, where there are problems such as long instruction information configuration time and large memory access overhead.

Method used

A fusion method based on data flow architecture is proposed. By selecting a fusion scheme based on the product of embedding in Attention and the input sequence length, transposed data is pre-passed to reduce configuration instruction time and memory access overhead, and the calculation steps of Attention are fused into two kernel functions with high degree of multiplexing.

Benefits of technology

It effectively reduces the instruction information configuration time and memory access overhead in Attention calculation, and improves the efficiency of data flow architecture chips in processing large-scale model inference calculations in different scales.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940434A_ABST
    Figure CN119940434A_ABST
Patent Text Reader

Abstract

The invention proposes an attention mechanism fusion method and device based on a data stream architecture accelerator, and the method comprises a method for accelerating Attention calculation on a GPDPU accelerator, the method selects a fusion scheme according to the product of the dimension of embedding in Attention and the length of an input sequence, and for the calculation with the small dimension, the fusion scheme is selected according to the product of the dimension of the embedding in the Attention and the length of the input sequence. All operations are fused in the same kernel function in a mode of pre-introducing transposed data, so that the time and memory access overhead of configuration instructions are reduced, input data blocks are introduced into a memory of a temporary storage data cache SPM for calculation for calculation with relatively large dimensions, and the calculation efficiency is improved. The Attention calculation step is fused into two kernel functions with very high multiplexing degree, so that the configuration time of instruction information is shortened.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence (AI) and high performance computing (HPC), and in particular to a method, device, electronic device, computer-readable storage medium and computer program product for fusion of an attention mechanism based on a data flow architecture. Background Art

[0002] Attention mechanism Attention is one of the core mechanisms of modern deep learning models (especially models in natural language processing, computer vision and multimodal tasks). In recent years, accelerating the Attention calculation in the Transformer model is an important area in deep learning research, especially in optimizing hardware efficiency and sequence length extension.

[0003] For example, Flash Attention is a specially optimized Attention implementation that relies on video memory access efficiency and computational flow optimization and is carefully designed for GPU hardware. It divides Attention calculations into small blocks, avoiding the writing and reading of the video memory of the complete Attention matrix, and reduces the waste of memory access time through operator fusion. However, it is mainly designed for GPUs and is difficult to port to data flow architecture chips. Sparse Attention reduces the computational complexity of traditional global Attention and reduces the complexity through sparse strategies, but its sparse mode needs to be designed according to specific tasks and data, and has poor versatility, and is also difficult to use on data flow chips DPU.

[0004] Nowadays, the embedding dimensions of Attention in large models range from 64 to 12288. Data flow chips need to be able to handle reasoning of large models of different sizes. It takes a lot of time for data flow chips to configure instruction information for each kernel function, and some calculations in Attention require matrix transposition, which may require multiple data transmissions and consume a considerable amount of memory access time. At the same time, the softmax nonlinear calculation in Attention requires special support from acceleration hardware. Although its computational complexity is only a small part compared to matrix multiplication, it requires multiple transmissions of all input values, which may incur huge overhead if not handled properly.

[0005] In view of this, there is an urgent need to provide a method that can handle Attention calculations of various scales, reuse kernel functions as much as possible to reduce instruction information configuration time, reduce unnecessary memory access overhead through fusion operations, and provide a method of transposing data that does not consume much memory to match the matrix multiplication and softmax calculations in Attention. Summary of the invention

[0006] The present invention is a GPDPU accelerator based on a data flow architecture. The data flow architecture processor GPDPU accelerator includes: main memory DRAM, data cache SPM, and PE array. In order to solve the above-mentioned multiple technical problems, namely (1) the GPDPU accelerator needs to be able to handle large-scale model inference calculations of different scales, (2) it takes a lot of time to configure instruction information for each kernel function in the GPDPU accelerator, and it is necessary to reduce the proportion of instruction information configuration time in attention calculation, (3) some calculations in Attention require matrix transposition, which may require multiple data transmissions, which will consume a considerable amount of memory access time, and it is necessary to reduce this memory access time. Ming proposed a method for accelerating Attention calculation on the GPDPU accelerator. This method selects a fusion scheme based on the product of the dimension of the embedding in Attention and the length of the input sequence. For calculations with smaller dimensions (such as 768), all operations are fused into the same kernel function by pre-transposing the data, thereby reducing the time for configuring instructions and memory access overhead. For calculations with larger dimensions (such as 4096), the input data is divided into blocks and passed into the memory of the temporary data cache SPM for calculation. The calculation steps of Attention are fused into two highly reused kernel functions to reduce the configuration time of instruction information.

[0007] In view of the shortcomings of existing technologies, such as Figure 6 As shown, the present invention proposes an attention mechanism fusion method based on a data flow architecture accelerator, which includes:

[0008] In the initial step, a data flow architecture accelerator for performing attention calculation is obtained, where the data flow architecture accelerator includes a main memory, a cache, and a processing unit; when the length of the input sequence matrix multiplied by the embedding dimension is lower than a threshold, a first fusion step is performed;

[0009] In the first fusion step, the input sequence matrix, the transposed matrix of the input sequence, the transposed matrix of the weight matrix Wq, the weight matrix Wk, and the transposed matrix Wv are transferred from the main memory to the cache, and the transposed matrix of the query matrix Q, the key matrix K, and the transposed matrix of the value matrix V are obtained by matrix multiplication. The processing unit then performs matrix multiplication to calculate the multiplication of the key matrix K and the transposed matrix of the query matrix Q to obtain the transposed matrix of S; the transposed matrix of S is sequentially divided by bit and the column-by-column softmax activation is calculated to obtain the transposed matrix of P; finally, the transposed matrix of the value matrix V is matrix multiplied with the transposed matrix of P to obtain the transposed matrix of O, and finally the transposed matrix of O is transposed and transferred back to the main memory to obtain the O matrix as the attention calculation result.

[0010] The attention mechanism fusion method based on the data flow architecture accelerator, wherein the initial step includes performing a second fusion step when the length of the input sequence matrix multiplied by the embedding dimension is greater than or equal to the threshold:

[0011] In the first block step, the input sequence is divided into blocks according to the length dimension of the input sequence to obtain subsequences, and the weight matrix W and the subsequence are written into the cache; a matrix T with a size of the subsequence length and the length of the input sequence is passed in. If the weight matrix W used to calculate this time is a K matrix or a V matrix, the matrix T is a 0 matrix; if it is used to calculate the query matrix Q and the S matrix this time, the T matrix is ​​the transpose of the K matrix;

[0012] The first calculation step is to pass the weight matrix W, the subsequence and the T matrix into the cache. When calculating the K or V matrix, the T matrix passed in is a 0 matrix. When calculating the Q matrix, the T matrix passed in is the transposed matrix of the K matrix. The weight matrix W and the subsequence are matrix multiplied to obtain the Q or K or V matrix. The K or V matrix is ​​multiplied with the matrix T to obtain an S matrix with all zeros. The K or V matrix block obtained this time and the S matrix with all zeros are passed out to the main memory. When calculating the Q matrix, Q and the matrix are multiplied. The S matrix block is obtained by multiplying the T matrix, and the Q matrix block obtained this time and the transposed S matrix block are transferred to the main memory; if the input sequence length is n, the first calculation step will be repeated 3×n / 64 times, the first n / 64 times are used to calculate the K matrix, the middle n / 64 times are used to calculate the V matrix, and the last n / 64 times are used to calculate the Q matrix and the S matrix. Only the non-zero S matrix obtained in the last n / 64 cycles is needed, and the zero S matrix obtained in the first 2×n / 64 cycles will be discarded. Repeat the first calculation step until the complete transposed S matrix is ​​stored in the main memory;

[0013] The second block division step is to divide the complete S matrix after transposition into intermediate matrix blocks and pass them into the cache, and at the same time pass in the transposed matrix of V;

[0014] The second calculation step is to receive the intermediate matrix block, perform bitwise division on it, and then perform column-wise softmax calculation to obtain the transposed matrix of P, and then multiply it with the transposed matrix of V to obtain a part of the transposed matrix of the O matrix and transpose it again and transfer it to the main memory. Repeat the second calculation step until the complete O matrix is ​​obtained in the main memory as the attention calculation result.

[0015] The attention mechanism fusion method based on the data flow architecture accelerator, wherein the threshold is the capacity of the cache.

[0016] The attention mechanism fusion method based on the data flow architecture accelerator, wherein the data flow architecture accelerator is a data flow chip DPU.

[0017] like Figure 7 As shown, the present invention also proposes an attention mechanism fusion device based on a data flow architecture accelerator, which includes:

[0018] The initial module obtains a data flow architecture accelerator for performing attention calculation, wherein the data flow architecture accelerator includes a main memory, a cache, and a processing unit; when the length of the input sequence matrix multiplied by the embedding dimension is lower than a threshold, the first fusion module is executed;

[0019] The first fusion module transfers the input sequence matrix, the transposed matrix of the input sequence, the transposed matrix of the weight matrix Wq, the weight matrix Wk, and the transposed matrix Wv from the main memory to the cache, and obtains the transposed matrix of the query matrix Q, the key matrix K, and the transposed matrix of the value matrix V through matrix multiplication. The processing unit then performs matrix multiplication to calculate the multiplication of the key matrix K and the transposed matrix of the query matrix Q to obtain the transposed matrix of S; the transposed matrix of S is sequentially divided by bit and the column-by-column softmax activation is calculated to obtain the transposed matrix of P, and finally the transposed matrix of the value matrix V is matrix multiplied with the transposed matrix of P to obtain the transposed matrix of O, and finally the transposed matrix of O is transposed and transferred back to the main memory to obtain the O matrix as the attention calculation result.

[0020] The attention mechanism fusion device based on the data flow architecture accelerator, wherein the initial module includes executing the first fusion module when the length of the input sequence matrix multiplied by the embedding dimension is greater than or equal to the threshold, that is, when the storage space occupied in the calculation process of the first fusion module does not exceed the cache capacity, and formulating it as The input sequence length is x, the embedding dimension is d, the number of attention heads is h, the number of bytes occupied by a single data is b, the cache capacity is C bytes, and the second fusion module is executed:

[0021] The first block module divides the input sequence into blocks according to the length dimension of the input sequence to obtain subsequences, writes the weight matrix W and the subsequence into the cache; passes in a matrix T with a size of the subsequence length and the length of the input sequence. If the weight matrix W used to calculate this time is a K matrix or a V matrix, the matrix T is a 0 matrix; if it is used to calculate the query matrix Q and the S matrix, the T matrix is ​​the transpose of the K matrix;

[0022] The first calculation module passes the weight matrix W, the subsequence and the T matrix into the cache. When calculating the K or V matrix, the T matrix passed in is a 0 matrix. When calculating the Q matrix, the T matrix passed in is the transposed matrix of the K matrix. The weight matrix W and the subsequence are matrix multiplied to obtain the Q or K or V matrix. The K or V matrix is ​​multiplied by the matrix T to obtain an S matrix with all zeros. The K or V matrix block obtained this time and the S matrix with all zeros are passed out to the main memory. When calculating the Q matrix, Q and the matrix are multiplied. The S matrix block is obtained by multiplying the T matrix, and the Q matrix block obtained this time and the transposed S matrix block are transferred to the main memory; if the input sequence length is n, the first calculation module will be repeated 3×n / 64 times, the first n / 64 times are used to calculate the K matrix, the middle n / 64 times are used to calculate the V matrix, and the last n / 64 times are used to calculate the Q matrix and the S matrix. Only the non-zero S matrix obtained in the last n / 64 cycles is needed, and the zero S matrix obtained in the first 2×n / 64 cycles will be discarded. Repeat the first calculation module until the complete transposed S matrix is ​​stored in the main memory;

[0023] The second block splitting module divides the complete S matrix after transposition into intermediate matrix blocks and passes them into the cache, and at the same time passes the transposed matrix of V;

[0024] The second calculation module receives the intermediate matrix block, performs bitwise division on it, and then performs column-wise softmax calculation to obtain the transposed matrix of P, and then multiplies it with the transposed matrix of V to obtain a part of the transposed matrix of the O matrix and transpose it again and transfer it to the main memory. The second calculation module is repeatedly executed until the complete O matrix is ​​obtained in the main memory as the attention calculation result.

[0025] The attention mechanism fusion device based on the data flow architecture accelerator, wherein the threshold is the capacity of the cache; the data flow architecture accelerator is a data flow chip DPU.

[0026] The present invention proposes an electronic device, including an attention mechanism fusion device based on a data flow architecture accelerator, and the electronic device may be connected to an information display device, which is used to display the attention calculation results according to display parameters and attributes set by the user or through an artificial intelligence model.

[0027] The present invention proposes a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of the attention mechanism fusion method based on the data flow architecture accelerator are implemented.

[0028] The present invention proposes a computer program product, including a computer program, wherein when the computer program is executed by a processor, the steps of the attention mechanism fusion method based on the data flow architecture accelerator are implemented.

[0029] It can be seen from the above scheme that the advantages of the present invention are:

[0030] For attention calculations with a small product of the embedding dimension and the input sequence length, the five-step attention calculation can be fused into one operator to maximize the use of SPM storage space and reduce the consumption of memory access time and instruction configuration time. For attention calculations with a large product of the embedding dimension and the input sequence length, the input data is passed in blocks to save SPM storage space. The five-step attention calculation can be fused into two operators, and the softmax calculation in the second operator can be accelerated to the maximum extent by adjusting the arrangement of the output data of the first operator. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 This is a schematic diagram of the overall structure of the GPDPU accelerator;

[0032] Figure 2 This is a schematic diagram of the standard implementation of Attention calculation on GPDPU;

[0033] Figure 3 This is a schematic diagram of merging five kernel functions into one kernel function according to solution 1 of the present invention;

[0034] Figure 4 Schematic diagram of the first kernel function of solution 2 of the present invention;

[0035] Figure 5 This is a schematic diagram of the second kernel function of the second solution of the present invention;

[0036] Figure 6 is a flow chart of the method of the present invention;

[0037] Figure 7 It is a module diagram of the device of the present invention;

[0038] Figure 8 This is a schematic diagram of the structure of a first electronic device of the present invention;

[0039] Fig. 9 This is a schematic diagram of the application environment structure of the first electronic device of the present invention;

[0040] Fig.10 It is a schematic diagram of the structure of a second electronic device of the present invention.

[0041] Reference numerals:

[0042] A-First electronic device;

[0043] B-Attention mechanism fusion device based on data flow architecture accelerator;

[0044] C-data acquisition equipment;

[0045] D-information display device;

[0046] 1000 - second electronic device;

[0047] Ⅰ-computational unit;

[0048] II-ROM;

[0049] III-RAM;

[0050] IV-bus;

[0051] V-interface;

[0052] VI - input unit;

[0053] VII-output unit;

[0054] VIII- Storage medium;

[0055] Ⅸ-Communication unit. DETAILED DESCRIPTION

[0056] It should be noted that, in the present invention, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device.

[0057] Without more constraints, an element defined by the phrase "comprising a..." does not exclude the existence of other identical elements in the process, method, article or apparatus comprising the element.

[0058] The processor described in the present invention is the control center of the electronic device, which can be a processor or a general term for multiple processing elements. For example, it can be one or more central processing units (CPUs), or application specific integrated circuits (ASICs), or one or more integrated circuits configured to implement the embodiments of the present invention, such as one or more microprocessors (digital signal processors, DSPs), or one or more field programmable gate arrays (FPGAs).

[0059] Optionally, the processor can perform various functions of the electronic device by running or executing a software program stored in the memory, and calling data stored in the memory.

[0060] In a specific implementation, as an embodiment, the processor may include one or more CPUs. Each of these processors may be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). The processor here may refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions). Electronic devices may include: servers, desktop computers, laptops, smart phones, tablet computers, embedded computers, etc., wherein the embedded computers include vehicles and robots, etc.

[0061] The memory is used to store the software program for executing the solution of the present invention, and the execution is controlled by the processor. The specific implementation method can refer to the above method embodiment, which will not be repeated here.

[0062] It should be noted that the structure of the electronic device shown in the drawings of the present invention does not constitute a limitation thereto, and the actual knowledge structure recognition device may include more or fewer components than shown in the drawings, or a combination of certain components, or a different arrangement of components.

[0063] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware or any other combination. When implemented by software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center by wired (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more available media sets. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a tape), an optical medium (for example, a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state hard disk.

[0064] It should also be understood that the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. A and B can be singular or plural. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship, but it may also indicate an "and / or" relationship. Please refer to the context for specific understanding.

[0065] In the present invention, "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can be represented by: a, b, c, ab, ac, bc, or abc, where a, b, c can be single or multiple.

[0066] It should also be understood that in various embodiments of the present invention, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0067] In the several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of the units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0068] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0069] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0070] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage media include: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, etc., various media that can store program codes.

[0071] In order to achieve the above technical effects, the present invention proposes the following key technical points:

[0072] Key point 1: The Attention fusion scheme to be adopted is determined based on the product of the embedding dimension and the input sequence length. Technical effect: For smaller Attention calculations, the calculations can be integrated to the greatest extent, making full use of the storage space of SPM and reducing the round-trip transmission time of intermediate calculation results. For larger Attention calculations, five kernel functions are fused into two by reusing kernel functions.

[0073] Key point 2: By multiplexing the kernel functions, the five kernel functions are integrated to reduce the proportion of instruction information configuration time and memory access time. The technical effect is to reduce the time spent on transferring intermediate computing data back and forth between the main memory and SPM, and to minimize the time for transmitting instruction information. By transferring the input sequence into SPM in blocks, the storage pressure of SPM is greatly reduced.

[0074] Key point 3: Both solutions use the transposed data transmission method. The small model solution uses transposed transmission because the matrix calculation of Attention and softmax calculations require matrix transposition. By passing in the transposed input sequence (the input sequence, such as the vector feature matrix obtained after the text sequence is encoded, or the vector feature sequence obtained after the image is processed) and the weight matrix (the weight parameter matrix obtained during the model training process, W q / W k / W v ) can directly calculate the required transposed matrix without transmitting the intermediate results, which can maximize the degree of fusion; the larger model solution uses transposed transmission in two places. The first one is also to make the kernel function more integrated, and the second one is to cooperate with the subsequent softmax calculation to fully utilize the acceleration effect of SIMD parallelization.

[0075] The present invention accelerates the overall calculation time of Attention on the DPU, mainly by reducing the time spent on data transmission between the main memory and SPM during the Attention calculation process through operator fusion, and the time spent on these calculation processes themselves is not reduced. In the calculation process of attention, the intermediate calculation results need to be transposed. According to the conventional method, they need to be transposed from the SPM to the main memory and then transferred back to the SPM from the main memory. The way to transpose the matrix on the DPU is to call the transfer function between the main memory and the SPM. This method will spend more time on transmitting data, and it will also take a lot of time to re-input the instruction information into the SPM. The method mentioned in the present invention of directly calculating the transposed matrix of the intermediate result by inputting the transposed matrix of the input sequence can avoid such time consumption.

[0076] In order to make the above features and effects of the present invention more clearly and understandably described, embodiments are given below and described in detail with reference to the accompanying drawings. This specification discloses one or more embodiments that include the features of the present invention. The disclosed embodiments are only for illustration. The scope of protection of the present invention is not limited to the disclosed embodiments, and the present invention is defined by the attached claims.

[0077] The present invention accelerates the attention mechanism calculation, and optimizes the problems that the GPDPU accelerator needs to process attention mechanism calculations of different scales, the time-consuming transmission of instruction information in the GPDPU accelerator, and the long memory access time required for data transposition.

[0078] In order to make the above features and effects of the present invention more clearly understood, actual examples are given below and detailed descriptions are given in conjunction with the accompanying drawings.

[0079] like Figure 1 As shown, the overall structure of the GPDPU accelerator used in the present invention. GPDPU is a coarse-grained data flow structure accelerator, which uses the data flow concept and adopts a flow control combined execution mode internally. Traditional control flow execution accelerators drive execution with instructions, while in this structure, the data flow direction of the data flow graph in the calculation is used to drive execution. Compared with traditional structures, this structure is more suitable for accelerating scientific calculations with high parallelism. The GPDPU core consists of six parts: main memory, a microcontroller Micro Controller, a 4*4 running array PE array, an instruction cache Cbuf and a data cache SPM, and a transmission network. The main memory is used to store configuration information, instructions, data, etc. transmitted from the CPU; the instruction cache Cbuf is used to cache the instructions to be executed by the execution unit PE, which are transmitted from the main memory; the data cache SPM is used to cache the input data, weight data, and calculated output data and intermediate result data required by the execution unit; the 4*4 PE array completes the entire calculation process; the transmission network is responsible for data transmission between the main memory and the cache, and between the cache and the PE array. The execution array consists of 16 PEs, which are interconnected through a mesh network. Each PE contains an instruction cache, a register stack, a four-stage pipeline unit, a router, and a microcontroller inside the PE. The microcontroller is used to configure the instruction information for each of the 16 PEs, control the execution order of the PEs, and transfer the results back to the main memory after the calculation is completed.

[0080] like Figure 2 As shown, this is a schematic diagram of the standard implementation of Attention calculation on GPDPU. In this standard implementation, for example Figure 2As shown on the left, five kernel functions need to be called, namely: matrix multiplication 1 (the matrix scale of the multiplication is [input_length, embedding_size] and [embedding_size, embedding_size / h]), matrix multiplication 2 (the matrix scale of the multiplication is [input_length, embedding_size / h] and [embedding_size / h, input_length]), bitwise division, row softmax, matrix multiplication 3 (the matrix scale of the multiplication is [input_length, input_length] and [input_length, embedding_size / h]), where input_length is the length of the input sequence, i.e. the number of characters. number, embedding_size is the dimension of each input character after being converted to a vector representation, h represents the use of the h-head attention mechanism, each kernel function depends on the previous calculation results, and the data dependency determines that the five kernel functions need to be executed serially. After each kernel function is calculated, the calculation result needs to be transferred from the SPM back to the main memory, and then from the main memory to the SPM for the calculation of the next kernel function. Therefore, memory access is one of the bottlenecks for accelerating attention calculations. If you can use registers to save intermediate results and omit memory access under the premise of ensuring data dependency, or store the intermediate results back in the SPM, a storage device with lower consumption, instead of the main memory, you can save or reduce the storage and access overhead of this part of the intermediate result data. This takes advantage of the data storage characteristics of main memory access overhead > SPM access overhead > inter-PE memory access overhead > PE register access overhead in GPDPU.

[0081] The implementation steps of the present invention are introduced below. The input sequence is actually a matrix whose shape is the input sequence length × embedding dimension. First, determine whether the product of the embedding dimension and the input sequence length is less than If the conditions are met, then option 1 is enabled; otherwise, option 2 is enabled.

[0082] Scenario 1:

[0083] The above five kernel functions are combined into one, such as Figure 3As shown in the figure, first, the input sequence matrix, the transposed matrix of the input sequence, the transposed matrix of the Wq weight matrix, the transposed matrix of the Wk weight matrix, and the transposed matrix of the Wv weight matrix are transferred from the main memory to the SPM, and the transposed matrix of the query matrix (Q matrix), the key matrix (K matrix), and the transposed matrix of the value matrix (V matrix) are calculated by matrix multiplication 1. Then, the transposed matrix of K and Q is multiplied by matrix multiplication 2 to obtain the transposed matrix of S (similarity matrix), and then the bitwise division and column-wise softmax activation function are performed on them in turn to obtain the transposed matrix of P. Finally, the transposed matrix of V and the transposed matrix of P (attention weight matrix) are matrix multiplied by 3 to obtain the transposed matrix of O. Finally, the transposed matrix of O (output matrix) is transposed and transferred back to the main memory to obtain the O matrix. The O matrix is ​​the output obtained by the standard Attention implementation. Therefore, the correct calculation result can be obtained through Scheme 1.

[0084] This solution only requires one kernel function, that is, it only needs to transfer instruction information from the main memory to the SPM once, which can greatly save the overall time required for Attention.

[0085] This solution only transfers data from the main memory to the SPM once. The following calculates the difference in the amount of data transferred between this solution and the standard implementation.

[0086] The total amount of data transmitted by the standard implementation of Attention (single head):

[0087] 1) Pass in the input sequence: input_length*embedding_size

[0088] 2) Pass in Wq matrix, Wk matrix, Wv matrix: embedding_size*embedding_size / h*3

[0089] 3) Output K, Q, V matrices: input_length*embedding_size / h*3

[0090] 4) Pass in the Q matrix and the transposed matrix of K: input_length*embedding_size / h*2

[0091] 5) Output and then input S matrix: input_length*input_length*2

[0092] 6) Output and then input S' matrix: input_length*input_length*2

[0093] 7) Pass out and then pass in the P matrix, pass in the transposed matrix of V: ​​input_length*input_length*2+input_length*embedding_size / h

[0094] 8) Output O matrix: input_length*embedding_size / h

[0095] Total 6*input_length^2+(7 / h+1)*input_length*embedding_size+3 / h*embedding_size^2

[0096] The total amount of data transmitted by Attention solution 1 (single head):

[0097] 1) Input sequence and transpose of input sequence: input_length*embedding_size*2

[0098] 2) Pass in the transpose of the Wq weight matrix, the Wk weight matrix, and the Wv weight matrix: embedding_size*embedding_size / h*3

[0099] 3) Output O matrix: input_length*embedding_size / h

[0100] Total(2+1 / h)*input_length*embedding_size+3 / h*embedding_size^2

[0101] It can be seen that compared with the standard implementation, this solution can save 6*input_length^2+(6 / h-1)input_length*embedding_size data transmission time for each head calculation.

[0102] This solution directly calculates the transposed matrix of S, eliminating the step of transferring the S matrix back to the main memory and then transposing it to the SPM, making it possible to merge the Attention calculation into one operator on the GPDPU.

[0103] Scenario 2:

[0104] When the product of the embedding dimension and the input sequence length is less than 2^20, the above five kernel functions can be fused into two. At this time, the matrix multiplication calculation in Attention requires a large storage space. For example, for input_length = embedding_size = 4096, the calculated data type is fp16 and there are 32 attention heads. The maximum amount of data that needs to be passed into SPM in the standard implementation is 35MB, which is much larger than the capacity of SPM. Therefore, the input sequence needs to be passed in blocks.

[0105] 1) Kernel function 1

[0106] Kernel function 1 is Figure 4 As shown, Figure 4 The left parts of the two figures above and below are different, indicating that they calculate different subsequences of the input sequence.

[0107] Data block and input: The weight matrix of a single head is small enough and therefore not block-based. The input_length dimension of the input sequence is block-based in units of 64, and the embedding_size dimension is not divided. The weight matrix W (Wq or Wk or Wv matrix) and the block-based input sequence are passed into SPM. The size of the weight matrix is ​​[embedding_size, length_per_head], length_per_head is generally 64 or 128, and the size of the block-based input sequence is [64, embedding_size]. In addition, a matrix T of size [length_per_head, input_length] must be passed in. Matrix T is actually the transpose of the K matrix used to obtain the S matrix (and is also calculated by kernel function 1). If the kernel function is used to calculate the K or V matrix, the matrix T passed in is a 0 matrix. If the kernel function is used to calculate the Q matrix and the S matrix, the T matrix passed in should be the transpose of the K matrix.

[0108] The input_length dimension of the input sequence is divided into blocks of 64. The number 64 is also a relatively small number selected in consideration of the capacity limitation of SPM. According to the embedding_size and the number of attention heads of the current mainstream large model, it is also set to 128. The present invention is not limited to this. In order to facilitate calculation in DPU, the optimal setting is a multiple of 64. Let embedding_size be d, the number of blocks be x, the number of attention heads be h, the number of bytes of the calculated data type be b, and the SPM capacity be C bytes, then That's it.

[0109] Calculation and result transmission: Matrix multiplication is performed on the weight matrix and the input sequence after slicing, that is, the matrix of size [64, embedding_size] is multiplied by the matrix of size [embedding_size, length_per_head] to obtain the Q or K or V matrix of size [64, length_per_head]. When the kernel function calculates the K or V matrix this time, the T matrix passed in is a 0 matrix. The K or V matrix is ​​multiplied by the 0 matrix to obtain a matrix of all 0s, and the K or V matrix block obtained this time and the S matrix of 0 are transmitted to the main memory; the Q matrix is ​​the same as the K / V matrix, which is obtained by matrix multiplication of the weight matrix and the input sequence after slicing. When the kernel function calculates the Q matrix this time, the T matrix passed in is the transpose of the K matrix. The transpose of the Q matrix and the K matrix is ​​multiplied to obtain part of the S matrix, and the Q matrix block and the S matrix block obtained this time are transmitted to the main memory. It should be noted that the S matrix block needs to be transposed and transmitted back to the main memory to match the subsequent softmax calculation.

[0110] The ultimate goal of kernel function 1 is to calculate the S matrix. It integrates the two kernel functions in the standard implementation of Attention when the input sequence must be divided into blocks. First, kernel function 1 is called to calculate the K matrix. The T matrix passed in is the 0 matrix. Each time kernel function 1 is called, a K matrix block of [64, length_per_head] is obtained. The kernel function is called input_length / 64 times to obtain the complete K matrix. After obtaining the complete K matrix, kernel function 1 is called to calculate the Q matrix. At this time, the T matrix passed in is the transposed matrix of the K matrix. Each time kernel function 1 is called, a Q matrix block of [64, length_per_head] and an S matrix of [64, input_length] are obtained. The kernel function 1 is called input_length / 64 times to obtain the complete Q matrix and the complete S matrix. The calculation process of matrix multiplication is to continuously accumulate on the same S matrix. Each time the kernel function 1 is called, a subsequence of a different input series is calculated. When all subsequences of the input sequence are calculated once, the complete S matrix will be obtained.

[0111] 2) Kernel function 2

[0112] Kernel function 2 is Figure 5 As shown, Figure 5 The left parts of the two figures above and below are different, and the input is a different part of the ST matrix.

[0113] Data block division and input: After kernel function 1 calculates the complete transpose of the S matrix of size [input_length, input_length] multiple times, it is divided into matrix blocks of [input_length, 256] and passed to SPM. In addition, the transposed matrix of V needs to be passed in. Here, the transposed matrix of S is divided into 256 by column to facilitate the use of SIMD (single instruction multiple data) to parallelize the softmax calculation.

[0114] Calculation and result transmission: First, perform bitwise division on the transposed matrix of S, then perform column-wise softmax calculation on it to obtain the transposed matrix of P, and multiply it with the transposed matrix of V passed in in advance, that is, multiply the matrix of size [length_per_head, input_length] with the matrix of size [input_length, 256] to obtain a part of the transposed matrix of O of size [length_per_head, 256]. Finally, transpose and transmit the partial transposed matrix of O to obtain the partial matrix of O.

[0115] Each call to kernel function 2 gets a block of O matrix of [256, length_per_head], and kernel function 2 is used input_length / 256 times to get the complete O matrix. The complete O matrix is ​​[input_length, length_per_head], and each calculation gets the [256, length_per_head] part, so calculating input_length / 256 will get the complete O matrix.

[0116] The following is a system embodiment corresponding to the above method embodiment. This embodiment can be implemented in conjunction with the above embodiment. The relevant technical details mentioned in the above embodiment are still valid in this embodiment. In order to reduce repetition, they are not repeated here. Accordingly, the relevant technical details mentioned in this embodiment can also be applied in the above embodiment.

[0117] like Figure 7 As shown, the present invention also proposes an attention mechanism fusion device based on a data flow architecture accelerator, which includes:

[0118] The initial module obtains a data flow architecture accelerator for performing attention calculation, wherein the data flow architecture accelerator includes a main memory, a cache, and a processing unit; when the length of the input sequence matrix multiplied by the embedding dimension is lower than a threshold, the first fusion module is executed;

[0119] The first fusion module transfers the input sequence matrix, the transposed matrix of the input sequence, the transposed matrix of the weight matrix Wq, the weight matrix Wk, and the transposed matrix Wv from the main memory to the cache, and obtains the transposed matrix of the query matrix Q, the key matrix K, and the transposed matrix of the value matrix V through matrix multiplication. The processing unit then performs matrix multiplication to calculate the multiplication of the key matrix K and the transposed matrix of the query matrix Q to obtain the transposed matrix of S; the transposed matrix of S is sequentially divided by bit and the column-by-column softmax activation is calculated to obtain the transposed matrix of P, and finally the transposed matrix of the value matrix V is matrix multiplied with the transposed matrix of P to obtain the transposed matrix of O, and finally the transposed matrix of O is transposed and transferred back to the main memory to obtain the O matrix as the attention calculation result.

[0120] The attention mechanism fusion device based on the data flow architecture accelerator, wherein the initial module includes when the length of the input sequence matrix multiplied by the embedding dimension is greater than or equal to the threshold (b× Execute the second fusion module:

[0121] The first block module divides the input sequence into blocks according to the length dimension of the input sequence to obtain subsequences, writes the weight matrix W and the subsequence into the cache; passes in a matrix T with a size of the subsequence length and the length of the input sequence. If the weight matrix W used to calculate this time is a K matrix or a V matrix, the matrix T is a 0 matrix; if it is used to calculate the query matrix Q and the S matrix, the T matrix is ​​the transpose of the K matrix;

[0122] The first calculation module passes the weight matrix W, the subsequence and the T matrix into the cache. When calculating the K or V matrix, the T matrix passed in is a 0 matrix. When calculating the Q matrix, the T matrix passed in is the transposed matrix of the K matrix. The weight matrix W and the subsequence are matrix multiplied to obtain the Q or K or V matrix. The K or V matrix is ​​multiplied by the matrix T to obtain an S matrix with all zeros. The K or V matrix block obtained this time and the S matrix with all zeros are passed out to the main memory. When calculating the Q matrix, Q and the matrix are multiplied. The S matrix block is obtained by multiplying the T matrix, and the Q matrix block obtained this time and the transposed S matrix block are transferred to the main memory; if the input sequence length is n, the first calculation module will be repeated 3×n / 64 times, the first n / 64 times are used to calculate the K matrix, the middle n / 64 times are used to calculate the V matrix, and the last n / 64 times are used to calculate the Q matrix and the S matrix. Only the non-zero S matrix obtained in the last n / 64 cycles is needed, and the zero S matrix obtained in the first 2×n / 64 cycles will be discarded. Repeat the first calculation module until the complete transposed S matrix is ​​stored in the main memory;

[0123] The second block splitting module divides the complete S matrix after transposition into intermediate matrix blocks and passes them into the cache, and at the same time passes the transposed matrix of V;

[0124] The second calculation module receives the intermediate matrix block, performs bitwise division on it, and then performs column-wise softmax calculation to obtain the transposed matrix of P, and then multiplies it with the transposed matrix of V to obtain a part of the transposed matrix of the O matrix and transpose it again and transfer it to the main memory. The second calculation module is repeatedly executed until the complete O matrix is ​​obtained in the main memory as the attention calculation result.

[0125] The attention mechanism fusion device based on the data flow architecture accelerator, wherein the threshold is the capacity of the cache; the data flow architecture accelerator is a data flow chip DPU.

[0126] like Figure 8 As shown, the present invention further proposes a first electronic device A in another embodiment, including the attention mechanism fusion device based on the data flow architecture accelerator.

[0127] like Fig. 9 As shown, the first electronic device A can also be connected to the data acquisition device C and the information display device D through a wired or wireless information transmission scheme. The data acquisition device C is used to collect the vector feature matrix obtained after the text sequence is encoded, or the vector feature sequence obtained after the image is processed. The information display device D is used to display the attention calculation results or text semantic recognition results or image classification results obtained by the analysis of the present invention.

[0128] The information display device D can sort and process the data output by the first electronic device A based on the information display mechanism to improve the readability of the data output by the first electronic device A. The information display mechanism can be manually preset, for example, the data output by the first electronic device A is visually displayed, which can be based on the display parameters and / or attributes set by the user. The display parameters can be, for example, the display data range, and the display attributes can be, for example, the display font, color, whether to scroll and play, etc. The user is presented with the key information specified by the user, and the user can understand the information more promptly without having to access the secondary page or scroll the page, saving the user's operation. Or the information display mechanism can be an artificial intelligence AI display model, which can learn the user's key information based on the user's previous usage habits, such as viewing time, number of clicks, number of edits, etc., and then automatically present rich and necessary key information to the user.

[0129] The present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a readable storage medium. When the computer program is executed by a processor, the computer can execute the attention mechanism fusion method based on the data flow architecture accelerator provided by the above methods.

[0130] The present invention also proposes a storage medium VIII in another embodiment for storing a computer program for executing the attention mechanism fusion method based on the data flow architecture accelerator. It should be understood that the storage medium in the embodiment of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct RAM bus random access memory (DR RAM).

[0131] Fig.10 A schematic block diagram of a second electronic device 1000 that can be used to implement an embodiment of the present invention is shown. The second electronic device 1000 electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The second electronic device 1000 can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples, and are not intended to limit the implementation of the present invention described and / or required herein. The second electronic device 1000 may be the same or different from the first electronic device A.

[0132] The second electronic device 1000 includes a computing unit I, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory II (ROM) or a computer program loaded from a storage medium VIII into a random access memory (RAM) III. In RAM III, various programs and data required for the operation of the device 1000 can also be stored. The computing unit I, ROM II, and RAM III are connected to each other via a bus IV. An input / output (I / O) interface V is also connected to the bus IV.

[0133] A plurality of components in the second electronic device 1000 are connected to the I / O interface V, including: an input unit VI, such as a keyboard, a mouse, etc.; an output unit VII, such as various types of displays, speakers, etc.; a storage medium VIII, such as a disk, an optical disk, etc.; and a communication unit IX, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit IX allows the second electronic device 1000 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0134] The computing unit I may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of computing unit I include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit I performs the various methods and processes described above, such as method steps S1-S2. For example, in some embodiments, the method may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage medium VIII. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 1000 via ROM II and / or a communication unit IX. When the computer program is loaded into RAM III and executed by the computing unit I, one or more steps of the method described above may be performed. Alternatively, in other embodiments, the computing unit I may be configured to execute the method in any other appropriate manner (e.g., by means of firmware).

[0135] Although the embodiments of the present invention have been disclosed as above, they are not limited to the applications listed in the specification and the implementation modes, and they can be fully applied to various fields suitable for the present invention. For those familiar with the art, additional modifications can be easily implemented. Therefore, without departing from the general concept defined by the claims and the scope of equivalents, the present invention is not limited to the specific details and the illustrations shown and described herein.

Claims

1. A method for fusion of attention mechanism based on data flow architecture accelerator, characterized in that: include: An initial step is to obtain a data flow architecture accelerator for performing attention calculation, wherein the data flow architecture accelerator includes a main memory, a cache, and a processing unit; When the length of the input sequence matrix multiplied by the embedding dimension is lower than the threshold, the first fusion step is performed; In the first fusion step, the input sequence matrix, the transposed matrix of the input sequence, the transposed matrix of the weight matrix Wq, the weight matrix Wk, and the transposed matrix Wv are transferred from the main memory to the cache, and the transposed matrix of the query matrix Q, the key matrix K, and the transposed matrix of the value matrix V are obtained by matrix multiplication. The processing unit then performs matrix multiplication to calculate the multiplication of the key matrix K and the transposed matrix of the query matrix Q to obtain the transposed matrix of S; the transposed matrix of S is sequentially divided by bit and the column-by-column softmax activation is calculated to obtain the transposed matrix of P; finally, the transposed matrix of the value matrix V is matrix multiplied with the transposed matrix of P to obtain the transposed matrix of O, and finally the transposed matrix of O is transposed and transferred back to the main memory to obtain the O matrix as the attention calculation result.

2. The attention mechanism fusion method based on the data flow architecture accelerator according to claim 1, characterized in that: This initial step includes performing a second fusion step when the length of the input sequence matrix multiplied by the embedding dimension is greater than or equal to the threshold: The first block step is to block the input sequence according to the length dimension of the input sequence to obtain subsequences, and write the weight matrix W and the subsequences into the cache; Pass in a matrix T of the size of the subsequence length and the length of the input sequence. If the weight matrix W used to calculate this time is a K matrix or a V matrix, the matrix T is a 0 matrix; if it is used to calculate the query matrix Q and the S matrix, the T matrix is ​​the transpose of the K matrix; The first calculation step is to pass the weight matrix W, the subsequence and the T matrix into the cache. When calculating the K or V matrix, the passed T matrix is ​​a 0 matrix. When calculating the Q matrix, the passed T matrix is ​​the transposed matrix of the K matrix. The weight matrix W and the subsequence are matrix multiplied to obtain the Q or K or V matrix. The K or V matrix is ​​multiplied with the matrix T to obtain the S matrix of all 0s, and the K or V matrix block obtained this time and the S matrix of 0 are transferred to the main memory; when the Q matrix is ​​calculated, Q is multiplied with the T matrix to obtain the S matrix block, and the Q matrix block obtained this time and the transposed S matrix block are transferred to the main memory; if the input sequence length is n, the first calculation step will be repeated 3×n / subsequence length times, the first n / subsequence length times are used to calculate the K matrix, the middle n / subsequence length times are used to calculate the V matrix, and the last n / subsequence length times are used to calculate the Q matrix and the S matrix; The second block division step is to divide the complete S matrix after transposition into intermediate matrix blocks and pass them into the cache, and at the same time pass in the transposed matrix of V; The second calculation step is to receive the intermediate matrix block, perform bitwise division on it, and then perform column-wise softmax calculation to obtain the transposed matrix of P, and then multiply it with the transposed matrix of V to obtain a part of the transposed matrix of the O matrix and transpose it again and transfer it to the main memory. Repeat the second calculation step until the complete O matrix is ​​obtained in the main memory as the attention calculation result.

3. The attention mechanism fusion method based on the data flow architecture accelerator according to claim 1, characterized in that: The threshold is the capacity of the cache.

4. The attention mechanism fusion method based on the data flow architecture accelerator according to claim 1, characterized in that: The data flow architecture accelerator is a data flow chip DPU.

5. An attention mechanism fusion device based on a data flow architecture accelerator, characterized in that: include: An initialization module obtains a data flow architecture accelerator for performing attention calculation, the data flow architecture accelerator comprising a main memory, a cache, and a processing unit; When the length of the input sequence matrix multiplied by the embedding dimension is below the threshold executing a first fusion module; The first fusion module transfers the input sequence matrix, the transposed matrix of the input sequence, the transposed matrix of the weight matrix Wq, the weight matrix Wk, and the transposed matrix Wv from the main memory to the cache, and obtains the transposed matrix of the query matrix Q, the key matrix K, and the transposed matrix of the value matrix V through matrix multiplication. The processing unit then performs matrix multiplication to calculate the multiplication of the key matrix K and the transposed matrix of the query matrix Q to obtain the transposed matrix of S; the transposed matrix of S is sequentially divided by bit and the column-by-column softmax activation is calculated to obtain the transposed matrix of P, and finally the transposed matrix of the value matrix V is matrix multiplied with the transposed matrix of P to obtain the transposed matrix of O, and finally the transposed matrix of O is transposed and transferred back to the main memory to obtain the O matrix as the attention calculation result.

6. The attention mechanism fusion device based on the data flow architecture accelerator according to claim 5, characterized in that: The initial module includes executing the second fusion module when the length of the input sequence matrix multiplied by the embedding dimension is greater than or equal to the threshold: A first block module divides the input sequence into blocks according to the length dimension of the input sequence to obtain subsequences, and writes the weight matrix W and the subsequences into the cache; Pass in a matrix T of the size of the subsequence length and the length of the input sequence. If the weight matrix W used to calculate this time is a K matrix or a V matrix, the matrix T is a 0 matrix; if it is used to calculate the query matrix Q and the S matrix, the T matrix is ​​the transpose of the K matrix; The first calculation module passes the weight matrix W, the subsequence and the T matrix into the cache. When calculating the K or V matrix, the T matrix passed in is a 0 matrix. When calculating the Q matrix, the T matrix passed in is the transposed matrix of the K matrix. The weight matrix W and the subsequence are matrix multiplied to obtain the Q or K or V matrix. The K or V matrix is ​​multiplied by the matrix T to obtain the S matrix of all 0s, and the K or V matrix block obtained this time and the S matrix of 0 are transferred to the main memory. When the Q matrix is ​​calculated, Q is multiplied by the T matrix to obtain the S matrix block, and the Q matrix block obtained this time and the transposed S matrix block are transferred to the main memory. If the length of the input sequence is n, the first calculation step will be repeated 3×n / subsequence length times, the first n / subsequence length times are used to calculate the K matrix, the middle n / subsequence length times are used to calculate the V matrix, and the last n / subsequence length times are used to calculate the Q matrix and the S matrix. The first calculation step is repeated until the complete S matrix after transposition is stored in the main memory. The second block module divides the complete S matrix after transposition into intermediate matrix blocks and passes them into the cache, and at the same time passes the transposed matrix of V; The second calculation module receives the intermediate matrix block, performs bitwise division on it, and then performs column-wise softmax calculation to obtain the transposed matrix of P, and then multiplies it with the transposed matrix of V to obtain a part of the transposed matrix of the O matrix and transpose it again and transfer it to the main memory. The second calculation module is repeatedly executed until the complete O matrix is ​​obtained in the main memory as the attention calculation result.

7. The attention mechanism fusion device based on the data flow architecture accelerator according to claim 5, characterized in that: The threshold is the capacity of the cache; the data flow architecture accelerator is a data flow chip DPU.

8. An electronic device, characterized in that: It includes an attention mechanism fusion device based on a data flow architecture accelerator as described in claims 5-7, and the electronic device is connected to an information display device, which is used to display the attention calculation result with display parameters and attributes set by the user or through an artificial intelligence model.

9. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the attention mechanism fusion method based on a data flow architecture accelerator as described in any one of claims 1-4.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the attention mechanism fusion method based on the data flow architecture accelerator described in any one of claims 1-4 are implemented.

Citation Information

Patent Citations

  • GPU-based video SAR echo simulation parallel implementation method

    CN107229051A

  • Cholesky decomposition acceleration calculation method and system based on data stream architecture

    CN115391731A

  • Multi-head attention mechanism fusion calculation distribution method based on acceleration processor

    CN116431562A

  • Data stream design method compatible with multi-head self-attention and convolution calculation

    CN119003955A

  • Tensor processing

    US20230359697A1

Cited By

  • Multi-input-output Transform model chip architecture and calculation method

    CN121072630A