An acceleration method for matrix multiplication, an electronic device, and a storage medium

By converting the matrix data storage type in the GPU to match the expected storage type of the matrix multiplication accumulator, and using high-width bit data transfer instructions, the problem of low efficiency of matrix multiplication data handling is solved, and more efficient data transfer and accurate calculation results are achieved.

CN119988811BActive Publication Date: 2025-07-04METAX INTEGRATED CIRCUITS (SHANGHAI) CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510480430.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-07-04
Estimated Expiration
2045-04-17

AI Technical Summary

Technical Problem

In graphics processors (GPUs), data handling efficiency of matrix multiplication is low, especially when converting between different storage types, it is easy to cause errors in calculation results, and it is difficult for the prior art to efficiently transfer data.

Method used

By obtaining the expected storage type in the matrix multiplication accumulator MMA, converting the matrix data storage type in the intermediate register to match the expected storage type, and directly transferring the data using the high-width bit data transfer instruction to reduce the number of transfers.

Benefits of technology

It improves data handling efficiency, reduces the number of data transfer instructions, and ensures the accuracy of matrix multiplication calculation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988811B_ABST
    Figure CN119988811B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of chip design, and particularly to an acceleration method for matrix multiplication, an electronic device and a storage medium. The method includes obtaining an expected storage type of a first input matrix block in a matrix multiplication accumulator (MMA); obtaining an actual storage type of matrix data storing the first input matrix block in a memory; when the expected storage type is different from the actual storage type, reading the matrix data from the memory and storing it in an intermediate register, converting the matrix data stored in the intermediate register according to the actual storage type into matrix data stored according to the expected storage type to obtain expected matrix data; storing the expected matrix data stored in the intermediate register in a shared memory; taking out corresponding matrix data from the shared memory through a high-width data transfer instruction and loading it into a register for calculation and processing by the MMA. The present invention can directly use the high-width data transfer instruction to transfer data, so as to achieve the purpose of improving the data transfer efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of chip design, and particularly to an acceleration method for matrix multiplication, an electronic device, and a storage medium. Background Art

[0002] In a graphics processing unit (GPU), matrix multiplication is one of the core operations in high-performance computing applications. A matrix multiply-accumulate (MMA) is a component in computer architecture used to efficiently perform matrix multiplication and accumulation operations. The core of matrix multiplication is to calculate the sum of products of corresponding elements of two matrices. The matrix multiply-accumulator usually adopts a pipeline design, and the matrix multiplication operation process is decomposed into multiple stages, including data reading, multiplication operation, accumulation operation, etc. Each stage is completed in different clock cycles, so that different operation stages can process different data elements in parallel at the same time, improving the operation efficiency. To accelerate the implementation of matrix multiplication, GPUs usually adopt a block processing strategy, that is, the input matrices A and B are divided into smaller sub-matrices, for example, divided into multiple sub-matrices of M×K and K×N sub-matrices. These sub-matrix data are initially stored in global memory and then read into shared memory or directly into registers for efficient matrix operations. When the data is stored in shared memory, it needs to be read from shared memory into registers again for the matrix multiply-accumulator (MMA) to perform operations. However, users have various considerations for the storage method of matrices in memory, including NT type storage, NN type storage, TN type storage, and TT type storage, where NT type storage means that matrix A is stored in row-major order and matrix B is stored in column-major order, and NN type storage means that both matrix A and matrix B are stored in row-major order. However, the data sorting method read from shared memory or registers needs to meet the input conditions of MMA. Since MMA can only arrange data in a specific thread data manner when processing data, directly using the high-width bit lds (load from share memory) data transfer instruction will result in incorrect processing results; among them, the lds instruction is used to load data from shared memory into registers. Since the data transfer process when storing data in registers is very time-consuming in the GPU, there is an urgent need for a method to improve the transfer efficiency. Summary of the Invention

[0003] In view of the above technical problems, the technical solution adopted by the present invention is: an acceleration method for matrix multiplication, the method comprising the following steps:

[0004] S100, obtaining the expected storage type of the first input matrix block in the matrix multiply-accumulator MMA.

[0005] S200, obtain the actual storage type of the matrix data storing the first input matrix block in the memory.

[0006] S300, when the expected storage type of the first input matrix block is different from the actual storage type, read the matrix data from the memory and store it in an intermediate register, and convert the matrix data stored in the intermediate register according to the actual storage type into matrix data stored according to the expected storage type through a hardware instruction to obtain the expected matrix data.

[0007] S400, store the expected matrix data stored in the intermediate register in the shared memory.

[0008] S500, extract the corresponding matrix data from the shared memory through a data transfer instruction and load it into the input register specifically providing matrix data for the MMA.

[0009] In addition, the present invention also provides a non-transitory computer-readable storage medium, in which at least one instruction or at least one program segment is stored, and the at least one instruction or the at least one program segment is loaded and executed by a processor to implement the above method.

[0010] In addition, the present invention also provides an electronic device, including a processor and the above non-transitory computer-readable storage medium.

[0011] The present invention has at least the following beneficial effects:

[0012] An acceleration method for matrix multiplication, an electronic device and a storage medium provided by an embodiment of the present invention, by converting the matrix data already stored in an intermediate register into matrix data expected by the MMA, and further making the data stored in the shared memory be the expected matrix data. When loading the data in the shared memory into the register of the MMA, Q data transfer instructions with high-width bits can be directly used to transfer the data. When the amount of data transferred each time increases and the value of T increases, the value of Q decreases, that is, the number of data transfer instructions is reduced, and at the same time, the conversion step of S300 ensures that the operation result of the matrix multiplier will not be incorrect, achieving the purpose of improving the data transfer efficiency. Description of the Drawings

[0013] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0014] Figure 1 A flowchart of an acceleration method for matrix multiplication provided by an embodiment of the present invention;

[0015] Figure 2 A schematic diagram of the arrangement mode of threads provided by an embodiment of the present invention;

[0016] Figure 3 Based on Figure 2 A comparison schematic diagram of matrix elements before and after executing corresponding hardware instructions in the arrangement mode;

[0017] Figure 4 A schematic diagram of another arrangement mode of threads provided by an embodiment of the present invention;

[0018] Figure 5 Based on Figure 4 A comparison schematic diagram of matrix elements before and after executing corresponding hardware instructions in the arrangement mode;

[0019] Figure 6 A schematic diagram of another arrangement mode of threads provided by an embodiment of the present invention;

[0020] Figure 7 Based on Figure 6 A comparison schematic diagram of matrix elements before and after executing corresponding hardware instructions in the arrangement mode. Detailed implementation manners

[0021] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts shall fall within the protection scope of the present invention.

[0022] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present invention have the same meanings as those commonly understood by those skilled in the art.

[0023] Please refer to Figure 1 , which shows an acceleration method for matrix multiplication. The method includes the following steps:

[0024] S100, obtain the expected storage type of the first input matrix block in the matrix multiplication accumulator MMA.

[0025] Among them, the first input matrix block is the first D-level matrix block, and the second input matrix block is the second D-level matrix block. D is the maximum division level at which the input matrices in matrix multiplication are divided into matrix blocks, and the value range of D is from 1 to D. It should be noted that there are two input matrices in the matrix multiplication of MMA. To improve the efficiency of matrix multiplication, for the matrix multiplication operation of the first matrix A and the second matrix B, according to the standard matrix multiplication rule, the first matrix A is divided into multiple first matrix blocks {A1, A2, …, A i , …, A L}, A i is the i-th first matrix block, and the value range of i is from 1 to L; the second matrix B is also divided into multiple second matrix blocks {B1, B2, …, B i , …, B L}, B i is the i-th second matrix block, and the value range of i is from 1 to L; the first matrix block A i is multiplied by the second matrix block B i to obtain an intermediate result, and then the final result is calculated based on the intermediate result. In another embodiment, A i will also be further divided into multiple first and second-level matrix blocks {a i1 , a i2 , …, a ir , …, a iL}, a ir is the r-th first and second-level matrix block in A i , and similarly, B will also be correspondingly divided into multiple second and second-level matrix blocks. Embodiments of other divisions of matrix blocks with more levels also fall within the protection scope of the present invention.

[0026] The conversion steps provided in the embodiments of the present invention applicable to the first input matrix block are equally applicable to the second input matrix block.

[0027] In one implementation, the storage type of each input matrix block in memory is row-major storage or column-major storage.

[0028] Among them, MMA includes two input matrices, and the calculation of the two input matrices in MMA follows the standard matrix multiplication rule.

[0029] In one embodiment, the combined storage methods of the two input matrices of MMA are NT type storage, NN type storage, TN type storage, and TT type storage. Here, N and T are used to represent the storage types of the corresponding input matrices respectively, where N refers to row-major order storage and T refers to column-major order storage. For example, when the combined storage method of the two input matrices of MMA is NT type storage, it means that one of the two input matrices is stored in row-major order and the other input matrix is stored in column-major order; when the combined storage method of the two input matrices of MMA is NN type storage, both input matrices are stored in row-major order; and so on, which will not be elaborated here.

[0030] In one embodiment, the expected storage type of the first input matrix block of MMA is row-major order storage, and the expected storage type of the second input matrix block is column-major order storage. At this time, the expected combined storage method is NT type storage. The expected storage type of the first input matrix block of MMA is column-major order storage, and the storage type of the second input matrix block is row-major order storage. At this time, the expected combined storage method is TN type storage. The expected storage types of the first input matrix block and the second input matrix block of MMA are both row-major order storage or both column-major order storage. At this time, the expected combined storage method is NN type storage or TT type storage. Combinations of other storage types expected for the first input matrix block and the second input matrix block of MMA also fall within the protection scope of the present invention.

[0031] S200, obtain the actual storage type of the matrix data stored in the memory for the first input matrix block.

[0032] In one embodiment, the actual storage type of the first input matrix block is row-major order storage or column-major order storage. Among them, the actual storage type is the storage type configured by the user, which may be configured based on considerations of chip performance or storage space.

[0033] S300, when the expected storage type of the first input matrix block is different from the actual storage type, read the matrix data from the memory and store it in the intermediate register, and convert the matrix data stored in the intermediate register according to the actual storage type into matrix data stored according to the expected storage type through hardware instructions to obtain the expected matrix data.

[0034] Among them, the purpose of the embodiment of the present invention is to directly use the data transfer instruction with high width to transfer data. If the order of the matrix elements read currently is correct, it can transfer more matrix data through one data transfer instruction and ensure the correct calculation result of matrix multiplication. Compared with the data transfer instruction with low width, the number of transfers of the data transfer instruction with high width is greatly reduced, achieving the purpose of improving the data transfer efficiency. However, when the expected storage type is different from the actual storage type, if the storage type is not converted, it is easy to cause errors in matrix multiplication. For example, the storage type of the first input matrix block originally transferred is row-major storage, and each data transfer instruction transfers 32-bit data, while MMA expects matrix data in column-major storage; at this time, if the data transfer instruction transfers 64-bit data each time, it will cause an error in the calculation result of the matrix multiplier. In order to directly use the data transfer instruction with high width and ensure the correct calculation result of the matrix multiplier, it is necessary to convert the storage type of the matrix data in the intermediate register according to S300.

[0035] When reading data from the memory to the intermediate register, the data storage type in the intermediate register is the same as that in the memory, which is the actual storage type; in order to make the data storage type in the shared memory be the expected storage type, before transferring the matrix data in the intermediate register to the shared memory, it is necessary to convert the matrix data stored in the intermediate register according to the actual storage type into matrix data stored according to the expected storage type.

[0036] S400, store the expected matrix data stored in the intermediate register into the shared memory.

[0037] S500, take out the matrix data of the first input matrix block and the matrix data of the second input matrix block input to each MMA from the shared memory through Q data transfer instructions, and load the matrix data of the first input matrix block and the second input matrix block into the input register dedicated to providing matrix data for MMA; where Q satisfies: Q = U / (T × H), where U is the data volume of the first input matrix block and the second input matrix block, T is the data volume transferred by each data transfer instruction each time, and H is the number of threads in the thread group of each execution unit.

[0038] It should be noted that since the corresponding matrix data has been converted into the matrix data expected by MMA in the intermediate register, the data stored in the shared memory is already the expected matrix data. When loading the data in the shared memory into the register of MMA, the Q data transfer instructions with high width can be directly used to transfer data. When the data volume transferred each time increases, that is, when the value of T increases, the value of Q decreases, that is, the number of transfers of the data transfer instruction is reduced, and at the same time, the conversion step of S300 ensures that the operation result of the matrix multiplier will not be incorrect.

[0039] Among them, in S300, the matrix data stored in the intermediate register according to the actual storage type is converted into matrix data stored according to the desired storage type, and the conversion steps are configured according to the thread arrangement. In one implementation, the thread arrangement includes one thread reading the matrix data in N groups of intermediate registers, N threads reading the matrix data in one intermediate register, and one thread reading the matrix data in two intermediate registers. Other types of thread arrangements also fall within the protection scope of the present invention.

[0040] In one implementation, when an execution unit parallelly processes the matrix multiplications of N MMAs, the memory stores the matrix data of N×M×K matrix elements of N MMAs according to the actual storage type, where the matrix data of the N×M matrix elements is the matrix data of the first input matrix block of one MMA; the matrix data of N MMAs is parallelly read by a thread group of an execution unit, and the matrix data of each MMA is stored in N groups of intermediate registers. The conversion steps include: S310, obtaining the N×M matrix elements stored in the N groups of intermediate registers. S320, permuting the N×M matrix elements stored in the N groups of intermediate registers processed by the current thread through a hardware instruction, and converting the actual storage type of the N×M matrix elements into the desired storage type of the N MMAs to obtain the desired matrix data. Then, U in S500 is the data volume of M×K matrix elements.

[0041] In one implementation, the hardware instruction is a permutation instruction, and the permutation instruction is used to recombine the data of two intermediate registers. The combination method can be determined according to specific hardware design requirements. For example, it may be simple splicing, or it may be combined in a specific bit selection or cross manner. As an example, when the actual storage type is row-major storage, the desired storage type is column-major storage, and each intermediate register stores two matrix elements, and each row of matrix elements in the input matrix block is stored through two intermediate registers, the positions of the matrix elements stored in the intermediate registers are exchanged through the permutation instruction, so that the storage type of the matrix data in the intermediate register is the desired column-major storage.

[0042] In one implementation, N = 4, M = 4, each group includes two intermediate registers, each intermediate register stores 2 matrix elements, the actual storage type is row-major storage, and the desired storage type is column-major storage. Then, the permutation instruction is the perm instruction, and the matrix elements stored in the intermediate registers are permuted through the perm instruction.

[0043] In another embodiment, N = 4, M = 4, each group includes two intermediate registers, each intermediate register can store up to 32-bit data at most, and each intermediate register stores two 16-bit matrix elements. The actual storage type is row-major order storage, and the desired storage type is column-major order storage. Then, the matrix elements stored in the intermediate registers are permuted by the perm instruction.

[0044] As an example, please refer to Figure 2 , Figure 2 for the arrangement of threads and registers. In Figure 2 , T0 is thread 0, T1 is thread 1, and so on; r[0:1] are the intermediate registers r0 and r1, r[2:3], r[4:5], r[6:7] respectively represent r2 to r7, a total of 6 intermediate registers, and so on; each intermediate register stores 32-bit data. When each matrix element is 16-bit, the data volume is a matrix of 16 rows and 64 columns of 16-bit matrix elements, that is, a 16×(16×4) matrix. In Figure 3 , before the operation of the perm instruction, A16 and B16 are stored in r0, C16 and D16 are stored in r1, E16 and F16 are stored in r2, and so on, P16 and Q16 are stored in r7. Before the operation of the perm instruction, the matrix data is in row storage type. The matrix data and itself are used as the two inputs of the per instruction for conversion, and the storage type of the converted matrix data is the desired column-major order storage. Finally, 64 bits, that is, 4 matrix elements, are fetched by the data transfer instruction and put into the shared memory.

[0045] In another embodiment, N = 4, M = 4, each group includes two intermediate registers, each intermediate register can store up to 32-bit data at most, and each intermediate register stores four 8-bit matrix elements. The actual storage type is row-major order storage, and the desired storage type is column-major order storage. Then, the matrix data of the N×M matrix elements stored in the intermediate registers and itself are used as the two inputs of the perm instruction for the first permutation to obtain the permuted matrix data; then, the permuted matrix data and itself are used as the two inputs of the perm instruction again for the second permutation to obtain the desired matrix data.

[0046] As an example, when each matrix element is 16-bit and the data transfer instruction is 64-bit, each data transfer instruction transfers 4 matrix elements. When each matrix element is 16-bit and the data transfer instruction is 128-bit, each data transfer instruction transfers 8 matrix elements.

[0047] In one embodiment, when the memory stores matrix data of T×M matrix elements according to the actual storage type, and the matrix data of the T×M matrix elements is the matrix data of the first input matrix block of an MMA, and T threads read and each intermediate register stores M matrix elements, S300 further includes: respectively and concurrently reading the continuously stored matrix elements in the memory by T threads, and storing the M matrix elements read by each thread into an intermediate register; the conversion step includes: S310, obtaining the T×M matrix elements stored in T groups of intermediate registers. The hardware instructions in S320 include a swap instruction (shfl instruction) and a permutation instruction. The matrix elements stored in the intermediate registers processed by adjacent threads are swapped through the swap instruction to obtain the swapped matrix elements; the swapped matrix elements and the matrix elements before swapping are permuted through the permutation instruction to obtain the desired matrix data. Then U in S500 is the data volume of T×M matrix elements. Among them, the swap instruction is used to swap the matrix elements stored in the intermediate registers corresponding to two threads.

[0048] In one embodiment, T = 2, M = 2, and each intermediate register stores two matrix elements in half precision; then the hardware instructions include a swap instruction (shfl instruction) and a permutation instruction.

[0049] As an example, please refer to Figure 4 , Figure 4 for another arrangement of threads and registers. In Figure 4 , thread T0 and thread T1 read two matrix elements b0 and b1 in half precision and store them into an intermediate register, and are also used to read two matrix elements b2 and b3 in another intermediate register. When each intermediate register stores 32-bit data and each matrix element is 16-bit, its data volume is a matrix of 16 rows and 16 columns of 16-bit matrix elements, that is, a 16×16 matrix, where T0:b0 indicates that the data in the intermediate register read by thread 0 is b0. Figure 5 FIG. is a comparison schematic diagram of matrix elements before and after executing the corresponding hardware instructions. In Figure 5 , A16 and B16 before executing the shlf instruction correspond to Figure 4 T0:b0 and T0:b1 in Figure 4T1:b0 and T1:b1 in it. When the matrix elements stored in the intermediate registers corresponding to T0 and T1 are exchanged by the shlf instruction, the exchanged matrix elements are obtained; after the exchange, the matrix elements C16 and D16 correspond to T0, and the matrix elements A16 and B16 correspond to T1; then the perm instruction is used to exchange the exchanged matrix elements and the original matrix elements before the exchange to obtain the desired matrix data; in the desired matrix data, two half-precision matrix elements A16 and C16 stored in the intermediate register correspond to T0, and two half-precision matrix elements B16 and D16 stored in the intermediate register correspond to T1, that is, the row storage type of the original A16 and B16 is converted to the column storage type. Finally, two half-precision matrix elements are fetched by the data transfer instruction and placed in the shared memory.

[0050] In one embodiment, the maximum precision stored in the intermediate register is 32 bit, two matrix elements are stored in each intermediate register, and each matrix element is 16-bit data.

[0051] In one embodiment, the maximum precision stored in the intermediate register is 32 bit, four matrix elements are stored in each intermediate register, and each matrix element is 8-bit data. Then the step of permutation by the hardware instruction in S320 further includes: exchanging the matrix elements stored in the intermediate registers corresponding to two adjacent groups of threads by the exchange instruction to obtain T×4 exchanged matrix elements; performing permutation on the T×4 exchanged matrix elements and the T×4 matrix elements before the exchange by the perm instruction to obtain the processed matrix data. The processed matrix data and itself are permuted again to obtain the desired matrix data, and the data transfer instruction in S500 is a 32-bit instruction.

[0052] In one embodiment, when the memory stores matrix data of T×M matrix elements according to the actual storage type, and the matrix data of the T×M matrix elements is the matrix data of the first input matrix block of an MMA; when T threads read the matrix data of K intermediate registers and each intermediate register stores M / K matrix elements, S300 further includes: sequentially reading M matrix elements continuously stored in the memory by N threads; the conversion step includes: S310, obtaining the T×M matrix elements stored in N groups of intermediate registers; S320, the hardware instructions include swap instructions, permutation instructions, and move instructions (mov instructions). By the swap instructions, the matrix elements stored in the intermediate registers processed by adjacent threads are swapped to obtain the swapped matrix elements; by the permutation instructions, the swapped matrix elements and the matrix elements before swapping are permuted to obtain the permuted matrix data; by the swap instructions, the permuted matrix data is swapped again to obtain the matrix elements after the second swap; by the move instructions, the matrix elements after the second swap and the permuted matrix data are moved to obtain the desired matrix data. Then, in S500, U is the data volume of T×M matrix elements.

[0053] In one embodiment, T = 4, M = 4, and K = 2.

[0054] As an example, please refer to Figure 6 , Figure 6 for another arrangement of threads and registers. In Figure 6 , T0:b0, T0:b1, T0:b2, T0:b3 respectively represent that thread T0 processes matrix elements b0, b1, b2, and b3, and so on. When each matrix element is 16 bits, its data volume is a matrix of 16 rows and 16 columns of 16-bit matrix elements, that is, a 16×16 matrix. Figure 7 is a comparison schematic diagram of matrix elements before and after executing the corresponding hardware instructions. In Figure 7 , in the matrix before the first execution of the shlf instruction, A16, B16, C16, and D16 corresponding to T0 respectively correspond to Figure 6b0, b1, b2, and b3 in it. When the matrix elements stored in the intermediate registers corresponding to T0 and T1 are swapped by the shlf instruction, the swapped matrix elements are obtained; after swapping, the matrix elements corresponding to T0 are E16, F16, G16, and H16, and the matrix elements corresponding to T1 are A16, B16, C16, and D16, and so on. Then, the perm instruction is used to permute the swapped matrix elements and the original matrix elements before swapping to obtain the permuted matrix data; in the permuted matrix data, T0 corresponds to A16, E16, C16, and G16, and T1 corresponds to B16, F16, D16, and H16; then the shlf instruction and the mov instruction are executed again to convert the matrix data into the desired column storage type. Finally, 64 bits, that is, 4 matrix elements, are fetched through the data transfer instruction and placed into the shared memory.

[0055] An embodiment of the present invention also provides a non-transitory computer-readable storage medium, which can be disposed in an electronic device to store at least one instruction or at least one program segment related to a method for implementing a method in the method embodiment. The at least one instruction or the at least one program segment is loaded and executed by the processor to implement the method provided in the above embodiment.

[0056] An embodiment of the present invention also provides an electronic device, including a processor and the aforementioned non-transitory computer-readable storage medium.

[0057] An embodiment of the present invention also provides a computer program product, which includes program code. When the program product runs on an electronic device, the program code is used to cause the electronic device to execute the steps in the method according to various exemplary embodiments of the present invention described above in this specification.

[0058] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above.

[0059] Although some specific embodiments of the present invention have been described in detail by way of examples, those skilled in the art should understand that the above examples are only for illustration and not for limiting the scope of the present invention. Those skilled in the art should also understand that various modifications can be made to the embodiments without departing from the scope and spirit of the present invention. The scope of the present invention disclosed is defined by the appended claims.

Claims

1. An acceleration method for matrix multiplication, characterized in that, The method includes the following steps: S100, obtaining the expected storage type of the first input matrix block in the matrix multiplication accumulator (MMA); S200, obtaining the actual storage type of the matrix data storing the first input matrix block in the memory; S300, when the expected storage type of the first input matrix block is different from the actual storage type, reading the matrix data from the memory and storing it in an intermediate register, and converting the matrix data stored in the intermediate register according to the actual storage type into matrix data stored according to the expected storage type through a hardware instruction to obtain the expected matrix data; S400, storing the expected matrix data stored in the intermediate register into the shared memory; S500, taking out the matrix data of the first input matrix block and the matrix data of the second input matrix block input to each MMA from the shared memory through Q data transfer instructions, and loading the matrix data of the first input matrix block and the second input matrix block into the input register dedicated to providing matrix data for the MMA; where Q satisfies: Q = U / (T × H), where U is the data volume of the first input matrix block and the second input matrix block, T is the data volume transferred by each data transfer instruction each time, and H is the number of threads in the thread group of each execution unit; When an execution unit performs matrix multiplication of N MMAs in parallel, the memory stores the matrix data of N × M × K matrix elements of N MMAs according to the actual storage type, where the matrix data of M × K matrix elements is the matrix data of the first input matrix block of one MMA; the matrix data of N MMAs is read in parallel by the thread group of one execution unit, and the matrix data of each MMA is stored in N groups of intermediate registers; then the conversion step in S300 includes: S310, obtaining the N × M matrix elements stored in N groups of intermediate registers; S320, permuting the N × M matrix elements stored in the N groups of intermediate registers processed by the current thread through a hardware instruction, and converting the actual storage type of the N × M matrix elements into the expected storage type of N MMAs to obtain the expected matrix data; Then U in S500 is the data volume of M × K matrix elements.

2. The method according to claim 1, wherein N = 4, M = 4, each group includes two intermediate registers, each intermediate register can store a maximum of 32-bit data, and each intermediate register stores two 16-bit matrix elements. The actual storage type is row-major order storage, and the expected storage type is column-major order storage; then the permutation instruction is the perm instruction, and the matrix elements stored in the intermediate register are permuted through the perm instruction.

3. The method according to claim 1, wherein Where N = 4, M = 4, each group includes two intermediate registers, each intermediate register can store up to 32-bit data, and each intermediate register stores four 8-bit matrix elements. The actual storage type is row-major order storage, and the expected storage type is column-major order storage. Then, the matrix data of the N×M matrix elements stored in the intermediate register and itself are used as the two inputs of the perm instruction for the first permutation to obtain the permuted matrix data. Then, the permuted matrix data and itself are used as the two inputs of the perm instruction again for the second permutation to obtain the expected matrix data.

4. The method according to claim 1, wherein When the memory stores the matrix data of T×M matrix elements according to the actual storage type, and the matrix data of the T×M matrix elements is the matrix data of the first input matrix block of an MMA, when T threads read the matrix data of the same intermediate register and each intermediate register stores M matrix elements, S300 further includes: T threads respectively and concurrently read the continuously stored matrix elements in the memory, and each thread reads M matrix elements and stores them in an intermediate register; the conversion step includes: S310, obtaining the T×M matrix elements stored in T groups of intermediate registers; The hardware instructions in S320 include swap instructions and permutation instructions. The matrix elements stored in the intermediate registers processed by adjacent threads are swapped through the swap instructions to obtain the swapped matrix elements; the swapped matrix elements and the matrix elements before swapping are permuted through the permutation instructions to obtain the expected matrix data; Then, U in S500 is the data volume of T×M matrix elements.

5. The method according to claim 4, wherein Where T = 2, M = 2, and each intermediate register stores two half-precision matrix elements. Then, the hardware instructions include swap instructions and permutation instructions, and each data transfer instruction is used to fetch all the matrix elements in an intermediate register. Among them, the swap instruction is used to swap the matrix elements stored in the intermediate registers corresponding to two threads.

6. The method according to claim 4, characterized in that The maximum precision stored in the intermediate register is 32-bit, each intermediate register stores four matrix elements, and each matrix element is 8-bit data; Then, the step of permutation through the hardware instructions in S320 further includes: swapping the matrix elements stored in the intermediate registers corresponding to adjacent groups of threads through the swap instructions to obtain T×4 swapped matrix elements; the T×4 swapped matrix elements and the T×4 matrix elements before swapping are permuted through the permutation instructions to obtain the processed matrix data; the processed matrix data and itself are permuted again to obtain the expected matrix data.

7. The method according to claim 1, characterized in that, When the memory stores the matrix data of T×M matrix elements according to the actual storage type, where the matrix data of the T×M matrix elements is the matrix data of the first input matrix block of an MMA; when T threads read the matrix data of K intermediate registers and each intermediate register stores M / K matrix elements, S300 further includes: T threads sequentially read M continuously stored matrix elements in the memory; the conversion step includes: S310, obtain the T×M matrix elements stored in N sets of intermediate registers; S320, the hardware instructions include swap instructions, permutation instructions, and move instructions. Use the swap instructions to swap the matrix elements stored in the intermediate registers processed by adjacent threads to obtain the swapped matrix elements; use the permutation instructions to permute the swapped matrix elements and the matrix elements before swapping to obtain the permuted matrix data; use the swap instructions to swap the permuted matrix data again to obtain the matrix elements after the second swap; use the move instructions to move the matrix elements after the second swap and the permuted matrix data to obtain the desired matrix data; Then in S500, U is the data volume of T×M matrix elements.

8. A non-transitory computer-readable storage medium storing at least one instruction or at least one program segment, characterized in that, The at least one instruction or the at least one segment of program is loaded and executed by a processor to implement the method according to any one of claims 1-7.

9. An electronic device, characterized in that, It includes a processor and the non-transitory computer-readable storage medium described in claim 8.

Citation Information

Patent Citations

  • Message cache management method and device and network processor

    CN114157619A

  • Data transmission method and high-performance computing device

    CN116932239A

  • Method and apparatus for matrix calculation acceleration

    CN117501250A

  • Data migration method and device based on DPU, electronic equipment and storage medium

    CN118796738A

  • Hardware circuit, data migration method, chip, and electronic device

    WO2022227563A1