Processing method for matrix multiplication
By adjusting the scheduling order of thread blocks in matrix multiplication operation, we ensure that the second matrix is fully loaded into the cache and the first matrix is partially loaded into the cache, which solves the problem of low cache hit rate and improves the operation speed of matrix multiplication.
Patent Information
- Application Number
- CN202311460984.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-02
- Publication Date
- 2025-07-08
AI Technical Summary
When performing matrix multiplication operations in the prior art, since the first matrix cannot be loaded into the cache, the cache hit rate is low, which reduces the operation speed.
By obtaining the logical coordinate sequence of matrix blocks of the initial mapping relationship and the target mapping relationship, the scheduling order of the thread blocks is adjusted so that the second matrix is all loaded into the cache, the first matrix is partially loaded into the cache, and is scheduled in the order of the ordinal number of the thread blocks from small to large.
Improve the cache hit rate of matrix multiplication operations and improve the operation speed.
Smart Images

Figure CN120277308A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electrical digital data processing, and in particular, to a processing method for matrix multiplication. Background Art
[0002] In the prior art, when performing matrix multiplication operations (i.e., the operation of multiplying the first matrix by the second matrix) on two matrices using multiple thread blocks, the thread blocks are often scheduled in a fixed scheduling order without considering the scales of the above two matrices. For example, in a GPU in the prior art, the thread blocks are scheduled in a scheduling order of first scheduling the thread blocks on the x-axis and then scheduling the thread blocks on the y-axis, where the positive direction of the x-axis corresponds to the direction in which the number of rows of the matrix increases from small to large, and the positive direction of the y-axis corresponds to the direction in which the number of columns of the matrix increases from small to large; if the storage space occupied by the first matrix is larger than the storage space size of the cache, then when performing matrix multiplication operations on the first matrix and the second matrix, since the first matrix cannot be fully loaded into the cache, there will be more cache misses, reducing the speed of performing matrix multiplication operations. Summary of the Invention
[0003] The object of the present invention is to provide a processing method for matrix multiplication to improve the speed of performing matrix multiplication operations.
[0004] According to the present invention, there is provided a processing method for matrix multiplication, the method comprising the following steps:
[0005] S100, obtain a first matrix A and a second matrix B to be subjected to matrix multiplication operations, A is an m×k matrix, B is a k×n matrix, m is the number of rows of A, k is the number of columns of A, and n is the number of columns of B.
[0006] S200, obtain a logical coordinate sequence P of matrix blocks when mapping according to a preset initial mapping relationship, P=(p1, p2, …, p i , …, p u ), p i is the logical coordinate of C i , C i is a matrix block having a mapping relationship with blo i when mapping according to a preset initial mapping relationship, blo i is the i-th scheduled thread block when performing matrix multiplication operations on A and B, p i =(p i,x , p i,y ), p i,x is the x logical coordinate of C i , p i,y is the C iThe y logical coordinate, where the value range of i is from 1 to u, and u is the number of target thread blocks which are the thread blocks for performing matrix multiplication on A and B.
[0007] S300, if m > n and sto B <sto cache <sto A , then go to S400; sto B is the storage space size occupied by A, sto cache is the storage space size of the cache, sto A is the storage space size occupied by B.
[0008] S400, obtain the logical coordinate sequence P' of the matrix block when mapping according to the target mapping relationship, P'=(p'1, p'2,..., p' i ,..., p' u ), p' i is the logical coordinate of C' i and C' i is the matrix block having a mapping relationship with blo i when mapping according to the target mapping relationship, p' i =(p' i,x , p' i,y ), p' i,x is the x logical coordinate of C' i , p' i,y is the logical coordinate of C' i and p' e,all is less than β e+e0,all , β e,all is the ordinal number when blo e,all is scheduled for matrix multiplication on A and B, blo e,all is the thread block having a mapping relationship with mat e,all in the target mapping relationship, mat e,all is the matrix block for recording the matrix multiplication result of the e-th row in A and B, β e+e0,all is the ordinal number when blo e+e0,all is scheduled for matrix multiplication on A and B, blo e+e0,all is the thread block having a mapping relationship with mat e+e0,all in the target mapping relationship, mat e+e0,all is the matrix block for recording the matrix multiplication result of the (e + e0)-th row in A and B; the value range of e is from 1 to m, and e0 is a preset first row number threshold, e0 ≥ 1.
[0009] S500, schedule the target thread blocks in ascending order of the corresponding ordinal numbers, where the thread block with the corresponding ordinal number i is used to execute the logical coordinate p'i For the matrix multiplication operation of matrix blocks, when performing matrix multiplication on A and B, B is fully loaded into the cache and A is partially loaded into the cache.
[0010] Compared with the prior art, the present invention has at least the following beneficial effects:
[0011] When performing matrix multiplication on the first matrix and the second matrix, the present invention first obtains the logical coordinate sequence P of matrix blocks when mapping according to a preset initial mapping relationship. The order of the logical coordinates of different matrix blocks in P is arranged according to the scheduling order of thread blocks having a mapping relationship with the matrix blocks in the initial mapping relationship. That is, the i-th logical coordinate in P is the logical coordinate of the matrix block having a mapping relationship with the i-th scheduled thread block in the initial mapping relationship. Secondly, the present invention judges the scales of the first matrix and the second matrix. If the conditions that the number of rows of the first matrix is greater than the number of columns of the second matrix, the storage space occupied by the second matrix is less than the storage space of the cache, and the storage space occupied by the first matrix is greater than the storage space of the cache are satisfied, then the logical coordinate sequence P' of matrix blocks when mapping according to the target mapping relationship is obtained. The i-th logical coordinate in P' is the logical coordinate of the matrix block having a mapping relationship with the i-th scheduled thread block in the target mapping relationship. In the target mapping relationship of the present invention, the ordinal number of the thread block having a mapping relationship with the matrix block for the multiplication result of the smaller rows of the first matrix and the second matrix is smaller, and the ordinal number of the thread block having a mapping relationship with the matrix block for the multiplication result of the larger rows of the first matrix and the second matrix is larger. And when performing matrix multiplication on the first matrix and the second matrix, the second matrix is fully loaded into the cache and the first matrix is partially loaded into the cache. The present invention schedules the thread blocks in ascending order of the ordinal numbers of the thread blocks. Thus, during the process of performing matrix multiplication on the first matrix and the second matrix, the second matrix will be fully hit by the cache, and when multiplying the parts of the same row in the first matrix with different columns in the second matrix, the parts of the same row in the first matrix will also be repeatedly hit by the cache, thereby increasing the probability of cache hit when performing matrix multiplication on the first matrix and the second matrix. Since the reading speed of the cache is relatively high, the speed of the multiplication operation is thus increased. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0013] Figure 1 It is a flowchart of a processing method for matrix multiplication provided by an embodiment of the present invention. Detailed implementation manners
[0014] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0015] Embodiment 1
[0016] According to this embodiment, as Figure 1 shown, a processing method for matrix multiplication is provided, and the method includes the following steps:
[0017] S100. Obtain a first matrix A and a second matrix B to be subjected to matrix multiplication operation. A is an m×k matrix, B is a k×n matrix, m is the number of rows of A, k is the number of columns of A, and n is the number of columns of B.
[0018] S200. Obtain a logical coordinate sequence P of matrix blocks when mapping according to a preset initial mapping relationship, P=(p1, p2, …, p i , …, p u ), p i is the logical coordinate of C i , C i is a matrix block having a mapping relationship with blo i when mapping according to the preset initial mapping relationship, blo i is the i-th scheduled thread block when performing matrix multiplication operation on A and B, p i =(p i,x , p i,y ), p i,x is the x logical coordinate of C i , p i,y is the y logical coordinate of C i , and the value range of i is from 1 to u, where u is the number of target thread blocks, and the target thread block is the thread block for performing matrix multiplication operation on A and B.
[0019] In the preset initial mapping relationship of this embodiment, there is a mapping relationship between the thread block and the matrix block that satisfies the equality of the coordinates of the thread block and the logical coordinates of the matrix block. Optionally, the thread blocks are scheduled in the scheduling order of first scheduling the thread blocks on the x-axis and then scheduling the thread blocks on the y-axis, and the coordinate sequence of the thread blocks is ((0, 0), (1, 0), …, (u x -1, 0), (0, 1), (1, 1), …, (u x -1, 1), …, (0, u y -1), (1, uy -1),…,(u x -1,u y -1)), where the i-th coordinate corresponds to the coordinates of the i-th scheduled thread block, and the logical coordinate sequence P of the matrix block is ((0,0),(1,0),…,(u x -1,0),(0,1),(1,1),…,(u x -1,1),…,(0,u y -1),(1,u y -1),…,(u x -1,u y -1)).
[0020] Specifically, u = u x ×u y , u x is the number of thread blocks included in the thread grid for matrix multiplication of A and B in the x direction, u x = m / t x , t x is the number of rows corresponding to each thread block for matrix multiplication of A and B preset, u y is the number of thread blocks included in the thread grid for matrix multiplication of A and B in the y direction, u y = n / t y , t y is the number of columns corresponding to each thread block for matrix multiplication of A and B preset.
[0021] S300, if m > n and sto B <sto cache <sto A , then enter S400; sto B is the storage space size occupied by A, sto cache is the storage space size of the cache, sto A is the storage space size occupied by B.
[0022] Optionally, the cache is the L2 cache of the GPU.
[0023] S400, obtain the logical coordinate sequence P' of the matrix block when mapping according to the target mapping relationship, P' = (p'1, p'2, …, p' i ,…, p' u ), p' i is the logical coordinate of C' i , C' i is the matrix block having a mapping relationship with blo i when mapping according to the target mapping relationship, p' i = (p'i,x , p' i,y ), p' i,x is the x logical coordinate of C', p' i of C', p' i,y is the y logical coordinate of C'; in the target mapping relationship, β i is less than β e,all where β e+e0,all is the ordinal number when performing matrix multiplication on A and B for blo e,all is scheduled, blo e,all is the thread block that has a mapping relationship with mat in the target mapping relationship e,all mat e,all is the matrix block used to record the matrix multiplication result of the e-th row in A and B, β e,all is the ordinal number when performing matrix multiplication on A and B for blo e+e0,all is scheduled, blo e+e0,all is the thread block that has a mapping relationship with mat e+e0,all in the target mapping relationship e+e0,all mat e+e0,all is the matrix block used to record the matrix multiplication result of the (e + e0)-th row in A and B; the value range of e is from 1 to m, e0 is a preset first row number threshold, and e0 ≥ 1.
[0024] It should be understood that when performing matrix multiplication on A and B, the ordinal number of the i-th scheduled thread block is i.
[0025] In this embodiment, one matrix block is used to record the matrix multiplication result of the t x rows of A and the t y columns of B, and e0 = λ1 × t x , where λ1 is a preset first multiple, λ1 ≥ 1 and λ1 is an integer.
[0026] Optionally, when e0 = t x , p' i is equal to q(i) logical coordinates in P. If 1 ≤ i ≤ u x , q(i) = (i - 1) × u y + 1; if i > u x , q(i) = (%(i / u x ) - 1) × u y + floor(i / u x ), where %() is the remainder function and floor() is the floor function.
[0027] Specifically, in the target mapping relationship, β e,g is less than β e,g+g0 , β e,gWhen performing matrix multiplication on A and B, blo e,g The scheduled ordinal number, blo e,g In the target mapping relationship, with mat e,g The thread block with a mapping relationship, mat e,g Is a matrix block used to record the matrix multiplication result of the e-th row in A and the g-th column in B, β e,g+g0 When performing matrix multiplication on A and B, blo e,g+g0 The scheduled ordinal number, blo e,g+g0 In the target mapping relationship, with mat e,g+g0 The thread block with a mapping relationship, mat e,g+g0 Is a matrix block used to record the matrix multiplication result of the e-th row in A and the g+g0-th column in B; the value range of g is from 1 to n, and g0 is a preset first column number threshold, g0≥1.
[0028] Specifically, g0 = λ1×t y .
[0029] Optionally, in the target mapping relationship, if λ1 = 1, the ordinal number of the thread block with a mapping relationship to the matrix block is positively correlated with the x logical coordinate of the matrix block; for matrix blocks with the same x logical coordinate, the ordinal number of the thread block with a mapping relationship to the matrix block is positively correlated with the y logical coordinate of the matrix block. For example, the ordinal number of the thread block with a mapping relationship to the matrix block with the smallest x logical coordinate is less than the ordinal number of the thread block with a mapping relationship to the matrix block with the second smallest x logical coordinate, the ordinal number of the thread block with a mapping relationship to the matrix block with the second smallest x logical coordinate is less than the ordinal number of the thread block with a mapping relationship to the matrix block with the third smallest x logical coordinate, and so on; for matrix blocks with the same x logical coordinate, the ordinal number of the thread block with a mapping relationship to the matrix block with the smallest y logical coordinate is less than the ordinal number of the thread block with a mapping relationship to the matrix block with the second smallest y logical coordinate, the ordinal number of the thread block with a mapping relationship to the matrix block with the second smallest y logical coordinate is less than the ordinal number of the thread block with a mapping relationship to the matrix block with the third smallest y logical coordinate, and so on.
[0030] Optionally, in the target mapping relationship, if λ1 = 2, the ordinal number of the thread block having a mapping relationship with the two-row matrix block with the smallest x logical coordinate is less than the ordinal number of the thread block having a mapping relationship with the two-row matrix block with the second smallest x logical coordinate, the ordinal number of the thread block having a mapping relationship with the two-row matrix block with the second smallest x logical coordinate is less than the ordinal number of the thread block having a mapping relationship with the two-row matrix block with the third smallest x logical coordinate, and so on; for any of the two-row matrix blocks, the ordinal number of the thread block having a mapping relationship with the four matrix blocks with the smallest y logical coordinate is less than the ordinal number of the thread block having a mapping relationship with the four matrix blocks with the second smallest y logical coordinate, the ordinal number of the thread block having a mapping relationship with the four matrix blocks with the second smallest y logical coordinate is less than the ordinal number of the thread block having a mapping relationship with the four matrix blocks with the third smallest y logical coordinate, and so on; for any of the four matrix blocks, the ordinal number of the thread block having a mapping relationship with the matrix block having the smallest ordinal number of the corresponding thread block in the initial mapping relationship is the smallest, the ordinal number of the thread block having a mapping relationship with the matrix block having the second smallest ordinal number of the corresponding thread block in the initial mapping relationship is the second smallest, the ordinal number of the thread block having a mapping relationship with the matrix block having the second largest ordinal number of the corresponding thread block in the initial mapping relationship is the second largest, and the ordinal number of the thread block having a mapping relationship with the matrix block having the largest ordinal number of the corresponding thread block in the initial mapping relationship is the largest. The corresponding thread block in the initial mapping relationship is the thread block having a mapping relationship with the matrix block in the initial mapping relationship.
[0031] Optionally, in the target mapping relationship, if λ1≥3, the ordinal number of the thread block having a mapping relationship with the λ1-row matrix block with the smallest x logical coordinate is less than the ordinal number of the thread block having a mapping relationship with the λ1-row matrix block with the second smallest x logical coordinate, the ordinal number of the thread block having a mapping relationship with the λ1-row matrix block with the second smallest x logical coordinate is less than the ordinal number of the thread block having a mapping relationship with the λ1-row matrix block with the third smallest x logical coordinate, and so on; for any of the λ1-row matrix blocks, the ordinal number of the thread block having a mapping relationship with the λ1×λ1 matrix blocks with the smallest y logical coordinate is less than the ordinal number of the thread block having a mapping relationship with the λ1×λ1 matrix blocks with the second smallest y logical coordinate, the ordinal number of the thread block having a mapping relationship with the λ1×λ1 matrix blocks with the second smallest y logical coordinate is less than the ordinal number of the thread block having a mapping relationship with the λ1×λ1 matrix blocks with the third smallest y logical coordinate, and so on; for any λ1×λ1 matrix blocks, the ordinal number of the thread block having a mapping relationship with the λ1 matrix blocks with the smallest x logical coordinate is less than the ordinal number of the thread block having a mapping relationship with the λ1 matrix blocks with the second smallest x logical coordinate, the ordinal number of the thread block having a mapping relationship with the λ1 matrix blocks with the second smallest x logical coordinate is less than the ordinal number of the thread block having a mapping relationship with the λ1 matrix blocks with the third smallest x logical coordinate, and so on; for any of the λ1 matrix blocks, the ordinal number of the thread block having a mapping relationship with the matrix block with the smallest y logical coordinate is less than the ordinal number of the thread block having a mapping relationship with the matrix block with the second smallest y logical coordinate, the ordinal number of the thread block having a mapping relationship with the matrix block with the second smallest y logical coordinate is less than the ordinal number of the thread block having a mapping relationship with the matrix block with the third smallest y logical coordinate, and so on.
[0032] S500, schedule the target thread blocks in ascending order of the corresponding ordinal numbers, where the thread block with the corresponding ordinal number i is used to perform the matrix multiplication operation of the matrix block with the logical coordinate p’ i When performing the matrix multiplication operation on A and B, B is fully loaded into the cache and A is partially loaded into the cache.
[0033] In this embodiment, the order of scheduling the thread blocks remains unchanged, and the logical coordinates of the matrix blocks also remain unchanged. What changes is the mapping relationship between the thread blocks and the matrix blocks; schedule the target thread blocks in ascending order of the ordinal numbers. If the coordinate sequence of the thread blocks is ((0,0),(1,0),…,(u x -1,0),(0,1),(1,1),…,(u x -1,1),…,(0,u y -1),(1,u y -1),…,(u x -1,u y-1)), where the i-th coordinate corresponds to the coordinate of the i-th scheduled thread block. Then, the thread block with coordinates (0, 0) is scheduled first, followed by the thread block with coordinates (1, 0), and so on. Finally, the thread block with coordinates (u x -1, u y -1) is scheduled.
[0034] As a first specific embodiment, A is a 10×2 matrix, B is a 2×4 matrix, sto B <sto cache <sto A , t x = t y = 2, u x = 5, u y = 2. The number of thread blocks for performing matrix multiplication on A and B is 10, P = ((0, 0), (1, 0), (2, 0), (3, 0), (4, 0), (0, 1), (1, 1), (2, 1), (3, 1), (4, 1)). Among them, the matrix block with logical coordinates (0, 0) is used to record the matrix multiplication result of the first and second rows of A and the first and second columns of B. The matrix block with logical coordinates (1, 0) is used to record the matrix multiplication result of the third and fourth rows of A and the first and second columns of B. The matrix block with logical coordinates (2, 0) is used to record the matrix multiplication result of the fifth and sixth rows of A and the first and second columns of B. The matrix block with logical coordinates (3, 0) is used to record the matrix multiplication result of the seventh and eighth rows of A and the first and second columns of B. The matrix block with logical coordinates (4, 0) is used to record the matrix multiplication result of the ninth and tenth rows of A and the first and second columns of B. The matrix block with logical coordinates (0, 1) is used to record the matrix multiplication result of the first and second rows of A and the third and fourth columns of B. The matrix block with logical coordinates (1, 1) is used to record the matrix multiplication result of the third and fourth rows of A and the third and fourth columns of B. The matrix block with logical coordinates (2, 1) is used to record the matrix multiplication result of the fifth and sixth rows of A and the third and fourth columns of B. The matrix block with logical coordinates (3, 1) is used to record the matrix multiplication result of the seventh and eighth rows of A and the third and fourth columns of B. The matrix block with logical coordinates (4, 1) is used to record the matrix multiplication result of the ninth and tenth rows of A and the third and fourth columns of B. e0 = 2, P' = ((0, 0), (2, 0), (4, 0), (1, 1), (3, 1), (1, 0), (3, 0), (0, 1), (2, 1), (4, 1)).
[0035] As a second specific embodiment, A is a 12×2 matrix, B is a 2×4 matrix, sto B <sto cache <sto A , t x = t y = 2, ux = 6, u y = 2, the number of thread blocks for performing matrix multiplication on A and B is 12, P = ((0,0),(1,0),(2,0),(3,0),(4,0),(5,0),(0,1),(1,1),(2,1),(3,1),(4,1),(5,1)), where the matrix block with logical coordinates (0,0) is used to record the matrix multiplication result of the 1st and 2nd rows of A and the 1st and 2nd columns of B, the matrix block with logical coordinates (1,0) is used to record the matrix multiplication result of the 3rd and 4th rows of A and the 1st and 2nd columns of B, the matrix block with logical coordinates (2,0) is used to record the matrix multiplication result of the 5th and 6th rows of A and the 1st and 2nd columns of B, the matrix block with logical coordinates (3,0) is used to record the matrix multiplication result of the 7th and 8th rows of A and the 1st and 2nd columns of B, the matrix block with logical coordinates (4,0) is used to record the matrix multiplication result of the 9th and 10th rows of A and the 1st and 2nd columns of B, the matrix block with logical coordinates (5,0) is used to record the matrix multiplication result of the 11th and 12th rows of A and the 1st and 2nd columns of B, the matrix block with logical coordinates (0,1) is used to record the matrix multiplication result of the 1st and 2nd rows of A and the 3rd and 4th columns of B, the matrix block with logical coordinates (1,1) is used to record the matrix multiplication result of the 3rd and 4th rows of A and the 3rd and 4th columns of B, the matrix block with logical coordinates (2,1) is used to record the matrix multiplication result of the 5th and 6th rows of A and the 3rd and 4th columns of B, the matrix block with logical coordinates (3,1) is used to record the matrix multiplication result of the 7th and 8th rows of A and the 3rd and 4th columns of B, the matrix block with logical coordinates (4,1) is used to record the matrix multiplication result of the 9th and 10th rows of A and the 3rd and 4th columns of B, the matrix block with logical coordinates (5,1) is used to record the matrix multiplication result of the 11th and 12th rows of A and the 3rd and 4th columns of B, e0 = 4, P' = ((0,0),(1,0),(4,0),(5,0),(2,1),(3,1),(2,0),(3,0),(0,1),(1,1),(4,1),(5,1)).
[0036] When performing matrix multiplication on the first matrix and the second matrix in this embodiment, first obtain the logical coordinate sequence P of matrix blocks when mapping according to a preset initial mapping relationship. The order of the logical coordinates of different matrix blocks in P is arranged according to the scheduling order of the thread blocks having a mapping relationship with the matrix blocks in the initial mapping relationship. That is, the i-th logical coordinate in P is the logical coordinate of the matrix block having a mapping relationship with the i-th scheduled thread block in the initial mapping relationship. Secondly, this embodiment judges the scales of the first matrix and the second matrix. If the conditions that the number of rows of the first matrix is greater than the number of columns of the second matrix, the storage space occupied by the second matrix is less than the storage space of the cache, and the storage space occupied by the first matrix is greater than the storage space of the cache are satisfied, then obtain the logical coordinate sequence P' of matrix blocks when mapping according to the target mapping relationship. The i-th logical coordinate in P' is the logical coordinate of the matrix block having a mapping relationship with the i-th scheduled thread block in the target mapping relationship. In the target mapping relationship of this embodiment, the ordinal number of the thread block having a mapping relationship with the matrix block for the multiplication result of the smaller rows in the first matrix and the second matrix is smaller, and the ordinal number of the thread block having a mapping relationship with the matrix block for the multiplication result of the larger rows in the first matrix and the second matrix is larger. And when performing matrix multiplication on the first matrix and the second matrix, the second matrix is fully loaded into the cache and the first matrix is partially loaded into the cache. This embodiment schedules the thread blocks in ascending order of the ordinal numbers of the thread blocks. Thus, during the process of performing matrix multiplication on the first matrix and the second matrix, the second matrix will be fully hit by the cache, and when multiplying different columns in the second matrix with a part of the same row in the first matrix, the part of the same row in the first matrix will also be repeatedly hit by the cache, thereby increasing the probability of cache hit when performing matrix multiplication on the first matrix and the second matrix. Since the reading speed of the cache is relatively high, the speed of the multiplication operation is increased.
[0037] Embodiment Two
[0038] The above Embodiment One is applicable to the scenario where m>n and sto B <sto cache <sto A This embodiment is applicable to the scenario where sto B >sto cache and sto A >sto cache of the scenario.
[0039] S300 further includes: If sto B >sto cache and sto A >sto cache , then enter S600.
[0040] S600, obtain the logical coordinate sequence P of the matrix block when mapping according to the first mapping relationship 0 , P 0 = (p 0 1, p 0 2, …, p 0 i , …, p 0 u ), p 0 i is the logical coordinate of C 0 i , C 0 i is the matrix block that has a mapping relationship with blo i when mapping according to the target mapping relationship, p 0 i = (p 0 i,x , p 0 i,y ), p 0 i,x is the x logical coordinate of C 0 i , p 0 i,y is the y logical coordinate of C 0 i ; in the first mapping relationship, β f,h is less than β f+f0,h , β f,h is less than β f,h+h0 , β f,h is the ordinal number when blo f,h is scheduled for the matrix multiplication operation of A and B, blo f,h is the thread block that has a mapping relationship with mat f,h in the first mapping relationship, mat f,h is the matrix block used to record the matrix multiplication result of the f-th row in A and the h-th column in B, β f+f0,h is the ordinal number when blo f+f0,h is scheduled for the matrix multiplication operation of A and B, blo f+f0,h is the thread block that has a mapping relationship with mat f+f0,h in the first mapping relationship, mat f+f0,h is the matrix block used to record the matrix multiplication result of the (f + f0)-th row in A and the h-th column in B; β f,h+h0 is the ordinal number when blo f,h+h0 is scheduled for the matrix multiplication operation of A and B, blo f,h+h0 is the thread block that has a mapping relationship with mat f,h+h0 in the first mapping relationship, mat f,h+h0It is a matrix block for recording the result of matrix multiplication of the f-th row in A and the (h + h0)-th column in B; the value range of f is from 1 to m, f0 is a preset second row number threshold, f0 ≥ 1, the value range of h is from 1 to n, and h0 is a preset second column number threshold, h0 ≥ 1.
[0041] In this embodiment, a matrix block is used to record the result of matrix multiplication operation of the t x rows of A and the t y columns of B, f0 = λ2 × t x , where λ2 is a preset second multiple, λ2 ≥ 2 and λ2 is an integer.
[0042] In this embodiment, h0 = λ2 × t y .
[0043] Optionally, λ2 = 2, λ2 = 3 or λ2 = 4.
[0044] Optionally, in the first mapping relationship, if λ2 = 2, then the ordinal number of the thread block having a mapping relationship with the two columns of matrix blocks with the smallest y logical coordinates is less than the ordinal number of the thread block having a mapping relationship with the two columns of matrix blocks with the second smallest y logical coordinates, the ordinal number of the thread block having a mapping relationship with the two columns of matrix blocks with the second smallest y logical coordinates is less than the ordinal number of the thread block having a mapping relationship with the two columns of matrix blocks with the third smallest y logical coordinates, and so on; for any of the two columns of matrix blocks, the ordinal number of the thread block having a mapping relationship with the four matrix blocks with the smallest x logical coordinates is less than the ordinal number of the thread block having a mapping relationship with the four matrix blocks with the second smallest x logical coordinates, the ordinal number of the thread block having a mapping relationship with the four matrix blocks with the second smallest x logical coordinates is less than the ordinal number of the thread block having a mapping relationship with the four matrix blocks with the third smallest x logical coordinates, and so on; for any of the four matrix blocks, the ordinal number of the thread block having a mapping relationship with the matrix block with the smallest ordinal number of the corresponding thread block in the initial mapping relationship is the smallest, the ordinal number of the thread block having a mapping relationship with the matrix block with the second smallest ordinal number of the corresponding thread block in the initial mapping relationship is the second smallest, the ordinal number of the thread block having a mapping relationship with the matrix block with the second largest ordinal number of the corresponding thread block in the initial mapping relationship is the second largest, and the ordinal number of the thread block having a mapping relationship with the matrix block with the largest ordinal number of the corresponding thread block in the initial mapping relationship is the largest. The corresponding thread block in the initial mapping relationship is the thread block having a mapping relationship with the matrix block in the initial mapping relationship.
[0045] Optionally, in the first mapping relationship, if λ2≥3, then the ordinal number of the thread block having a mapping relationship with the λ1 column matrix block with the smallest y logical coordinate is less than the ordinal number of the thread block having a mapping relationship with the λ1 column matrix block with the second smallest y logical coordinate, and the ordinal number of the thread block having a mapping relationship with the λ1 column matrix block with the second smallest y logical coordinate is less than the ordinal number of the thread block having a mapping relationship with the λ1 column matrix block with the third smallest y logical coordinate, and so on; for any one of the λ1 column matrix blocks, the ordinal number of the thread block having a mapping relationship with the λ1×λ1 matrix blocks with the smallest x logical coordinate is less than the ordinal number of the thread block having a mapping relationship with the λ1×λ1 matrix blocks with the second smallest x logical coordinate, and the ordinal number of the thread block having a mapping relationship with the λ1×λ1 matrix blocks with the second smallest x logical coordinate is less than the ordinal number of the thread block having a mapping relationship with the λ1×λ1 matrix blocks with the third smallest x logical coordinate, and so on; for any one of the λ1×λ1 matrix blocks, the ordinal number of the thread block having a mapping relationship with the λ1 matrix blocks with the smallest x logical coordinate is less than the ordinal number of the thread block having a mapping relationship with the λ1 matrix blocks with the second smallest x logical coordinate, and the ordinal number of the thread block having a mapping relationship with the λ1 matrix blocks with the second smallest x logical coordinate is less than the ordinal number of the thread block having a mapping relationship with the λ1 matrix blocks with the third smallest x logical coordinate, and so on; for any one of the λ1 matrix blocks, the ordinal number of the thread block having a mapping relationship with the matrix block with the smallest y logical coordinate is less than the ordinal number of the thread block having a mapping relationship with the matrix block with the second smallest y logical coordinate, and the ordinal number of the thread block having a mapping relationship with the matrix block with the second smallest y logical coordinate is less than the ordinal number of the thread block having a mapping relationship with the matrix block with the third smallest y logical coordinate, and so on.
[0046] S700, schedule the target thread blocks in ascending order of the corresponding ordinal numbers, where the thread block with the corresponding ordinal number i is used to perform the matrix multiplication operation of the matrix block with the logical coordinate p 0 i When performing the matrix multiplication operation on A and B, both A and B are partially loaded into the cache.
[0047] In this embodiment, the scheduling order of the thread blocks remains unchanged, and the logical coordinates of the matrix blocks also remain unchanged. What changes is the mapping relationship between the thread blocks and the matrix blocks; schedule the target thread blocks in ascending order of the ordinal numbers. If the coordinate sequence of the thread blocks is ((0,0),(1,0),…,(u x -1,0),(0,1),(1,1),…,(u x -1,1),…,(0,u y -1),(1,u y -1),…,(u x -1,u y-1)), where the i-th coordinate corresponds to the coordinate of the i-th scheduled thread block. Then, the thread block with the coordinate (0, 0) is scheduled first, followed by the thread block with the coordinate (1, 0), and so on. Finally, the thread block with the coordinate (u x -1, u y -1) is scheduled.
[0048] In this embodiment, if sto B and sto A are both small, A and B can be loaded into the cache simultaneously. Then, steps S600 - S700 are not executed, and the following steps are performed: The target thread blocks are scheduled in ascending order of the corresponding ordinal numbers. Among them, the thread block with the corresponding ordinal number i is used to perform the matrix multiplication operation of the matrix block with the logical coordinate p i . When performing the matrix multiplication operation on A and B, both A and B are fully loaded into the cache.
[0049] As a first specific embodiment, A is a 12×2 matrix, B is a 2×12 matrix, sto B > sto cache and sto A > sto cache , t x = t y = 2, u x = 6, u y= 6, the number of thread blocks for performing matrix multiplication on A and B is 36, P = ((0,0),(1,0),(2,0),(3,0),(4,0),(5,0),(0,1),(1,1),(2,1),(3,1),(4,1),(5,1),(0,2),(1,2),(2,2),(3,2),(4,2),(5,2),(0,3),(1,3),(2,3),(3,3),(4,3),(5,3),(0,4),(1,4),(2,4),(3,4),(4,4),(5,4),(0,5),(1,5),(2,5),(3,5),(4,5),(5,5)), where the matrix block with logical coordinates (0,0) is used to record the matrix multiplication result of the 1st and 2nd rows of A and the 1st and 2nd columns of B, the matrix block with logical coordinates (1,0) is used to record the matrix multiplication result of the 3rd and 4th rows of A and the 1st and 2nd columns of B, the matrix block with logical coordinates (2,0) is used to record the matrix multiplication result of the 5th and 6th rows of A and the 1st and 2nd columns of B, the matrix block with logical coordinates (3,0) is used to record the matrix multiplication result of the 7th and 8th rows of A and the 1st and 2nd columns of B, the matrix block with logical coordinates (4,0) is used to record the matrix multiplication result of the 9th and 10th rows of A and the 1st and 2nd columns of B, the matrix block with logical coordinates (5,0) is used to record the matrix multiplication result of the 11th and 12th rows of A and the 1st and 2nd columns of B, and so on, the matrix block with logical coordinates (0,5) is used to record the matrix multiplication result of the 1st and 2nd rows of A and the 11th and 12th columns of B, the matrix block with logical coordinates (1,5) is used to record the matrix multiplication result of the 3rd and 4th rows of A and the 11th and 12th columns of B, the matrix block with logical coordinates (2,5) is used to record the matrix multiplication result of the 5th and 6th rows of A and the 11th and 12th columns of B, the matrix block with logical coordinates (3,5) is used to record the matrix multiplication result of the 7th and 8th rows of A and the 11th and 12th columns of B, the matrix block with logical coordinates (4,5) is used to record the matrix multiplication result of the 9th and 10th rows of A and the 11th and 12th columns of B, the matrix block with logical coordinates (5,5) is used to record the matrix multiplication result of the 11th and 12th rows of A and the 11th and 12th columns of B, λ2 = 2, f0 = 4, h0 = 4, P 0= ((0,0),(1,0),(4,0),(5,0),(2,1),(3,1),(2,0),(3,0),(0,1),(1,1),(4,1),(5,1),(0,2),(1,2),(4,2),(5,2),(2,3),(3,3),(2,2),(3,2),(0,3),(1,3),(4,3),(5,3),(0,4),(1,4),(4,4),(5,4),(2,5),(3,5),(2,4),(3,4),(0,5),(1,5),(4,5),(5,5)).
[0050] As a second specific embodiment, A is a 12×2 matrix, B is a 2×12 matrix, sto B >sto cache and sto A >sto cache , t x = t y = 2, u x = 6, u y= 6, the number of thread blocks for performing matrix multiplication on A and B is 36, P = ((0,0),(1,0),(2,0),(3,0),(4,0),(5,0),(0,1),(1,1),(2,1),(3,1),(4,1),(5,1),(0,2),(1,2),(2,2),(3,2),(4,2),(5,2),(0,3),(1,3),(2,3),(3,3),(4,3),(5,3),(0,4),(1,4),(2,4),(3,4),(4,4),(5,4),(0,5),(1,5),(2,5),(3,5),(4,5),(5,5)), where the matrix block with logical coordinates (0,0) is used to record the matrix multiplication result of the 1st and 2nd rows of A and the 1st and 2nd columns of B, the matrix block with logical coordinates (1,0) is used to record the matrix multiplication result of the 3rd and 4th rows of A and the 1st and 2nd columns of B, the matrix block with logical coordinates (2,0) is used to record the matrix multiplication result of the 5th and 6th rows of A and the 1st and 2nd columns of B, the matrix block with logical coordinates (3,0) is used to record the matrix multiplication result of the 7th and 8th rows of A and the 1st and 2nd columns of B, the matrix block with logical coordinates (4,0) is used to record the matrix multiplication result of the 9th and 10th rows of A and the 1st and 2nd columns of B, the matrix block with logical coordinates (5,0) is used to record the matrix multiplication result of the 11th and 12th rows of A and the 1st and 2nd columns of B, and so on, the matrix block with logical coordinates (0,5) is used to record the matrix multiplication result of the 1st and 2nd rows of A and the 11th and 12th columns of B, the matrix block with logical coordinates (1,5) is used to record the matrix multiplication result of the 3rd and 4th rows of A and the 11th and 12th columns of B, the matrix block with logical coordinates (2,5) is used to record the matrix multiplication result of the 5th and 6th rows of A and the 11th and 12th columns of B, the matrix block with logical coordinates (3,5) is used to record the matrix multiplication result of the 7th and 8th rows of A and the 11th and 12th columns of B, the matrix block with logical coordinates (4,5) is used to record the matrix multiplication result of the 9th and 10th rows of A and the 11th and 12th columns of B, the matrix block with logical coordinates (5,5) is used to record the matrix multiplication result of the 11th and 12th rows of A and the 11th and 12th columns of B, λ2 = 3, f0 = 6, h0 = 6, P 0= ((0,0),(3,0),(0,1),(3,1),(0,2),(3,2),(1,0),(4,0),(1,1),(4,1),(1,2),(4,2),(2,0),(5,0),(2,1),(5,1),(2,2),(5,2),(0,3),(3,3),(0,4),(3,4),(0,5),(3,5),(1,3),(4,3),(1,4),(4,4),(1,5),(4,5),(2,3),(5,3),(2,4),(5,4),(2,5),(5,5)).
[0051] When performing matrix multiplication on the first matrix and the second matrix in this embodiment, first obtain the logical coordinate sequence P of matrix blocks when mapping according to a preset initial mapping relationship. The order of arrangement of the logical coordinates of different matrix blocks in P is arranged according to the scheduling order of thread blocks having a mapping relationship with the matrix blocks in the initial mapping relationship. That is, the i-th logical coordinate in P is the logical coordinate of the matrix block having a mapping relationship with the i-th scheduled thread block in the initial mapping relationship. Secondly, the present invention judges the scales of the first matrix and the second matrix. If the condition that the storage space occupied by the first matrix and the storage space occupied by the second matrix are both greater than the storage space of the cache is satisfied, then obtain the logical coordinate sequence P of matrix blocks when mapping according to the first mapping relationship 0 , P 0 The i-th logical coordinate in it is the logical coordinate of the matrix block having a mapping relationship with the i-th scheduled thread block in the first mapping relationship. In the first mapping relationship of this embodiment, the ordinal number of the thread block having a mapping relationship with the matrix block of the multiplication result of the smaller rows in the first matrix and the smaller columns in the second matrix is smaller, and the ordinal number of the thread block having a mapping relationship with the matrix block of the multiplication result of the larger rows in the first matrix and the smaller columns in the second matrix is larger. The ordinal number of the thread block having a mapping relationship with the matrix block of the multiplication result of the smaller rows in the first matrix and the larger columns in the second matrix is larger. And when performing matrix multiplication on the first matrix and the second matrix, the first matrix and the second matrix are partially loaded into the cache. Thus, when matrix multiplication operations are performed on thread blocks with adjacent scheduling orders during the process of performing matrix multiplication on the first matrix and the second matrix, the elements obtained in the first matrix and the second matrix are relatively consistent, so that the probability of the elements in the first matrix and the second matrix being cached hits will increase. Since the reading speed of the cache is relatively high, the speed of the multiplication operation is thus improved.
[0052] Although some specific embodiments of the present invention have been described in detail by way of examples, those skilled in the art should understand that the above examples are only for illustration purposes and not for limiting the scope of the present invention. Those skilled in the art should also understand that various modifications can be made to the embodiments without departing from the scope and spirit of the present invention. The scope of the present invention is defined by the appended claims.
Claims
1. A processing method for matrix multiplication, characterized in that, The method includes the following steps: S100, obtaining a first matrix A and a second matrix B to perform matrix multiplication operation, where A is an m×k matrix, B is a k×n matrix, m is the number of rows of A, k is the number of columns of A, and n is the number of columns of B; S200, obtain the logical coordinate sequence P of the matrix blocks when mapping according to the preset initial mapping relationship, P = (p1, p2, …, p i , …, p u ), where p i is the logical coordinate of C i , and C i is the matrix block that has a mapping relationship with blo i when mapping according to the preset initial mapping relationship. blo i is the i-th scheduled thread block during the matrix multiplication operation of A and B, and p i = (p i,x , p i,y ), where p i,x is the x logical coordinate of C i , and p i,y is the y logical coordinate of C i . The value range of i is from 1 to u, where u is the number of target thread blocks, and the target thread blocks are the thread blocks for the matrix multiplication operation of A and B; S300, if m > n and sto B <sto cache <sto A , then go to S400; sto B is the storage space size occupied by A, sto cache is the storage space size of the cache, sto A is the storage space size occupied by B; S400, obtain the logical coordinate sequence P' of the matrix block when mapping according to the target mapping relationship, P' = (p'1, p'2, …, p' i , …, p' u ), where p' i is the logical coordinate of C' i , and C' i is the matrix block that has a mapping relationship with blo i when mapping according to the target mapping relationship. p' i = (p' i,x , p' i,y ), where p' i,x is the x logical coordinate of C' i , and p' i,y is the y logical coordinate of C' i . In the target mapping relationship, β e,all is less than β e+e0,all . β e,all is the ordinal number when blo e,all is scheduled for the matrix multiplication operation of A and B. blo e,all is the thread block that has a mapping relationship with mat e,all in the target mapping relationship. mat e,all is the matrix block used to record the matrix multiplication result of the e-th row in A and B. β e+e0,all is the ordinal number when blo e+e0,all is scheduled for the matrix multiplication operation of A and B. blo e+e0,all is the thread block that has a mapping relationship with mat e+e0,all in the target mapping relationship. mat e+e0,all is the matrix block used to record the matrix multiplication result of the (e + e0)-th row in A and B. The value range of e is from 1 to m, and e0 is a preset first row number threshold, e0 ≥ 1; S500 schedules target thread blocks in ascending order according to the corresponding ordinal numbers, where the thread block with the corresponding ordinal number i is used to perform matrix multiplication operations on the matrix block with logical coordinates p’ i When performing matrix multiplication on A and B, B is fully loaded into the cache and A is partially loaded into the cache.
2. The processing method for matrix multiplication according to claim 1, wherein u = u x × u y , u x is the number of thread blocks included in the thread grid for performing matrix multiplication on A and B in the x direction, u x = m / t x , t x is the number of rows corresponding to each thread block for performing matrix multiplication on A and B as preset, u y is the number of thread blocks included in the thread grid for performing matrix multiplication on A and B in the y direction, u y = n / t y , t y is the number of columns corresponding to each thread block for performing matrix multiplication on A and B as preset.
3. The processing method for matrix multiplication according to claim 2, characterized in that, e0 = λ1 × t x , where λ1 is a preset first multiple, λ1 ≥ 1 and λ1 is an integer.
4. The processing method for matrix multiplication according to claim 2, wherein, P = ((0, 0), (1, 0), …, (u x - 1, 0), (0, 1), (1, 1), …, (u x - 1, 1), …, (0, u y - 1), (1, u y - 1), …, (u x - 1, u y - 1)).
5. The processing method for matrix multiplication according to claim 4, wherein When e0 = t x , p’ i is equal to q(i) logical coordinates in P. If 1 ≤ i ≤ u x , q(i) = (i - 1) × u y + 1; if i > u x , q(i) = (%(i / u x ) - 1) × u y + floor(i / u x ), where %() is the remainder function and floor() is the floor function.
6. The processing method for matrix multiplication according to claim 2, wherein β in the target mapping relationship e,g less than β e,g+g0 , β e,g is the ordinal number of blo e,g scheduled during the matrix multiplication operation of A and B, and blo e,g is the thread block having a mapping relationship with mat e,g in the target mapping relationship, and mat e,g is the matrix block used to record the matrix multiplication result of the e-th row in A and the g-th column in B, and β e,g+g0 is the ordinal number of blo e,g+g0 scheduled during the matrix multiplication operation of A and B, and blo e,g+g0 is the thread block having a mapping relationship with mat e,g+g0 in the target mapping relationship, and mat e,g+g0 is the matrix block used to record the matrix multiplication result of the e-th row in A and the (g + g0)-th column in B; the value range of g is from 1 to n, and g0 is a preset first column number threshold, where g0 ≥ 1.
7. The processing method for matrix multiplication according to claim 6, characterized in that, g0 = λ1 × t y .