Apparatus, method, and medium for distributing processing threads for matrix-matrix multiplication
By performing adaptive threading and contention mitigation techniques on the K dimension in matrix-matrix multiplication, the problems of low cache utilization and slow processing in the prior art are solved, and more efficient matrix multiplication processing is achieved.
Patent Information
- Application Number
- CN202110393705.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-10-29
- Filing Date
- 2021-04-13
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2041-04-13
AI Technical Summary
In the prior art, when performing matrix-matrix multiplication, it is difficult to effectively utilize cache, resulting in slower processing speed.
By performing adaptive threading in the K dimension, the matrix block size is dynamically adjusted to improve cache memory utilization and reduce memory contention between processing threads through contention mitigation engines.
Improve the processing efficiency of matrix-matrix multiplication, reduce processing time, and optimize cache utilization.
Smart Images

Figure CN114428936B_ABST
Abstract
Description
Technical Field
[0001] Apparatus, method, and medium for distributing processing threads for matrix-matrix multiplication are disclosed. Background Art
[0002] A given computer system may include a specialized math library that provides routines to support arithmetic operations in a variety of engineering, data mining, digital signal processing, data analysis, and machine learning applications. One such routine may implement a version of the Generalized Matrix-Matrix Multiplication (GEMM) algorithm for performing matrix-matrix multiplication. For some applications, the matrices involved in matrix-matrix multiplication may be relatively large (e.g., a given matrix may have tens of thousands or hundreds of thousands of rows and columns, or more), resulting in a relatively large number of floating-point multiplication operations for each matrix-matrix multiplication. Summary of the Invention
[0003] A non-transitory storage medium storing machine-readable instructions that, when executed by a machine, cause the machine to: access data representing a first dimension of a first matrix, a second dimension of a second matrix, and a third dimension shared by the first matrix and the second matrix; determine a plurality of candidate factorizations of the first matrix, wherein the candidate factorizations differ in size relative to each other along the first dimension and the third dimension; determine an associated fitness value for each candidate factorization of the plurality of candidate factorizations; select, based on the fitness value, a candidate factorization from the plurality of candidate factorizations to provide a selected candidate factorization; and provide data representing the selected candidate factorization such that processing threads are assigned to submatrices of the first matrix, wherein the processing threads determine a third matrix based on the multiplication of the first matrix and the second matrix.
[0004] An apparatus for matrix-matrix multiplication includes: a processor; and a memory for storing instructions that, when executed by the processor, cause the processor to: perform threading of a first matrix along a first dimension of the first matrix and a second dimension of the matrix, wherein the threading represents a block size of the first matrix for processing threads to be assigned to a multiplication algorithm for determining a third matrix representing the product of the first matrix and a second matrix, the block size including a first block size along the first dimension and a second block size along the second dimension, and the second matrix shares the second dimension with the first matrix; and provide data representing the first block size and the second block size to the multiplication algorithm.
[0005] A method for matrix-matrix multiplication, comprising: at least one hardware processor accessing data representing a first matrix, a second matrix, a first block size, and a second block size, wherein the first block size represents a sub-matrix decomposition size of the first matrix along a first dimension, the second block size represents a sub-matrix decomposition size of the first matrix along a second dimension, and the second matrix shares the second dimension with the first matrix; and the at least one hardware processor multiplying the first matrix and the second matrix to determine a third matrix, wherein the multiplying includes allocating processing threads to sub-matrices of the first matrix based on the first block size and the second block size. Description of the Drawings
[0006] Figure 1 is a schematic diagram of a computer system according to an example embodiment.
[0007] Figure 2 is depicting by a Figure 1 computer system according to an example embodiment of the thread decomposition engine and contention mitigation engine used by the process flowchart.
[0008] Figure 3 is an illustration of a nested processing loop of a general matrix-matrix multiplication (GEMM) algorithm according to an example embodiment.
[0009] Figure 4 is an illustration of matrix-matrix multiplication processing performed by K groups of processing threads according to an example embodiment.
[0010] Figure 5A and 5B shows a flowchart of a process for performing adaptive K-dimensional threading according to an example embodiment.
[0011] Figure 6 is a flowchart depicting a process for determining a sub-block size for a processing thread sub-loop according to an example embodiment.
[0012] Figure 7 is a flowchart depicting a process for determining whether to recommend using a local temporary buffer to mitigate processing thread contention according to an example embodiment.
[0013] Figure 8 is an illustration of a non-transitory storage medium storing machine-executable instructions according to an example embodiment, the machine-executable instructions when executed by a machine cause the machine to provide data representing decomposition for processing thread allocation.
[0014] Figure 9 is a schematic diagram of a device according to an example embodiment, the device including a processor that provides data representing block sizes for threading by a multiplication algorithm.
[0015] Figure 10 is a flowchart depicting a process of multiplying matrices according to an example embodiment. Detailed Embodiment
[0016] The multiplication of two matrices (referred to herein as “matrix-matrix multiplication” or “matrix multiplication”) can be performed in a computer system having one or more multi-core central processing unit (CPU) semiconductor packages (or “chips”). The CPU semiconductor package can include multiple CPU processing cores that can access local on-chip memory. Additionally, the CPU semiconductor package can employ a non-uniform memory access (NUMA) architecture. Generally, a NUMA architecture recognizes that the access time of a processing node to local memory is faster than the access time to non-local memory. Thus, in a NUMA architecture, processing cores can be grouped according to NUMA domains, such that the processing cores of each NUMA domain use local memory access to perform most of their computations. A NUMA domain can have one or more cache levels, enabling more efficient access to local on-chip memory by storing recently accessed data in faster cache memory.
[0017] For matrix-matrix computations, the computer system can employ a general matrix-matrix multiplication (GEMM) algorithm that relies on different processing threads (on corresponding NUMA nodes and sockets) for performing different parts of the multiplication. Matrix-matrix multiplication can involve multiplying relatively large matrices that contain tens of thousands, or hundreds of thousands (or more) of rows and columns.
[0018] To accommodate such computationally intensive operations, a process called “threading” can be used to distribute different parts of the matrix-matrix multiplication processing workload among the processing threads of the computer system. In this context, a “processing thread” (or “thread”) refers to a unit of machine-executable instructions assigned to a processing core of the computer system. Processing threads can be executed in parallel (i.e., simultaneously) by the processing cores assigned to them.
[0019] In the context of the present disclosure, “threading” refers to distributing the processing workload of matrix-matrix multiplication among different processing threads of the computer system. Threading can, for example, partition the input matrices involved in matrix-matrix multiplication and assign the resulting partitions to corresponding processing threads.
[0020] One type of threading is "M / N" threading, where matrix partitioning is performed along the M dimension and the N dimension to form corresponding M×N sub-matrices or blocks, and these sub-matrices or blocks are assigned to different processing threads. In this article, the "M dimension" refers to the row dimension of the input matrix A of the matrix-matrix multiplication, and the "N dimension" refers to the column dimension of the input matrix B of the matrix-matrix multiplication. The definition of matrix-matrix multiplication is as follows: A×B = C, where "C" represents the output matrix, or the result of the matrix-matrix multiplication. Therefore, the dimension of the output matrix C is M×N, or M rows multiplied by N columns. The input matrix A and the input matrix B share a common dimension, which is referred to as the "K dimension" in this article. More specifically, the K dimension is the column dimension of the input matrix A and the row dimension of the input matrix B.
[0021] The challenge with using M / N threading is that the result data units corresponding to the sub-matrices may be too small to effectively utilize the cache (e.g., the cache line boundary of the last-level cache (LLC) of the CPU). Therefore, due to underutilization of the cache, there may not be enough cached data to amortize the cost of moving data from the main system memory, resulting in slower processing of the matrix-matrix multiplication.
[0022] According to the example embodiments described herein, threading is performed on the K dimension, i.e., the dimension shared by the input matrices. More specifically, according to the example embodiments, a thread decomposition with an adaptive K-dimension chunking engine (referred to as the "thread decomposition engine" herein) provides processing thread assignments to the GEMM algorithm. The corresponding sub-matrices or sub-blocks are sized to improve cache memory utilization, and generally, reduce the processing time of the matrix-matrix multiplication. According to the example embodiments, the thread decomposition engine assigns processing threads to sub-matrices or blocks based on a two-dimensional matrix partitioning, where one dimension is the K dimension and the other dimension is either the M dimension or the N dimension.
[0023] The K-dimension-based threading described herein can be one of two forms: "M / K threading", which refers to the threading that assigns processing threads to the respective matrix blocks obtained by partitioning along the M dimension and the K dimension; and "N / K threading", which refers to the threading that assigns processing threads to the matrix blocks obtained by partitioning the matrix along the N dimension and the K dimension. The M / K or N / K processing thread assignments result in block sizes that are closer to or correspond to the optimal cache block size, i.e., block sizes that result in more cache hits and better utilization of the cache memory when performing matrix-matrix multiplication compared to M / N threading. Additionally, according to the example embodiments, the thread decomposition engine performs partitioning along the K dimension in an adaptive manner that takes into account the optimal cache block size.
[0024] Also as described herein, according to an example embodiment, a contention mitigation engine provides further information to the GEMM algorithm, which reduces (or eliminates) memory contention among processing threads when processing matrix-matrix multiplication. More specifically, according to an example embodiment, the contention mitigation engine further subdivides or partitions along the M dimension and / or the N dimension to create sub-blocks of each M×N block of the output matrix C. Each processing thread can "sub-loop" on the corresponding sub-block. As a result of the sub-blocks and sub-loops, memory resource contention among "K groups" of processing threads can be reduced. In this case, the "K groups" of processing threads refer to all processing threads that contribute to a given M×N block of the output matrix C.
[0025] In addition, according to an example embodiment, as a further measure to reduce memory resource contention, the contention mitigation engine advises whether the K groups of processing threads should each use a local temporary buffer (i.e., a "scratchpad" buffer) to determine the results of the corresponding M×N block of the output matrix C. If the K groups of processing threads use local temporary buffers, then each processing thread in the K groups uses a local temporary buffer to determine its contribution to the M×N block of the output matrix C, and the contributions from the local temporary buffers can then be transferred to the final output buffer of the output matrix C by adding and updating.
[0026] Referring Figure 1 , as a more specific example, according to some embodiments, the computer system 100 may include a GEMM engine 110. In this case, the "GEMM engine 110" represents a software-driven engine in which multiple parallel processing threads 112 jointly apply the GEMM algorithm to perform matrix-matrix multiplication. In conjunction with Figure 1 Referring Figure 2 , for the example embodiments described herein, matrix-matrix multiplication refers to the multiplication of the input matrix A 210 and the input matrix B 214 to determine the output matrix C 220. In other words, the GEMM engine 110 determines the product described by the following matrix-matrix multiplication: A×B = C.
[0027] The input matrix A 210 has rows along the row dimension M (i.e., M rows indexed by M indices) and columns along the K dimension (i.e., K columns indexed by K indices). The input matrix B 214 has rows along the K dimension (i.e., K rows indexed by K indices) and columns along the N dimension (i.e., N columns indexed by N indices). Thus, the input matrix A 210 and the input matrix B 214 share the common dimension K. The output matrix C 220 has rows along the M dimension (i.e., M rows indexed by M indices) and columns along the N dimension (i.e., N columns indexed by N indices).
[0028] To configure the GEMM engine 110 ( Figure 1)To perform matrix-matrix multiplication, the GEMM engine 110 receives data 250 which, according to an example embodiment, represents the allocation of processing threads and, as further described herein, can represent other parameters that further enhance matrix-matrix multiplication processing. According to an example embodiment, at least a portion of the data 250 is provided by a thread decomposition engine having an adaptive K-threading engine 114 (referred to herein as the "thread decomposition engine 114"). As described herein, the thread decomposition engine 114 uses certain criteria to analyze the row and column sizes of matrices 210, 214, and 220 to provide parameters (referred to herein as "BSM", "BSN", and "BSK") that represent the matrix partitioning for the allocation of processing threads 112. According to an example embodiment, partitioning is performed along the K dimension and either along the N dimension or the M dimension.
[0029] The BSM parameter (referred to herein as the "BSM block size") represents the number of rows in each matrix partition or block along the M dimension. The BSN parameter (referred to herein as the "BSN block size") represents the number of columns in each matrix block along the N dimension. The BSK parameter (referred to herein as the "BSK block size") represents the number of rows / columns in the matrix partition or block along the K dimension (depending on whether it is the input matrix A 210 or the input matrix B 214).
[0030] According to an example embodiment, for a given matrix-matrix multiplication, the thread decomposition engine 114 performs M / K threading (corresponding to matrix partitioning along the M and K dimensions) or N / K threading (corresponding to matrix partitioning along the N and K dimensions) to determine the matrix blocks assigned to the processing threads 112. For N / K threading, the thread decomposition engine 114 partitions the input matrix B 214 into blocks of size BSK×BSN (i.e., each block has a size of BSK rows by BSN columns); and each BSK×BSN block of the input matrix B 214 is assigned to a different corresponding processing thread 112. More specifically, the thread decomposition engine 114 provides the following block sizes as a result of N / K threading: the BSN block size that represents the number of columns in each matrix block along the N dimension; the BSK block size that represents the number of rows / columns in each matrix block along the K dimension; and the default BSM block size, since the partitioning for work sharing does not occur in the M dimension for N / K threading. Each processing thread 112 processes the operations of the matrix-matrix multiplication related to the BSK×BSN block of the input matrix B 214 assigned to it to obtain the corresponding BSM×BSN block of the output matrix 220.
[0031] For M / K threading, the thread decomposition engine 114 partitions the input matrix A 210 into blocks, and each BSM×BSK block of the input matrix A 210 is assigned to a different corresponding processing thread 112. More specifically, the thread decomposition engine 114 provides the following block sizes for M / K threading: the BSM block size representing the number of rows of each matrix block along the M dimension; the BSK block size representing the number of rows / columns of each matrix block along the K dimension; and the default BSN block size, since partitioning for work sharing does not occur in the N dimension for M / K threading. Each processing thread 112 processes the operations of matrix-matrix multiplication related to the BSM×BSK block of the input matrix A 210 assigned to it to obtain the corresponding BSM×BSN block of the output matrix 220.
[0032] As a more specific example, the thread decomposition engine 114 can apply N / K threading to determine the assignment of 16 processing threads 112. For this example, the matrices can have the following dimensions: the input matrix A 210 can have 96 rows along the M dimension and 3840 columns along the K dimension, represented by the M×K dimension of "(96×3840)"; the input matrix B 214 can have the K×N dimension of (3840×192); and the output matrix C 220 can have the corresponding M×N dimension of (96×192). As further described herein, the thread decomposition engine 114 can estimate these dimensions and determine that the BSM, BSN, and BSK block sizes are 96, 48, and 960 respectively. In this way, the 16 processing threads 112 of this example are assigned to 16 48×960 blocks of the input matrix B.
[0033] Combined Figure 1 Referencing Figure 3 According to some embodiments, the GEMM engine 110 processes matrix-matrix multiplication following nested processing loops. In this regard, Figure 3 depicts the outer m c loop 304, which refers to the processing in the M dimension. For M / K threading, when the processing thread 112 indexes along the K dimension in the inner corresponding k c loop 308 and along the N dimension in the inner n c loop 312 respectively, a given processing thread 112 processes the operations of the BSM×BSK block assigned to it. For N / K threading, when the processing thread 112 indexes along the M dimension in the corresponding m c loop 304 and along the K dimension in the k c loop 308 respectively, a given processing thread 112 processes the operations of the BSN×BSK block assigned to it.
[0034] Figure 3 Also depicted are the following three nested innermost loops, which are part of the macro kernel of the GEMM engine 110: mr Loop 324, n r Loop 328 and k r Loop 332. In Loops 324, 328, and 332, for a given BSM×BSN block of the output matrix C 220( Figure 2 ), the processing thread 112 iterates in the M dimension, N dimension, and K dimension, respectively. The sub-loops introduce two processing loops outside the macro-kernel of the GEMM engine 110: an m_sub loop 316 corresponding to the sub-loop in the M dimension and an n_sub loop 320 corresponding to the sub-loop in the N dimension. More specifically, in conjunction with Figure 1 Reference Figure 2 , according to an example embodiment, the contention mitigation engine 118 of the computer system 100 subdivides the BSM and BSN block sizes into corresponding sub-matrices or sub-blocks, each having BSM sub and BSN sub dimensions. The GEMM engine 110 includes the following nested loops to execute the sub-loops: an m_sub loop 316 to sub-loop over sub-blocks of size BSM sub of each matrix block, and an n_sub loop 320 to sub-loop over sub-blocks of size BSN sub of each matrix block. According to an example embodiment, BSM sub and BSN sub can be BSMR and BSNR, respectively. As further described herein, according to an example embodiment, when the processing threads 112 jointly process a BSM×BSN block of the output matrix C 220, the sub-loops can prevent memory contention among the K groups of processing threads 112.
[0035] In addition to providing BSM sub and BSN subBeyond the size being part of the data 250, according to an example embodiment, the contention mitigation engine 118 can provide advice to the GEMM engine 110 on whether to use a temporary local output buffer. In this case, the "temporary output buffer" refers to a "scratchpad" buffer that can be used by each of the K groups of processing threads that jointly process a particular BSM×BSN block of the output matrix C 220. Instead of each processing thread 112 in the K groups of processing threads 112 updating the BSM×BSN block in the final result buffer, the processing threads 112 instead update the temporary output buffer. Additionally, according to an example embodiment, the temporary local output buffer is optimally aligned with the cache boundary. When the K groups of processing threads 112 have completed processing the BSM×BSN block of the output matrix C 220 and the temporary local output buffer stores the complete result of the sub-block of the output matrix, then the contents of the temporary buffer can be transferred (e.g., via the update routine and add-update of the GEMM engine 110) to the final output buffer of the matrix C 220.
[0036] Referring back Figure 1 , according to many possible embodiments, the GEMM engine 110, the thread decomposition engine 114, and the contention mitigation engine 118 can be provided by any of several different computer architectures. For Figure 1 an example embodiment, the computer system 100 includes one or more central processing units (CPUs) 102 (e.g., a CPU semiconductor package or "chip"), where each CPU 102 includes a set of processing cores 101. According to some embodiments, several processing cores 101 on a given CPU 102 can form corresponding NUMA domains; and each CPU 102 can have multiple NUMA domains.
[0037] According to a further example embodiment, processing cores other than central processor processing cores can be used. In this way, according to a further embodiment, the processing cores 101 can be graphics processing unit (GPU) cores, field programmable gate arrays (FPGAs), node accelerator cores, etc.
[0038] According to an example embodiment, the computer system 100 can be formed by one or more physical machines, where each physical machine is made of or formed by actual software and actual hardware. In this way, according to an example embodiment, each physical machine can include one or more CPUs 102 and memory. The memory of the physical machines jointly passes through the memory 106 in Figure 2is represented in. As an example, the memory 106 may store machine-executable instructions 108 that, when executed by the CPU 102, form one or more software components described herein, such as the GEMM engine 110, the thread decomposition engine 114, and the contention mitigation engine 118.
[0039] According to an example embodiment, the memory 106 may store data, such as, by way of example, data representing the input matrix A; data representing the input matrix B; data representing the intermediate and final results of the output matrix C; data representing the temporary local output buffer; data representing the parameters described herein, such as block size, sub-block size, etc.; data representing the block allocation of the processing threads 112; and so on.
[0040] Generally, the memory 106 is a non-transitory storage medium, which may be formed by a semiconductor storage device, a memristor-based storage device, a magnetic storage device, a phase change storage device, a combination of storage devices corresponding to one or more of these storage technologies, and so on. In addition, the memory 106 may be a volatile memory, a non-volatile memory, or a combination of different storage types, such as volatile memory and / or non-volatile memory.
[0041] The physical machine of the computer system 100 may take many different forms, such as one or more rack-mounted modules, one or more server blades, desktop computers, laptop computers, tablet computers, smartphones, wearable computers, and so on. Depending on the particular embodiment, the GEMM engine 110, the thread decomposition engine 114, and the contention mitigation engine 118 may be formed by the entire physical machine, multiple physical machines, or a portion thereof. In addition, according to some embodiments, the GEMM engine 110, the thread decomposition engine 114, and the contention mitigation engine 118 may include and / or correspond to one or more virtual components or one or more virtual environments of an actual physical machine, such as one or more virtual machines, one or more containers, and so on.
[0042] According to an example embodiment, the thread decomposition engine 114 performs Figure 5A and 5B the depicted process 500 for determining either the BSM and BSK block sizes (for M / K threading) or the BSN and BSK block sizes (for N / K threading). Generally, the determination of the block size has two parts: a first part, where the thread decomposition engine 114 determines a candidate matrix decomposition (block 502) and the BSK block size (block 504); and a second part (i.e., the remainder of the process 500), where the thread decomposition engine 114 performs an iterative process that estimates the candidate matrix decomposition for determining the remaining BSM block size (for M / K threading) or the BSN block size (for N / K threading).
[0043] Turning now to a specific example, according to some embodiments, the thread decomposition engine 114 may determine the BSK block size as follows (in accordance with Figure 5A block 504). First, the thread decomposition engine 114 initially defines the BSM, BSN, and BSK block sizes as follows (where the subscript "init" indicates the initial state):
[0044] BSM init = min(BSM opt , M) Equation 1
[0045] BSN init = min(BSN opt , N) and Equation 2
[0046] BSK init = min(BSK opt , K) Equation 3
[0047] In these equations, "BSM opt ", "BSN opt ", and "BSK opt " represent the optimal values of the BSM, BSN, and BSK block sizes, respectively; "min()" represents the minimum function, i.e., the function that selects the minimum value from a tuple of multiple values. In this case, the optimal block size refers to the number of rows or columns corresponding to a data unit that is aligned or nearly aligned with the cache line boundary. Thus, according to an example embodiment, the optimal block size may be a function of, for example, the cache architecture (e.g., the architecture of the last-level cache (LLC)) and the size of the matrix entries.
[0048] For M / K threading, the thread decomposition engine 114 may determine several parameters for deriving the BSK block size. These parameters include the number of blocks of size BSN in the N dimension (referred to as "nblk"); the number, referred to as "mblk_max", of blocks of size BSMR in the M dimension, where "BSMR" represents the minimum possible sub-block size of the output matrix C (i.e., M rows); the re-determined BSM block size; and the number of blocks of size BSM in the M dimension (referred to as "mblk"):
[0049] nblk = ceil(N / BSN) Equation 4
[0050] mblk_max = ceil(M / BSMR) Equation 5
[0051] BSM = min(BSM init , ceil(mblk_max / div_m)*BSMR), and Equation 6
[0052] mblk = ceil(M / BSM) Equation 7
[0053] In these equations, ceil() represents the application of the ceiling or limit function; "div_m" represents the number of threads in the M dimension; "N" represents the number of columns of the input matrix B; and "M" represents the number of rows of the input matrix A.
[0054] According to an example embodiment, preferably the number of M blocks is an integer multiple of the number of threads in the M dimension; and according to this preference, the thread decomposition engine 114 estimates the following equation:
[0055] if (mblk > div_m) mblk = ceil(mblk / div_m) * div_m.
[0056] Then the thread decomposition engine determines the corresponding number of threads in the K dimension (referred to herein as "div_k") as follows:
[0057] div_k = maxth / div_m Equation 8
[0058] where "maxth" represents the total number of processing threads.
[0059] For N / K threading, according to an example embodiment, the thread decomposition engine 114 determines the number of mblk as follows:
[0060] mblk = ceil(M / BSM) Equation 9
[0061] Furthermore, according to an example embodiment, for N / K threading, the thread decomposition engine 114 determines the number of blocks of BSNR size in the N dimension (referred to herein as "nblk_max"), where "BSNR" represents the minimum possible sub-block size along the N dimension of the output matrix C; the BSN block size; and the number of blocks of BSN size in the N dimension (referred to herein as "nblk") as follows:
[0062] nblk_max = ceil(N / BSNR) Equation 10
[0063] BSN = min(BSN init , ceil(nblk_max / div_n) * BSNR, and Equation 11
[0064] nblk = ceil(N / BSN) Equation 12
[0065] According to an example embodiment, preferably the number of N blocks is an integer multiple of the number of threads in the N dimension; and thus, according to an example embodiment, the thread decomposition engine 114 estimates the following equation:
[0066] If (nblk > div_n) nblk = ceil(nblk / div_n) * div_n.
[0067] Then, according to the exemplary embodiment, the thread decomposition engine 114 determines the number of div_k for threading in the K dimension as follows:
[0068] div_k = maxth / div_n Equation 13
[0069] According to the exemplary embodiment, regardless of whether M / K threading or N / K threading is used, the thread decomposition engine 114 dynamically adjusts the value of the BSK block size for efficiency optimization. For an inappropriate / copied GEMM kernel (i.e., a GEMM kernel that copies the input matrices A and B into a buffer for consecutive memory access), the thread decomposition engine 114 determines the optimal resident size of the packing buffer for the input matrices A and B as follows:
[0070] Packing Buffer Size = BSK opt * (BSM opt + BSN opt ) Equation 14
[0071] Using the optimal resident size of the packing buffer, the thread decomposition engine 114 can determine the maximum possible BSK block size as follows:
[0072] BSK = BSK opt * (BSM opt + BSN opt ) / (BSM + BSN) Equation 15
[0073] BSK = ceil(BSK / BSKR) * BSKR Equation 16
[0074] BSK = min(BSK, K) and Equation 17
[0075] kblk = ceil(K / BSK) Equation 18
[0076] where "BSKR" represents the unroll factor of the innermost K loop 324 ( Figure 3 ) in the macro kernel. Unrolling is a loop transformation technique that improves performance by completing multiple iterations in each loop trip, and the "unroll factor" refers to the number of iterations in each loop trip.
[0077] According to the exemplary embodiment, the thread decomposition engine 114 can then adjust the BSK block size as follows. For M / K threading, the thread decomposition engine 114 estimates the following equation:
[0078] if (div_k > kblk) AND if (nblk * kblk % div_k != 0) kblk = div_k.
[0079] If the above equation is true, then according to the example embodiment, the thread decomposition engine 114 uses the optimal BSK determined as described above because there are sufficient N blocks for equal work distribution.
[0080] For N / K threading, the following equation is determined:
[0081] If (div_k > kblk) AND if (mblk * kblk % div_k != 0) kblk = div_k.
[0082] As described below, if the equation is true, then the thread decomposition engine 114 calculates another BSK block size. Otherwise, the optimal BSK block size can be used because there are sufficient M blocks for equal work distribution.
[0083] If the thread decomposition engine 114 determines that the optimal BSK block size cannot be used, then, according to the example embodiment, the thread decomposition engine 114 recalculates the BSK block size as follows:
[0084] BSK = min(BSK, ceil(ceil(K / BSKR) / kblk) * BSKR) Equation 19
[0085] The thread decomposition engine 114 can then determine the number of threads in the K dimension as follows:
[0086] kblk = ceil(K / BSK) Equation 20
[0087] Still referring to Figure 5A , according to the example embodiment, the thread decomposition engine 114 determines (block 502) candidate matrix decompositions using all possible divisors of the number of threads divided between the M dimension and the K dimension (for M / K threading) or between the N dimension and the K dimension (for N / K threading). The thread decomposition engine 114 determines (block 508) the conditional values for each candidate matrix decomposition based on the BSM, BSN, and BSK block sizes. Note that, according to the example embodiment, if M / K threading is used, then the default BSN block size is used; if N / K threading is used, then the default BSM block size is used.
[0088] According to block 512, the thread decomposition engine 114 normalizes the condition values. In this case, "normalizing" a particular condition value means determining the maximum value of the condition values for each candidate matrix decomposition and normalizing the condition values based on that maximum value. This in turn effectively weights the condition values for a particular condition value category. The thread decomposition engine 114 then determines a fitness value for each candidate matrix decomposition based on the normalized condition values (e.g., the thread decomposition engine 114 adds together the normalized condition values) according to block 516.
[0089] Reference Figure 5B , according to an example embodiment, the thread decomposition engine 114 classifies or ranks the candidate matrix decompositions based on the respective individual fitness values of the candidate matrix decompositions (block 520). For example, according to some embodiments, the thread decomposition engine 114 ranks the candidate matrix decompositions in ascending order according to the fitness values of the candidate matrix decompositions, such that the candidate matrix decomposition with the lowest respective fitness value is at the top of the ranking, i.e., is the highest-ranked candidate matrix decomposition.
[0090] The thread decomposition engine 114 then uses the ranked candidate matrix decompositions to perform an iterative process (e.g., the iterative process set forth in blocks 524 and 528) for selecting a particular candidate matrix decomposition based on the ranking. More specifically, according to an example embodiment, the thread decomposition engine 114 selects (block 524) the next candidate matrix decomposition based on the ranking. For example, for an initial selection, the thread decomposition engine 114 selects the highest-ranked candidate matrix decomposition. As further described herein, the thread decomposition engine 114 then determines (decision block 528) whether the candidate matrix decomposition is acceptable based on one or more selection criteria. If the candidate matrix decomposition is not acceptable, then, according to an example embodiment, the thread decomposition engine 114 returns to block 524 to select the next candidate matrix decomposition that is second-highest ranked (relative to the previous selection) and continues to perform decision block 528 to determine whether the candidate matrix decomposition is acceptable.
[0091] Based on one or more selection criteria, the thread decomposition engine 114 ultimately selects an acceptable particular candidate matrix decomposition and, according to block 532, then transmits data representing the block size of the selected matrix decomposition to the GEMM engine 110.
[0092] As an example of the condition values that the thread decomposition engine 114 can use to determine the respective fitness values (according to Figure 5A blocks 508, 512, and 516), according to an example embodiment, for M / K threading, the thread decomposition engine 114 can determine one or more (or all) of the following seven condition values for a candidate matrix decomposition for determining the fitness value:
[0093] Condition value 1: The absolute difference between the number of sub - blocks in the K dimension and the smallest such value (preferably the number of blocks set using BSK).
[0094] Condition value 2: The absolute difference between the BSN and BSK block sizes.
[0095] Condition value 3: The normalized upper - bound calculation of the ratio of the number of threads in the M dimension to the maximum number of decomposable blocks of the input matrix B available for shared packing.
[0096] Condition value 4: The remainder when the local size in the M dimension is divided by the BSM block size.
[0097] Condition value 5: BSM opt and the difference in BSM block sizes.
[0098] Condition value 6: If the BSM block size is strictly less than the BSM init block size, the difference between BSM and the BSM init block size.
[0099] Condition value 7: If the BSK block value is not equal to the BSK init block size, add 1; if the BSM block size is not equal to the BSM init block size (strict inequality), also add 1.
[0100] According to the example implementation, for N / K threading, the thread decomposition engine 114 can determine one or more (or all) of the following eight condition values for candidate matrix decomposition to determine the fitness value:
[0101] Condition value 1: The absolute difference between the number of sub - blocks in the N dimension and the smallest such value (preferably the number of sub - blocks set using BSN).
[0102] Condition value 2: The absolute difference between the BSN and BSK block sizes.
[0103] Condition value 3: If the number of threads in the N dimension is an integer multiple of the number of processing cores per NUMA node, it is 0, otherwise it is 1.
[0104] Condition value 4: If the number of threads in the N dimension is less than or equal to the number of processing cores per NUMA node, it is 0, otherwise it is 1.
[0105] Condition value 5: The absolute difference between the number of threads in the N dimension and the number of processing cores per LLC.
[0106] Condition value 6: The normalized upper - bound calculation of the ratio of the number of threads in the N dimension to the maximum number of decomposable blocks of the input matrix A available for shared packing.
[0107] Condition value 7: The BSN block size and BSNinit The difference in block sizes, if the difference is strictly less than BSN init Block size,.
[0108] Condition value eight: If the BSK block size is not equal to BSK init block size, then add 1, if the BSN block size is not equal to BSN init block size (strict inequality), also add 1.
[0109] According to an example embodiment, the thread decomposition engine 114 then normalizes (relative to the maximum value) each condition value in all possible candidate matrix decompositions to provide a corresponding weighted contribution; and then, according to an example embodiment, the thread decomposition engine 114 sums the condition values of each candidate matrix decomposition to determine the corresponding single fitness value of the candidate matrix decomposition. Next, the thread decomposition engine 114 classifies or sorts the candidate matrix decompositions in ascending order (e.g., according to the sorting of Figure 5B box 520) and, if the following conditions are met, selects the candidate matrix decomposition with the smallest corresponding fitness value.
[0110] According to an example embodiment, for M / K threading, the thread decomposition engine 114 determines whether a particular candidate decomposition (identified by the sorting) is acceptable (e.g., performing the Figure 5B decision box 528) based on whether both of the following two conditions are met:
[0111] Condition one: The number of threads in the K dimension > 1; and
[0112] Condition two: The BSK block size is less than or equal to the BSM block size, or ceil(BSN / BSNR) is less than the total number of threads.
[0113] According to an example embodiment, for N / K threading, if both of the following two condition values are met, the thread decomposition engine 114 selects the candidate matrix decomposition (identified by the sorting) as acceptable (e.g., performing the Figure 5B decision box 528):
[0114] Condition one: The number of threads in the K dimension > 1; and
[0115] Condition value two: The BSK block size is less than or equal to the BSN block size, or ceil(BSM / BSMR) is less than the total number of threads.
[0116] As discussed above in connection with Figure 5A and 5B if, for a given candidate matrix decomposition identified by the sorting, the condition values are not met, then the next highest candidate matrix decomposition is selected and the corresponding condition values are re-evaluated.
[0117] Figure 4is an example illustration 400 of M / K threading, and specifically, illustrates partitioning an input matrix A 210 into corresponding blocks 410, where each BSM×BSK block 410 is assigned to a specific processing thread 112( Figure 1 ). Note that although example illustration 400 depicts the BSN block size as equal to N (i.e., the number of columns of input matrix B 214 and the number of columns of output matrix C 220), according to further example embodiments, the BSN block size can be less than N. Note that although in Figure 4 the blocks 410 are depicted as being square, the blocks 410 can be rectangular because the BSM and BSK block sizes can be different. A particular K group of processing threads 112 (such as the K group of processing threads corresponding to the blocks 410 of row 412 of Figure 4 ) together determine the corresponding BSM×BSN block 420 of output matrix C 220. In other words, the processing threads corresponding to each block 410 of row 412 compute a portion of the BSM×BSN block 420 of output matrix C 220. For example, the processing thread 112 assigned to block 410-1 of row 412, for example, performs a matrix multiplication operation based on block 410-1 and block 414-1 of input matrix B 214 to produce a corresponding contribution to block 420 of output matrix C 220. As another example, the processing thread 112 assigned to block 410-2 performs a matrix multiplication operation based on block 410-2 and block 414-2 of input matrix B 214 to produce an additional contribution to block 420 of output matrix C 220. Because the processing threads 112 in a given K group update the same shared BSM×BSN block of output matrix C, there may be memory resource contention among these processing threads 112.
[0118] According to an example embodiment, the contention mitigation engine 118( Figure 1 ) determines an optimal decomposition of the blocks of output matrix 220 in the M dimension and / or N dimension in a manner that allows the processing threads 112 in a given K group to more efficiently perform matrix-matrix multiplication operations by accessing smaller data blocks. Further subdivision along the M dimension and / or N dimension allows each processing thread 112 to "sub-loop" over the corresponding sub-blocks. Thus, as part of the sub-loop, each processing thread 112 processes its associated sub-blocks one at a time according to Figure 3 the loops 316 and 320 of
[0119] Figure 6 Illustrates a process 600 used by the contention mitigation engine 118 to further decompose the M dimension and / or N dimension according to an example embodiment. In conjunction with Figure 1 reference Figure 6, according to an example embodiment, the contention mitigation engine 118 determines (block 630) candidate submatrix factorizations by considering all divisors of the number of threads in a given set of K. More specifically, according to an example embodiment, the contention mitigation engine 118 may determine (block 634) a set of conditional values for each candidate submatrix factorization and determine a fitness value for each set of conditional values. For example, according to an example embodiment, the contention mitigation engine 118 may determine the fitness value for each candidate submatrix factorization based on the following conditional values:
[0120] Conditional value one: If the BSM sub sub-block size is greater than or equal to the BSN sub block size (i.e., preferably the M dimension is greater than the N dimension), add 1.
[0121] Conditional value two: If the BSN sub sub-block size is equal to the BSN opt block size (i.e., preferably the optimal cache block size for the N dimension), add 1.
[0122] In accordance with block 638, the contention mitigation engine 118 may classify or sort (e.g., sort the submatrix factorizations based on the fitness value determined from the conditional values) the candidate submatrix factorizations based on the associated conditional values. The contention mitigation engine 118 may then select (block 642) a submatrix factorization based on the sorting. For example, according to some embodiments, the contention mitigation engine 118 may sort the candidate submatrix factorizations in descending order based on the respective fitness values of the candidate submatrix factorizations, and the contention mitigation engine 118 may then select the candidate submatrix factorization having the largest respective fitness value determined from the two conditional values as described above.
[0123] The contention mitigation engine 118 may then determine (decision block 650) whether a particular matrix factorization thread is only in the K dimension, and if so, then determine (decision block 654) whether to use M / K threading. For M / K threading (i.e., the "yes" branch of decision block 654), the contention mitigation engine 118 transmits (block 662) to the GEMM engine 110 data representing a sub-loop only in the M dimension, where the M dimension is subdivided according to an M dimension sub-block size (BSM sub ) equal to the BSMR block size. Otherwise, if the threading occurs only in the K dimension and N / K threading is used (the "no" branch of decision block 662), then in accordance with block 656, the contention mitigation engine 118 transmits to the GEMM engine 110 data representing a sub-loop only in the N dimension, where the N dimension is according to an N dimension sub-block size (BSN sub ) equal to the BSNR block size. If the threading occurs in a dimension other than the K dimension (the "no" branch of decision block 650), then in accordance with process 600, the contention mitigation engine 118 transmits (block 668) to the GEMM engine 110 data representing the BSNsub and BSM sub values of data.
[0124] According to some embodiments, the contention mitigation engine 118 may recommend whether each group of K processing threads uses a temporary local buffer to obtain a corresponding BSM×BSN block of the output matrix C, such that when the group of processing threads K finishes processing the block, the content of the temporary local buffer can be transferred to the output buffer of the output matrix C. According to some embodiments, the local temporary buffer is aligned with the cache boundary, and the updates to the temporary local buffer are essentially additive.
[0125] The contention mitigation engine 118 may perform Figure 7 the process 700 depicted to determine whether to recommend using a temporary local output buffer. According to an example embodiment, the process 700 involves the contention mitigation engine 118 recommending the use of a temporary local output buffer if both branches of a two-way branch test are satisfied. Otherwise, the contention mitigation engine 118 does not recommend using a temporary local output buffer. In connection with Figure 1 reference Figure 7 , according to an example embodiment, for the first branch, the contention mitigation engine 118 determines whether any one of the BSM block size (decision box 704), the BSN block size (decision box 708), or the BSK block size (decision box 712) is respectively equal to or greater than the optimal block sizes BSM opt , BSN opt and BSK opt . If so, then the first branch has been satisfied, and for the second branch, the contention mitigation engine 118 determines (decision box 720) whether the number of threads in the K dimension is less than a predetermined threshold number (e.g., the threshold number "8"). If the number of threads in the K dimension is greater than the predetermined threshold, then according to an example embodiment, both branches have been satisfied, and the contention mitigation engine 118 recommends using the local temporary buffer and transmits data with the recommendation (box 724). Otherwise, if either branch is not satisfied, then according to an example embodiment, as depicted in box 716, the contention mitigation engine 118 transmits data indicating not to use the local temporary output buffer to the GEMM engine 110.
[0126] Referring back to Figure 1, according to some embodiments, the GEMM engine 110 may have one or more of the following load balancing features. The GEMM engine 110 adjusts the number of threads for the common blocks that cooperate to pack either matrix A and / or B according to the fair sharing principle to help improve group synchronization. If a local temporary output buffer with sub-loops is used, then the GEMM engine 110 calculates the GEMM solution for a specific sub-block and uses a non-blocking loop on the update function. If the GEMM engine 110 uses a local temporary output buffer without sub-loops, then the GEMM engine 110 calculates the GEMM solution and then immediately calls the update function for a specific sub-block. If the GEMM engine 110 uses sub-loops without a local temporary output buffer, then within the loop, the GEMM engine 110 locally packs the input matrix A or B and waits for a lock to calculate the GEMM solution for updating the output matrix C. If the GEMM engine 110 neither uses sub-loops nor a local temporary output buffer, then the GEMM engine 110 waits for a lock to calculate the GEMM solution for updating the output matrix C.
[0127] As a specific example of the output of the thread decomposition engine 114 and the contention mitigation engine 118, for 16 processing threads 112, the engines 114 and 118 may determine the block sizes, sub-block sizes, and other parameters for the following matrices: matrix A (192×3840), matrix B (3840×96), and matrix C (192×96). For this example, assume that the central processing unit 102 has 64 processing cores, 4 NUMA domains, and 4 processing cores share the last-level cache (LLC). For this example, also assume that N / K threading is used.
[0128] For this example, the thread decomposition engine 114 estimates the following M×N×K candidate decompositions: 1×1×16, 1×2×8, 1×4×4, 1×8×2, and 1×16×1, although 1×16×1 is invalid (because there is no partitioning in the N dimension). Considering the selection criteria, the thread decomposition engine 114 selects the 1×1×16 decomposition, so the threading will be only in the K dimension. For this example, the thread decomposition engine 114 selects BSM, BSN, and BSK block sizes of 192 / 96 / 240.
[0129] For this example, the contention mitigation engine 118 determines to sub-loop in the M dimension with blocks of size BSMR. Additionally, for this example, the contention mitigation engine 118 determines that the use of a local temporary output buffer is not recommended.
[0130] As another example, for the above example computer system, the thread decomposition engine 114 and the contention mitigation engine 118 may process the following matrices: matrix A (96×3840), matrix B (3840×192), and matrix C (96×192). For this example, assume that N / K threading is used.
[0131] For this example, the thread decomposition engine 114 estimates the following candidate M×N×K decompositions: 1×1×16, 1×2×8, 1×4×4, 1×8×2, and 1×16×1, although 1×16×1 is invalid (because no threading or decomposition occurs along the N dimension). Based on the selection criteria, the thread decomposition engine 114 selects the 1×4×4 decomposition such that threading will occur along the N dimension and the K dimension. For this example, the thread decomposition engine 114 selects BSM / BSN / BSK blocks to be 96 / 48 / 960.
[0132] For this second example, the contention mitigation engine 118 sub-loops along the M dimension and the N dimension. Additionally, for this second example, the contention mitigation engine 118 recommends using a local temporary output buffer for efficiency.
[0133] Reference Figure 8 , according to an example embodiment, the non-transitory storage medium 800 stores machine-readable instructions 804 that, when executed by a machine, cause the machine to access data representing a first dimension of a first matrix, a second dimension of a second matrix, and a third dimension shared by the first matrix and the second matrix. The instructions 804, when executed by the machine, further cause the machine to determine a plurality of candidate decompositions of the first matrix, where the candidate decompositions differ in size relative to each other along the first dimension and the third dimension. The instructions 804, when executed by the machine, also cause the machine to determine an associated fitness value for each candidate decomposition of the plurality of candidate decompositions, and select a candidate decomposition from the plurality of candidate decompositions based on the fitness value to provide the selected candidate decomposition. The instructions 804, when executed by the machine, further cause the machine to provide data representing the selected candidate decomposition such that processing threads are assigned to sub-matrices of the first matrix. The processing threads determine a third matrix based on the multiplication of the first matrix and the second matrix.
[0134] Reference Figure 9 , according to an example embodiment, the apparatus 900 includes a processor 910 and a memory 904 storing instructions 908. The instructions 908, when executed by the processor 910, cause the processor 910 to perform threading of a first matrix along a first dimension of the first matrix and a second dimension of the matrix. The threading represents the block sizes of the first matrix for processing threads assigned to a multiplication algorithm for determining a third matrix representing the product of the first matrix and the second matrix. The block sizes include a first block size along the first dimension and a second block size along the second dimension. The second matrix shares the second dimension with the first matrix. The instructions 908, when executed by the processor 910, cause the processor 910 to provide data representing the first block size and the second block size to the multiplication algorithm.
[0135] Reference Figure 10, according to an example embodiment, technique 1000 includes at least one hardware processor that accesses (block 1002) data representing a first matrix, a second matrix, a first block size, and a second block size. The first block size represents the sub - matrix decomposition size of the first matrix along a first dimension; the second block size represents the sub - matrix decomposition size of the first matrix along a second dimension. The second matrix shares the second dimension with the first matrix. Technique 1000 includes the hardware processor multiplying the first matrix and the second matrix (block 1004) to determine a third matrix. The multiplication includes allocating processing threads to sub - matrices of the first matrix based on the first block size and the second block size.
[0136] According to an example embodiment, a plurality of candidate decomposition block sizes are determined along a third dimension such that the plurality of candidates are each decomposed along the third dimension according to the block size. A specific advantage is that it can improve the processing efficiency of matrix - matrix multiplication.
[0137] According to an example embodiment, the block size can be determined based on the characteristics of the cache memory used by the processing threads to determine the third matrix. A specific advantage is that it can improve the processing efficiency of matrix - matrix multiplication.
[0138] According to an example embodiment, for each fitness value, a plurality of conditions associated with the candidate decomposition are determined; and each fitness value is determined based on the plurality of conditions. A specific advantage is that it can improve the processing efficiency of matrix - matrix multiplication.
[0139] According to an example embodiment, for each fitness value, each condition among the plurality of conditions is normalized to provide a plurality of normalized condition values. The plurality of normalized condition values are combined to determine the fitness value. A specific advantage is that it can improve the processing efficiency of matrix - matrix multiplication.
[0140] According to an example embodiment, the fitness values are sorted to provide a corresponding ranking of the plurality of candidate decompositions. The selected candidate decomposition is selected based on the ranking. A specific advantage is that it can improve the processing efficiency of matrix - matrix multiplication.
[0141] According to an example embodiment, a first candidate decomposition is selected based on the ranking; and it is determined that the first candidate decomposition does not meet the selection criteria. A second candidate decomposition other than the first candidate decomposition is selected based on the ranking; and it is determined that the second candidate decomposition meets the selection criteria. The second candidate decomposition is determined as the selected candidate decomposition. A specific advantage is that it can improve the processing efficiency of matrix - matrix multiplication.
[0142] According to an example embodiment, the selected candidate decomposition corresponds to a third matrix that is subdivided into a plurality of submatrices. Each submatrix is associated with a set of processes, and a submatrix decomposition of the plurality of submatrices is determined. Data representing the submatrix decomposition is provided such that each processing thread performs a sub-loop based on the submatrix decomposition. A specific advantage is that the processing efficiency of matrix-matrix multiplication can be improved.
[0143] According to an example embodiment, based on the characteristics of the selected candidate decomposition, it is determined whether to recommend using a buffer to store the preliminary results of the multiplication. The routine updates the third matrix using an addition update from the content stored in the buffer. A specific advantage is that the processing efficiency of matrix-matrix multiplication can be improved.
[0144] Although the present disclosure is described with respect to a limited number of embodiments, those skilled in the art who benefit from the present disclosure will understand many modifications and variations therefrom. The appended claims are intended to cover all such modifications and variations.
Claims
1. A non-transitory storage medium storing machine-readable instructions that, when executed by a machine, cause the machine to: access data representing a first dimension of a first matrix, a second dimension of a second matrix, and a third dimension shared by the first matrix and the second matrix; determine a plurality of candidate decompositions of the first matrix, wherein the candidate decompositions differ in size relative to each other along the first dimension and the third dimension; determine an associated fitness value for each candidate decomposition of the plurality of candidate decompositions; select, based on the fitness values, a candidate decomposition from the plurality of candidate decompositions to provide a selected candidate decomposition; and provide data representing the selected candidate decomposition such that a processing thread is assigned to a sub-matrix of the first matrix, wherein the processing thread determines a third matrix based on a multiplication of the first matrix and the second matrix.
2. The storage medium according to claim 1, wherein the instructions, when executed by the machine, further cause the machine to determine a block size for the plurality of candidate decompositions such that each candidate decomposition of the plurality of candidate decompositions is decomposed along the third dimension according to the block size.
3. The storage medium according to claim 2, wherein the instructions, when executed by the machine, further cause the machine to determine the block size based on characteristics of a cache memory used by the processing thread in the determination of the third matrix.
4. The storage medium according to claim 1, wherein the instructions, when executed by the machine, further cause the machine to: for each fitness value of the fitness values: determine a plurality of conditions for the associated candidate decomposition; and determine each fitness value based on the plurality of conditions.
5. The storage medium according to claim 4, wherein the instructions, when executed by the machine, further cause the machine to: for each fitness value: normalize each condition of the plurality of conditions to provide a plurality of normalized condition values; and combine the plurality of normalized condition values to determine each fitness value.
6. The storage medium according to claim 1, wherein the instructions, when executed by the machine, further cause the machine to: sort the fitness values to provide a corresponding sorting of the plurality of candidate decompositions; and select the selected candidate decomposition based on the sorting.
7. The storage medium according to claim 6, wherein the instructions, when executed by the machine, further cause the machine to: select a first candidate decomposition based on the sorting; determine that the first candidate decomposition does not meet a selection criterion; select a second candidate decomposition other than the first candidate decomposition based on the sorting; determine that the second candidate decomposition meets the selection criterion; and determine that the second candidate decomposition is the selected candidate decomposition.
8. The storage medium according to claim 1, wherein The selected candidate decomposition corresponds to the third matrix that is subdivided into a plurality of sub-matrices, where each of the plurality of sub-matrices is associated with a set of processes of a plurality of processes, and the instructions, when executed by the machine, further cause the machine to: Determine a plurality of candidate sub-matrix decompositions, where the candidate sub-matrix decompositions are different in size relative to each other along the first dimension and the second dimension; Determine an associated second fitness value for each candidate sub-matrix decomposition of the plurality of candidate sub-matrix decompositions; Based on the second fitness value, select a candidate sub-matrix decomposition from the plurality of candidate sub-matrix decompositions to provide a selected candidate sub-matrix decomposition; And Provide data representing the selected candidate sub-matrix decomposition so that each processing thread of the processing threads performs a sub-loop based on the selected sub-matrix decomposition.
9. The storage medium according to claim 1, wherein, The selected candidate decomposition corresponds to the third matrix that is subdivided into a plurality of sub-matrices, where each of the plurality of sub-matrices is associated with a set of processes of a plurality of processes, and the instructions, when executed by the machine, further cause the machine to: Determine a sub-matrix decomposition of the plurality of sub-matrices; and Provide data representing the sub-matrix decomposition so that each processing thread of the processing threads performs a sub-loop based on the sub-matrix decomposition.
10. The storage medium according to claim 1, wherein, The instructions, when executed by the machine, further cause the machine to determine whether to recommend using a buffer to store preliminary results of the multiplication based on at least one feature of the selected candidate decomposition, where a routine updates the third matrix using an addition update from the content stored in the buffer.
11. The storage medium according to claim 10, wherein, The instructions, when executed by the machine, further cause the machine to determine whether to recommend using the buffer based on at least one of the following: a comparison of a first block size of the candidate decomposition along the first dimension with a first cache-optimal block size along the first dimension, a comparison of a second block size of the candidate decomposition along the second dimension with a second cache-optimal block size along the second dimension, or a comparison of a third block size of the candidate decomposition along the third dimension with a third cache-optimal block size along the third dimension.
12. The storage medium according to claim 1, wherein: The first matrix includes an M×K matrix, and the second matrix includes a K×N matrix; The first dimension corresponds to the M dimension of the M×K matrix; The second dimension corresponds to the N dimension of the K×N matrix; The third dimension corresponds to the K dimension of the M×K matrix and the K dimension of the K×N matrix; The multiplication includes (M×K)(K×N) matrix multiplication; And Determining the plurality of candidate decompositions of the first matrix includes M / K threading.
13. The storage medium according to claim 1, wherein: The second matrix includes an M×K matrix, and the first matrix includes a K×N matrix; The first dimension corresponds to the N dimension of the K×N matrix; The second dimension corresponds to the M dimension of the M×K matrix; The third dimension corresponds to the K dimension of the M×K matrix and the K dimension of the K×N matrix; The multiplication includes (M×K)(K×N) matrix multiplication; and Determining the plurality of candidate decompositions of the first matrix includes N / K threading.
14. An apparatus for matrix-matrix multiplication, comprising: a processor; and a memory for storing instructions that, when executed by the processor, cause the processor to: Perform threading of the first matrix along a first dimension of the first matrix and a second dimension of the matrix, where the threading represents the block size of the first matrix for the processing threads to be assigned to a multiplication algorithm for determining a third matrix representing the product of the first matrix and a second matrix, the block size includes a first block size along the first dimension and a second block size along the second dimension, and the second matrix shares the second dimension with the first matrix; and Provide data representing the first block size and the second block size to the multiplication algorithm.
15. The apparatus according to claim 14, wherein, The multiplication algorithm includes a Generalized Matrix-Matrix Multiplication (GEMM) algorithm.
16. The apparatus according to claim 14, wherein, The instructions, when executed by the processor, further cause the processor to: Determine the second block size; and Estimate candidate matrix decompositions, each of the candidate matrix decompositions being decomposed along the second dimension according to the second block size, where the estimation is for determining the second block size.
17. The apparatus according to claim 14, wherein, The threading corresponds to the third matrix being subdivided into a plurality of sub-matrices, where each of the plurality of sub-matrices is associated with a set of threads among the plurality of threads, and the instructions, when executed by the processor, further cause the processor to: Determine the sub-matrix decomposition of the plurality of sub-matrices; and Provide data representing the sub-matrix decomposition so that each of the plurality of threads performs a sub-loop based on the sub-matrix decomposition.
18. The apparatus according to claim 14, wherein, The instructions, when executed by the processor, further cause the processor to determine whether to recommend using a buffer to store preliminary results of the multiplication based on at least one of the block sizes, where a routine updates the third matrix using an addition update from the content stored in the buffer.
19. A method for matrix-matrix multiplication, comprising: At least one hardware processor accesses data representing a first matrix, a second matrix, a first block size, and a second block size, where the first block size represents the sub-matrix decomposition size of the first matrix along a first dimension, the second block size represents the sub-matrix decomposition size of the first matrix along a second dimension, and the second matrix shares the second dimension with the first matrix; and The at least one hardware processor multiplies the first matrix and the second matrix to determine a third matrix, wherein the multiplication includes allocating processing threads to sub-matrices of the first matrix based on the first block size and the second block size.
20. The method according to claim 19, wherein, the multiplication includes each processing thread of the processing threads sub-looping on sub-blocks of the sub-matrix allocated to that processing thread.
Citation Information
Patent Citations
Generalized acceleration of matrix multiply accumulate operations
CN108874744A
Data processing method and device, electronic equipment and storage medium
CN111158874A