Optimization method for sparse-dense matrix multiplication parallel algorithm based on CSR characteristics
By designing the CSR sparse matrix data buffer and load balancing strategy in sparse dense matrix multiplication parallel calculation, the problems of load imbalance and frequent discontinuous memory fetching are solved, and more efficient computing performance is achieved.
Patent Information
- Application Number
- CN202111355302.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-16
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2041-11-16
AI Technical Summary
The existing sparse and dense matrix multiplication parallel computing has load imbalance and frequent discontinuous memory fetching problems, resulting in inefficient computing.
By designing a load balancing strategy for parallel data preloading of CSR sparse matrix data buffer and buffer, the load balancing of dense matrix computing tasks is optimized, and a temporary storage design of partial accumulation sum is adopted during the calculation process to reduce unnecessary memory access.
It improves the calculation efficiency of sparse and dense matrix multiplication, reduces the number of discontinuous memory accesses, and improves the utilization rate of computer resources.
Smart Images

Figure CN114048035B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of computer parallel computing and computer architecture, and in particular to an optimization method for a sparse-dense matrix multiplication parallel algorithm based on CSR characteristics. Background Art
[0002] Graph-based machine learning algorithms, especially graph neural networks, are widely used in link prediction, classification, social networks and other fields. Compared with traditional neural networks, they have extremely strong feature expression capabilities and have set off a research boom in academia and industry. However, in reality, the training of larger graph neural networks may take several days or even dozens of days, which greatly limits the continued development of graph neural networks. Therefore, the acceleration of graph neural network training is particularly important. In addition, in the training of real graph neural networks, graphs are often sparse, and convolution algorithms are generally based on sparse dense matrix multiplication (SpMM). Therefore, the acceleration of training is mainly aimed at sparse dense matrix multiplication.
[0003] Compressed Sparse Row (CSR) is a data format that compresses and stores sparse matrices in row order. Graphs in practical applications are often sparse. If they are stored directly in the form of matrices, a large amount of storage space is required, and during calculations, many invalid calculations will occur due to the presence of a large number of zero elements, wasting computer resources. The CSR format only stores non-zero elements in the graph, which greatly improves storage efficiency, avoids invalid calculations, and improves computing efficiency.
[0004] Most of the current sparse-dense matrix multiplications are directly parallelized. As the amount of concurrency increases, different threads will swap data into the cache when calculating their own data, which will cause the operating system to have a large number of discontinuous memory accesses and frequent redundant data exchanges between the memory and the cache, reducing the computing efficiency. At the same time, direct parallelism is difficult to predict the load of each thread, especially for some irregular sparse matrices, the load between individual threads often has a large difference, and the efficiency of the entire system will be subject to the thread with the largest load. Summary of the invention
[0005] In view of the shortcomings of the prior art, the present invention aims to provide an optimization method for the sparse dense matrix multiplication parallel algorithm based on the CSR characteristics. By taking the parallel version of sparse dense matrix multiplication as the basis and proposing a buffer data preloading strategy, the load balancing of threads during calculation is improved to improve the computational efficiency of sparse dense matrix multiplication.
[0006] In order to achieve the above object, the technical solution adopted by the present invention is as follows:
[0007] The optimization method for the sparse-dense matrix multiplication parallel algorithm based on CSR characteristics includes the design of CSR sparse matrix data buffer, load balancing of buffer parallel data preloading, load balancing of dense matrix calculation tasks on threads, and temporary storage design of partial accumulation of calculation results; which also includes:
[0008] S1 finds the length n of the row with the most non-zero elements in the given sparse matrix;
[0009] S2 allocates space of length n to the buffer;
[0010] S3 enters the main calculation loop to determine whether each row of the sparse matrix has been calculated. If yes, the calculation is completed, otherwise it enters step S4;
[0011] S4 specifies the preloaded data range based on the rowPtr array;
[0012] S5 distributes the data evenly to the corresponding threads according to the distribution formula and the number of available threads, and loads the data into the buffer;
[0013] S6 sets the synchronization node and waits for all threads to finish loading their respective data;
[0014] S7 calculates the start and end column numbers of the dense matrix in the current thread computing task according to the current thread number and the load balancing distribution formula;
[0015] Each S8 thread takes the sparse matrix data from the buffer and performs a dot multiplication operation with the corresponding data in the dense matrix;
[0016] The intermediate accumulation sums during the dot product calculation process of S9 are stored in the temporary register;
[0017] S10 Each thread writes the final result of the calculation into the result matrix;
[0018] S11 returns to step S3.
[0019] It should be noted that the CSR sparse matrix data buffer design can adaptively configure the buffer size according to the size of the sparse matrix. Through the CSR sparse matrix data buffer, the data required by the thread during calculation can be directly obtained from the cache buffer, reducing the number of discontinuous memory accesses and improving computing efficiency.
[0020] It should be noted that the load balancing of the buffer parallel data preloading is able to divide a row of data of the CSR sparse matrix into uniform data blocks according to the number of row elements of different sparse matrices, and assign them to different parallel threads to load into the buffer, so as to reduce the biggest bottleneck of data loading.
[0021] It should be noted that the load balancing of threads by the dense matrix calculation task is to divide the core operation of matrix multiplication. Since each row of the sparse matrix needs to be dot-multiplied with each column of the dense matrix, and the dot multiplications are independent of each other, the calculation amount of a single thread can be reduced through parallel optimization; among them, the size of the data block can be adaptively adjusted by column according to the size of different dense matrices, and the data blocks can be evenly distributed to the corresponding threads for calculation, so as to reduce the time spent on the entire operation.
[0022] It should be noted that the temporary storage design of the partial cumulative sum of the calculation results is to design a temporary register during the calculation process, store the intermediate calculation results in the temporary register first, and then write them into the result matrix after the calculation is completed, which can effectively reduce some unnecessary memory access overhead and improve parallel computing performance.
[0023] The beneficial effect of the present invention is that by analyzing some memory access problems existing in the current SpMM parallel algorithm, buffer design is used as an entry point and load balancing is used as a means to achieve the purpose of improving computing performance and improving computer resource utilization. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 Schematic diagram of the CSR storage format of the sparse matrix in the present invention;
[0025] Figure 2 A schematic diagram of the sparse-dense matrix multiplication based on the buffer in the present invention;
[0026] Figure 3 Schematic diagram of buffer design and data preloading strategy in the present invention;
[0027] Figure 4 It is a schematic diagram of the dense matrix data load balancing strategy in the present invention;
[0028] Figure 5 It is a flowchart of the optimization method for the sparse-dense matrix multiplication parallel algorithm based on the CSR characteristics in the present invention. DETAILED DESCRIPTION
[0029] The present invention will be further described below in conjunction with the accompanying drawings. It should be noted that this embodiment is based on the technical solution and provides a detailed implementation method and specific operation process, but the protection scope of the present invention is not limited to this embodiment.
[0030] The following is a brief description of the abbreviations and key terms used in the embodiments of the present invention:
[0031] SpMM (Sparse-Dense Matrix-Matrix Multiplication): Matrix multiplication is often used as the basic operation of convolution. In practical applications of graph neural networks, the feature matrix often exhibits great sparsity, while the convolution kernel is dense. Therefore, the convolution operation in graph neural networks is often expressed as an operation between a sparse feature matrix and a dense convolution kernel. SpMM is one of the most basic and important ones.
[0032] CSR (Compressed Sparse Row): CSR is a compressed storage format for sparse matrices. This format only requires three types of data to express: offset, column number, and value. The row offset indicates the position of the starting element of a row in the val array. Since CSR only stores non-zero elements in the matrix, and graphs in practical applications are often sparse, using the CSR format can significantly improve storage efficiency and computing efficiency.
[0033] Cache: refers to the memory that is physically located between the CPU and the memory and can exchange data at high speed. It can exchange data between the CPU and the memory to solve the access rate matching problem between the CPU and the memory. Modern computer caches are often divided into multiple levels. The closer the cache is to the CPU, the smaller the capacity and the faster the access rate. In essence, cache is a memory that provides a copy of the data in the memory for CPU calculation.
[0034] The present invention is an optimization method for a sparse-dense matrix multiplication parallel algorithm based on CSR characteristics, including a CSR sparse matrix data buffer design, load balancing of buffer parallel data preloading, load balancing of dense matrix calculation tasks on threads, and temporary storage design of partial cumulative sums of calculation results; wherein, it also includes:
[0035] S1 finds the length n of the row with the most non-zero elements in the given sparse matrix;
[0036] S2 allocates space of length n to the buffer;
[0037] S3 enters the main calculation loop to determine whether each row of the sparse matrix has been calculated. If yes, the calculation is completed, otherwise it enters step S4;
[0038] S4 specifies the preloaded data range based on the rowPtr array;
[0039] S5 distributes the data evenly to the corresponding threads according to the distribution formula and the number of available threads, and loads the data into the buffer;
[0040] S6 sets the synchronization node and waits for all threads to finish loading their respective data;
[0041] S7 calculates the start and end column numbers of the dense matrix in the current thread computing task according to the current thread number and the load balancing distribution formula;
[0042] Each S8 thread takes the sparse matrix data from the buffer and performs a dot multiplication operation with the corresponding data in the dense matrix;
[0043] The intermediate accumulation sums during the dot product calculation process of S9 are stored in the temporary register;
[0044] S10 Each thread writes the final result of the calculation into the result matrix;
[0045] S11 returns to step S3.
[0046] It should be noted that the CSR sparse matrix data buffer design can adaptively configure the buffer size according to the size of the sparse matrix. Through the CSR sparse matrix data buffer, the data required by the thread during calculation can be directly obtained from the cache buffer, reducing the number of discontinuous memory accesses and improving computing efficiency.
[0047] It should be noted that the load balancing of the buffer parallel data preloading is able to divide a row of data of the CSR sparse matrix into uniform data blocks according to the number of row elements of different sparse matrices, and assign them to different parallel threads to load into the buffer, so as to reduce the biggest bottleneck of data loading.
[0048] It should be noted that the load balancing of threads by the dense matrix calculation task is to divide the core operation of matrix multiplication. Since each row of the sparse matrix needs to be dot-multiplied with each column of the dense matrix, and the dot multiplications are independent of each other, the calculation amount of a single thread can be reduced through parallel optimization; among them, the size of the data block can be adaptively adjusted by column according to the size of different dense matrices, and the data blocks can be evenly distributed to the corresponding threads for calculation, so as to reduce the time spent on the entire operation.
[0049] It should be noted that the temporary storage design of the partial cumulative sum of the calculation results is to design a temporary register during the calculation process, store the intermediate calculation results in the temporary register first, and then write them into the result matrix after the calculation is completed, which can effectively reduce some unnecessary memory access overhead and improve parallel computing performance.
[0050] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be described in more detail below in conjunction with the accompanying drawings of the present invention.
[0051] Figure 1An example of CSR compression storage of a sparse matrix is given. The left side of the figure is the logical representation of the sparse matrix, and it can be seen that there are a large number of zero element values in the matrix. The right side is the CSR format data actually stored in the computer, including offset, column number, and value.
[0052] Convolution is the most important operator in neural networks. In the subdivision neighborhood of graph neural networks, feature maps often have high sparsity. Storing and calculating in the form of compressed matrices can improve the training efficiency of graph neural networks.
[0053] In the example, ag represents the non-zero value in the matrix, which is usually a numerical value in actual calculation. The rowPtr array in physical storage is the row offset, indicating the position of the starting element of a row in the val array. The colInd array indicates the column number where the value is located, and the val array stores the actual value. The colInd array is the same length as the val array, which is longer than the rowPtr array. The last element in the rowPtr array indicates the end of the val array.
[0054] In the actual design of the present invention, the buffer size is set to the maximum difference of all adjacent elements in rowPtr, that is, the number of non-zero elements in the row with the most non-zero elements in the sparse matrix. The above method can reduce one traversal of the matrix. The buffer only stores a row of data in the colInd and val arrays at a time. Since the rowPtr data is not involved in the actual parallel computing, it is not stored in the buffer.
[0055] Figure 2 This is a schematic diagram of the buffer-based SpMM of the present invention, where x represents a non-zero element in a sparse matrix, and the small rectangular block between the sparse and dense matrices represents a buffer for storing preloaded data. The key data in the sparse matrix will be stored in the buffer, and when subsequent threads execute parallel computing tasks, they can directly obtain data from the buffer.
[0056] SpMM is an important way to implement convolution algorithms. However, the current parallelization solution used for the CSR storage format will become more and more obvious with the increase in concurrency, causing some data access discontinuity problems, which will become a bottleneck for SpMM computing efficiency.
[0057] In traditional parallel algorithms, since the variable i of parallel threads is the same in the inner loop, that is, they operate on the same row, when some threads calculate faster and some threads calculate slower, some threads will access earlier elements while some threads will access later elements, causing different threads to repeatedly load the same redundant data, causing the operating system to frequently exchange data between the cache and memory, affecting performance.
[0058] In order to solve this problem, the present invention designs an adaptive buffer between the sparse matrix and the dense matrix based on the characteristics of the CSR storage format. The size of the buffer is determined by the maximum value of the difference between all adjacent elements in rowPtr, so the time complexity is only O(m), where m is the number of rows of the sparse matrix.
[0059] The thread loads the data in the sparse matrix into the buffer, and each subsequent thread obtains the data from the buffer and performs its own computing task, avoiding repeated loading of the same data and improving the cache hit rate.
[0060] The buffer space is first allocated in the memory space. Since the buffer address will be frequently hit during subsequent data loading, the operating system will swap the data block where the buffer is located into the cache to reduce the number of memory accesses for subsequent operations.
[0061] Figure 3 It is a schematic diagram of the detailed design of the buffer and the data preloading strategy. Since the height of the matrix in this example is 4, the data in the buffer will be loaded four times. As shown in the figure, each load will accurately load the non-zero elements of a row in the sparse matrix into the buffer according to the rowPtr array.
[0062] During a data loading process, the preloaded data will be evenly distributed to different threads for loading through load balancing (indicated by different depths of color in the figure) to improve data loading efficiency. The distribution formula is: Where s is the number of non-zero elements in the row.
[0063] When preloading data, synchronization is required to wait for all threads to complete loading the data they need to load before subsequent calculations can be performed.
[0064] Figure 4 The figure is a diagram of the load balancing strategy for dense matrix computing data. The computing data in the dense matrix is divided by columns. The reason is that in matrix multiplication, the dot product operation is performed on the rows of matrix A and the columns of matrix B. Therefore, this division method can ensure that the principle of spatial locality can be used to achieve continuity of memory access during calculation.
[0065] exist Figure 4 In the example, the first two columns are calculated by thread 0, the next two columns are calculated by thread 1, and the last column is calculated by thread 2. As shown in the figure, they are divided by different colors. The division formula is Where cols is the width of the dense matrix.
[0066] In actual calculations, the thread will take a row of data from the buffer and perform a dot multiplication operation on the row of data and the data block allocated in the dense matrix.
[0067] The intermediate accumulated results generated during the dot multiplication process will be placed in the temporary register, and when the calculation of an element in the final result matrix is completed, the value will be written into the result matrix.
[0068] In fact, we only care about the final calculation result, not the intermediate results. This method can ensure correctness while avoiding frequent writing to memory, thereby improving calculation efficiency.
[0069] Figure 5 This is a flowchart of the optimization algorithm for parallel computing of sparse and dense matrix multiplication based on CSR storage using buffers and load balancing. It describes the calculation process of matrix multiplication for a given sparse matrix (CSR format) and a dense matrix:
[0070] 1) Find the length n of the row with the most non-zero elements in a given sparse matrix.
[0071] 2) Allocate space of length n to the buffer.
[0072] 3) Enter the main calculation loop to determine whether each row of the sparse matrix has been calculated. If so, the calculation is completed, otherwise enter step 4).
[0073] 4) Specify the preloaded data range based on the rowPtr array.
[0074] 5) Evenly distribute the data to the corresponding threads according to the allocation formula and the number of available threads, and load the data into the buffer.
[0075] 6) Set the synchronization node and wait for all threads to finish loading their respective data.
[0076] 7) Based on the current thread number and the load balancing distribution formula, find the starting and ending column numbers of the dense matrix in the current thread calculation task
[0077] 8) Each thread takes the sparse matrix data from the buffer and performs a dot multiplication operation with the corresponding data in the dense matrix.
[0078] 9) The intermediate accumulation and storage during the dot product calculation are stored in the temporary register.
[0079] 10) Each thread writes the final result of the calculation into the result matrix.
[0080] 11) Return to step 3).
[0081] For those skilled in the art, various corresponding changes can be made according to the above technical solutions and concepts, and all of these changes should be included in the protection scope of the claims of the present invention.
Claims
1. Optimization method for sparse-dense matrix multiplication parallel algorithm based on CSR characteristics, It is characterized in that It includes the design of CSR sparse matrix data buffer, load balancing of buffer parallel data preloading, load balancing of dense matrix calculation tasks on threads, and temporary storage design of partial accumulation of calculation results; it also includes: S1 finds the length n of the row with the most non-zero elements in the given sparse matrix; S2 allocates space of length n to the buffer; S3 enters the main calculation loop to determine whether each row of the sparse matrix has been calculated. If yes, the calculation is completed, otherwise it enters step S4; S4 specifies the preloaded data range based on the rowPtr array; S5 distributes the data evenly to the corresponding threads according to the distribution formula and the number of available threads, and loads the data into the buffer; S6 sets the synchronization node and waits for all threads to finish loading their respective data; S7 calculates the start and end column numbers of the dense matrix in the current thread computing task according to the current thread number and the load balancing distribution formula; Each S8 thread takes the sparse matrix data from the buffer and performs a dot multiplication operation with the corresponding data in the dense matrix; The intermediate accumulation sums during the dot product calculation process of S9 are stored in the temporary register; S10 Each thread writes the final result of the calculation into the result matrix; S11 returns to step S3.
2. The optimization method for sparse-dense matrix multiplication parallel algorithm based on CSR characteristics according to claim 1, It is characterized in that The CSR sparse matrix data buffer design can adaptively configure the size of the buffer according to the size of the sparse matrix. Through the CSR sparse matrix data buffer, the data required by the thread during calculation can be directly obtained from the cache buffer, reducing the number of discontinuous memory accesses and improving computing efficiency.
3. According to the optimization method of the parallel algorithm of sparse and dense matrix multiplication based on CSR characteristics in claim 1, It is characterized in that The load balancing of the buffer parallel data preloading is to divide a row of CSR sparse matrix data into uniform data blocks according to the number of row elements of different sparse matrices, and distribute them to different parallel threads to load into the buffer, so as to reduce the biggest bottleneck of data loading.
4. The optimization method for the sparse-dense matrix multiplication parallel algorithm according to claim 1, Features: The load balancing of threads by the dense matrix calculation task is to divide the core operation of matrix multiplication. Since each row of the sparse matrix needs to be dot-multiplied with each column of the dense matrix, and the dot multiplications are independent of each other, the calculation amount of a single thread can be reduced through parallel optimization; among them, the size of the data block can be adaptively adjusted by column according to the size of different dense matrices, and the data blocks can be evenly distributed to the corresponding threads for calculation, so as to reduce the time spent on the entire operation.
5. The optimization method for the sparse-dense matrix multiplication parallel algorithm according to claim 1, Features: The temporary storage design of the partial cumulative sum of the calculation results is to design a temporary register during the calculation process, store the intermediate calculation results in the temporary register first, and then write them into the result matrix after the calculation is completed, which can effectively reduce some unnecessary memory access overhead and improve parallel computing performance.
Citation Information
Patent Citations
Workload optimization in a multi-processor system executing sparse-matrix vector multiplication
EP2657842A1
Sparse matrix multiplication method using GPU
KR101400577B1