Tensor Core-based diagonal sparse matrix-vector product solving method

By using BDIA format and Warp-level collaborative optimization on Tensor Core, the storage and computing bottlenecks of traditional sparse matrix format on the GPU are solved, and efficient sparse matrix-vector multiplication is achieved, which improves computing performance and storage efficiency.

CN120296294APending Publication Date: 2025-07-11INST OF SOFTWARE - CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510256022.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The traditional sparse matrix storage format cannot efficiently perform sparse matrix-vector multiplication on modern hardware accelerators such as GPU and Tensor Core, resulting in wasted storage space, poor computing performance and discontinuity of memory access, affecting computing efficiency.

Method used

The diagonal sparse matrix-vector product solution method based on Tensor Core is used to convert the sparse matrix into a diagonal sparse matrix using the BDIA format, and matrix-vector multiplication is performed through Warp-level collaboration and Tensor Core registers, and precision compensation is performed in combination with TF32 mode.

Benefits of technology

The sparse matrix processing speed has been significantly improved, the computing efficiency has been improved by more than 8 times, the storage efficiency has been improved, the memory access continuity has been improved, and the computing performance has been greatly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296294A_ABST
    Figure CN120296294A_ABST
Patent Text Reader

Abstract

The invention discloses a Tensor Core-based diagonal sparse matrix-vector product solving method, and belongs to the technical field of special hardware accelerators. The method comprises the following steps: acquiring a sparse matrix in a BDIA format; wherein the sparse matrix in the BDIA format is obtained by converting a diagonal sparse matrix; preparing an input vector and an output vector, and transferring the input vector, the output vector and the sparse matrix in the BDIA format to a GPU global memory; starting configuration of the CUDA kernel is set; after it is determined that each warp is in a row section in an output vector, the input vector and the sparse matrix in the BDIA format are divided, and vector blocks and matrix blocks are obtained; loading diagonal blocks of the matrix blocks and the vector blocks from a global memory to a Tensor Core register through cooperation in warp; each warp uses a Tensor Core register to execute matrix-vector multiplication, and a final vector result corresponding to the warp is obtained; and writing the final vector result corresponding to each warp into an output vector. According to the invention, storage and transmission quantity can be reduced, and operation speed and energy efficiency are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of dedicated hardware accelerators, and particularly relates to a method for solving diagonal sparse matrix-vector products based on Tensor Core. Background Art

[0002] The calculation of sparse matrix-vector multiplication (SpMV) has wide technical applications and is suitable for occasions that require efficient linear algebra calculations such as engineering / scientific computing, graphics processing, and big data analysis. For example, when applied to the engineering / scientific computing scenario, it will involve specific scenarios such as partial differential equation (PDE) solving, quantum chemistry calculation, and machine learning.

[0003] Partial differential equation (PDE) solving: In the fields of material simulation, fluid mechanics, structural mechanics, electronic structure calculation, etc., the finite difference method (FDM), finite volume method (FVM), or finite element method (FEM) is commonly used to discretize PDEs to obtain large-scale linear equations with diagonal or block diagonal sparse structures.

[0004] Quantum chemistry calculation: In electronic structure or quantum chemistry simulations, it is often necessary to repeatedly solve matrices with specific sparse structures to obtain information such as energy band and energy level distributions.

[0005] Machine learning: In graph neural networks (GNNs) or high-dimensional sparse data processing, diagonal block sparse matrices (or block banded matrices) may also appear, and efficient sparse matrix multiplication is required to accelerate the training or inference process.

[0006] For example, the Graph Convolutional Network (GCN), as an important model of deep learning for graph data, is widely used in fields such as recommendation systems, social network analysis, and knowledge graphs. In the field of recommendation systems, modern personalized recommendation systems often use the GCN model to perform message passing on the sparse graph constructed from "user-item" interaction data, so as to capture the potential correlation features between users and items and provide more accurate and personalized recommendation results. For example, in the movie recommendation scenario, GCN can use the rating or click relationship between users and movies for embedding learning, and then achieve fast interest prediction and real-time recommendation. In the field of social network analysis, the relationships between users often form ultra-large-scale and sparse social graphs. Through iterative message passing, GCN can effectively mine user group structures, community partitions, and potential friend recommendations, thus supporting personalized services and precision marketing in social networks. In the field of knowledge graphs, entities and relationships can be regarded as nodes and edges, with scales often reaching from millions to hundreds of millions, and the structure is highly sparse. Through GCN, entity representations can be learned to complete tasks such as relationship prediction, entity linking, and semantic retrieval, significantly improving the intelligence level of question-answering systems or search engines.

[0007] However, as the scale of graph data continues to increase (the number of nodes or edges can reach tens of millions to hundreds of millions), GCN faces dual challenges of storage and computational efficiency in the actual deployment process. Specifically, in the "message passing" step of GCN, large-scale sparse matrix-vector multiplication (SpMV) or sparse matrix-dense matrix multiplication (SpMM) needs to be repeatedly performed. Traditional sparse matrix storage formats (such as CSR, ELL, etc.) often cannot efficiently execute such sparse operations on deep learning hardware accelerators (such as GPUs, especially Tensor Cores), resulting in the following problems.

[0008] 1) A large number of zero elements are filled, wasting video memory and bandwidth. When traditional sparse matrix formats process block diagonal matrices, a large number of zero fillings are often required. Although these zero fillings can ensure the structural consistency of the matrix, in sparse matrices, the vast majority of elements are zero. This not only causes traditional formats to allocate memory for each element to store these zero elements, wasting a large amount of storage space, but also affects the matrix operation efficiency because reading zero elements also occupies storage bandwidth.

[0009] 2) Poor computational performance. When traditional sparse matrix formats perform matrix-vector multiplication (SpMV), it usually involves a large number of index lookups and non-contiguous memory accesses, which poses great challenges to hardware parallel computing:

[0010] Computation limitation: In traditional formats, the multiplication of a sparse matrix and a dense vector relies on complex indexing operations and element-by-element calculations, which leads to inefficient parallel computing and makes it difficult to fully utilize the computing power of modern hardware (such as GPUs and Tensor Cores).

[0011] Hardware incompatibility: The design of traditional sparse matrix formats does not consider the parallel computing characteristics of modern hardware, especially the efficient matrix multiplication acceleration ability of Tensor Cores, and cannot fully utilize hardware acceleration resources, resulting in low computing efficiency.

[0012] 3) Discontinuous memory access pattern. Since traditional sparse matrix formats usually require frequent index lookups and memory jumps, it leads to discontinuous data access:

[0013] High cache miss rate: Due to discontinuous memory access, it is difficult for the hardware cache to effectively prefetch the required data, resulting in a high cache miss rate and increasing memory access latency;

[0014] Low computing efficiency: Due to the discontinuous memory access pattern, the data flow efficiency during the computing process is reduced, seriously affecting the computing performance.

[0015] In summary, when implementing the SpMV operation, traditional sparse matrix storage formats, such as Compressed Sparse Row (CSR) and External List Storage (ELL), mainly reduce storage space and memory access by recording the position information of non-zero elements, thereby improving computing efficiency. However, with the rise of dedicated hardware accelerators (such as Tensor Cores), these formats fail to fully exploit the parallel computing potential of the hardware and are thus limited in performance. Summary of the Invention

[0016] To address the above problems, the present invention proposes a method for solving the diagonal sparse matrix-vector product based on Tensor Core, which can reduce storage and transmission amounts and improve operation speed and operation energy efficiency.

[0017] To achieve the above object, the technical solution of the present invention includes the following content.

[0018] A method for solving the diagonal sparse matrix-vector product based on Tensor Core, the method comprising:

[0019] Obtain a sparse matrix in BDIA format; wherein, the sparse matrix in BDIA format is obtained by converting a diagonal sparse matrix;

[0020] Prepare an input vector and an output vector, and transfer the input vector, output vector, and sparse matrix in BDIA format to the GPU global memory;

[0021] Set the startup configuration of the CUDA kernel;

[0022] After determining the row segments of each warp in the output vector, partition the input vector and the sparse matrix in BDIA format to obtain vector blocks and matrix blocks; wherein, the matrix block is an 8×8 matrix block in FP32 format, and the vector block is an 8×1 vector block in FP32 format.

[0023] Through within-warp cooperation, load the diagonal blocks of the matrix block and the vector block from global memory into Tensor Core registers.

[0024] Each warp uses Tensor Core registers to perform matrix-vector multiplication to obtain the final vector result corresponding to the warp.

[0025] Write the final vector result corresponding to each warp into the output vector y.

[0026] Furthermore, the sparse matrix in BDIA format consists of an offset array and a block value array. The offset array records the positions of the diagonals of each block in the entire diagonal sparse matrix, and the block value array is used to store all the element values of the corresponding diagonal block.

[0027] Furthermore, the launch configuration of the CUDA kernel includes: configuring the thread block size, the number of computing thread blocks, and launching the kernel.

[0028] Furthermore, through within-warp cooperation, loading the diagonal blocks of the matrix block from global memory into Tensor Core registers includes:

[0029] Allocate and register a 16×8 minimum computing unit in the Tensor Core register area. The minimum computing unit includes: register 0 at the upper left position, register 1 at the lower left position, register 2 at the upper right position, and register 3 at the lower right position.

[0030] Transfer the 8×8 matrix block in FP32 format to register 0 and register 2, and use the inline assembly instruction cvt.rna.satfinite.tf32.f32 to convert the 8×8 matrix block in FP32 format to an 8×8 matrix block in TF32 format.

[0031] Calculate the low-precision bits lost when converting the 8×8 matrix block in FP32 format to an 8×8 matrix block in TF32 format to obtain the first lost precision, and after converting the first lost precision to TF32 format, store it in register 1 and register 3.

[0032] Furthermore, through within-warp cooperation, loading the diagonal blocks of the vector block from global memory into Tensor Core registers includes:

[0033] Transfer the 8×1 vector block in FP32 format to the first column of register 0 and register 1, and use the inline assembly instruction cvt.rna.satfinite.tf32.f32 to convert the 8×1 vector block in FP32 format into an 8×1 vector block in TF32 format;

[0034] Calculate the low-precision loss when the 8×1 vector block in FP32 format is converted into an 8×1 vector block in TF32 format to obtain the second loss precision, and after converting the second loss precision into TF32 format, store it in the eighth column of register 1 and register 3.

[0035] Furthermore, each warp uses Tensor Core registers to perform matrix-vector multiplication to obtain the final vector result corresponding to this warp, including:

[0036] Each thread performs matrix multiplication by calling the mma.sync.aligned.m16n8k8.row.col.f32.tf32.tf32.f32 instruction twice to obtain a 16×16 result matrix; among them, the 16×16 result matrix consists of the upper-left block of high-bit multiplied by high-bit, the upper-right block of high-bit multiplied by low-bit, the lower-left block of low-bit multiplied by high-bit, and the lower-right block of low-bit multiplied by low-bit;

[0037] Extract the first column of the upper-left block, the upper-right block, and the lower-left block and merge them to obtain the final vector result calculated by this thread;

[0038] Accumulate the final vector results calculated by the threads within the warp to obtain the final vector result corresponding to this warp.

[0039] Furthermore, the output vector y is the calculation result of the diagonal sparse matrix-vector product in the field of partial differential equation solving, the field of quantum chemistry calculation, or the field of machine learning.

[0040] An optimization method for accelerating the GCN model based on the Tensor Core block diagonal sparse matrix format, the method includes:

[0041] Step S1: Generate the graph structure of the training data and calculate the adjacency matrix of this graph structure Among them, the graph structure includes several categories of nodes;

[0042] Step S2: Generate the node feature matrix H (0) ;

[0043] Step S3: The graph neural network is based on the node feature matrix H (0) and the adjacency matrix Calculate the predicted scores of node pairs, and calculate the loss based on the predicted scores and the true scores of the node pairs; where each node pair consists of nodes of different categories, and the graph neural network is based on the node feature matrix H (0) and the adjacency matrix When calculating the predicted scores of node pairs, the product result between the node feature matrix and the adjacency matrix is obtained based on the method for solving the diagonal sparse matrix-vector product based on Tensor Core according to any one of claims 1 to 6

[0044] Step S4: Perform backpropagation based on the loss L to complete the training of the graph convolutional network in this round

[0045] Step S5: Repeat Step S1 - Step S4 until the graph convolutional network reaches a specific metric on the validation set, and then obtain the trained graph convolutional network

[0046] An electronic device, the electronic device includes: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, it implements the method for solving the diagonal sparse matrix-vector product based on Tensor Core according to any one of the above or implements the optimization method of the accelerated GCN model in the block diagonal sparse matrix format based on Tensor Core as described above

[0047] A computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, they implement the method for solving the diagonal sparse matrix-vector product based on Tensor Core according to any one of the above or implement the optimization method of the accelerated GCN model in the block diagonal sparse matrix format based on Tensor Core as described above

[0048] Compared with the prior art, the present invention has at least the following beneficial effects

[0049] 1) Optimized BDIA format based on Tensor Core characteristics. Based on the existing BDIA format, the present invention redesigns the block size in combination with the characteristics of NVIDIA Tensor Core, selects block layouts of 8x8 and 16x16, so as to make the most of the computing power of Tensor Core, thereby reducing zero padding, improving storage efficiency, and at the same time making matrix operations on Tensor Core hardware more efficient, breaking through the bottleneck between storage and computing in traditional formats, and significantly improving the processing speed of diagonal sparse matrices

[0050] 2) Accelerating with Tensor Core for the first time in the BDIA field. The present invention is the first technical solution to introduce Tensor Core to accelerate matrix-vector multiplication during the BDIA format processing. By applying the efficient matrix operation of Tensor Core to the BDIA format, a revolutionary performance improvement is achieved. Especially in complex operations, the calculation time is significantly shortened, and the overall operation efficiency is improved.

[0051] 3) Warp-level optimization and efficient use of shared memory. The present invention reconstructs the CUDA kernel, optimizes thread allocation and cooperation through Warp-level operations, reduces thread divergence, and makes full use of shared memory to promote data processing efficiency. By optimizing kernel operations and data distribution, more efficient data loading and calculation are achieved, further reducing latency and improving the overall calculation efficiency and throughput.

[0052] 4) Balanced strategy of high precision and high performance. Combining with the TF32 mode of Tensor Core, the triple floating-point simulation (3TF) strategy is used for precision compensation to balance high performance and calculation precision. That is to say, while maintaining high calculation precision, the present invention achieves a leapfrog improvement in performance through hardware acceleration, and the speed is increased by more than 8 times compared with the traditional method, which is of great significance to fields such as scientific computing and graphics processing. Description of the Drawings

[0053] Figure 1 Flowchart of the method for solving the diagonal sparse matrix-vector product based on Tensor Core.

[0054] Figure 2 BDIA data structure diagram.

[0055] Figure 3 Thread group wrap task division diagram.

[0056] Figure 4 Flowchart of the optimization method for accelerating the GCN model based on the block diagonal sparse matrix format of Tensor Core. Detailed Description of the Invention

[0057] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below through specific embodiments in combination with the accompanying drawings.

[0058] Tensor Core acceleration is an NVIDIA hardware unit optimized for deep learning. It uses 16-bit floating-point numbers to quickly perform tensor operations, thereby accelerating matrix operations. The present invention utilizes the characteristics of Tensor Core and the BDIA (Blocked Diagonal) format to implement the calculation of solving the diagonal sparse matrix-vector product.

[0059] The method for solving diagonal sparse matrix-vector multiplication based on Tensor Core of the present invention, as Figure 1 shown, includes the following steps 1 to 8.

[0060] Step 1: Obtain a sparse matrix in BDIA format, which is converted from a diagonal sparse matrix.

[0061] As Figure 2 shown, in the BDIA format, an offset array (offsets) is used to record the positions of each block diagonal in the entire matrix, and a block value array (block_values) is used to store all the element values of the corresponding diagonal block. Different from the traditional diagonal storage (DIA), BDIA divides the matrix into multiple small blocks, each small block corresponding to a certain range of rows and columns on a certain diagonal of the matrix. Therefore, these small blocks can be batch-multiplied by Tensor Core on the GPU side, thus creating a prerequisite for improving the throughput of sparse matrix-vector multiplication (SpMV).

[0062] Among them, when constructing the offset array, the host side records the start and span information of each diagonal according to the diagonal positions of the non-zero blocks in the diagonal sparse matrix, and stores the start and span information of the diagonal in the offsets array for quickly locating the corresponding diagonal block on the GPU side.

[0063] When constructing the block value array, 8×8 small blocks in the matrix are extracted (filled with zeros or skipped if there is no data) according to the offsets and stored sequentially in the block value array.

[0064] Finally, ensure that the sizes of the offset array and the block value array correspond correctly, and record the row and column sizes of each block to assist in the subsequent adaptation and configuration of Tensor Core registers.

[0065] Step 2: Prepare the input / output vectors and transfer the input / output vectors and the sparse matrix in BDIA format to the GPU global memory.

[0066] In the GCN scenario, the input vector x can represent a column (or a mini-batch) of the node feature matrix, with a size matching the number of rows of the matrix. The size of the output vector is the same as that of the input vector x and is initially 0, storing the SPMV result.

[0067] In this embodiment, the offsets, block_values, x, and y are copied to the GPU global memory through the CUDA API.

[0068] Step 3: Set the startup configuration of the CUDA kernel.

[0069] The launch configuration of the CUDA kernel includes configuring the thread block size, the number of computing thread blocks, and launching the kernel, etc.

[0070] Thread block size (BLOCK_SIZE): Usually choose 128 or 256, and the warp (32 threads) is the basic scheduling unit.

[0071] Number of computing thread blocks (gridDim): If the number of rows of the matrix is row, then let gridDim = row / 32 (each warp calculates 32 output elements).

[0072] Launch the kernel: Call spmv_bdia_kernel<<<gridDim,blockDim>>>(offsets,block_values,x,y,...), and pass in relevant parameters.

[0073] Step 4: After determining the row section of each warp in the output vector, perform partitioning on the input vector and the sparse matrix in BDIA format to obtain vector blocks and matrix blocks; among them, the matrix block is an 8*8 matrix block in FP32 format, and the vector block is an 8*1 vector block in FP32 format.

[0074] 1) Calculate the global thread ID: In the kernel function, after obtaining global_tid through blockIdx.x and threadIdx.x, get warp ID = global_tid / 32 and lane ID = global_tid % 32.

[0075] 2) Determine the warp division of labor: The warp ID determines the row section [warpId * 32, warpId * 32 + 31] in the output vector, so as to ensure that the division of labor of different warps does not overlap. Figure 3 As shown in the task division diagram of the thread group (warp), each warp is responsible for calculating a section of the output vector (for example, 32 result elements). The threads in the warp work together, sequentially read the diagonal blocks of the matrix (and their corresponding vector blocks) in the loop, and complete block-level matrix multiplication through Tensor Core.

[0076] Step 5: Through cooperation within the warp, load the diagonal blocks of the matrix block and the vector block from global memory to the TensorCore register.

[0077] 1) Determine the diagonal blocks to be accessed.

[0078] The warp traverses the offsets, compares the row section it is responsible for, and identifies the index position of the relevant 8×8 matrix block in block_values.

[0079] 2) Load matrix blocks and vector blocks.

[0080] The loading of matrix blocks and vector blocks does not use shared memory. Instead, each thread within a warp directly reads 8×8 block data and the corresponding input vector sub-block x_sub from the GPU global memory and stores them in registers (or local variables) for subsequent operations, thus ensuring that each thread has a clear division of labor and avoiding read-write conflicts.

[0081] Specifically, the present invention first allocates and registers a Tensor Core register area, loads small matrix blocks and small vector blocks into the registers, and then performs the transfer of matrix blocks and vector blocks. Among them, the transfer method of matrix blocks is that the 8×8 matrix block is mapped to the upper half (registers 0 and 2) of 16×8, and the cvt.rna.satfinite.tf32.f32 inline assembly instruction is used to convert FP32 to TF32; and the lost low precision is also converted to TF32 and stored in the lower half (registers 1, 3), forming high and low matrix blocks.

[0082] The transfer method of vector blocks is to store the 8×1 vector in the first column of registers 0 and 1 of the 8×16 minimum unit; the cvt.rna.satfinite.tf32.f32 is also used to convert FP32 to TF32; and the lost precision is placed in the eighth column of registers 2 and 3, forming high and low vectors.

[0083] Step 6: Each warp uses the Tensor Core register to perform matrix-vector multiplication to obtain the final vector result corresponding to this warp.

[0084] The present invention performs matrix multiplication through GPU inline assembly instructions, calculates the corresponding result vector block, and finally writes it back to the global memory.

[0085] 1) Each thread calls the mma.sync instruction to perform matrix multiplication of the minimum calculation unit between vector blocks and matrix blocks to obtain the final vector result of this thread.

[0086] The present invention first performs matrix multiplication by calling the mma.sync.aligned.m16n8k8.row.col.f32.tf32.tf32.f32 instruction twice to obtain a 16×16 result matrix, where this 16×16 matrix consists of 4 8×8 sub-blocks: upper left (high×high), upper right (high×low), lower left (low×high), lower right (low×low).

[0087] Then, extract the first columns of the upper-left, upper-right, and lower-left 3 blocks and combine them to obtain the final vector result. Among them, the lower-right (low bit × low bit) is ignored because the magnitude is too small (e^-20), which does not affect the accuracy and can reduce redundant accumulation.

[0088] 2) Based on the final vector results of each thread, obtain the final vector result corresponding to the warp.

[0089] When generating the final vector result corresponding to the warp, thread synchronization should be ensured, that is, after each thread within the warp completes the operation on the accumulator, the output is written back.

[0090] Step 7: Write the final vector result corresponding to each warp into the output vector y.

[0091] Write the final vector result corresponding to each warp into y[out_row], where out_row is the row interval responsible for the warp; if it is not in this interval, it is not written to avoid competition.

[0092] Step 8: Return and post-process the output vector y.

[0093] If further processing or visualization of the output vector y is required on the CPU side, call the CUDA API to copy the output vector y back to the host.

[0094] In summary, the present invention uses Tensor Core for acceleration in the BDIA field, thereby providing a significant improvement in the operation speed. The calculation efficiency is significantly improved compared with the traditional method, and it is applicable to the calculation of diagonal sparse matrix-vector products in fields such as partial differential equation (PDE) solving, quantum chemistry calculation, and machine learning.

[0095] Next, taking the graph neural network training of the "personalized movie recommendation system" in the field of machine learning as an example, the application of the present invention will be specifically described. In this embodiment, the method for solving the diagonal sparse matrix-vector product based on Tensor Core is introduced into the forward propagation process of the graph neural network, so as to obtain an optimized method for accelerating the GCN model in a specific block diagonal sparse matrix format based on Tensor Core. As Figure 4 described, the optimized method for accelerating the GCN model in the block diagonal sparse matrix format based on Tensor Core includes the following steps S1 to S5.

[0096] Step S1: Generate the graph structure of the training data and calculate the adjacency matrix of the graph structure Among them, the graph structure includes several types of nodes.

[0097] In this embodiment, the present invention randomly divides the rating data of the personalized movie recommendation system into a training set (80%), a validation set (10%), and a test set (10%) so as to evaluate the recommendation effect by calculating the error (such as RMSE) between the predicted rating and the true rating during validation and testing. Among them, the data of the personalized movie recommendation system usually comes from the interaction behaviors between users and movies, such as user ratings, viewing history, click behaviors, etc.

[0098] After obtaining the training set, this embodiment first loads the MovieRating-10K dataset from the training set data. This dataset contains the interaction data (ratings from 1 to 5 stars) of 10,000 users and 5,000 movies. Then, users and movies are regarded as nodes in the graph, and the interaction behaviors between users and movies form the edges in the graph. Since the data structure of this graph shows sparsity and only some users rate some movies, an adjacency matrix A in the form of a sparse matrix is obtained. Among them, the shape of the adjacency matrix A is 15,000×15,000 (15,000 is the total number of nodes). Finally, the adjacency matrix A is normalized and a self-loop is added to generate the adjacency matrix which will be used as the input for the graph convolution operation in the GCN model. At this time, the adjacency matrix shows a block-diagonal sparse structure, where the block structure corresponds to the local connections between user groups and movie categories.

[0099] In addition, due to the sparsity of the adjacency matrix in the recommendation system, the present invention stores it in the 8×8 blocked BDIA (Blocked Diagonal) format. The BDIA format is particularly suitable for the local connectivity in large-scale graph data.

[0100] Step S2: Generate the node feature matrix H (0) ,

[0101] The present invention initializes 32-dimensional feature vectors for each user and movie node (which can be randomly initialized or use some prior information). These feature vectors form the initial node feature matrix H (0) , where the node feature matrix H (0) has a size of 15,000×32.

[0102] Step S3: The graph neural network calculates the predicted rating of the node pair based on the node feature matrix H (0) and the adjacency matrix and calculates the loss according to this predicted rating and the true rating of the node pair; among them, each node pair is composed of nodes of different categories, and the graph neural network is based on the node feature matrix H (0) and the adjacency matrix When calculating the predicted scores of node pairs, the method for solving the diagonal sparse matrix-vector product based on Tensor Core provided by the present invention is used to implement the product calculation between the node feature matrix and the adjacency matrix.

[0103] Step S3.1: Based on the node feature matrix H (0) Perform graph convolution calculation to obtain the node feature matrix H (w) .

[0104] The GCN network of the present invention includes two layers. The input feature dimension of the first layer of GCN is 32, and the output dimension is 64, so that the node feature matrix can be obtained, where σ is ReLU, and W (l) is the corresponding trainable parameter matrix. The input feature dimension of the second layer of GCN is 64, and the output dimension is 16, so that the node feature matrix

[0105] For each epoch, the present invention performs forward propagation in the GCN network. Among them, in the matrix multiplication and the matrix multiplication process, the BDIA+Tensor Core acceleration operator of the present invention will be called to replace the traditional CSRSPMV operation, thus greatly improving the efficiency.

[0106] Step S3.2: Make predictions based on the node feature matrix H (w) to obtain the predicted scores of node pairs.

[0107] To convert the output dimension 16 into the prediction of movie scores, a simple scoring prediction module (such as a fully connected layer + activation function) can be connected after the GCN output or the inner product form can be used to obtain the predicted score of node pair i

[0108]

[0109] where h u and h m are the output representations of the user and movie nodes obtained based on the node feature matrix H (2) respectively, represents the concatenation or splicing operation, and MLP represents the multi-layer perceptron.

[0110] Step S3.3: Based on the predicted scores of node pairs and the true scores of the node pairs, obtain the loss L.

[0111] In this embodiment, the mean square error (MSE) is used as the loss function to obtain the loss L.

[0112]

[0113] Among them, N represents the number of node pairs, and y i represents the true score of node pair i.

[0114] In one embodiment, the present invention uses the Adam optimizer (Adaptive Moment Estimation), which can accelerate model convergence and improve training stability by adaptively adjusting the learning rate of each parameter. The update rule of the Adam optimizer is as follows:

[0115]

[0116] where θ t and θ t+1 are the model parameters obtained in the t-th round and the (t + 1)-th round of training respectively, and are the estimates of the momentum and variance of the gradient during the t-th round of training respectively, η is the learning rate, and ∈ is a small constant to avoid division by zero errors.

[0117] Step S4: Perform backpropagation based on the loss L to complete the training of the graph convolutional network in this round.

[0118] The present invention calculates the gradient based on the above loss L and updates each layer of W (0) W (1) according to the gradient. Among them, the Adam algorithm can adaptively adjust the learning rate.

[0119] Step S5: Repeat Step S1 - Step S4 until the graph convolutional network reaches a specific metric on the validation set, and then obtain the trained graph convolutional network.

[0120] After completing the training of the graph convolutional network in each round, it is judged whether the loss of the graph convolutional network has no obvious decrease or reaches the maximum epoch. If so, the training of the graph neural network is completed. Then, calculate the RMSE or corresponding metric of the graph neural network on the validation set. If the root mean square error (RMSE) or corresponding metric is not satisfactory, repeat the training; otherwise, output the trained graph convolutional network. The calculation formula of RMSE is as follows:

[0121]

[0122] It can be seen that in the GCN model, the adjacency matrix of a graph usually exhibits a sparse structure, where most elements are zero, and the non-zero elements are concentrated near the main diagonal or in local block regions. Nodes in a graph usually have local connectivity, meaning that most nodes are only connected to a small number of neighbor nodes. Therefore, the non-zero elements in the adjacency matrix tend to be distributed near the diagonal or in local block regions. This structure poses challenges to the storage and calculation of sparse matrices. In this embodiment, a method for solving the diagonal sparse matrix-vector product based on Tensor Core is applied to the message passing optimization in the GCN model prediction task. By optimizing the storage structure of the adjacency matrix in the GCN task and combining the hardware characteristics of Tensor Core, the performance of the GCN model is improved in multiple aspects, mainly reflected in the reduction of GPU video memory occupancy and the acceleration of training time.

[0123] Specifically, through the BDIA+Tensor Core acceleration operator, not only is the training time of each epoch significantly reduced, with a speed increase of 1.5 - 2.0 times compared to the traditional CSR, the memory occupancy can be reduced by 30% - 50%, the training and convergence speed is increased by 10% - 20%, the speed of generating prediction results is accelerated, the real-time response ability of the system and the user experience are improved, but also the TF32+ high-low bit compensation strategy achieves efficient parallelism without sacrificing the FP32 accuracy, and there is almost no difference from the FP32 method in terms of indicators such as Top-N recommendation and RMSE.

[0124] In addition, since more training rounds can be performed or a larger data volume can be processed after acceleration, indicators such as the Top-10 recommendation accuracy and nDCG@10 tend to be maintained or even slightly improved, achieving a better personalized experience.

[0125] Moreover, the 8×8 block BDIA format proposed in the present invention effectively solves the problems in the above-mentioned traditional sparse matrix formats, especially for the characteristics of sparse graph data in the GCN model, and has significant advantages. If the block size is set too large, there are too many zero elements in the matrix, and the register resources of Tensor Core cannot be fully utilized, resulting in waste of memory bandwidth and storage space and a decrease in computing efficiency. If the block is set too small, the parallel computing ability of Tensor Core cannot be effectively utilized, reducing the utilization efficiency of hardware resources.

[0126] In summary, for the common adjacency matrices in GCN, the connections of many nodes are only limited to local neighbors, so the sparse matrix presents a block-diagonal sparse structure. Traditional sparse matrix formats (such as CSR) must fill in a large number of zero elements, while the BDIA format stores the matrix in blocks and only stores the non-zero elements in the block-diagonal area, avoiding the redundant storage of a large number of zero elements. After adopting the BDIA format, the sparse structure of the adjacency matrix can be efficiently represented, and the storage requirement is reduced by 30% - 50%, providing a more efficient storage solution for large-scale processing of graph data.

[0127] The BDIA format enables the adjacency matrix to be stored in a block-diagonal manner, reducing unnecessary access to zero elements and improving the locality of memory access. Due to the block-diagonal structure of the sparse matrix, data access becomes more continuous, thus optimizing the parallelism during the calculation process.

[0128] Tensor Core supports efficient matrix multiplication. Through the MMA instruction, the speed of dense matrix calculation can be significantly improved. The storage method of the BDIA format is highly compatible with the parallel computing ability of Tensor Core, enabling the calculation process to fully utilize the hardware acceleration resources. The matrix multiplication operation of each 8×8 block perfectly matches the parallel execution ability of Tensor Core, significantly enhancing the message passing speed in the GCN model.

[0129] Since the minimum computing unit of Tensor Core is a 16×8 matrix, in this embodiment, the 8×8 matrix block is divided into two parts. The upper half occupies register 0 and register 2 of Tensor Core, and the in-line assembly instruction cvt.rna.satfinite.tf32.f32 is used to convert the FP32-precision data into TF32-precision data for correct acceleration using Tensor Core. The lost precision is ensured by converting the data into TF32 format and storing it in register 1 and register 3, so that the high-order and low-order data can be processed in parallel. Finally, these four 8×8 matrix blocks are combined to obtain a 16×16 matrix. The results of the upper left, upper right, and lower left are all extracted for the final calculation, while the low-order * low-order result in the lower right is discarded because its precision is too small (about e-20) and does not affect the final prediction result.

[0130] After adopting the BDIA format, the matrix data can be stored in a block-aligned manner and supports batch loading, which makes the memory access pattern more continuous. Such a storage method is highly compatible with the memory access characteristics of the GPU, effectively reducing the cache miss rate and improving the memory access efficiency.

[0131] Due to the continuity of the data, memory access becomes more efficient, reducing memory jumps and index lookup operations, thereby reducing latency during the calculation process and enhancing the overall performance of the system.

[0132] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the present disclosure. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and examples are only considered exemplary, and the present disclosure is not limited to the exact structures already described and shown in the drawings, and various modifications and changes can be made without departing from its scope.

Claims

1. A method for solving diagonal sparse matrix-vector products based on Tensor Core, characterized in that, The method includes: obtaining a sparse matrix in BDIA format; wherein, the sparse matrix in BDIA format is obtained by converting a diagonal sparse matrix; Preparing an input vector and an output vector, and transferring the input vector, the output vector, and the sparse matrix in BDIA format to the GPU global memory; Setting the startup configuration of the CUDA kernel; After determining the row section of each warp in the output vector, performing partitioning on the input vector and the sparse matrix in BDIA format to obtain a vector block and a matrix block; wherein, the matrix block is an 8×8 matrix block in FP32 format, and the vector block is an 8×1 vector block in FP32 format; Through within-warp cooperation, loading the diagonal block of the matrix block and the vector block from the global memory to the Tensor Core register; Each warp uses the Tensor Core register to perform matrix-vector multiplication to obtain the final vector result corresponding to this warp; Writing the final vector result corresponding to each warp to the output vector.

2. The method for solving diagonal sparse matrix-vector product based on Tensor Core according to claim 1, wherein The sparse matrix in BDIA format consists of an offset array and a block value array. The offset array records the positions of the diagonals of each block in the entire diagonal sparse matrix, and the block value array is used to store all the element values of the corresponding diagonal block.

3. The method for solving the diagonal sparse matrix-vector product based on Tensor Core according to claim 1, wherein The startup configuration of the CUDA kernel includes: configuring the thread block size, the number of computing thread blocks, and starting the kernel.

4. The method for solving diagonal sparse matrix-vector product based on Tensor Core according to claim 1, wherein Through within-warp cooperation, loading the diagonal block of the matrix block from the global memory to the Tensor Core register includes: Allocating and registering a 16×8 minimum computing unit in the Tensor Core register area. The minimum computing unit includes: register 0 at the upper left position, register 1 at the lower left position, register 2 at the upper right position, and register 3 at the lower right position; Transferring the 8×8 matrix block in FP32 format to register 0 and register 2, and using the inline assembly instruction cvt.rna.satfinite.tf32.f32 to convert the 8×8 matrix block in FP32 format to an 8×8 matrix block in TF32 format; Calculating the low-precision bits lost when converting the 8×8 matrix block in FP32 format to an 8×8 matrix block in TF32 format to obtain a first lost precision, and after converting this first lost precision to TF32 format, storing it in register 1 and register 3.

5. The method for solving the diagonal sparse matrix-vector product based on Tensor Core according to claim 4, wherein Through within-warp cooperation, loading the diagonal block of the vector block from the global memory to the Tensor Core register includes: Transferring the 8×1 vector block in FP32 format to the first column of register 0 and register 1, and using the inline assembly instruction cvt.rna.satfinite.tf32.f32 to convert the 8×1 vector block in FP32 format to an 8×1 vector block in TF32 format; Calculating the low-precision bits lost when converting the 8×1 vector block in FP32 format to an 8×1 vector block in TF32 format to obtain a second lost precision, and after converting this second lost precision to TF32 format, storing it in the eighth column of register 1 and register 3.

6. The method for solving diagonal sparse matrix-vector product based on Tensor Core according to claim 5, wherein Each warp uses Tensor Core registers to perform matrix-vector multiplication to obtain the final vector result corresponding to that warp, including: Each thread performs matrix multiplication by calling the mma.sync.aligned.m16n8k8.row.col.f32.tf32.tf32.f32 instruction twice to obtain a 16×16 result matrix; wherein, the 16×16 result matrix consists of an upper-left block of high-order multiplied by high-order, an upper-right block of high-order multiplied by low-order, a lower-left block of low-order multiplied by high-order, and a lower-right block of low-order multiplied by low-order; Extract the first columns of the upper-left block, the upper-right block, and the lower-left block and combine them to obtain the final vector result calculated by this thread; Accumulate the final vector results calculated by the threads within the warp to obtain the final vector result corresponding to that warp.

7. The method for solving diagonal sparse matrix-vector product based on Tensor Core according to claim 5, wherein The output vector y is the calculation result of a diagonal sparse matrix-vector product in the field of partial differential equation solving, the field of quantum chemistry calculation, or the field of machine learning.

8. An optimization method for accelerating the GCN model based on the block diagonal sparse matrix format of Tensor Core, characterized in that, The method includes: Step S1: Generate the graph structure of the training data and calculate the adjacency matrix of the graph structure Among them, the graph structure includes nodes of several categories; Step S2: Generate the node feature matrix H (0) ; Step S3: The graph neural network is based on the node feature matrix H (0) and the adjacency matrix calculate the predicted scores of node pairs, and calculate the loss based on the predicted scores and the true scores of the node pairs; wherein, each node pair consists of nodes of different categories, and the graph neural network is based on the node feature matrix H (0) and the adjacency matrix When calculating the predicted scores of node pairs, the product result between the node feature matrix and the adjacency matrix is obtained based on the method for solving the diagonal sparse matrix-vector product based on Tensor Core according to any one of claims 1 to 6; Step S4: Perform backpropagation based on the loss L to complete the training of the graph convolutional network in this round; Step S5: Repeat Step S1 - Step S4 until the graph convolutional network reaches a specific metric on the validation set, and then obtain the trained graph convolutional network.

9. An electronic device, characterized in that, The electronic device includes: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, it implements the method for solving the diagonal sparse matrix-vector product based on Tensor Core according to any one of claims 1-7 or implements the optimization method for accelerating the GCN model in the block diagonal sparse matrix format based on Tensor Core according to claim 8.

10. A computer-readable storage medium, characterized in that, Computer program instructions are stored on the computer-readable storage medium, and when the computer program instructions are executed by the processor, they implement the method for solving the diagonal sparse matrix-vector product based on Tensor Core according to any one of claims 1-7 or implement the optimization method for accelerating the GCN model in the block diagonal sparse matrix format based on Tensor Core according to claim 8.