Accelerator and acceleration method based on Top-K sparse moment vector multiplication
By using data preprocessing, sparse matrix coding and FPGA architecture design methods in Top-K sparse matrix vector multiplication, the problem of low performance of Top-K SpMV algorithm in the prior art is solved, and significant computing efficiency and performance improvement are achieved.
Patent Information
- Application Number
- CN202510166895.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-12-09
- Filing Date
- 2025-02-14
- Publication Date
- 2025-06-10
AI Technical Summary
When performing Top-K sparse matrix vector multiplication, the prior art faces bottlenecks in efficient memory transmission and data storage formats, resulting in poor performance, especially the problem of random access and handling of data in the sorting stage has not been effectively solved.
An accelerator based on Top-K sparse matrix vector multiplication is proposed. Through data preprocessing and sparse matrix coding, the data structure is optimized by steps such as quantization, sparse reconstruction and shuffling, reducing memory transmission and calculation amount, and designing a quantitative computing core and a full-precision computing core in the FPGA architecture to improve computing efficiency.
The overall computing efficiency of the Top-K SpMV algorithm is significantly improved, and the end-to-end performance is improved by 153.3 times, 2.5 times and 3.4 times the baseline performance of CPU, GPU and FPGA respectively, significantly reducing the latency and computational complexity of random access of memory.
Smart Images

Figure CN120123629A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to an accelerator and an acceleration method based on Top-K sparse matrix-vector multiplication. Background Art
[0002] The Top-K Sparse Matrix-Vector Multiplication (SpMV) operator is an optimized linear algebra operation designed to efficiently select the K most relevant items for each user or entity from the result of the multiplication of a sparse matrix and a vector. Specifically, the Top-K SpMV operator not only performs the traditional Sparse Matrix-Vector Multiplication (SpMV), but also further integrates the function of Top-K selection to ensure that the calculation result can be directly used to generate personalized recommendation lists or priority rankings in other application scenarios.
[0003] The Top-K SpMV operator is a key operator in fields such as recommendation systems, artificial intelligence, and graph analysis, and its performance has a significant impact on the overall efficiency of these applications. In practical applications, the Sparse Embeddings technology is widely used. For example, in the recommendation system of an e-commerce platform, the Sparse Embeddings technology is used to process user behavior and product features and provide real-time recommendation services for users. Due to its interpretability and the characteristic of reducing computational complexity, the Sparse Embeddings technology has also been widely used in fields such as social networks, image retrieval and classification, natural language processing, and graph analysis.
[0004] In a recommendation system, the Top-K SpMV operator is used to quickly calculate the similarity between users and products, so as to select the K products that best match the user's preferences. The recommendation system must respond in real time to the ever-changing user behavior to provide personalized recommendation results. The performance of the system response is crucial because any delay will directly affect the user experience, and thus affect the user's satisfaction and retention rate. Studies have shown that a 1-second increase in system latency will result in a 13% decrease in user conversion rate.
[0005] In a social network, the Top-K SpMV operator can efficiently identify the K friends or communities most relevant to a specific user. This process involves the calculation of a sparse adjacency matrix. Especially in the case of a large and dynamically changing user relationship network, the performance improvement can significantly shorten the latency of information dissemination, ensure that users receive social information and recommendations in a timely manner, and thus enhance the interactivity.
[0006] In image retrieval and classification applications, the Top-K SpMV operator is used to quickly find the K images most relevant to a given query from a large number of image features. Since image data is usually high-dimensional and sparse, the efficiency of the algorithm directly affects the real-time performance of retrieval. The improvement of end-to-end performance can ensure that users can get results almost immediately after uploading a query image, which enhances the user experience.
[0007] In NLP tasks, the Top-K SpMV operator is used to calculate the similarity between texts to quickly retrieve the K documents or sentences most relevant to a specific text. For many applications, such as automatic question-answering systems and information retrieval, the system must respond to user queries in milliseconds. An efficient Top-K SpMV operator can ensure that the system quickly returns accurate results when processing large-scale corpora, thereby improving the overall efficiency of the system.
[0008] In the field of graph analysis, Top-K SpMV can accelerate the calculation of sparse adjacency matrices, quickly find the K neighbors most relevant to a certain node, and thus improve the analysis efficiency of graph-structured data. For retrieval-based generative models, Top-K SpMV can efficiently retrieve the documents most relevant to a query from a large corpus, improving the accuracy and relevance of the generated content.
[0009] In practical scenarios, the scale of sparse matrices usually reaches the order of millions or even tens of millions of rows, and the non-zero elements are usually very randomly distributed. The highly random data access patterns, large-scale streaming computing requirements, and strong real-time constraints pose huge challenges to traditional cache-based CPU and GPU architectures.
[0010] Specifically, the Top-K SpMV algorithm implemented based on the traditional CPU and GPU architectures usually involves two-stage operations of SpMV and Sort. The Sort stage usually involves a large amount of random data access and transfer, resulting in more than 60% of the latency. At the same time, there is an inherent mismatch between the regular architecture and the sparsity of the matrix, and the dense memory random access accounts for more than 60% of the total computing latency of SpMV. The current FPGA architecture design for Top-K SpMV still follows the traditional optimization method for the SpMV operator, and generally uses Float-32 data for full-scale SpMV calculations. Although some studies have introduced quantization techniques to relieve the pressure on the transmission bandwidth, in order to ensure the calculation accuracy, the data bit width still remains at a relatively high level, such as 20-bit. However, this high data bit width usage strategy leads to a significant bottleneck: within a single clock cycle, the 512-bit high-bandwidth memory transfer bit width can only carry at most 15 non-zero element transfers. Existing technologies tend to use the Coordinate (COO) or Compressed Sparse Row (CSR) format to store sparse matrices. The COO data format does not perform any form of compression on the coordinates of non-zero elements, resulting in low utilization of the transmission bandwidth and large memory occupancy overhead. Although the CSR data format compresses the row coordinates of non-zero elements, its variable-length row index array representation is only suitable for row-level acceleration and fails to fully utilize the acceleration potential at the non-zero element level; using a fixed-length row index representation can parallelize the optimization at the non-zero element granularity, but it causes waste of the transfer bit width, which exacerbates the problem of low utilization of the transmission bandwidth.
[0011] CN117495651A discloses a GPU acceleration optimization method based on sparse matrix-vector multiplication, which is characterized by including: obtaining a sparse matrix and a first multiplicand vector in the memory space, reordering the sparse matrix and the first multiplicand vector to obtain a first matrix and a second multiplicand vector; compressing the first matrix to obtain a second matrix; determining the number of rows allocated to each warp according to the number of non-zero elements in each row of the second matrix to obtain an allocation result, and performing warp allocation processing on the second matrix according to the allocation result to obtain a number of warps with allocation completed; multiplying the non-zero elements in each row of the number of warps with allocation completed by the second multiplicand vector for dot product calculation to generate a calculation result. This technical solution optimizes the performance of the Top-K SpMV algorithm in the SpMV calculation stage through means such as reordering, matrix compression, and warp allocation, but it fails to completely solve the problem of random data access and transfer in the Sort stage. Especially when dealing with matrices with high sparsity or uneven distribution of non-zero elements, the high latency of memory random access has not been fundamentally improved. In addition, this technical solution does not fully consider the characteristics of the FPGA architecture, such as data bit-width optimization and fine-grained parallel computing, nor does it propose an efficient solution for the Top-K selection stage, and lacks a cross-platform adaptive optimization strategy.
[0012] In summary, the inefficient memory transfer problem remains the key bottleneck restricting the performance of the Top-K SpMV acceleration system. How to optimize the memory transfer mechanism and data storage format is a problem that needs to be further solved.
[0013] In addition, on the one hand, there are differences in the understanding of those skilled in the art; on the other hand, although the applicant studied a large number of literatures and patents when making this invention, all details and contents are not listed in detail due to space limitations. However, this does not mean that this invention does not possess the features of these prior arts. On the contrary, this invention already possesses all the features of the prior arts, and the applicant reserves the right to add relevant prior arts in the background art. Summary of the Invention
[0014] The Top-K sparse matrix-vector multiplication (SpMV) algorithm executed under traditional CPU and GPU architectures generally includes two steps: SpMV operation and sorting. The frequent random data access and transmission during the sorting process account for more than 60% of the time delay. In addition, due to the mismatch between the regular hardware design and the sparse matrix structure, intensive memory random access occupies more than 60% of the total delay of SpMV calculation. Currently, the FPGA design solutions for Top-K SpMV still follow the traditional optimization path for SpMV operations, and mostly use 32-bit floating-point numbers for the complete SpMV operation. Although some studies have tried to reduce the bandwidth pressure through quantization techniques, in order to maintain sufficient computing accuracy, the data width usually remains at a relatively high level, such as 20 bits. This relatively high data width strategy introduces a key limitation: within a single clock cycle, a 512-bit wide memory can only process the transmission of up to 15 non-zero elements.
[0015] The current approach tends to use the coordinate (COO) or compressed sparse row (CSR) format to store sparse matrix information. The COO data format does not compress the positions of non-zero elements at all, resulting in inefficient bandwidth usage and large memory requirements. Although the CSR data format compresses the rows where non-zero elements are located, its variable-length row index array can only provide acceleration at the row level and does not fully exploit the parallel potential at the non-zero element level. If a fixed-length row index representation method is adopted, although parallel optimization can be achieved at the non-zero element level, this will lead to unnecessary waste of transmission bit width, thereby further reducing the bandwidth utilization rate.
[0016] In view of the deficiencies of the prior art, the present invention provides an accelerator based on Top-K sparse matrix-vector multiplication from the first aspect, including a first acceleration unit and a second acceleration unit. The first acceleration unit is used for data preprocessing and sparse matrix encoding, and the second acceleration unit is used for performing Top-K sparse matrix-vector multiplication to obtain the top K results with the largest values in the operation result of multiplying a sparse matrix by a dense vector; wherein, the first acceleration unit includes a preprocessing module and an encoding module; the preprocessing steps of the preprocessing module include quantization, sparse reconstruction, and shuffling; the quantization step based on the sparse matrix is used to reduce the broadband requirement for data transmission; the non-zero elements with less impact on sorting are pruned based on the sparse reconstruction algorithm to reduce redundant memory transmission and computational amount; the shuffling operation is performed on the sparse matrix after sparse reconstruction to avoid the Top-K aggregation phenomenon; the encoding module is used to re-encode the preprocessed sparse matrix.
[0017] By dividing the task into two stages: preprocessing and computing, the present invention significantly improves the overall computing efficiency. The preprocessing module optimizes the data structure through quantization, sparse reconstruction, and shuffling steps, reducing unnecessary memory transfers and computational amounts, thereby accelerating the subsequent computing process. The second acceleration unit focuses on efficiently executing the SpMV operation and can quickly locate the results of the top K maximum values, enhancing the speed and accuracy of result extraction. Compared with the baselines of the best-practice CPU, GPU, and FPGA, the present invention achieves a significant effect of increasing the end-to-end performance by 153.3 times, 2.5 times, and 3.4 times respectively.
[0018] According to a preferred embodiment, the steps of sparse matrix encoding include: re-encoding the preprocessed sparse matrix to form an Ultra-CSR data format or a Random-CSR data format, while keeping the sparse structure of the elements within each row unchanged, recombining the sparse matrix to form a new sparse matrix with a completely random row arrangement. This encoding method not only retains the internal structure of the original sparse matrix but also breaks the original pattern by randomly reorganizing the rows, reducing access conflicts and improving storage and access efficiency. The Ultra-CSR and Random-CSR data formats are particularly suitable for processing large-scale sparse matrices, capable of reducing storage space occupancy and increasing the reading speed while maintaining high precision.
[0019] According to a preferred embodiment, the second acceleration unit includes a quantization computing core, a full-precision computing core, a reader, and an arbiter; the quantization computing core is connected to the reader, the reader is connected to the arbiter, and the arbiter is connected to the full-precision computing core; the quantization computing core is used to execute the sparse matrix-vector multiplication and update the Top-16 row indexes maintained by each, forming the final Top-512 row indexes; the reader is used to read the full-precision data from the corresponding channels according to the row indexes output by the quantization computing core and pack them into a format suitable for processing by the full-precision computing core; the arbiter is used to merge the multiplexed data into one path in a polling manner and send it to the full-precision computing core; the full-precision computing core extracts the full-precision data packet and performs the SpMV calculation.
[0020] The present invention is designed in this way to achieve the full-process automated management from the preliminary screening (Top-16 row indexes) to the final precise calculation (Top-512 row indexes). The quantization computing core can quickly identify key data, the reader and the arbiter can ensure the efficient management and scheduling of the data stream, and the full-precision computing core can complete the final precise calculation, which greatly improves the response speed and computing power of the entire system.
[0021] According to a preferred embodiment, the quantization method includes: multiplying all non-zero elements in the sparse matrix proportionally by a constant and then rounding to make the numerical distribution fall within the range of [-32, 31]. This method effectively compresses the data range, reduces the storage requirement and transmission bandwidth, while maintaining sufficient precision. It enables the use of a smaller data type for representation in subsequent calculations, further accelerating the calculation speed and reducing energy consumption.
[0022] According to a preferred embodiment, the steps of sparse reconstruction include: calculating the original precision; setting the precision loss tolerance and the initial pruning ratio; in each iterative pruning process, sorting the elements in the same row for each row vector in the sparse matrix and removing elements according to the pruning ratio of the current iteration round, where the elements are those with smaller absolute values; determining whether the precision meets the requirement; in the case where the precision meets the requirement, increasing the pruning ratio and continuing to sort the elements in the same row; in the case where the precision does not meet the requirement, decreasing the pruning ratio and determining whether the ratio has been searched. In the case where the ratio has been searched, a pruned sparse matrix is formed; if it is determined that the ratio has not been searched, re-sorting the elements in the same row is performed.
[0023] The steps of sparse reconstruction significantly reduce the computational amount and memory occupation by intelligently removing elements that have less impact on the final result, while ensuring the accuracy and reliability of the result. This method is particularly applicable to application scenarios that require balancing precision and efficiency, such as large-scale data analysis and machine learning model training.
[0024] According to a preferred embodiment, the steps of sparse matrix encoding include: when the sparsity of the coefficient embedding in the sparse matrix is greater than 1%, encoding the sparse matrix from the COO data format to the Ultra-CSR data format; when randomly accessing the full-precision sparse matrix, encoding the sparse matrix from the COO data format to the Random-CSR data format.
[0025] This strategy can dynamically adapt to different types of sparse matrices, maximize the utilization of hardware resources, reduce the storage overhead and access latency. For matrices with high sparsity, Ultra-CSR provides a compact storage solution; while for scenarios with frequent random access, Random-CSR provides a more flexible access mode, improving the overall performance.
[0026] According to a preferred embodiment, the quantization calculation core includes a first decoder, a first element-wise multiplier, a first aggregator, and a Top-16 updater. The first decoder is used to parse the input data; the first element-wise multiplier is used to perform the element-wise multiplication between the sparse matrix and the vector; the first aggregator is used to summarize the multiplication results to generate partial sums; the Top-16 updater is used to update the row indices of the most important 16 non-zero elements according to the calculation results.
[0027] The combined action of the components inside the quantization calculation core ensures the effective parsing of input data, fast multiplication operations, and the efficient aggregation and update of results. In particular, the design of the Top-16 updater can track the most important non-zero elements in real time, greatly improving the efficiency of screening Top-K results.
[0028] According to a preferred embodiment, the full-precision calculation core includes a second decoder, a second element-wise multiplier, and a second aggregator. The second decoder is used to parse the data packets read and compressed by the reader; the second element-wise multiplier is used to optimize the SpMV performance by replicating the full-precision dense vector values multiple times in the URAM; the second aggregator is used to aggregate the dot product results of the elements in the same row.
[0029] The full-precision calculation core significantly improves the performance of SpMV calculations by optimizing the parsing and multiplication operations of data packets, especially by accelerating dot product calculations through multiple replications in the URAM. This enables the system to maintain high precision while still having excellent calculation speed and efficiency.
[0030] The present invention provides, from a second aspect, an acceleration method for Top-K sparse matrix-vector multiplication. The method includes: performing data preprocessing and sparse matrix encoding on the sparse matrix; performing Top-K sparse matrix-vector multiplication to obtain the top K results with the largest values in the operation result of multiplying the sparse matrix by the dense vector; wherein, the preprocessing steps of the preprocessing module include quantization, sparse reconstruction, and shuffling; reducing the broadband requirement for data transmission based on the quantization step of the sparse matrix; pruning non-zero elements that have less impact on sorting based on the sparse reconstruction algorithm to reduce redundant memory transmission and computational amount; performing a shuffling operation on the sparse matrix after sparse reconstruction to avoid the Top-K aggregation phenomenon; and re-encoding the preprocessed sparse matrix.
[0031] The method of the present invention optimizes the data structure of the sparse matrix through steps such as quantization, sparse reconstruction, and shuffling, reducing unnecessary memory transmission and computational amount. Quantization reduces the data transmission bandwidth requirement, reducing the data volume while maintaining a certain precision. Sparse reconstruction reduces redundant calculations by pruning non-critical elements and avoids the aggregation phenomenon of Top-K results through the shuffling operation, improving the uniformity of result distribution.
[0032] According to a preferred embodiment, the steps of sparse reconstruction include: calculating the original precision; setting the precision loss tolerance and the initial pruning ratio; in each iterative pruning process, sorting the elements in the same row for each row vector in the sparse matrix, and removing elements according to the pruning ratio of the current iteration round, where the elements are those with smaller absolute values; determining whether the precision meets the requirement; in the case where the precision meets the requirement, increasing the pruning ratio and continuing to sort the elements in the same row; in the case where the precision does not meet the requirement, decreasing the pruning ratio and determining whether the ratio has been searched. In the case where it has been searched, a pruned sparse matrix is formed; if it is determined that the ratio has not been searched, sorting the elements in the same row again.
[0033] Different from the prior art, the present invention adopts the basic principle method of approximate calculation and introduces multiple optimization techniques such as quantization, unstructured pruning, multi-core approximation, expanding the approximation range, and fusing the SpMV and Sort operators. The present invention introduces a preprocessing method for sparse matrices, which combines quantization, pruning, and Shuffle operations to optimize the data processing flow. This method aims to effectively improve the storage and transmission efficiency of sparse matrices, thereby reducing power consumption and increasing the inference speed; the present invention adopts a two-stage Top-K SpMV algorithm, which not only has high scalability but also proposes Ultra-CSR and Random-CSR data formats for streaming and randomly accessed sparse matrices, further improving the data transmission and access efficiency; in terms of CPU implementation, the present invention uses the MKL library and heap sorting algorithm for efficient single-core and multi-core calculations. Especially in a multi-core environment, through physical core binding technology, the resource utilization rate is maximized, and the calculation speed is significantly increased; in GPU implementation, the present invention combines INT6 quantization and ReSparse pruning technologies to greatly reduce the data transmission volume and calculation burden, improving the throughput and calculation performance. Compared with traditional methods, the present invention shows more significant end-to-end performance advantages when processing large-scale sparse matrices, especially in data scenarios containing millions to tens of millions of rows. This makes the present invention more practical in actual applications and can effectively meet the requirements of low power consumption and high inference speed. Brief Description of the Drawings
[0034] Figure 1 is a schematic diagram of the simplified module connection relationship of the accelerator provided by the present invention;
[0035] Figure 2 is a schematic flowchart of the pruning method for low-influence elements of the sparse matrix provided by the present invention;
[0036] Figure 3 is a schematic diagram of the Ultra-CSR sparse matrix data format provided by the present invention;
[0037] Figure 4It is a schematic diagram of the Random-CSR sparse matrix data format provided by the present invention;
[0038] Figure 5 It is a schematic diagram of the result of the quantization calculation core Q-Core of the accelerator provided by the present invention.
[0039] Figure 6 It is a logical schematic diagram of the reader of the accelerator provided by the present invention.
[0040] Figure 7 It is a logical schematic diagram of the arbiter of the accelerator provided by the present invention.
[0041] Figure 8 It is a schematic diagram of the structure of the full-precision calculation core F-Core of the accelerator provided by the present invention.
[0042] List of reference numerals
[0043] 100: First acceleration unit; 110: Preprocessing module; 120: Encoding module; 130: DRAM; 200: Second acceleration unit; 210: High-bandwidth memory; 220: Quantization calculation core; 221: First decoder; 222: First element-wise multiplier; 223: First aggregator; 224: Top-16 updater; 230: Full-precision calculation core; 231: Second decoder; 232: Second element-wise multiplier; 233: Second aggregator; 240: Reader; 250: Arbiter. Detailed implementation manners
[0044] The following is a detailed description with reference to the accompanying drawings.
[0045] Currently, the FPGA architecture design for Top-K SpMV generally continues the optimization method of the traditional SpMV operator and generally uses the Float-32 data format for the complete SpMV calculation. Although some research has introduced quantization techniques to relieve the pressure on the transmission bandwidth, in order to ensure the calculation accuracy, the data bit width still remains at a relatively high level, such as 20-bit. This strategy of high data bit width brings significant problems: within one clock cycle, the transmission bandwidth of the 512-bit high-bandwidth memory 210 can only support the transmission of at most 15 non-zero elements.
[0046] Existing storage technologies mainly rely on the Coordinate (COO) or Compressed Sparse Row (CSR) format to represent sparse matrices. The COO data format does not compress the coordinates of non-zero elements, resulting in low utilization of transmission bandwidth and increased memory occupancy. Although the CSR data format compresses the row indices of non-zero elements, its variable-length row index array is only suitable for row-level acceleration and fails to fully utilize the parallel potential at the non-zero element level. Although using a fixed-length row index representation can achieve parallel optimization at the non-zero element level, it wastes transmission bandwidth and further exacerbates the problem of low bandwidth utilization.
[0047] To address the deficiencies of the prior art, the present invention provides an accelerator based on Top-K sparse matrix-vector multiplication and an acceleration method thereof. The present invention can also provide an electronic device equipped with the accelerator of the present invention. The present invention can also provide a processor capable of running the encoded program of the acceleration method of the present invention. The present invention can also provide a storage medium for storing the encoded program of the acceleration method of the present invention.
[0048] Embodiment 1
[0049] This embodiment provides an accelerator based on Top-K sparse matrix-vector multiplication. The accelerator can run the acceleration method of the present invention. Preferably, the Top-K sparse matrix-vector multiplication (which can also be written as sparse matrix-vector multiplication) (SpMV) in the present invention refers to the process of obtaining the top K results with the largest values in the operation result of multiplying a sparse matrix by a dense vector.
[0050] The accelerator based on Top-K sparse matrix-vector multiplication of the present invention includes a first acceleration unit 100 and a second acceleration unit 200. The first acceleration unit 100 and the second acceleration unit 200 form a heterogeneous operating environment. The first acceleration unit 100 is a CPU, and the second acceleration unit 200 can be one of physical hardware such as a central processing unit (CPU), a field programmable gate array (FPGA), a graphics processing unit (GPU), an application specific integrated circuit (ASIC), a neural network processor (NPU), in-memory computing, etc.
[0051] Preferably, Figure 1 One of the hardware combination modes of the accelerator based on Top-K sparse matrix-vector multiplication of the present invention is shown to illustrate the acceleration method of the present invention. Figure 1The accelerator shown has a first acceleration unit 100 as a central processing unit (CPU); the second acceleration unit 200 is a field-programmable gate array (FPGA). There are various ways to connect the central processing unit (CPU) and the field-programmable gate array (FPGA), such as connecting through a PCI Express (PCIe) bus. Specifically, the FPGA is usually inserted into the motherboard in the form of a PCIe interface card, and the CPU communicates with it through the PCIe slot on the motherboard. For another example, the CPU and the FPGA can be connected through a dedicated high-speed parallel or serial interface. These interfaces can be optimized according to the requirements of specific applications to provide a very high data transfer rate. For another example, the CPU and the FPGA are integrated to form a system-on-chip (SoC). In this architecture, the CPU and the FPGA are on the same chip, sharing memory resources and can communicate directly through an internal bus. For another example, in an FPGA containing an embedded CPU, the AXI bus can be used as the main communication channel between the CPU and the FPGA logic block. For another example, the CPU and the FPGA can also communicate with each other through Ethernet or other network protocols. This method is suitable for application scenarios that require remote access or cooperation across multiple nodes. For another example, if the CPU and the FPGA can access a common memory area, they can directly exchange data within this area. This can be achieved through physically shared RAM or through memory mapped to the address spaces of both parties. The method of shared memory can simplify the data transfer process, but requires an appropriate synchronization mechanism to ensure data consistency.
[0052] Preferably, during the execution of a computing task, the CPU is used for operations such as data preprocessing and sparse matrix encoding. Preferably, as Figure 1 shown, the CPU is configured with a preprocessing module 110 and an encoding module 120. The preprocessing module 110 is used to execute preprocessing steps. Preferably, the preprocessing steps include quantization, sparse reconstruction (ReSparse), and shuffling steps. The encoding module 120 is used to encode the sparse matrix. The encoding module 120 can implement two efficient sparse matrix encoding methods. The preprocessing module 110 and the encoding module 120 significantly optimize the data representation form, providing strong support for the subsequent computing process of the FPGA: on the one hand, effectively reducing the redundant computation amount, and on the other hand, significantly improving the memory transfer efficiency, thereby ensuring that the computing performance of the entire system is maximally improved. After the FPGA finishes the calculation of Top-K SpMV, the CPU reads back the calculation result through the PCIe bus and displays it to the user.
[0053] Preferably, a CPU memory, that is, DRAM 130, is also set inside the CPU. A high-bandwidth memory 210 (HBM) is also set at the connection end of the FPGA.
[0054] Preferably, as the core computing unit of the entire accelerator system, the FPGA is used to execute the core computing task of Top-K sparse matrix-vector multiplication (SpMV) under the unified scheduling of the CPU. In the FPGA architecture design, through the high customization of programmable gate circuits, a multi-core parallel computing unit is integrated, specifically including 32 quantization computing cores 220 (Q-Core) and 1 full-precision computing core 230 (F-Core). To optimize the computing data stream and achieve efficient cooperation among multiple cores, the FPGA of the present invention improves and provides a new data path design: 32 readers 240 and 1 arbiter 250 are configured. Preferably, each quantization computing core 220 (Q-Core) is configured with a reader 240.
[0055] Preferably, as Figure 1 shown, the reader 240 is arranged between the quantization computing core 220 (Q-Core) and the arbiter 250. The reader 240 is used to efficiently obtain full-precision data from the memory according to the Top-K vector index calculated by the quantization computing core 220 (Q-Core). The arbiter 250 is connected to the full-precision computing core 230 (F-Core) and provides input data for the full-precision computing core 230 (F-Core). Preferably, the arbiter 250 is used to realize the intelligent scheduling and integration of multiple data streams. The arbiter 250 merges the parallelly read data streams into a single data stream as the input of the full-precision computing core 230 (F-Core).
[0056] Further preferably, the CPU and the FPGA achieve efficient transmission of data and control signals through a PCIe (Peripheral Component Interconnect Express) interface. Specifically, data transmission mainly includes two directions: one is to transmit the encoded sparse matrix from the CPU to the high-bandwidth memory 210 (HBM) on the FPGA; the other is to transmit the processing results of the full-precision computing core 230 (F-Core) back to the CPU for further analysis or display by the user. Specifically, the CPU is connected to the high-bandwidth memory 210 through a PCIe interface. The high-bandwidth memory 210 is directly connected to the FPGA through the AXI (Advanced eXtensible Interface) communication protocol, forming a tight storage-computation coupling architecture. The high-bandwidth memory 210 provides the FPGA with large-capacity and high-bandwidth storage support, significantly improving the read / write efficiency of sparse matrices and vector data. Inside the FPGA, the quantization computing core 220 (Q-Core) and the full-precision computing core 230 (F-Core) can directly access the data in the high-bandwidth memory 210, effectively reducing data transmission latency and avoiding bandwidth bottlenecks. In addition, the upstream and downstream modules inside the FPGA achieve fully pipelined data stream processing through optimized wiring design, further improving the overall computing performance.
[0057] Preferably, the quantization computing core 220 (Q-Core) is used to process the SpMV calculation of partial rows of the sparse matrix and the vector, and select the Top-16 row indices therefrom. Each quantization computing core 220 (Q-Core) is equipped with a dedicated quantization unit and a Top-K selection unit, which can efficiently perform the SpMV calculation on the quantized data and dynamically select the non-zero elements most likely to enter the Top-K. 32 quantization computing cores 220 (Q-Core) work in parallel, significantly improving the throughput and efficiency of the SpMV calculation.
[0058] The full-precision computing core 230 (F-Core) is used to process tasks that require high-precision floating-point operations. The full-precision computing core 230 (F-Core) plays an auxiliary role in the Top-K SpMV calculation, and is used to perform further precise calculations on the Top-16 row indices selected by the quantization computing core 220 (Q-Core) to ensure the accuracy of the final result.
[0059] Further preferably, as Figure 1As shown, the correct transmission of the data stream between the reader 240 and the quantization calculation core 220 (Q-Core) is achieved through the wiring resources inside the FPGA. The reader 240 reads the full-precision data from the high-bandwidth memory 210 according to the Top-K vector index of the quantization calculation core 220 (Q-Core) and transfers it to the arbiter 250. The arbiter 250 arbitrates the data from 32 readers 240 in a polling manner, merges them into one path, and then transfers them to the full-precision calculation core 230 (F-Core). After the full-precision calculation core 230 (F-Core) completes the full-precision calculation, it outputs the final Top-K result.
[0060] Preferably, after the preprocessing module 110 in the CPU receives the sparse matrix, the operations on the sparse matrix are as follows.
[0061] S1: Execute the data preprocessing steps, that is, perform the preprocessing operations of quantizing, sparse reconstruction (ReSparse), and shuffling on the sparse matrix.
[0062] S11: Quantize the sparse matrix into integers with a 6-bit width, thereby significantly reducing the bandwidth requirements for data transmission.
[0063] In the current FPGA architecture design practice of the Top-K sparse matrix vector multiplication (SpMV), most solutions still rely on traditional SpMV operator optimization strategies and usually use 32-bit floating-point numbers (Float-32) for all calculations. Although there are technical attempts to reduce the pressure on the transmission bandwidth by introducing quantization methods, due to the need to maintain calculation accuracy, the data width still remains at a relatively high level, such as 20 bits. However, this high-data-width design method leads to a significant problem: even with a high-speed memory with a width of 512 bits, at most 15 non-zero element data transmissions can be processed in each clock cycle, which severely restricts the calculation efficiency and resource utilization.
[0064] To solve the above bottleneck, the present invention proposes an innovative method - to achieve the Top-K vector index statistics based on integer quantization by expanding the approximation interval. The core of this method is to use a 6-bit integer format for quantization, thereby greatly reducing the representation of non-zero elements from the traditional 32 bits to only 6 bits. This improvement not only greatly reduces the bandwidth required for data transmission but also improves the resource utilization rate because smaller data units can support more parallel processing tasks under the same hardware resources.
[0065] Preferably, the quantization method is: multiply all non-zero elements in the sparse matrix by a constant in equal proportion, and then take the integer so that the numerical distribution falls within the interval [-32, 31].
[0066] The quantization formula is expressed as:
[0067] Q A = clip(round(S × A), -2 B-1 , 2 B-1 -1).
[0068] Among them, A represents the original sparse matrix; A max represents the largest element in the sparse matrix, and A min represents the smallest element in the sparse matrix. B represents the number of bits used for quantization, S represents the quantization scale. After the original sparse matrix A is multiplied by the quantization scale S, it is truncated by round() and clip(), and scaled to the interval [-31, 32]. Q A represents the quantized sparse matrix.
[0069] This quantization step brings two significant advantages:
[0070] First, at the transmission level of the high-bandwidth memory 210, the data channel with a width of 512 bits can be improved from originally only transmitting 15 non-zero elements to transmitting 24, which means the bandwidth utilization rate is increased by 60%. At the same time, the computing parallelism of the quantization computing core 220 (Q-Core) is also increased by 60% accordingly. This not only optimizes the data processing speed but also significantly enhances the overall computing efficiency.
[0071] Second, integer operations consume fewer hardware resources than floating-point operations. Therefore, while the computing parallelism is increased, the consumption of hardware resources can actually be reduced instead of resulting in a decrease in resource utilization. This effective utilization of resources creates conditions for the improvement of the accelerator clock frequency, thereby further enhancing the overall computing performance of the accelerator.
[0072] Preferably, perform the SpMV calculation for all rows on the sparse matrix after quantization and sparse reconstruction, and count the Top-Ks values during this process, where Ks > K. For example, when K = 100 is set, the Ks value obtained by statistics can reach 512. This method not only improves the computing accuracy but also ensures the efficiency and reliability when processing large-scale data sets.
[0073] Here, the quantized and sparsely reconstructed sparse matrix is used to reduce the bandwidth pressure for accessing the sparse matrix; the total number of memory accesses for accessing the sparse matrix is reduced by reducing the number of bits occupied by each non-zero element in the sparse matrix and the number of non-zero elements in the sparse matrix.
[0074] First, on multiple matrix datasets with different sparsities and scales, calculate the exact Top-K SpMV results respectively. These datasets cover a wide range of sparse matrix instances to ensure the general applicability of the results. Then, use the INT6 quantization method to preprocess these matrices. During the preprocessing, not only the data bitwidth of each non-zero element is reduced, but also the quantized matrix is ensured to maintain sufficient precision for effective Top-K calculation.
[0075] Then, count the approximate results of Top-K under different K values. By analyzing the results under different K values, when Ks is set to 512, it can provide result quality comparable to the exact Top-K calculation while ensuring computational efficiency. Especially in the case of K = 100, this setting can effectively capture the most important non-zero elements, thus ensuring the accuracy and reliability of the final result.
[0076] Finally, calculate the error between the approximate result and the exact result under different K values, and use recall as the evaluation metric:
[0077]
[0078] Among them, TopKs {approximate} represents the set of Top-K vector indices obtained through approximate calculation (such as Top-K SpMV based on Ks = 512). TopK {true} represents the set of Top-K vector indices obtained through exact calculation. TopKs {apprbximate} ∩TopK {true} represents the number of indices commonly included in the approximate result and the exact result. K is the target Top-K value (e.g., K = 100). Through experiments, it is found that when Ks = 512, the accuracy and resource consumption of the approximate result reach the best balance. Specifically, the approximate calculation accuracy of multiple datasets is 100%, and the FPGA resource utilization rate remains within a reasonable range.
[0079] S12: Prune the non-zero elements that have less impact on sorting based on the ReSparse algorithm, thereby effectively reducing redundant memory transfers and computational amounts.
[0080] Existing Top-K Sparse Matrix-Vector Multiplication (SpMV) accelerators and their acceleration methods usually do not introduce the concept of pruning based on the granularity of non-zero elements to reduce unnecessary computational effort. In the absence of such a fine-grained pruning mechanism, there are often significant differences in the absolute values of non-zero elements in the sparse matrix. Specifically, those non-zero elements with relatively small absolute values have extremely limited impact on the final sorting result and can almost be ignored. Therefore, sending these low-impact non-zero elements to the FPGA for calculation not only wastes valuable computing resources but may also reduce the overall computing efficiency.
[0081] In view of this limitation, the present invention proposes a new algorithm for ReSparse. This algorithm can intelligently identify and zero out those non-zero elements that contribute less to the sorting process. In this way, the ReSparse algorithm effectively reduces the total number of non-zero elements that need to be transmitted and processed, thereby optimizing the data transmission efficiency and improving the computing performance.
[0082] In addition, the ReSparse algorithm not only reduces the redundant computational effort but also ensures that the key information is retained, thus guaranteeing the accuracy and reliability of the final sorting result. This method provides new ideas and technical means for efficiently processing large-scale sparse matrices, especially suitable for resource-constrained hardware platforms such as FPGAs.
[0083] The specific operation steps of the ReSparse algorithm are as Figure 2 shown.
[0084] S121: Calculate the original precision.
[0085] Directly calculate the Top-K SpMV on the original sparse matrix. Since the value of K varies in different scenarios, to improve the robustness of the present invention, preferably, K is taken as 1, 8, 16, 32, 50, 75, and 100, and the calculation results are used as the baseline.
[0086] Preferably, since multiple K values are used to measure the precision of the Top-K SpMV algorithm, when verifying the precision loss, a unified measurement index key_metric is required to determine the search convergence, that is, calculate the original precision. This measurement index is the accumulation of multiple Top-K precisions.
[0087] The calculation formula for the measurement index is:
[0088] K = [1, 8, 16, 32, 50, 75, 100].
[0089] Among them, key_metric represents a measurement index for determining search convergence; len(K) represents the number of values of K, which is 7 in this example; precision i represents the precision when K takes the value of K i .
[0090] The present invention provides an optimization algorithm for original precision. This optimization algorithm is used to ensure the fairness of calculation by presetting different weights under different importance indexes (Top-K). The core of this optimization algorithm is to optimize the storage and calculation efficiency of the sparse matrix by adjusting the sparsification ratio of the sparse matrix.
[0091] Specifically, this optimization algorithm uses the bisection method to adjust the current sparsification ratio in each iteration cycle. When initializing the iteration, if the sparsification ratio is set to 1.0, that is, no sparsification is performed. At the same time, the optimization algorithm will determine a precision threshold according to the accuracy loss tolerance (accuracy_threshold) set by the user to ensure the accuracy of the calculation result.
[0092] In subsequent iterations, this optimization algorithm will calculate the matching ratio (prec) between the actual importance index (Top-K) and the approximate calculation result (Top-512), and evaluate the precision of the optimization algorithm in combination with the preset weight. If the current sparsification ratio has been tested and the measurement index (key_metric) of the previous iteration reaches the preset key performance threshold (key_threshold), it is considered that the optimization algorithm has converged. At this time, the optimization algorithm will terminate and return the sparsification ratio (keep_percent) of the previous iteration as the optimal sparsification ratio.
[0093] Through this optimization process, the optimization algorithm can effectively determine the best sparsification ratio on the premise of meeting the precision requirements, thereby improving the performance of sparse matrix multiplication operations.
[0094] Assume that the true calculation result of the i-th row is y i = A·x. After pruning using the sparse reconstruction method, the calculation result becomes:[[]]
[0095] Among them, is the sparsified row vector, which retains kp × nnz(A i ) non-zero elements. kp represents the retention ratio of non-zero elements retained during the pruning process. nnz(A i ) represents the matrix A iNumber of NonZeros. This technical term comes from the requirements of sparse matrix storage formats. Since most elements in a sparse matrix are zeros, storing only non-zero elements can significantly save storage space and improve computational efficiency.
[0096] Assume that among the elements set to zero, the maximum absolute value is M, and the maximum absolute value element in vector x is ||x|| ∞ , then in the worst case, the computational error of the i-th row is:
[0097]
[0098] where δ i represents the computational error of the i-th row; and represents the set of indices of non-zero elements retained, kp represents the retention ratio of non-zero elements retained during the pruning process. A ij represents the non-zero element in the i-th row and j-th column. S represents the set containing all non-zero element indices in matrix A.
[0099] To ensure the accuracy of the sorting result, it is necessary to ensure that: y K -δ i ≥y K+1 .
[0100] where y K represents the minimum value in the Top-K rows, and y K+1 represents the maximum value in the non-Top-K rows.
[0101] By adjusting the non-zero element retention ratio kp, the sorting error brought by the sparse reconstruction method can be controlled to ensure that the error meets the application requirements. On 5 public datasets and 6 synthetic large-scale coefficient matrices, with an average compression rate of 35.7%, the sparse reconstruction method causes at most a 0.87% accuracy loss.
[0102] S122: Set the accuracy loss tolerance and set the initial pruning ratio.
[0103] The accuracy loss tolerance accThresh represents the lowest value of the ratio of the test accuracy of the sparse matrix after sparse reconstruction to the original accuracy. Using the accuracy loss tolerance, the lower limit key_thresh of the threshold of the unified measurement index key_metric can be calculated, which represents the lowest computational accuracy that needs to be ensured after pruning.
[0104] key_thresh = key_metric * accThresh.
[0105] Preferably, the precision loss tolerance is set to 99.9%. It should be noted that for different sparse matrices, the parameter value of the precision loss tolerance will be different, and this parameter value can float in the range of 90% - 99.99%. 99.9% is a relatively universal value obtained after testing on multiple data sets.
[0106] Based on the search principle of the binary search algorithm, the initial non-zero element retention ratio kp should be set to 0.5.
[0107] S123: According to the set precision loss tolerance and retention ratio, search for the highest pruning ratio at this precision loss tolerance based on the binary search algorithm.
[0108] In the process of optimizing the sparse matrix processing, a binary search algorithm is used to determine the optimal non-zero element retention ratio. This process involves multiple rounds of pruning iterations, and in each round of iteration, the sparse matrix is pruned based on the currently estimated retention ratio. Specifically, in each iteration, first, approximate calculations are performed using the sparse matrix pruned in the current round, and the results are compared with the preset minimum precision requirement. If the matrix at the current pruning level fails to meet the precision standard, the algorithm will accordingly adjust the search range; conversely, if the precision is met, the search interval for the retention ratio is further refined.
[0109] The iteration process continues until two key convergence conditions are met: First, the value of the retention ratio needs to be accurate to one decimal place and no longer change; second, the sparse matrix after pruning must be able to reach or exceed the minimum precision requirement. Once these two conditions are simultaneously met, the algorithm terminates the iteration and outputs the finally pruned sparse matrix as the processing result.
[0110] Through this iterative and step-by-step approximation method, not only the computational efficiency is improved, but also the accuracy and reliability of the results are ensured. This method is particularly effective for resource-constrained hardware platforms because it can minimize unnecessary computational volume while ensuring computational precision, thereby improving the overall performance and response speed.
[0111] As Figure 3 shown, in each iteration pruning process, the elements in each row vector of the sparse matrix are sorted in the same row, and elements are removed according to the pruning ratio of the current iteration round. Here, the elements are those with smaller absolute values.
[0112] As Figure 2 shown, when the pruning is completed, the Top-K SpMV is recalculated using the pruned sparse matrix, and a precision test is performed to obtain the key_metric, the same measure of precision for the current round i , that is, to determine whether the precision meets the requirements. If key_metrici > key_thresh, that is, if the judgment result of accuracy compliance is yes, it indicates that at the current pruning ratio, the accuracy still remains at a relatively high level. At this time, the non-zero element retention ratio kp is reduced through the binary search algorithm, that is, the pruning ratio is increased, and the sorting of elements in the same row continues. If key_metric i < key_thresh, that is, if the judgment result of accuracy compliance is no, the non-zero element retention ratio kp is increased, that is, the pruning ratio is reduced, and the lower limit of the search interval is updated to the current pruning ratio. If this ratio has been searched, this value is the maximum pruning ratio that can be achieved under the condition of accuracy compliance. In each round of search, the elements of each row of the sparse matrix need to be sorted to determine which elements are eliminated.
[0113] Finally, the highest pruning ratio that can be achieved without exceeding the accuracy loss tolerance is obtained, so as to maximize the pruning effect while ensuring the performance of the sparse matrix, and improve the efficiency and performance of the model.
[0114] S13: To avoid the Top-K aggregation phenomenon, a shuffling operation is performed on the sparse matrix after sparse reconstruction to ensure the uniformity of data distribution.
[0115] Shuffle refers to the process of randomly rearranging a group of elements. In computer programs, the shuffle algorithm is used to generate a random permutation of a sequence so that each element has an equal probability of appearing in any position.
[0116] Specifically, each row of the sparse matrix is regarded as an independent unit, and by randomly shuffling its row order, a new sparse matrix with a completely random row permutation is recombined, while keeping the sparse structure of the elements within each row unchanged.
[0117] S2: Encode the preprocessed sparse matrix into Ultra-CSR data format and Random-CSR data format.
[0118] When dealing with large-scale sparse matrices, due to the huge amount of data and uneven distribution, accessing memory (memory access) has become a significant bottleneck. Therefore, designing an efficient sparse matrix storage layout is particularly important for optimizing data storage and access efficiency. Currently, there are mainly two commonly used sparse matrix storage formats: Coordinate Format (COO) and Compressed Sparse Row (CSR).
[0119] Although the COO data format is very suitable for streaming reading and parallel computing because it allows direct sequential reading of non-zero elements and their corresponding row and column indices, it does not perform any form of compression on non-zero elements. For example, in the case of a 512-bit width, it can only carry data for 5 non-zero elements. In contrast, the CSR data format improves storage efficiency by compressing the column indices of each row. However, it uses variable-length arrays to represent row indices, which introduces data dependencies and makes the architecture sensitive to the distribution of non-zero elements at the row level, thus losing some acceleration opportunities.
[0120] To meet the requirements of streaming parallel computing, the BS-CSR (Block Sparse CSR) format emerged. This format uses a fixed-length representation for compressed row indices to reduce data dependencies and simplify the access pattern. However, this approach may lead to redundant zero-padding, resulting in a waste of bit width.
[0121] Therefore, although data formats such as COO, CSR, and BS-CSR have their respective advantages, they also have their limitations. To overcome these limitations, the present invention proposes a new encoding method, aiming to balance the relationship between storage density and computational efficiency while minimizing performance losses caused by data structures. By continuously improving the storage format of sparse matrices, the overall performance in large-scale data processing tasks can be further enhanced.
[0122] S21: Construct the Ultra-CSR data format.
[0123] Figure 3 In, x and y respectively represent the x and y coordinates of the non-zero element. val represents the value of the non-zero element.
[0124] As Figure 3 shown, the sparsity of Sparse Embeddings is usually greater than 1%, generally between 2% and 30%. Take a sparse matrix with 500 columns and a sparsity of 2% as an example. In this case, there are approximately 10 non-zero elements in each row. When using the BS-CSR data format, 10 bits are used to represent the y coordinate and 32 bits are used to represent the non-zero element value val. P (the maximum number of non-zero elements that can be carried by a 512-bit bandwidth) needs to meet the following conditions:
[0125] That is, P = 11, and each ptr occupies 4 bits. ptr refers to a single pointer value in the storage format of the sparse matrix.
[0126] Compared with the COO data format, the BS-CSR (Block Sparse CSR) data format significantly improves the ability of the high-bandwidth memory 210 to transfer non-zero elements, achieving a 2.2-fold improvement and being able to transfer up to 11 non-zero elements at a time. However, among these 11 non-zero elements, there are usually only two line breaks, which means that there are only two valid elements in the ptr array, and the remaining nine positions are used to fill with zero values, resulting in a waste of 36-bit bitwidth. Here, the ptr array is an array composed of multiple Ptr values and is used to store the starting position information of non-zero elements in each row or each block.
[0127] When directly encoding the quantization design into the BS-CSR data format, this problem of bitwidth waste caused by zero padding becomes more serious. If 6 bits are used to represent the non-zero element value Qval, the parameter P needs to satisfy:
[0128] Calculated according to this formula, P = 24, that is, each ptr occupies 5 bits. In this case, for 24 non-zero elements, there are usually three line breaks, so there are only three valid elements in the ptr array, which results in 105 bits for zero padding, plus 8 bits of unused space, wasting a total of 113-bit bitwidth, accounting for approximately 22.1% of the total bitwidth of the high-bandwidth memory 210.
[0129] To more effectively utilize this part of the wasted bitwidth and thus alleviate the memory access bottleneck, the present invention proposes an Ultra-CSR encoding method. This method aims to solve the problem of bitwidth waste in the BS-CSR data format. The Ultra-CSR data format uses 10 bits to represent the y coordinate, 6 bits to represent the non-zero element value Qval, and uses 1 bit for each non-zero element to indicate whether it has a line break relative to the previous non-zero element. Such a design scheme not only retains the regularity of the BS-CSR data format at the non-zero element level, facilitating parallel processing, but also enables the fixed-width high-bandwidth memory 210 to transfer up to 30 non-zero elements simultaneously. Compared with the BS-CSR design, the bandwidth utilization rate is doubled. The Ultra-CSR data format is an optimization technology for sparse matrix storage, and it improves the storage efficiency and computational performance through a series of specific steps.
[0130] First, each non-zero element is assigned a 1-bit identifier to indicate whether the element marks the start of a new row. This design helps to maintain the row-major arrangement order of the data and ensures the regularity of the data.
[0131] Secondly, this identification method effectively reduces storage redundancy. The extremely efficient 1-bit compressed row index occupies less bitwidth. Since this identification method eliminates the extra row index information and simplifies the calculation process of the row index, the starting position of each row can be directly determined through the identifier.
[0132] In addition, the Ultra-CSR data format is conducive to hardware acceleration and parallel processing because it allows for quick positioning to the start of each row, enabling parallel processing of multiple rows of data.
[0133] The Ultra-CSR data format also uses a 10-bit bitwidth to identify the column coordinates where non-zero elements are located. Such a bitwidth selection not only takes into account the sparsity of the sparse matrix but also meets the requirements of medium-scale matrices commonly found in practical applications.
[0134] Finally, to reduce the burden of data transmission, the numerical values of non-zero elements are represented in the INT6 format with a 6-bit bitwidth, achieving a balance between precision and bandwidth through quantization techniques.
[0135] Generally speaking, through these carefully designed steps, the Ultra-CSR data format optimizes the storage and computational efficiency of the sparse matrix, making it more suitable for hardware acceleration and parallel processing in modern computing environments.
[0136] S22: Construct the Random-CSR data format.
[0137] Figure 4 In it, x and y represent the x and y coordinates of non-zero elements respectively. Val represents the value of non-zero elements.
[0138] As Figure 4 shown, different from the streaming access pattern of the quantized matrix in the first stage, the access pattern of the full-precision sparse matrix in the second stage is row-random. This change brings two major challenges to the layout method of the sparse matrix.
[0139] Firstly, in the second stage, data needs to be read in parallel from 32 channels, and the arrival of this data is random and does not follow the expected order. Therefore, in the encoding stage, a synchronization mechanism needs to be introduced to ensure that the read data packets can be correctly decompressed. Secondly, although the element bitwidth of the full-precision sparse matrix is relatively high, and a data format similar to the COO data format can synchronize the row index information, the limited memory capacity of the high-bandwidth memory 210 still requires a data format with a high compression ratio to ensure that hundreds of millions of non-zero elements can be accommodated. For this reason, the present invention designs a storage format suitable for random access, the Random-CSR data format, for the full-precision sparse matrix.
[0140] Preferably, the Random-CSR data format adopts a page table-like design to address the random access requirements of the second-phase full-precision sparse matrix and make full use of the high parallelism and bandwidth of the high-bandwidth memory 210.
[0141] The Random-CSR data format contains two main data structures: rowPtr and package. Among them, the rowPtr structure data is responsible for storing the metadata information of each row, acting as a page table.
[0142] Specifically, the rowPtr table, i.e., rowPtr i stores the packet number start of the first non-zero element in the i-th row i , and the number of packets spanned by this row, length i . These information enable the system to quickly locate the starting position and span range of the non-zero elements in this row by reading the fields in the rowPtr table when randomly accessing the full-precision matrix.
[0143] On the other hand, the design of the package structure data draws on the Ultra-CSR data format, dividing the full-precision sparse matrix into multiple packets. Each packet contains 10 non-zero elements, and the storage method of these non-zero elements retains the efficient encoding strategy of Ultra-CSR. Specifically, a 1-bit new_row_flag is used in the package structure data to identify whether the current non-zero element changes rows compared to the previous non-zero element, maintaining the row-major order. For the column coordinates of non-zero elements, a 10-bit wide y coordinate is used to compress the data volume.
[0144] Different from the Ultra-CSR data format, the non-zero element values in the package structure data are stored using 32-bit to store FP32-precision numerical values, ensuring the accuracy of full-precision calculations. In the specific data access process, the accelerator first reads the corresponding rowPtr from the rowPtr table according to the row index i i , and then, based on start i and length i fields, locates and reads the corresponding multiple packets, packagestart i , from the package structure data. The non-zero elements contained in the packets are parsed and input into the corresponding modules to ensure efficient processing of the random access requirements of the full-precision sparse matrix.
[0145] This design not only improves the storage and access efficiency of the system for large-scale sparse matrices but also provides strong support for parallel computing in a high-bandwidth environment.
[0146] S3: Compute the approximate Top-K SpMV based on heterogeneous multi-cores.
[0147] The computing cores include soft cores and hard cores. A soft core refers to a processor core implemented using the programmable logic resources within an FPGA. These cores are not pre-fabricated physical entities but functional modules defined by HDL (Hardware Description Language, such as Verilog or VHDL) code. A hard core is a fixed-function processor directly embedded in the FPGA chip. They have a fixed physical structure, similar to an ASIC (Application-Specific Integrated Circuit), and usually offer higher performance and lower power consumption. A solid core is a form between a soft core and a hard core. It is usually an IP core submitted to the user in the form of a netlist file, which has been synthesized but can still be customized to a certain extent. The advantage of a solid core is that it is more difficult to be pirated than a soft core while retaining a certain degree of flexibility.
[0148] Preferably, as Figure 1 shown, in the FPGA of the present invention, 32 quantization computing cores 220 (Q-Cores), 1 full-precision computing core 230 (F-Core), 32 readers 240, and 1 arbiter 250 are soft cores, and they together constitute an approximate Top-K SpMV accelerator based on heterogeneous multi-cores.
[0149] S31: Perform the quantized approximate Top-K SpMV operation based on multiple quantization computing cores 220 (Q-Cores).
[0150] In the present invention, 32 quantization computing cores 220 (Q-Cores) serve as the first computing engine dedicated to processing quantized sparse matrix data, responsible for parallelly executing the preliminary sparse matrix-vector multiplication (SpMV) calculation and updating the Top-16 row indices they each maintain to form the final Top-512 row indices, as Figure 5 shown.
[0151] Preferably, four main modules are integrated inside the first engine: the first decoder 221, the first element-wise multiplier 222, the first aggregator 223, and the Top-16 updater 224.
[0152] The first decoder 221 is used to parse the input data for data preprocessing. The first decoder 221 decompresses the received sparse matrix data packet in Ultra-CSR data format and stores the decompressed data in the local memory for subsequent computing operations. The Ultra-CSR data format reintroduces data dependencies during the decoding phase. To overcome this challenge, a tree-structured adder is used to optimize the long dependencies, reducing the accumulated clock delay to log2n, thereby enabling the simultaneous parsing of the x coordinates of 30 non-zero elements within a single clock cycle.
[0153] The first element-wise multiplier 222 is used to perform an element-wise multiplication operation between a sparse matrix and a vector. The first element-wise multiplier 222 performs an element-wise multiplication operation between the dense vector Qx and the non-zeros extracted from the data packet. In order to perform random access of the highly reused dense vector Qx within a single clock cycle, it is pre-loaded into UltraRAM (Ultra Rapid RAM, URAM). Each data packet contains 30 non-zero elements. However, each URAM has only two read ports, which limits the parallel access capability.
[0154] To address this limitation, the dense vector Qx is replicated 15 times in the URAM to achieve simultaneous parallel access.
[0155] The core function of the first aggregator 223 is to sum up the multiplication results to generate partial sums. Specifically, the first aggregator 223 accumulates the element-wise multiplication results of the elements in the same row to calculate the partial sum of each row. When processing consecutive data packets, the first aggregator 223 retains the aggregation result of the last row in the previous data packet and dynamically determines whether to perform cross-packet aggregation operations based on the row flag indicator (flag) of the first element in the new data packet.
[0156] The Top-16 updater 224 updates the row indices of the most important 16 non-zero elements according to the calculation results. The Top-16 updater 224 compares the current Top-16 results of the quantization calculation core 220 (Q-Core) with the aggregation results of the new data packet to update the best row indices of the Top-16 updater 224. Each Top-16 updater 224 of the quantization calculation core 220 (Q-Core) on the field-programmable gate array FPGA maintains a merged comparison circuit implemented by a LUT to identify the minimum value in each clock cycle and compare it with the input value. If the input value exceeds the minimum value in the current Top-16, the value is updated in place, which effectively alleviates the intensive data movement usually associated with sorting algorithms. After the calculation is completed, the quantization calculation core 220 (Q-Core) calculates the Top-16 results of the sub-matrix it is responsible for. When K exceeds 100 and the accuracy of the Top-512 approximation decreases, the Top-16 updater 224 can be quickly adjusted to the Top-32 updater or higher by modifying the parameters.
[0157] S32: Execute an arbitration strategy for multi-channel polling.
[0158] This step requires two types of execution units: a reader 240 and an arbiter 250. The reader 240 is used to read full-precision data from the corresponding channel according to the row index output by the quantization calculation core 220 (Q-Core) and pack it into a format suitable for processing by the full-precision calculation core 230 (F-Core). The arbiter 250 merges multiplexed data into one path in a polling manner to avoid requests that cause the full-precision calculation core 230 (F-Core) to be unable to process. As shown in the reader 240 Figure 6 shown, a total of 32 are configured. Each reader 240 reads the FP32 sparse matrix stored in the Random-CSR data format from a specific channel using the Top-16 row indexes output by the quantization calculation core 220 (Q-Core). After the reader 240 retrieves the required data packets, it repackages them and includes the necessary tag information for subsequent decoding. The tag contains three fields: ch., offset, and end. The ch. uses 5 bits to identify the high-bandwidth memory 210 channel from which the data packet comes. The offset field uses 4 bits to indicate the relative position of the first element of the i-th row in the data packet. The end field uses 1 bit to indicate whether the i-th row has been fully read. The x field represents the row index of the first element of the current data packet. The flag field represents the line break flag for 10 non-zero elements. The y field records the column indexes of 10 non-zero elements. The value field records the full-precision values of 10 non-zero elements.
[0159] The output ports of each reader 240 have the same priority. When each reader 240 initiates a write request simultaneously, the arbiter 250 responds in sequence. As shown in Figure 7 shown, the arbiter 250 receives requests from 32 readers 240 through reader_req[31:0]. Each bit represents the request status of a reader 240. The arbiter 250 generates the next arbitration decision nxt_arb[31:0] based on the current state curr_arb[31:0] and the input requests, and updates the state under the control of the arb_en signal. The en signal is used to activate the arbiter 250, and the ready signal indicates whether the arbiter 250 is ready for the next operation. The valid signal is used to verify the validity of the arbitration result. The arbiter 250 ensures the fairness and efficiency of resource allocation by dynamically updating curr_arb[31:0] and monitoring reader_req[31:0], ensuring the complete processing of data from 32 readers 240.
[0160] S33: Perform the full-precision SpMV operation.
[0161] The full-precision calculation core 230 (F-Core), as the second calculation engine designed for full-precision floating-point (FP32) data, aims to complete the final accurate calculation. The full-precision calculation core 230 (F-Core) extracts full-precision data packets from the corresponding high-bandwidth memory 210 channels according to the Top-16 row indexes output by the quantization calculation core 220 (Q-Core) and performs SpMV calculations. After being parsed by the reader 240 and arbitrated by the arbiter 250, the full-precision data packets are input into the full-precision calculation core 230 (F-Core) for full-precision calculation.
[0162] As Figure 8 shown, the full-precision calculation core 230 (F-Core) includes a second decoder 231, a second element-wise multiplier 232, and a second aggregator 233. The second decoder 231 is used to decode the input data. The second element-wise multiplier 232 performs element-wise multiplication operations, and the second aggregator 233 is responsible for aggregating the results.
[0163] The second decoder 231 parses the data packets read and compressed by the reader 240, extracting the fields tag, x, y, and value. Since the position of the first floating-point element in the i-th row of the initial data packet is uncertain, the present invention uses flag and offset to determine which element in the data packet to start parsing from. The aggregator will later use the ch. and e (end) fields.
[0164] The second element-wise multiplier 232 has the same design concept as the quantization calculation core 220 (Q-Core), except that the multiplication unit in the full-precision calculation core 230 (F-Core) optimizes the SpMV performance by replicating the full-precision dense vector values 6 times in the URAM to ensure 10 non-zero parallel random accesses.
[0165] The second aggregator 233 aggregates the dot product results of the elements in the same row. The second aggregator 233 uses the ch. field parsed by the parser to map the multiplication results to the corresponding variables for accumulation. Since a single channel processes multiple rows of data, the end field helps to identify the line breaks within each channel.
[0166] As Figure 1As shown, in the FPGA, a heterogeneous accelerator is implemented based on HLS (High-Level Synthesis), using the Vitis 2023.2 toolchain and deployed on the Alveo U280 accelerator card platform. The accelerator includes 33 computing cores, 32 readers 240, and 1 arbiter 250. Among them, the computing cores include 32 quantization computing cores 220 (Q-Core) and 1 full-precision computing core 230 (F-Core). The quantization computing core 220 (Q-Core) is responsible for processing the sparse matrix data after INT6 quantization, mainly performing preliminary SpMV calculations and in-place updating of the top-16 row indices. The quantization computing core 220 (Q-Core) significantly improves the computing throughput by processing low-bitwidth INT6 data, reducing the computing complexity and data transmission burden. The full-precision computing core 230 (F-Core) is responsible for processing high-precision FP32 data to ensure the accuracy of the final calculation results of Top-K SpMV.
[0167] The reader 240 is responsible for reading full-precision data from 32 high-bandwidth memory 210 channels of the Alveo U280 accelerator card, and then integrating 32 data streams into one through the arbiter 250 and transmitting it to the full-precision computing core 230 (F-Core) for full-precision SpMV calculation. The design of the arbiter 250 optimizes data scheduling and transmission efficiency, ensuring efficient support for multi-channel data exchange in a high-bandwidth environment.
[0168] As Figure 6 shown, the quantization computing core 220 (Q-Core) is used to count the top-16 row indices in the SpMV results of the partial rows of the quantized sparse matrix (SpM) and the vector x. The architecture design of the quantization computing core 220 (Q-Core) includes multiple key components, aiming to efficiently process large-scale sparse data. The first decoder 221 is used to decode the row indices of non-zero elements in the data packet, decompress the sparse matrix data packet in Ultra-CSR data format, and store the decompressed data in the local memory for subsequent calculations. Since the Ultra-CSR data format introduces data dependencies during the decoding stage, that is, the x coordinate of the latter non-zero element depends on the x coordinate value of the previous non-zero element, it is challenging to parallelly decode 30 non-zero elements. Therefore, the decoding problem of the x coordinates of non-zero elements is transformed into parallelly counting how many line breaks there are before each non-zero element in a 512-bit-wide data packet.
[0169] The specific execution process for forming the Ultra-CSR data format is as follows: For non-zero elements with index i, first calculate the number of all 1s in the new array before index i, and then add this number to the maximum row number in the previous data packet as the x-coordinate of this non-zero element. The first element-wise multiplier 222 is a key component of the quantization calculation core 220 (Q-Core), and its main responsibility is to perform element-wise multiplication operations on the vector x and non-zero elements within the data packet.
[0170] To reduce the latency generated during random access to the vector x, before processing the sparse matrix SpM operation, the vector x is pre-transferred to the UltraRAM (URAM) memory of the quantization calculation core 220 (Q-Core). During the actual operation process, the data of the vector x can be directly read from the URAM, thus achieving efficient data access. Each data packet contains 30 non-zero elements. In the most unfavorable case, all these non-zero elements need to perform multiplication operations with the same element of the vector x. Since the URAM only provides two read ports, restricting the parallel access ability, the vector x needs to be copied 15 times in the URAM to ensure that 30 non-zero elements can be accessed and calculated simultaneously.
[0171] The Alevo U280 device is configured with 960 URAM units, and the capacity of each URAM is 4096x72 bits (bit), providing a total storage space of 270MB. In this application scenario, the size of the vector x is 1024x6 bits (bit), which means it can be completely stored within a single URAM. Even if 15 copies of the vector x are each replicated within 32 cores and each copy of x occupies a single URAM, only 480 URAMs are required to meet the demand. This method not only significantly improves the computing efficiency but also effectively utilizes the hardware resources.
[0172] The first aggregator 223 is responsible for aggregating the element-wise multiplication results in the same row. During the processing, it saves the aggregation result of the last row of the previous data packet and uses the new_row_flag flag of the first element of the new data packet to determine whether there are identical row elements that need to be aggregated between the two data packets. After being processed by the first aggregator 223, the quantization calculation core 220 (Q-Core) completes the basic operation of the sparse matrix-vector multiplication (SpMV).
[0173] The Top-16 updater 224 is used to track and record the optimal Top-16 row indices in real time. During the processing, the Top-16 updater 224 compares the current Top-16 results maintained by the quantization calculation core 220 (Q-Core) with the results aggregated from new data packets. If the newly aggregated results are better than the worst result in the current Top-16, the Top-16 updater 224 will synchronously update the Top-16 calculation results and their corresponding row indices.
[0174] To find the worst result within Top-16, the idea of merge sort is adopted to implement the argmin16() function, which introduces four data dependencies. To avoid these dependencies from breaking the pipeline interval (II), registers are inserted between three of the dependencies in the design, and they are processed in three clock cycles to ensure that the pipeline interval II = 1. Through this design, the quantization calculation core 220 (Q-Core) realizes efficient Top-K sparse matrix-vector multiplication (SpMV) calculation, greatly improving the performance of processing large-scale sparse data.
[0175] Embodiment 2
[0176] This embodiment is a further improvement of Embodiment 1, and the repeated content will not be elaborated.
[0177] In a large-scale online recommendation system, the system needs to process the interaction data between millions of users and a vast amount of commodities in real time. These data usually exist in the form of sparse matrices, and the distribution of non-zero elements is highly random. The Top-K SpMV operator is widely used in recommendation systems to calculate the similarity between users and commodities, so as to provide personalized commodity recommendations for users. To ensure that users obtain recommendation results during real-time access, the response time of the system must be controlled within the millisecond level.
[0178] However, the traditional Top-K SpMV algorithm has significant performance bottlenecks when processing large-scale sparse matrices. Due to the random access pattern of sparse matrices and the uneven distribution of non-zero elements, the traditional algorithms based on CPUs and GPUs are inefficient in terms of memory bandwidth and data transmission. Especially in the sorting stage, the latency of data movement and random access occupies more than 60% of the calculation time, resulting in low memory utilization and further increasing the calculation latency.
Claims
1. An accelerator based on Top-K sparse matrix vector multiplication, characterized in that: include: A first acceleration unit (100) is used for data preprocessing and sparse matrix encoding. The second acceleration unit (200) is used to perform Top-K sparse matrix-vector multiplication to obtain the top K results with the largest median value of the operation results of multiplying the sparse matrix and the dense vector; Wherein, the first acceleration unit (100) comprises a preprocessing module (110) and an encoding module (120); The preprocessing steps of the preprocessing module (110) include quantization, sparse reconstruction and shuffling; A sparse matrix-based quantization step to reduce the bandwidth requirements for data transmission; Based on the sparse reconstruction algorithm, non-zero elements with little impact on sorting are pruned to reduce redundant memory transmission and calculation; Shuffle the sparse matrix after sparse reconstruction to avoid Top-K aggregation; The encoding module (120) is used to re-encode the pre-processed sparse matrix.
2. The accelerator according to claim 1, characterized in that The steps of sparse matrix encoding include: The preprocessed sparse matrix is re-encoded into the Ultra-CSR data format or the Random-CSR data format. While keeping the sparse structure of the internal elements of each row unchanged, the sparse matrix is recombined to form a new sparse matrix with completely random row arrangement.
3. The accelerator according to claim 1 or 2, characterized in that: The second acceleration unit (200) includes a quantization computing core (220), a full-precision computing core (230), a reader (240) and an arbitrator (250); The quantization computing core (220) is connected to a reader (240), the reader (240) is connected to an arbitrator (250), and the arbitrator (250) is connected to a full-precision computing core (230); The quantization calculation core (220) is used to perform sparse matrix-vector multiplication and update the Top-16 row indexes maintained by each to form a final Top-512 row index; The reader (240) is used to read the full-precision data from the corresponding channel according to the row index output by the quantization computing core (220) and pack it into a format suitable for processing by the full-precision computing core (230); The arbiter (250) is used to merge multiple channels of data into one channel in a round-robin manner and send the data to the full-precision computing core (230); The full-precision computing core (230) extracts full-precision data packets and performs SpMV calculations.
4. The accelerator according to any one of claims 1 to 3, characterized in that: The quantization method includes: multiplying all non-zero elements in the sparse matrix by a constant in equal proportion, and then rounding them so that the value distribution falls within the interval [-32, 31].
5. The accelerator according to any one of claims 1 to 4, characterized in that: The sparse reconstruction step comprises: Calculate the original precision; Set the accuracy loss tolerance and the initial pruning ratio; In each iterative pruning process, the elements in the same row are sorted for each row vector in the sparse matrix, and elements are removed according to the pruning ratio of the current iteration round, and the elements are elements with smaller absolute values; Determine whether the accuracy meets the requirements; When the accuracy is met, the pruning ratio is increased and the elements in the same row are sorted. If the accuracy does not meet the requirement, the pruning ratio is reduced, and it is determined whether the ratio has been searched, and if it has been searched, a pruned sparse matrix is formed; If it is determined that the proportion has not been searched, the elements in the same row are re-sorted.
6. The accelerator according to any one of claims 1 to 5, characterized in that: The steps of sparse matrix encoding include: When the coefficient embedding sparsity of the sparse matrix is greater than 1%, the sparse matrix is encoded from the COO data format to the Ultra-CSR data format; In case of random access to full precision sparse matrices, the sparse matrix is encoded from the COO data format to the Random-CSR data format.
7. The accelerator according to any one of claims 1 to 6, characterized in that: The quantization calculation core (220) includes: A first decoder (221), for parsing input data; A first element-by-element multiplier (222) for performing element-by-element multiplication between a sparse matrix and a vector; a first aggregator (223) for aggregating the multiplication results to generate partial sums; The Top-16 updater (224) is used to update the row indexes of the 16 most important non-zero elements according to the calculation results.
8. The accelerator according to any one of claims 1 to 7, characterized in that: The full-precision computing core (230) includes: A second decoder (231) for parsing the data packet read and compressed by the reader (240); a second element-by-element multiplier (232) for optimizing SpMV performance by replicating the full-precision dense vector value multiple times in the URAM; The second aggregator (233) is used to aggregate the dot product results of the elements in the same row.
9. An acceleration method based on Top-K sparse moment vector multiplication, characterized in that: The method comprises: Perform data preprocessing and sparse matrix encoding on sparse matrices; Perform Top-K sparse matrix-vector multiplication to obtain the top K results with the largest median value of the multiplication results of the sparse matrix and the dense vector; Wherein, the preprocessing steps include quantization, sparse reconstruction and shuffling; A sparse matrix-based quantization step to reduce the bandwidth requirements for data transmission; Based on the sparse reconstruction algorithm, non-zero elements with little impact on sorting are pruned to reduce redundant memory transmission and calculation; Shuffle the sparse matrix after sparse reconstruction to avoid Top-K aggregation; Re-encode the preprocessed sparse matrix.
10. The method according to claim 9, characterized in that The sparse reconstruction step comprises: Calculate the original precision; Set the accuracy loss tolerance and the initial pruning ratio; In each iterative pruning process, the elements in the same row are sorted for each row vector in the sparse matrix, and elements are removed according to the pruning ratio of the current iteration round, and the elements are elements with smaller absolute values; Determine whether the accuracy meets the requirements; When the accuracy is met, the pruning ratio is increased and the elements in the same row are sorted. If the accuracy does not meet the requirement, the pruning ratio is reduced, and it is determined whether the ratio has been searched, and if it has been searched, a pruned sparse matrix is formed; If it is determined that the proportion has not been searched, the elements in the same row are re-sorted.
Citation Information
Patent Citations
GPU acceleration optimization method and device based on sparse matrix vector multiplication
CN117495651A
Cited By
Special accelerator for hierarchical greedy decoding algorithm
CN120671862A
Inference acceleration method and device
CN121706994A
Memory allocation method and device, computer equipment and storage medium
CN122044827A
Universal TopK computing device and method based on insertion merging
CN122450504A