Multi-mode sparse matrix vector multiplication accelerator and acceleration system
By designing a multi-mode sparse matrix vector multiplication accelerator, flexible switching of computing modes and diversity of data representation are achieved, and the problem of existing accelerators being limited to specific algorithms and format combinations is solved, and efficient computing performance and energy efficiency optimization under different conditions are achieved.
Patent Information
- Application Number
- CN202510221889.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-20
AI Technical Summary
Existing sparse matrix and vector multiplication accelerators are often limited to specific algorithms and format combinations, making it difficult to achieve optimal performance on different architectural configurations, input matrices, and vectors.
A multi-mode sparse matrix vector multiplication accelerator and acceleration system are designed to realize flexible switching of calculation modes and diversity of data representation through the combination of index calculation units, loading storage units, numerical registers and multiplication units.
By providing flexible data representation and calculation mode selection, efficient sparse matrix vector multiplication acceleration can be achieved under various matrix and vector characteristics, improving computing performance and reducing energy consumption.
Smart Images

Figure CN120179974A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of high-performance computing technology, and more particularly to a multi-mode sparse matrix-vector multiplication accelerator and an acceleration system. Background Art
[0002] In the field of high-performance computing, sparse matrix-vector multiplication (SpMV) is one of the key computations and is widely used in fields such as expert systems, big data analysis, and machine learning. Given the importance of sparse matrix-vector multiplication, custom accelerators for sparse matrices have been widely studied. These accelerators mainly face two key choices in design: matrix format and algorithm. There are theoretically multiple possible design combinations for these two design aspects. However, existing sparse matrix-vector multiplication accelerators are often limited to specific algorithm and format combinations.
[0003] When dealing with high-performance computing workloads, a single fixed design does not achieve optimal performance on all different architectural configurations, input matrices, and vectors. To fully utilize the characteristics of sparse matrices, the design of accelerators should fully consider hybrid design, adaptive optimization, cross-platform compatibility, and energy efficiency optimization. For this reason, the present invention proposes a multi-mode sparse matrix-vector multiplication acceleration system that can change modes according to the best data representation. By providing flexible data representation and calculation mode selection, as well as efficient input format conversion and data processing mechanisms, it is possible to achieve efficient sparse matrix-vector multiplication acceleration under various matrix and vector characteristics, thereby improving computing performance and reducing energy consumption.
[0004] Existing sparse matrix-vector multiplication accelerators are often limited to specific algorithm and format combinations. How to achieve the hybrid design, adaptive optimization, cross-platform compatibility, and energy efficiency optimization of accelerators is a technical problem that needs to be solved. Summary of the Invention
[0005] The technical task of the present invention is to address the above deficiencies by providing a multi-mode sparse matrix-vector multiplication accelerator and an acceleration system to solve the technical problem of how to achieve the hybrid design, adaptive optimization, cross-platform compatibility, and energy efficiency optimization of accelerators.
[0006] In a first aspect, a multi-mode sparse matrix-vector multiplication accelerator of the present invention includes an index calculation unit, a load / store unit, a numerical register, a multiply-accumulate unit, and on-chip memory;
[0007] The index calculation unit is configured to switch between a plurality of custom calculation modes based on an input mode selection instruction, and call the load / store unit and the multiply-accumulate unit based on the switched calculation mode, wherein the plurality of custom calculation modes include dense calculation, compressed matrix multiplication calculation, and dot matrix multiplication calculation;
[0008] Correspondingly, the load / store unit is used to read relevant data from the on-chip memory based on the switched computing mode and cache the relevant data into the numeric register;
[0009] Correspondingly, the multiply-accumulate unit is used to read the data cached in the numeric register based on the switched computing mode for calculation, obtain a calculation result, and store the calculation result in the numeric register;
[0010] Correspondingly, the load / store unit is used to store the calculation result stored in the numeric register into the on-chip memory.
[0011] Preferably, the mode selection instruction is generated by the runtime system that controls the accelerator and input into the accelerator. The runtime system determines the selected computing mode based on the service scenario to be processed and the required data, according to the dimension and density of the required data, and generates a corresponding mode selection instruction;
[0012] The runtime system has functions of device initialization, device management, memory management, device abstraction, interrupt response, error handling, and log recording.
[0013] Preferably, for a service scenario with fixed requirements, the runtime system is used to specify a computing mode based on the data volume required by the service scenario, generate a computing mode instruction, and pre-store the computing mode instruction locally in the runtime system. When the runtime system receives the requirements of the relevant service scenario, it inputs the pre-stored computing mode instruction into the index calculation unit;
[0014] Correspondingly, the index calculation unit is used to call the load / store unit and the multiply-accumulate unit based on the input computing mode instruction;
[0015] Correspondingly, the load / store unit is used to read relevant data from the on-chip memory based on the computing mode specified in the computing mode instruction and cache the relevant data into the numeric register;
[0016] Correspondingly, the multiply-accumulate unit is used to read the data cached in the numeric register based on the computing mode specified in the computing mode instruction for calculation, obtain a calculation result, and store the calculation result in the numeric register;
[0017] Correspondingly, the load / store unit is used to store the calculation result stored in the numeric register into the on-chip memory.
[0018] Preferably, for multiple custom computing modes, the index calculation unit is configured with a scheduling module corresponding to each computing mode. Each call module is used to perform index calculation for its corresponding computing mode, and call the load / store unit and the multiply-accumulate unit for data reading and calculation.
[0019] Preferably, if the switched computing mode is compressed matrix multiplication, the index calculation unit is used to calculate the column index of non-zero values, call the load / store unit to read the values corresponding to the column index from the on-chip memory, and send the values corresponding to the column index to the numeric register through the load / store unit. Correspondingly, the multiply-accumulate unit reads the values corresponding to the column index from the numeric register for calculation;
[0020] If the switched computing mode is dot matrix multiplication, the index calculation unit is used to determine the index position of non-zero values through leading non-zero detection logic, call the load / store unit to read the values corresponding to the index position from the on-chip memory, and send the values corresponding to the index position to the numeric register through the load / store unit. Correspondingly, the multiply-accumulate unit reads the values corresponding to the column index from the numeric register for calculation;
[0021] If the switched computing mode is dense computing, the index calculation unit is used to directly call the load / store unit, and the load / store unit directly reads the values in the matrix from the on-chip memory in sequence, sends the values to the numeric register. Correspondingly, the multiply-accumulate unit reads the values corresponding to the column index from the numeric register for calculation.
[0022] Preferably, if the switched computing mode is compressed matrix multiplication, the index calculation unit is used to perform the following operations to calculate the column index of non-zero values:
[0023] Obtain the row pointer: Read the starting pointer of the current row to determine the starting position of the row;
[0024] Sequentially read the column index array: Detect the column indexes of non-zero elements included in the current row by sequentially reading the column index array;
[0025] Extract non-zero values: According to the detected column indexes, call the load / store unit to fetch the corresponding non-zero values from the on-chip memory, and send the read non-zero values to the numeric register.
[0026] Preferably, if the switched computing mode is dot matrix multiplication, the index calculation unit is used to perform the following operations to determine the index position of non-zero values through leading non-zero detection logic:
[0027] Read the row bitmap: Read the bitmap of the matrix row, and each bit in the bitmap indicates whether the element at the corresponding position is non-zero;
[0028] Leading non-zero detection: Use the leading non-zero detection logic to sequentially check the bitmap to find the position of the next non-zero bit;
[0029] Return the column index: Use the position of the detected non-zero bit as the column index of the next non-zero value, call the load / store unit to read the value corresponding to the column index from the on-chip memory, and send the read value to the numeric register.
[0030] In a second aspect, a multi-mode sparse matrix-vector multiplication acceleration system according to the present invention includes an accelerator as described in any one of the first aspects, and the accelerator is controlled by a runtime system, which is designed and implemented in C++.
[0031] The multi-mode sparse matrix-vector multiplication accelerator and acceleration system of the present invention have the following advantages: It can switch the calculation mode according to the service scenario and the data volume of the required data. By providing flexible data representation and calculation mode selection, as well as efficient input format conversion and data processing mechanisms, it can achieve efficient sparse matrix-vector multiplication acceleration under various matrix and vector characteristics, thereby improving the calculation performance and reducing the energy consumption. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0033] The present invention will be further described below in conjunction with the drawings.
[0034] Figure 1 It is a structural block diagram of a multi-mode sparse matrix-vector multiplication accelerator for Embodiment 1. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0035] The present invention will be further described below in conjunction with the drawings and specific embodiments, so that those skilled in the art can better understand the present invention and implement it. However, the embodiments cited are not intended to limit the present invention. Without conflict, the embodiments of the present invention and the technical features in the embodiments can be combined with each other.
[0036] Embodiments of the present invention provide a multi-mode sparse matrix-vector multiplication accelerator and acceleration system, which are used to solve the technical problems of how to achieve the hybrid design, adaptive optimization, cross-platform compatibility and energy efficiency optimization of the accelerator.
[0037] Embodiment 1:
[0038] A multi-mode sparse matrix-vector multiplication accelerator according to the present invention includes an index calculation unit, a load / store unit, a numerical register, a multiply-accumulate unit, and on-chip memory.
[0039] The index calculation unit is used to switch between multiple custom calculation modes based on the input mode selection instruction, and call the load / store unit and the multiply-accumulate unit based on the switched calculation mode. Among them, the multiple custom calculation modes include dense calculation, compressed matrix multiplication calculation, and dot matrix multiplication calculation.
[0040] In this embodiment, the mode selection instruction is generated by the runtime system that controls the accelerator and input into the accelerator. The runtime system determines the selected calculation mode based on the business scenario to be processed and the required data, according to the dimension and density of the required data, and generates the corresponding mode selection instruction. Among them, the runtime system that controls the accelerator is designed and implemented in C++. The runtime system has functions such as device initialization, device management, memory management, device abstraction, interrupt response, error handling, and logging.
[0041] Correspondingly, the load / store unit is used to read relevant data from the on-chip memory based on the switched calculation mode and cache the relevant data into the numeric register.
[0042] Correspondingly, the multiply-accumulate unit is used to read the data cached in the numeric register based on the switched calculation mode for calculation, obtain the calculation result, and store the calculation result in the numeric register.
[0043] Correspondingly, the load / store unit is used to store the calculation result stored in the numeric register into the on-chip memory.
[0044] In this embodiment, for the multiple custom calculation modes, scheduling modules corresponding one-to-one to the calculation modes are configured in the index calculation unit. Each call module is used to perform index calculation for its corresponding calculation mode, and call the load / store unit and the multiply-accumulate unit for data reading and calculation.
[0045] If the switched calculation mode is compressed matrix multiplication calculation, the index calculation unit is used to calculate the column index of non-zero values, call the load / store unit to read the values corresponding to the column index from the on-chip memory, and send the values corresponding to the column index to the numeric register through the load / store unit. Correspondingly, the multiply-accumulate unit reads the values corresponding to the column index from the numeric register for calculation;
[0046] If the switched calculation mode is dot matrix multiplication calculation, the index calculation unit is used to determine the index position of non-zero values through the leading non-zero detection logic, call the load / store unit to read the values corresponding to the index position from the on-chip memory, and send the values corresponding to the index position to the numeric register through the load / store unit. Correspondingly, the multiply-accumulate unit reads the values corresponding to the column index from the numeric register for calculation;
[0047] If the switched computing mode is dense computing, the index calculation unit is used to directly call the load / store unit, and the load / store unit directly reads the values in the matrix from the on-chip memory in sequence, sends the values to the numeric register. Correspondingly, the multiply-accumulate unit reads the values corresponding to the column index from the numeric register for calculation.
[0048] As an improvement of this embodiment, for a business scenario with fixed requirements, the runtime system is used to specify the computing mode based on the data volume required by the business scenario, generate a computing mode instruction, and pre-store the computing mode instruction locally in the runtime system. When the runtime system receives the requirements of the relevant business scenario, it inputs the pre-stored computing mode instruction into the index calculation unit. That is, for the calculation with a fixed input matrix, the computing mode is determined in advance and stored in one of the dense format, compressed format, and bitmap format.
[0049] Correspondingly, the index calculation unit is used to call the load / store unit and the multiply-accumulate unit based on the input computing mode instruction.
[0050] Correspondingly, the load / store unit is used to read relevant data from the on-chip memory based on the computing mode specified in the computing mode instruction and cache the relevant data into the numeric register.
[0051] Correspondingly, the multiply-accumulate unit is used to read the data cached in the numeric register based on the computing mode specified in the computing mode instruction for calculation, obtain the calculation result, and store the calculation result in the numeric register.
[0052] Correspondingly, the load / store unit is used to store the calculation result stored in the numeric register into the on-chip memory.
[0053] For some specific applications, the workload may contain multiple matrices with different sparsities, and the mode can be dynamically selected at runtime.
[0054] The accelerator of this embodiment can switch among three modes: dense computing, compressed matrix multiplication computing, and dot matrix multiplication computing. The index calculation unit is responsible for calculating the calculation method of the non-zero value column index and fetching the value to be sent to the multiply-accumulate unit according to the index. In the compressed mode, the index calculation unit first obtains the row pointer and detects the column index of the non-zero elements contained in the row by sequentially reading the column index array; in the dot matrix mode, the index calculation unit uses the leading non-zero detection logic to sequentially check the bitmap of the matrix row and returns the position of the next non-zero bit as the column index of the next non-zero value; for dense computing, the accelerator runs efficiently by deactivating all modules under the index calculator and providing values in sequence.
[0055] Embodiment 2:
[0056] A multi-mode sparse matrix-vector multiplication acceleration system of the present invention includes the accelerator disclosed in Embodiment 1, which is controlled by a runtime system designed and implemented in C++.
[0057] The above has introduced in detail the multi-mode sparse matrix-vector multiplication accelerator and acceleration system provided by the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A multi-mode sparse matrix-vector multiplication accelerator, characterized in that: Includes index calculation unit, load storage unit, value register, multiply-add unit and on-chip memory; The index calculation unit is used to switch between multiple custom calculation modes based on the input mode selection instruction, and call the load storage unit and the multiplication and addition unit based on the switched calculation mode, wherein the multiple custom calculation modes include dense calculation, compressed matrix multiplication calculation and lattice matrix multiplication calculation; Correspondingly, the load storage unit is used to read relevant data from the on-chip memory based on the switched computing mode and cache the relevant data into the value register; Correspondingly, the multiplication and addition unit is used to read the data cached in the numerical register based on the switched calculation mode to perform calculation, obtain the calculation result, and store the calculation result in the numerical register; Correspondingly, the load storage unit is used to store the calculation results stored in the value register into the on-chip memory.
2. The multi-mode sparse matrix-vector multiplication accelerator according to claim 1, characterized in that: The mode selection instruction is generated by the runtime system that controls the accelerator and input into the accelerator. The runtime system determines the selected computing mode based on the business scenario to be processed and the required data, and the dimension and density of the required data, and generates the corresponding mode selection instruction; The runtime system has the functions of device initialization, device management, memory management, device abstraction, interrupt response, error handling and log recording.
3. The multi-mode sparse matrix-vector multiplication accelerator according to claim 2, characterized in that: For business scenarios with fixed requirements, the runtime system is used to specify the calculation mode based on the data volume required by the business scenario, generate calculation mode instructions, and pre-store the calculation mode instructions locally in the runtime system. When the runtime system receives the requirements of the relevant business scenario, it inputs the pre-stored calculation mode instructions into the index calculation unit; Correspondingly, the index calculation unit is used to call the load storage unit and the multiply-add unit based on the input calculation mode instruction; Correspondingly, the load storage unit is used to read relevant data from the on-chip memory based on the calculation mode specified in the calculation mode instruction and cache the relevant data into the value register; Correspondingly, the multiplication and addition unit is used to read the data cached in the numerical register to perform calculations based on the calculation mode specified in the calculation mode instruction, obtain the calculation results, and store the calculation results in the numerical register; Correspondingly, the load storage unit is used to store the calculation results stored in the value register into the on-chip memory.
4. The multi-mode sparse matrix-vector multiplication accelerator according to any one of claims 1 to 3, characterized in that: For multiple customized computing modes, the index computing unit is configured with scheduling modules corresponding to the computing modes one by one. Each calling module is used to perform index calculation for its corresponding computing mode, and calls the load storage unit and the multiplication and addition unit to read and calculate data.
5. The multi-mode sparse matrix-vector multiplication accelerator according to claim 4, characterized in that: If the switched calculation mode is compressed matrix multiplication calculation, the index calculation unit is used to calculate the column index of the non-zero value, call the load storage unit to read the value corresponding to the column index from the on-chip memory, send the value corresponding to the column index to the value register through the load storage unit, and correspondingly, the multiplication and addition unit reads the value corresponding to the column index from the value register for calculation; If the switched calculation mode is dot matrix multiplication calculation, the index calculation unit is used to determine the index position of the non-zero value through the leading non-zero detection logic, call the load storage unit to read the value corresponding to the index position from the on-chip memory, and send the value corresponding to the index position to the value register through the load storage unit. Correspondingly, the multiplication and addition unit reads the value corresponding to the column index from the value register for calculation; If the switched computing mode is intensive computing, the index computing unit is used to directly call the load storage unit, the load storage unit directly reads the values in the matrix from the on-chip memory in sequence and sends the values to the numerical register. Correspondingly, the multiplication and addition unit reads the value corresponding to the column index from the numerical register for calculation.
6. The multi-mode sparse matrix-vector multiplication accelerator according to claim 5, characterized in that: If the switched calculation mode is compressed matrix multiplication calculation, the index calculation unit is used to perform the following operations to calculate the column index of the non-zero value: Get row pointer: read the starting pointer of the current row to determine the starting position of the row; Sequentially read the column index array: By sequentially reading the column index array, the column index of the non-zero element contained in the current row is detected; Extract non-zero values: According to the detected column index, call the load storage unit to fetch the corresponding non-zero value from the on-chip memory and send the read non-zero value to the value register.
7. The multi-mode sparse matrix-vector multiplication accelerator according to claim 5, characterized in that: If the switched calculation mode is the dot matrix multiplication calculation, the index calculation unit is used to perform the following operations to determine the index position of the non-zero value through the leading non-zero detection logic: Read row bitmap: read the bitmap of the matrix row, each bit in the bitmap indicates whether the element at the corresponding position is non-zero; Leading non-zero detection: Use the leading non-zero detection logic to check the bitmap in sequence to find the position of the next non-zero bit; Return column index: Use the position of the detected non-zero bit as the column index of the next non-zero value, call the load storage unit to read the value corresponding to the column index from the on-chip memory, and send the read value to the value register.
8. A multi-mode sparse matrix-vector multiplication acceleration system, characterized in that: The accelerator comprises the accelerator as claimed in any one of claims 1 to 7, wherein the accelerator is controlled by a runtime system, and the runtime system is designed and implemented in C++.