Matrix multiplication optimization method and system based on RISC-V architecture

By converting sparse matrix into structured sparse matrix and designing custom vector index multiplication instructions, the performance bottleneck of sparse matrix multiplication on vector processors is solved, and more efficient data access and parallel computing is achieved.

CN120336686APending Publication Date: 2025-07-18SHANDONG INSPUR SCI RES INST CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510363862.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Traditional sparse matrix multiplication algorithms have performance bottlenecks in vector processors, including irregular data access patterns, excessive number of instructions, and parallel optimization of circular dependence and register pressure limit instruction-level.

Method used

Convert sparse matrix into a structured sparse matrix, store it in a vector register file, and store the column index in a scalar register file. Through interleaving execution and custom vector index multiplication and addition instructions to optimize loop expansion, custom vector index multiplication and addition instructions are designed and implemented to replace traditional vector loading instructions.

Benefits of technology

It improves the performance of sparse matrix multiplication, reduces index storage overhead, improves cache hit rate and instruction utilization, and realizes efficient parallel computing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336686A_ABST
    Figure CN120336686A_ABST
Patent Text Reader

Abstract

The invention discloses a matrix multiplication optimization method and system based on an RISC-V architecture, belongs to the technical field of machine learning, and aims to solve the technical problem of how to effectively improve the performance of sparse matrix multiplication. Comprising the following steps: constructing a sparse matrix adaptive to a machine learning model; the sparse matrix is converted into a structured sparse matrix, non-zero elements in the structured sparse matrix are stored in a vector register file, and column indexes corresponding to the non-zero elements are stored in a scalar register file; determining expansion factors of the internal circulation and the external circulation, and adjusting the expansion factors of the internal circulation; executing instructions in different loop iterations in a staggered manner in a staggered manner; designing and realizing a user-defined vector index multiply-add instruction; a custom vector index multiply-add instruction is executed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of machine learning, and more specifically to a matrix multiplication optimization method and system based on the RISC-V architecture. Background Art

[0002] With the booming development of machine learning applications, matrix multiplication has become a core computational operation. To improve computational efficiency and reduce storage overhead, sparse matrices are widely used in machine learning models. Sparse matrix multiplication can utilize zero-valued elements to reduce the amount of computation and storage space.

[0003] However, traditional sparse matrix multiplication algorithms have performance bottlenecks when executed on vector processors. Firstly, the data access pattern is irregular. The non-zero elements of the sparse matrix are unevenly distributed, resulting in an irregular data access pattern, which easily causes cache misses and reduces performance. Secondly, the number of instructions is excessive. Traditional algorithms require a large number of instructions to process sparse matrices, such as loading, storing, index calculation, etc., increasing the execution time. Thirdly, loop dependencies and register pressure limit instruction-level parallel optimization.

[0004] How to effectively improve the performance of sparse matrix multiplication is a technical problem that needs to be solved. Summary of the Invention

[0005] The technical task of the present invention is to address the above deficiencies and provide a matrix multiplication optimization method and system based on the RISC-V architecture to solve the technical problem of how to effectively improve the performance of sparse matrix multiplication.

[0006] In a first aspect, a matrix multiplication optimization method based on the RISC-V architecture according to the present invention includes the following steps:

[0007] Matrix construction: Construct a machine learning model based on the application scenario, and perform model training on the machine learning model based on the collected sample data. During the model training process, construct a sparse matrix adapted to the machine learning model;

[0008] Data processing: Convert the sparse matrix into a structured sparse matrix, store the non-zero elements of the structured sparse matrix in the vector register file, and store the column indices corresponding to the non-zero elements in the scalar register file;

[0009] Loop optimization: Determine the unfolding factors of the inner loop and the outer loop according to the size of the sparse matrix, the number of vector registers, and the characteristics of the structured sparse matrix, and adjust the unfolding factor of the inner loop;

[0010] Interleaved instruction execution: Interleave the instructions in different loop iterations by means of interleaved execution;

[0011] Custom Instruction: Design and implement a custom vector index multiply-add instruction, which is used to replace the traditional vector load instruction and perform vector index multiply-add operations;

[0012] Execute Custom Instruction: Execute the custom vector index multiply-add instruction.

[0013] Preferably, when the custom vector index multiply-add instruction is executed, the following operations are included:

[0014] Read the index value from the scalar register file;

[0015] Read the corresponding vector elements from the vector register file according to the index value;

[0016] Multiply the read vector elements by the elements in another vector register, and accumulate the results of the element multiplications into the target vector register.

[0017] Preferably, during data processing, decompose the sparse matrix into multiple sub-matrices, assign each sub-matrix to a core for processing, convert each sub-matrix into a structured sparse matrix, store the non-zero elements of the structured sparse matrix in the vector register file, and store the column indices corresponding to the non-zero elements in the scalar register file;

[0018] Correspondingly, when the custom instruction is executed, the index value of the custom vector index multiply-add instruction is passed between the cores through shared memory or a message mechanism.

[0019] Preferably, during data processing, convert the sparse matrix into a structured sparse matrix, store the non-zero elements of the structured sparse matrix in the vector register file, store the column indices corresponding to the non-zero elements in the scalar register file, and convert the structured sparse matrix into a form that is easy to implement in software, including compressed sparse row or compressed sparse column format;

[0020] Correspondingly, when the custom instruction is executed, the custom vector index multiply-add instruction is implemented through software simulation or compiler optimization.

[0021] In a second aspect, the present invention provides a matrix multiplication optimization system based on the RISC-V architecture, which is used to optimize matrix multiplication through a matrix multiplication optimization method according to any one of the first aspects. The system includes a matrix construction module, a data processing module, a loop optimization module, a custom instruction module, and an execute custom instruction module;

[0022] The matrix construction module is used to perform the following: construct a machine learning model based on the application scenario, train the machine learning model based on the collected sample data, and construct a sparse matrix adapted to the machine learning model during the model training process;

[0023] The data processing module is used to perform the following: convert a sparse matrix into a structured sparse matrix, store the non-zero elements of the structured sparse matrix in a vector register file, and store the column indices corresponding to the non-zero elements in a scalar register file;

[0024] The loop optimization module is used to perform the following: determine the unfolding factors of the inner loop and the outer loop according to the size of the sparse matrix, the number of vector registers, and the characteristics of the structured sparse matrix, and adjust the unfolding factor of the inner loop;

[0025] The interleaved execution instruction module is used to perform the following: interleaving the instructions in different loop iterations in an interleaved execution manner;

[0026] The custom instruction module is used to perform the following: design and implement a custom vector index multiply-add instruction, and the custom vector index multiply-add instruction is used to replace the traditional vector load instruction and perform vector index multiply-add operations;

[0027] The execution custom instruction module is used to perform the following: execute the custom vector index multiply-add instruction.

[0028] Preferably, when the custom vector index multiply-add instruction is executed, the execution custom instruction module is used to perform the following operations:

[0029] Read the index value from the scalar register file;

[0030] Read the corresponding vector element from the vector register file according to the index value;

[0031] Multiply the read vector element by the element in another vector register, and accumulate the result of the element multiplication into the target vector register.

[0032] Preferably, the data processing module is used to perform the following: decompose the sparse matrix into multiple sub-matrices, assign each sub-matrix to a core for processing, convert each sub-matrix into a structured sparse matrix, store the non-zero elements of the structured sparse matrix in a vector register file, and store the column indices corresponding to the non-zero elements in a scalar register file;

[0033] Correspondingly, when the custom instruction is executed, the index value of the custom vector index multiply-add instruction is passed between the cores through shared memory or a message mechanism.

[0034] Preferably, the data processing is used to convert the sparse matrix into a structured sparse matrix, store the non-zero elements of the structured sparse matrix in a vector register file, store the column indices corresponding to the non-zero elements in a scalar register file, and convert the structured sparse matrix into a form that is easy to implement in software, including compressed sparse row or compressed sparse column format;

[0035] Correspondingly, when executing a custom instruction, the custom instruction execution module is used to implement a custom vector index multiply-add instruction through software simulation or compiler optimization.

[0036] The matrix multiplication optimization method and system based on the RISC-V architecture of the present invention have the following advantages:

[0037] 1. Convert a sparse matrix into a structured sparse matrix. The structured sparse matrix has a regular data access pattern, which can reduce index storage overhead and improve cache hit rate;

[0038] 2. Determine the unfolding factors of the inner loop and the outer loop according to the matrix size, the number of vector registers, and the characteristics of the structured sparse matrix. Adopt an interleaved execution method to interleave the instructions in different loop iterations, reduce read-after-write dependencies, reduce register pressure, and improve instruction utilization;

[0039] 3. Design and implement a custom vector index multiply-add instruction. This instruction can perform vector index multiply-add operations, reduce the number of memory accesses, and improve data locality;

[0040] 4. Decompose the sparse matrix into multiple sub-matrices, and assign each sub-matrix to a core for processing. Each core can independently perform loop unfolding according to its own hardware resources, achieving higher parallelism. The index values of the custom instructions can be passed between the cores through shared memory or message passing mechanisms, realizing efficient parallel computing. Description of the Drawings

[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0042] The present invention will be further described below with reference to the drawings.

[0043] Figure 1 It is a flowchart of a matrix multiplication optimization method based on the RISC-V architecture for Embodiment 1;

[0044] Figure 2 It is a flowchart of the first improvement scheme of a matrix multiplication optimization method based on the RISC-V architecture for Embodiment 1;

[0045] Figure 3 It is a flowchart of the second improvement scheme of a matrix multiplication optimization method based on the RISC-V architecture for Embodiment 1. Detailed Embodiments

[0046] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, so that those skilled in the art can better understand the present invention and be able to implement it. However, the embodiments given are not intended to limit the present invention. Without conflict, the embodiments of the present invention and the technical features in the embodiments can be combined with each other.

[0047] The embodiments of the present invention provide a matrix multiplication optimization method and system based on the RISC-V architecture, which are used to solve the technical problem of how to effectively improve the performance of sparse matrix multiplication.

[0048] Embodiment 1:

[0049] In this embodiment, the relevant technical terms are described as follows:

[0050] Sparse matrix: A matrix in which most elements are zero. Compared with a dense matrix, a sparse matrix has higher storage efficiency and computational efficiency.

[0051] Block sparse matrix: A special form of a sparse matrix, in which the non-zero elements in the sparse matrix are organized into blocks according to certain rules. For example, a 2:4 block sparse matrix means that at most 2 non-zero elements exist in every 4 consecutive elements.

[0052] Structured sparse matrix: Usually refers to a sparse matrix with a clear pattern, which can include various sparse forms, such as block sparse matrices, list storage sparse matrices, etc.

[0053] Compressed sparse row: A commonly used storage format for sparse matrices, which stores the row information of the sparse matrix in a compressed manner. This format includes three arrays: a value array, a column index array, and a row pointer array.

[0054] Compressed sparse column: Another commonly used storage format for sparse matrices, which stores the column information of the sparse matrix in a compressed manner. This format includes three arrays: a value array, a row index array, and a column pointer array.

[0055] A matrix multiplication optimization method based on the RISC-V architecture of the present invention includes six steps: matrix construction, data processing, loop optimization, interleaved execution of instructions, custom instruction, and execution of custom instructions.

[0056] Step S100 Matrix construction: Build a machine learning model based on the application scenario, and train the machine learning model based on the collected sample data. During the model training process, build a sparse matrix adapted to the machine learning model.

[0057] Step S200 Data processing: Convert the sparse matrix into a structured sparse matrix, store the non-zero elements in the structured sparse matrix in the vector register file, and store the column indices corresponding to the non-zero elements in the scalar register file.

[0058] In this embodiment, this step converts the sparse matrix into a structured sparse matrix, such as a block sparse matrix. The structured sparse matrix has a regular data access pattern, which can reduce the index storage overhead and improve the cache hit rate. Then, the non-zero elements and the corresponding column indices of the structured sparse matrix are stored in the vector register file and the scalar register file respectively. The hybrid data placement can reduce the register name dependencies and simplify the loop unrolling.

[0059] Step S300: Loop optimization: Determine the unrolling factors of the inner loop and the outer loop according to the size of the sparse matrix, the number of vector registers, and the characteristics of the structured sparse matrix, and adjust the unrolling factor of the inner loop.

[0060] In this embodiment, this step optimizes the loop unrolling and determines the unrolling factors of the inner loop and the outer loop according to the matrix size, the number of vector registers, and the characteristics of the structured sparse matrix.

[0061] Step S400: Interleaved instruction execution: Interleave the instructions in different loop iterations through an interleaved execution method.

[0062] In this embodiment, an interleaved execution method is adopted to interleave the instructions in different loop iterations, reduce the read-after-write dependencies, reduce the register pressure, and improve the instruction utilization rate.

[0063] Step S500: Custom instruction design: Design and implement a custom vector index multiply-add instruction, which is used to replace the traditional vector load instruction and perform the vector index multiply-add operation.

[0064] In this embodiment, a custom vector index multiply-add instruction is designed and implemented. This instruction can perform the vector index multiply-add operation, reduce the number of memory accesses, and improve the data locality.

[0065] Step S600: Execute the custom instruction: Execute the custom vector index multiply-add instruction.

[0066] In this embodiment, when the custom vector index multiply-add instruction is executed, the following operations are included:

[0067] (1) Read the index value from the scalar register file;

[0068] (2) Read the corresponding vector element from the vector register file according to the index value;

[0069] (3) Multiply the read vector element by the element in another vector register, and accumulate the result of the element multiplication into the target vector register.

[0070] Based on the method disclosed in this embodiment, taking a 2:4 structured block sparse matrix as an example, the specific implementation method of the present invention is described as follows:

[0071] (1) Data preprocessing, including:

[0072] (1-1) Convert the sparse matrix into a 2:4 structured sparse matrix;

[0073] (1-2) According to the characteristics of the block sparse matrix, store the non-zero elements and column indices of each block in consecutive vector registers respectively to achieve data locality;

[0074] (2) Optimize loop unrolling, including:

[0075] (2-1) Unroll the inner loop and the outer loop, and the unrolling factors are 4 and 8 respectively;

[0076] (2-2) Due to the characteristics of the block sparse matrix, the unrolling factor of the inner loop can be adjusted according to the block size to make full use of the capacity of the vector register;

[0077] (3) Customize the vector index multiply-add instruction, including: using the customized vector index multiply-add instruction to replace the traditional vector load instruction to reduce the number of memory accesses. The index value of this instruction can be obtained through simple addition operations, thereby further improving the instruction efficiency.

[0078] As the first improved solution of this embodiment, during data processing, decompose the sparse matrix into multiple sub-matrices, allocate each sub-matrix to a core for processing, convert each sub-matrix into a structured sparse matrix, store the non-zero elements of the structured sparse matrix in the vector register file, and store the column indices corresponding to the non-zero elements in the scalar register file; correspondingly, when executing the custom instruction, the index value of the customized vector index multiply-add instruction is passed between each core through shared memory or message mechanism.

[0079] Based on the above first solution, the specific description of multi-core expansion implementation is given:

[0080] (1) Data preprocessing, including:

[0081] (1-1) Convert the sparse matrix into a structured sparse matrix;

[0082] (1-2) Decompose the sparse matrix into multiple sub-matrices, and allocate each sub-matrix to a core for processing. Each sub-matrix can adopt different structured sparse forms to adapt to different hardware resources;

[0083] (2) Optimize loop unrolling, including:

[0084] (2-1) Unroll the inner loop and the outer loop, and adjust the unrolling factor according to the matrix size and the number of vector registers;

[0085] (2-2) Each core can independently perform loop unrolling according to its own hardware resources to achieve a higher degree of parallelism;

[0086] (3) Customize the vector index multiply-add instruction, including: using the custom vector index multiply-add instruction to replace the traditional vector load instruction to reduce the number of memory accesses, and the index value of this instruction can be passed between each core through shared memory or message passing mechanism to achieve efficient parallel computing.

[0087] As the second improvement solution of this embodiment, during data processing, convert the sparse matrix into a structured sparse matrix, store the non-zero elements of the structured sparse matrix in the vector register file, store the column indexes corresponding to the non-zero elements in the scalar register file, and convert the structured sparse matrix into a form that is easy to implement in software, including compressed sparse row or compressed sparse column format; correspondingly, when executing the custom instruction, implement the custom vector index multiply-add instruction through software simulation or compiler optimization.

[0088] Based on the above improved solution two, give the specific description based on software implementation:

[0089] (1) Data preprocessing, including:

[0090] (1-1) Convert the sparse matrix into a structured sparse matrix;

[0091] (1-2) Convert the structured sparse matrix into a form that is easy to implement in software, such as compressed sparse row or compressed sparse column format;

[0092] (2) Optimize loop unrolling, including:

[0093] (2-1) Unroll the inner loop and the outer loop, and adjust the unrolling factor according to the matrix size and the number of vector registers;

[0094] (2-2) Loop unrolling can be automatically performed by the compiler or achieved through manual optimization;

[0095] (3) Customize the vector index multiply-add instruction, including: compile the code using the RISC-V GNU toolchain and add support for the custom vector index multiply-add instruction, and this instruction can be implemented through software simulation or compiler optimization.

[0096] Embodiment 2:

[0097] The matrix multiplication optimization system based on the RISC-V architecture of the present invention includes a matrix construction module, a data processing module, a loop optimization module, a custom instruction module, and an execution custom instruction module.

[0098] The matrix construction module is used to perform the following: constructing a machine learning model based on the application scenario, training the machine learning model based on the collected sample data, and constructing a sparse matrix adapted to the machine learning model during the model training process.

[0099] The data processing module is used to perform the following: converting the sparse matrix into a structured sparse matrix, storing the non-zero elements of the structured sparse matrix in the vector register file, and storing the column indices corresponding to the non-zero elements in the scalar register file.

[0100] In this embodiment, the data processing module is used to convert the sparse matrix into a structured sparse matrix, such as a block sparse matrix. The structured sparse matrix has a regular data access pattern, which can reduce the index storage overhead and improve the cache hit rate. Then, the non-zero elements and the corresponding column indices of the structured sparse matrix are respectively stored in the vector register file and the scalar register file. The mixed data placement can reduce the register name dependencies and simplify the loop unrolling.

[0101] The loop optimization module is used to perform the following: determining the unrolling factors of the inner loop and the outer loop according to the size of the sparse matrix, the number of vector registers, and the characteristics of the structured sparse matrix, and adjusting the unrolling factor of the inner loop.

[0102] In this embodiment, the loop optimization module optimizes the loop unrolling, and determines the unrolling factors of the inner loop and the outer loop according to the matrix size, the number of vector registers, and the characteristics of the structured sparse matrix.

[0103] The interleaved execution instruction module is used to perform the following: interleaving the instructions in different loop iterations in an interleaved execution manner.

[0104] In this embodiment, the interleaved execution instruction module interleaves the instructions in different loop iterations in an interleaved execution manner, reduces the read-after-write dependencies, reduces the register pressure, and improves the instruction utilization rate.

[0105] The custom instruction module is used to perform the following: designing and implementing a custom vector index multiply-add instruction, where the custom vector index multiply-add instruction is used to replace the traditional vector load instruction and perform the vector index multiply-add operation.

[0106] In this embodiment, a custom vector index multiply-add instruction is designed and implemented, which can perform the vector index multiply-add operation, reduce the number of memory accesses, and improve the data locality.

[0107] The custom instruction execution module is used to perform the following: execute a custom vector index multiply-add instruction.

[0108] In this embodiment, when the custom vector index multiply-add instruction is executed, the custom instruction execution module is used to perform the following operations:

[0109] (1) Read the index value from the scalar register file;

[0110] (2) Read the corresponding vector element from the vector register file according to the index value;

[0111] (3) Multiply the read vector element by the element in another vector register, and accumulate the result of the element multiplication into the target vector register.

[0112] The system of this embodiment can execute the method disclosed in Embodiment 1 to optimize matrix multiplication.

[0113] The above has introduced in detail the matrix multiplication optimization method and system based on the RISC-V architecture provided by the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A matrix multiplication optimization method based on the RISC-V architecture, characterized in that It includes the following steps: Matrix construction: Based on the application scenario, a machine learning model is constructed, and the machine learning model is trained based on the collected sample data. During the model training process, a sparse matrix adapted to the machine learning model is constructed; Data processing: Convert the sparse matrix into a structured sparse matrix, store the non-zero elements of the structured sparse matrix in the vector register file, and store the column indices corresponding to the non-zero elements in the scalar register file; Loop optimization: Determine the unfolding factors of the inner loop and the outer loop according to the size of the sparse matrix, the number of vector registers, and the characteristics of the structured sparse matrix, and adjust the unfolding factor of the inner loop; Interleaved execution of instructions: Interleave the instructions in different loop iterations in an interleaved execution manner; Custom instruction: Design and implement a custom vector index multiply-add instruction, which is used to replace the traditional vector load instruction and perform vector index multiply-add operations; Execute the custom instruction: Execute the custom vector index multiply-add instruction.

2. The matrix multiplication optimization method based on the RISC-V architecture according to claim 1, wherein When the custom vector index multiply-add instruction is executed, it includes the following operations: Read the index value from the scalar register file; Read the corresponding vector element from the vector register file according to the index value; Multiply the read vector element by the element in another vector register, and accumulate the result of the element multiplication into the target vector register.

3. The matrix multiplication optimization method based on the RISC-V architecture according to claim 1, wherein During data processing, decompose the sparse matrix into multiple sub-matrices, assign each sub-matrix to a core for processing, convert each sub-matrix into a structured sparse matrix, store the non-zero elements of the structured sparse matrix in the vector register file, and store the column indices corresponding to the non-zero elements in the scalar register file; Correspondingly, when the custom instruction is executed, the index value of the custom vector index multiply-add instruction is transmitted between the cores through shared memory or a message mechanism.

4. The matrix multiplication optimization method based on the RISC-V architecture according to claim 1, wherein During data processing, convert the sparse matrix into a structured sparse matrix, store the non-zero elements of the structured sparse matrix in the vector register file, store the column indices corresponding to the non-zero elements in the scalar register file, and convert the structured sparse matrix into a form that is easy to implement in software, including the compressed sparse row or compressed sparse column format; Correspondingly, when the custom instruction is executed, the custom vector index multiply-add instruction is implemented through software simulation or compiler optimization.

5. A matrix multiplication optimization system based on the RISC-V architecture, characterized in that, For matrix multiplication optimization implemented by a matrix multiplication optimization method based on the RISC-V architecture as described in any one of claims 1-4, the system includes a matrix construction module, a data processing module, a loop optimization module, a custom instruction module, and an execution custom instruction module; The matrix construction module is used to perform the following: Based on the application scenario, a machine learning model is constructed, and the machine learning model is trained based on the collected sample data. During the model training process, a sparse matrix adapted to the machine learning model is constructed; The data processing module is used to perform the following: Convert the sparse matrix into a structured sparse matrix, store the non-zero elements of the structured sparse matrix in the vector register file, and store the column indices corresponding to the non-zero elements in the scalar register file; The loop optimization module is used to perform the following: determine the unrolling factors of the inner loop and the outer loop according to the size of the sparse matrix, the number of vector registers, and the characteristics of the structured sparse matrix, and adjust the unrolling factor of the inner loop; The instruction interleaving execution module is used to perform the following: interleaving the instructions in different loop iterations by means of interleaved execution; The custom instruction module is used to perform the following: design and implement a custom vector index multiply-add instruction, which is used to replace the traditional vector load instruction and perform vector index multiply-add operations; The custom instruction execution module is used to perform the following: execute the custom vector index multiply-add instruction.

6. The matrix multiplication optimization system based on the RISC-V architecture according to claim 5, characterized in that, When the custom vector index multiply-add instruction is executed, the custom instruction execution module is used to perform the following operations: Read the index value from the scalar register file; Read the corresponding vector element from the vector register file according to the index value; Multiply the read vector element by the element in another vector register, and accumulate the result of the element multiplication into the target vector register.

7. The matrix multiplication optimization system based on the RISC-V architecture according to claim 5, wherein The data processing module is used to perform the following: decompose the sparse matrix into multiple sub-matrices, allocate each sub-matrix to a core for processing, convert each sub-matrix into a structured sparse matrix, and store the non-zero elements of the structured sparse matrix in the vector register file and store the column indices corresponding to the non-zero elements in the scalar register file; Correspondingly, when the custom instruction is executed, the index value of the custom vector index multiply-add instruction is passed between the cores through shared memory or a messaging mechanism.

8. The matrix multiplication optimization system based on the RISC-V architecture according to claim 5, wherein Data processing is used to convert the sparse matrix into a structured sparse matrix, store the non-zero elements of the structured sparse matrix in the vector register file, store the column indices corresponding to the non-zero elements in the scalar register file, and convert the structured sparse matrix into a form that is easy to implement in software, including compressed sparse row or compressed sparse column format; Correspondingly, when the custom instruction is executed, the custom instruction execution module is used to implement the custom vector index multiply-add instruction through software simulation or compiler optimization.

Citation Information

Cited By

  • Method and device for processing sparse matrix vector multiplication based on RISC-V

    CN121918880A

  • Method and apparatus for sparse matrix-vector multiplication based on RISC-V

    CN121918880B