Acceleration Method, Device and Medium for Sparse Matrix Multiplication Based on Double-Layer NOC
Through the sparse matrix multiplication acceleration method based on the two-layer NOC architecture, the problems of load imbalance and data path blocking in sparse matrix operations are solved, and more efficient sparse matrix multiplication operations are achieved.
Patent Information
- Application Number
- CN202510560388.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-04-30
AI Technical Summary
The existing sparse matrix multiplication acceleration units have problems such as unbalanced load, low hardware utilization and easy blockage of data paths, making it difficult to effectively improve the computing efficiency of sparse matrix.
The sparse matrix multiplication acceleration method based on the two-layer NOC architecture is adopted, and the triple set is generated by preprocessing and encoding the sparse matrix, and the independent control path and data path are used for parallel operations. The allocation and scheduling of multiplication and addition units are realized in combination with the lazy listening mode, and the data flow and storage architecture are optimized.
It realizes faster and reliable data paths, improves hardware utilization and computing efficiency, flexibility and scalability, and optimizes overall accelerator performance.
Smart Images

Figure CN120066449B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of on-chip interconnection networks, and relates to a sparse matrix multiplication acceleration method, device, and medium based on a double-layer NOC. Background Art
[0002] Currently, in the field of modern scientific computing, the research and development of sparse matrix multiplication hardware acceleration units have received increasing attention. A sparse matrix, as a special type of matrix, is characterized by a large number of zero elements in the matrix, while the effective data elements are relatively few. This data structure characteristic makes sparse matrices play an important role in high-dimensional operations. Especially in cutting-edge technology fields such as artificial intelligence (AI) and machine learning, sparse matrix operations are an indispensable part.
[0003] The reason why sparse matrices have attracted great interest from researchers is their potential acceleration performance. Since zero elements dominate in sparse matrices, when performing matrix operations, if these zero elements can be effectively skipped and only non-zero elements are calculated, the operation efficiency can be greatly improved. However, achieving this goal is not easy. The distribution of non-zero elements in sparse matrices is often uneven, which leads to problems such as load imbalance, low hardware utilization, limited acceleration ratio, and easy blockage of data paths when traditional matrix operation acceleration units process sparse matrices.
[0004] To solve these problems, researchers have begun to explore sparse matrix multiplication acceleration units based on the Network on Chip (NoC). As an advanced interconnection architecture, NoC builds multiple nodes inside the chip and connects these nodes through high-speed data paths, thereby achieving efficient communication between various modules inside the chip. Applying NoC to sparse matrix multiplication acceleration units can make full use of its advantages such as high bandwidth, low latency, and scalability to further optimize the performance of sparse matrix operations.
[0005] In the current design of sparse matrix multiplication acceleration units, the matrix is usually divided statically, and different regions of the matrix are assigned to specific operation units for parallel processing. At the same time, a crossbar switch is used to build a data path to achieve data transmission between various operation units. However, this design method often fails to achieve an ideal acceleration effect when dealing with sparse matrices. Because the distribution of non-zero elements in sparse matrices is uneven, the static matrix division method is likely to cause some operation units to be overloaded while other operation units are idle, thus reducing the hardware utilization. In addition, the crossbar switch is also likely to become a bottleneck during data transmission, resulting in data path blockage and further affecting the acceleration performance.
[0006] Therefore, how to design a multiplication acceleration unit that can efficiently process sparse matrices has become a key issue that needs to be solved in the current scientific computing field. The NoC-based sparse matrix multiplication acceleration unit is a new hardware architecture that is expected to solve this problem.
[0007] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present application, and therefore may include information that does not constitute the prior art known to ordinary technicians in the field. Summary of the invention
[0008] In order to provide a basic understanding of some aspects of the disclosed embodiments, a brief summary is given below. The summary is not an extensive review, nor is it intended to identify key / critical elements or delineate the scope of protection of these embodiments, but rather serves as a prelude to the detailed description that follows.
[0009] The disclosed embodiments provide a sparse matrix multiplication acceleration method, device, and medium based on a double-layer NOC, which solve the defects of existing acceleration units such as unbalanced load, low hardware utilization, limited acceleration ratio, and easy blockage of data paths, and realize the allocation, scheduling, and enabling of multiplication units and addition units, thereby optimizing the overall accelerator performance.
[0010] In some embodiments, the method comprises:
[0011] Preprocessing and encoding the input sparse matrix to generate a set of triples including row coordinates, column coordinates and non-zero data, and randomly storing the set of triples in an off-chip memory;
[0012] The weight matrix is stored in an on-chip memory bank in a row-first manner, and each row memory is independently connected to the on-chip interconnect network and maintains the corresponding network coordinates;
[0013] The control path independently transmits the preemption signal and the row memory activation signal, and the data path transmits the operation data in parallel; the control path initiates the multiplication unit preemption request according to the column priority of the input matrix, and the routing controller allocates the adjacent multiplication unit to each non-zero element of the input matrix based on the proximity priority principle; the data path sends the weight data to the target multiplication unit in a multicast manner.
[0014] The multiplication unit performs the product operation of the input element and the corresponding weight row data in parallel, generates intermediate results and stores them in the output buffer according to the column index;
[0015] The intermediate results are routed to the corresponding addition units based on the input matrix row indices, and accumulation operations are performed in a lazy listening mode until the sparse matrix multiplication result is obtained.
[0016] Preferably, during the process of storing the weight matrix in the on-chip memory bank in row-major order, the storage location is determined by the row and column indices of the weight matrix; the memory bank performs data access operations with row data as the finest granularity. Each row memory contains an independent data interface to access the on-chip interconnection network, acts as a node on the on-chip interconnection network to access the network, and maintains the on-chip interconnection network coordinates of each row memory based on the physical layout.
[0017] Preferably, the allocation and release of the multiplication units are maintained by the routing controller through a lookup table. For the elements of the input matrix to be operated on, after the allocation of the multiplication units is completed, the on-chip interconnection network node coordinates where the target multiplication units are located are encapsulated into data packets and enter the data path of the on-chip interconnection network to start transmission.
[0018] Preferably, the preemption signal is transmitted prior to the row memory activation signal.
[0019] Preferably, when sending the row weight data in the multicast mode, after receiving the data packet from the on-chip interconnection network, the multiplication unit compares the target row label with the source label of the data packet and only loads the matching data packet.
[0020] Preferably, the multiplication unit starts to perform multiplication operations in parallel immediately after the elements of the input matrix and the corresponding row weight data are both ready, completes all operations in a single cycle, and stores the operation results in the output buffer of the multiplication unit.
[0021] Preferably, the addition unit maintains a passive receiving state and only activates the accumulation operation when it detects that the row label in the input data packet matches. The bit width of each addition unit is equal to the depth of the output buffer of the multiplication unit. The number of addition units is equal to the number of rows of the input matrix, and the bit width of each addition unit is consistent with the number of columns of the weight matrix.
[0022] Preferably, the working process of the addition unit is as follows:
[0023] The routing controller generates initialization signals for the corresponding number of addition units according to the row dimension of the input matrix. Each addition unit configures an accumulator with a bit width equal to the depth of the output buffer of the multiplication unit and binds the row label of the input matrix;
[0024] After the multiplication unit completes the operation, it sends the operation results containing the row label to the addition unit corresponding to the row label for accumulation. After the accumulation is completed, it sends the row data of the final result matrix to off-chip storage and releases the resources of the addition unit.
[0025] In some embodiments, the device includes a processor and a memory storing program instructions. The processor is configured to execute the above-mentioned sparse matrix multiplication acceleration method based on the double-layer NOC when running the program instructions.
[0026] In some embodiments, the storage medium stores a computer program which, when executed by a processor, implements the sparse matrix multiplication acceleration method based on a double-layer NoC.
[0027] The sparse matrix multiplication acceleration method, device, and medium based on a double-layer NoC provided by the embodiments of the present disclosure can achieve the following technical effects:
[0028] First, by introducing a double-layer NoC architecture, this solution realizes a faster and more reliable data path. Traditional sparse matrix operation acceleration units usually rely on cross switches to build data paths, and this method is prone to data path blockage problems when processing sparse matrices, thus affecting the acceleration performance. However, the double-layer NoC architecture adopted in this solution realizes the optimization of the allocation, scheduling, and enabling of multiplication units and addition units by introducing an independent fast control NoC layer, thereby effectively avoiding data path blockage problems and improving the efficiency and reliability of data transmission.
[0029] Second, this solution designs an optimized data flow and storage architecture for the characteristics of sparse matrices. Traditional sparse matrix operation acceleration units often adopt the method of statically partitioning matrices, and this method is prone to problems such as load imbalance and low hardware utilization when processing sparse matrices. However, this solution preprocesses and encodes the input matrix and stores it in off-chip storage in the form of {row coordinates, column coordinates, non-zero data}, and issues operation requests according to the non-zero elements of the input matrix in column-major order, thereby realizing the optimization of the data flow of sparse matrix operations. At the same time, this solution also stores the weight matrix in the on-chip memory bank in row-major order and sends row data to all target multiplication units through multicast, further improving the efficiency of data reading and writing and the parallelism of operations.
[0030] The NoC-based sparse matrix accelerator proposed in this solution is more flexible and scalable in hardware implementation. Since the NoC architecture has advantages such as high bandwidth, low latency, and scalability, this solution can conveniently adjust the number of NoC nodes and connection methods according to actual needs, thereby realizing the acceleration of sparse matrix operations of different scales. At the same time, this solution can also be integrated and work collaboratively with other hardware functions, further expanding its application scope.
[0031] In summary, the sparse matrix multiplication acceleration unit based on a double-layer NoC proposed by the present invention shows significant beneficial effects in terms of data processing efficiency, hardware utilization, and the flexibility and scalability of hardware implementation, providing new ideas and methods for the development of the field of sparse matrix operation acceleration.
[0032] The above general description and the following description are only exemplary and explanatory, and are not used to limit this application. Description of the Drawings
[0033] One or more embodiments are exemplarily illustrated by corresponding drawings. These exemplary illustrations and the drawings do not limit the embodiments. Elements with the same reference numerals in the drawings are shown as similar elements. The drawings do not constitute a scale limitation, and wherein:
[0034] Figure 1 is a schematic diagram of the method flow of the present invention;
[0035] Figure 2 is a schematic diagram of the overall architecture of the sparse matrix multiplication acceleration unit of the present invention;
[0036] Figure 3 is a schematic diagram of the working process of the addition unit of the present invention;
[0037] Figure 4 is a schematic diagram of the device structure of the present invention. Detailed Embodiments
[0038] In order to be able to understand the features and technical content of the embodiments of the present disclosure in more detail, the implementation of the embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. The accompanying drawings are for reference and illustration only and are not used to limit the embodiments of the present disclosure. In the following technical description, for the sake of explanation, numerous details are provided to give a full understanding of the disclosed embodiments. However, one or more embodiments can still be implemented without these details. In other cases, well-known structures and devices may be shown in a simplified manner to simplify the drawings.
[0039] In the embodiments of the present disclosure, terms such as "first" and "second" in the specification and the above-mentioned drawings are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so as to implement the embodiments of the present disclosure described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion.
[0040] Unless otherwise specified, the term "plural" means two or more.
[0041] In the embodiments of the present disclosure, the character " / " indicates that the objects before and after are in an "or" relationship. For example, A / B means: A or B.
[0042] The term "and / or" is a description of the association relationship of an object and indicates that three relationships can exist. For example, A and / or B means: A or B, or, the three relationships of A and B.
[0043] The term "corresponding" may refer to an association relationship or a binding relationship. That A corresponds to B means that there is an association relationship or a binding relationship between A and B.
[0044] As shown in Figure 1 and Figure 2 , a sparse matrix multiplication acceleration method based on a double-layer NOC includes:
[0045] S1: Preprocess and encode the input sparse matrix to generate a set of triples containing row coordinates, column coordinates, and non-zero data, and randomly store the set of triples in off-chip memory.
[0046] As a refinement of the above embodiment, the input sparse matrix will be compressed and encoded first: for each non-zero element, store the triple of its {x coordinate, y coordinate, element value}, and ignore all zero elements. The non-zero element data will be randomly shuffled and stored in off-chip storage to avoid the high hot-spot communication pressure on some data paths caused by dense storage during matrix operations, thereby making the load of the NoC more balanced.
[0047] Specifically, the steps for generating triples are as follows:
[0048] Retrieve each element of the input matrix. If the number of non-zero elements is less than the number of zero elements and it is not a special permutation element (such as a diagonal matrix, etc.), then special encoding and pre-encoding are triggered, record the row coordinate Y, column coordinate X, and the element data itself of each non-zero element, and store them in external memory such as DRAM. The traditional storage method is to store all element values in order in the storage, without including row and column coordinates, but including many elements with a value of 0.
[0049] The NoC (Network on Chip) in this embodiment is the internal NoC for sparse matrix multiplication and does not serve any other hardware functions, so it can be regarded as a reliable data path. In addition, to improve hardware efficiency, the NoC adopts a double-layer data path: a control path and a data path, which are independent of each other, and the data path has a wider bandwidth. The independent control path is used to support the allocation of multiplication and addition units and can initiate data access and loading operations in the first place, reducing system latency.
[0050] S2: Store the weight matrix in the on-chip memory bank in row-major order, and each row memory is independently connected to the on-chip interconnection network and maintains the corresponding network coordinates.
[0051] As a refinement of the above embodiment, the storage location is determined by the row and column indices of the weight matrix during the process of storing the weight matrix in the on-chip memory bank in row-major order; specifically, for the weight matrix W of size , the elements of its first row {w(0,0), w(0,1), … w(0, )} will be continuously stored in the row memory with row label 1, and the remaining row data will be stored in their respective corresponding row memories in turn.
[0052] The on-chip memory group can be accessed through address and control signals, and its types include SRAM, sDRAM or register group. The row-first storage method is to store a whole row of data in adjacent locations first. Taking the baseaddr+offset storage method of SRAM as an example, the same Baseaddr is used in storing the data of a row of the weight matrix, and a comparison table of Baseaddr and row number is maintained.
[0053] The memory group performs data access operations at the finest granularity of row data. To improve read and write efficiency, each row memory contains an independent data interface to access the on-chip interconnect network, joins the network as an on-chip interconnect network node, and maintains the on-chip interconnect network coordinates of each row memory based on the physical layout.
[0054] S3: independently transmitting the preemption signal and the row memory activation signal through the control path, and transmitting the operation data in parallel through the data path; the control path initiates a multiplication unit preemption request according to the column priority of the input matrix, and allocates adjacent multiplication units to each non-zero element of the input matrix based on the proximity priority principle through the routing controller; the data path sends the weight data to the target multiplication unit in a multicast manner.
[0055] As a refinement of the above embodiment, for each non-zero element in the input matrix, a multiplication unit will be preempted. The multiplication unit preemption logic is implemented by the routing controller in the NoC. In the , the element at the i-th row and j-th column of the product matrix is equal to the sum of the products of the i-th row of the left matrix and the j-th column of the right matrix. Therefore, any element in the x-th column of the input matrix needs to be multiplied with every element in the x-th row of the weight matrix. In order to utilize the reusability of the weight data of this row, the off-chip storage initiates a multiplication unit preemption request to the NoC in a column-first manner. The specification of the multiplication unit is 1× . After receiving the preemption request for the x-th column data of the input matrix, the routing controller follows the proximity priority principle and preferentially allocates the multipliers close to the row memory of the x-th row of the weight matrix. The distance is obtained through NoC coordinate calculation. The allocation and release of the multiplication unit is maintained by the routing controller through a lookup table. For the elements to be calculated in the input matrix, after the multiplication unit allocation is completed, the NoC node coordinates where the target multiplication unit is located will be encapsulated into the data packet and enter the NoC data path to start transmission. At the same time, the preemption signal Occupy used to preempt the calculation unit and the enable signal Read_Row used to activate the corresponding row memory will enter the NoC control path to be transmitted at a faster speed and with lower blocking expectation. Occupy is transmitted before the Read_Row signal.
[0056] Among them, the row and column labels of the input elements in the input matrix are transmitted simultaneously with the Occupy signal. The NoC labels of all the multiplication units that require the weight data of this row in the current matrix sparse operation are transmitted simultaneously with Read_Row. When the row memory receives the Read_Row enable signal, it reads out the row weight data stored in the memory and sends the row data to all target multiplication units through multicast. Due to the proximity-first allocation principle of the routing controller, the target multiplication units are always clustered near the row memory, reducing the probability of path congestion and making the operands ready as soon as possible. After receiving the data packet from the NoC, the multiplication unit activated by Read_Row compares the target row label with the data packet source label. If they match, the data packet is unpacked and the row data is loaded; if they do not match, the current data packet is ignored.
[0057] It should be noted that in this embodiment, the NoC adopts a two-layer NoC architecture. The traditional NoC architecture usually has only one physical link. Even if multiple virtual channel FIFOs are designed, each FIFO needs to follow the arbitration logic to competitively use the physical link. Then, the data packet containing specific calculation data is usually larger, has a lower priority, and is more likely to be blocked compared to the data packet containing control signals / address signals. This solution designs two parallel and non-interfering physical links, namely, a control link with a relatively narrow bit width and a higher priority for transmitting address and control signals, and a data link with a relatively wide bit width. Both layers of the on-chip network have their own independent routers and actual circuit connections, and can perform data transmission work in parallel. The multiplier preemption signal and the row memory activation signal belong to control signals, and the effective bit width is 1. After being encapsulated by the data packet of the NoC network interface, they enter the control link for transmission.
[0058] S4: In the multiplication unit, the product operation of the input element and the corresponding weight row data is executed in parallel to generate an intermediate result and store it in the output buffer according to the column index.
[0059] As a refinement of the above embodiment, the multiplication unit immediately starts parallel multiplication operations as soon as the input matrix element and the corresponding weight row data are both ready, completes all operations in a single cycle, and stores the operation results in the output buffer of the multiplication unit. The depth of the output buffer is , and is indexed by the column label.
[0060] Since the input matrix is sparse, this embodiment only focuses on the non-zero elements thereof. Therefore, the multiplication with the weight row is a multiplication operation between a scalar and a vector. The x-th element in a certain row of the left matrix only interacts with the data in the x-th row of the right matrix, which is determined by the operation principle of matrix multiplication. The design of this part of the data link is to implement the multiplication operation between the non-zero elements and the corresponding rows of the weight matrix in a multiplication unit, improve the reusability of the weight matrix data, and at the same time, each row can perform its respective operations in parallel to improve efficiency. The specific multiplication operation depends on the multiplier design in the design, and it may be a signed / unsigned floating-point / scalar or other format multiplier. The output cache can be understood by analogy with the row memory on the weight matrix side and can be SRAM, sDRAM or register bank,
[0061] S5: Route the intermediate result to the corresponding adder unit based on the row label of the input matrix, and perform the accumulation operation through the lazy listening mode until the result of the sparse matrix multiplication is obtained.
[0062] As a refinement of the above embodiment, while processing the preemption of the multiplication unit and the reading of the row memory, the routing controller also sends an Accum adder unit initialization signal, and a total of adder units are initialized. The bit width of the adder unit is . Each adder unit contains the row label information of the input matrix and adopts the lazy listening working mode, that is, it does not actively initiate a read request, but passively receives the input data and performs accumulation and caching. The working process is as Figure 2 shown: When processing the calculation request for the j-th column of the input matrix, m multiplication units will be preempted for operation, where m is the number of non-zero elements in this column. Each multiplication unit maintains the matrix label information (i, j) of the corresponding element, where i < . After the multiplication operation unit with the matrix label (i, j) completes the operation, the operation result of 1×1× will be directly sent to the adder unit corresponding to the row label i for accumulation, and the multiplication unit will be released. And so on. When the operation results of all multiplication units are loaded into the corresponding adder units and the accumulation is completed, the operation result in each adder unit is the element value of each row in the final result matrix. The sparse matrix calculation is completed, and the result is sent to the off-chip storage and then all adder units are released, and the entire process is completed.
[0063] The line numbers in this step are the line numbers where the scalars are located in step S4. In this embodiment, the matrix multiplication of A×B = C is split. The value of the (x,y) element in C is the result of the dot product of the x-th row of A and the y-th column of B. The operation performed by the multiplier actually only completes the operation of one partial product in the dot product at a time, such as AB0. The other partial products such as A1B1 and A2B2 are completed by the parallel operations of other multipliers. It is necessary to enable the output buffer of the multiplier through the line number, and then transfer the partial products to the adder to perform addition.
[0064] Lazy listening means that there is a master-slave relationship between the multiplier corresponding to the line number and the adder. The data is actively sent by the multiplier and sent immediately after the operation is completed to release the multiplier resources. The adder receives passively and immediately performs an accumulation operation after receiving.
[0065] It should be noted that: the sparse matrix accelerator based on NoC proposed in this solution uses the NoC architecture to achieve a more reliable and faster data path, and at the same time uses a double-layer NoC architecture to further prevent the delay and performance degradation caused by data blocking. In addition, the data flow and storage architecture optimized specifically for sparse matrices have the advantages of higher reliability, higher acceleration, and lower cost compared with existing designs.
[0066] Combined Figure 4 As shown, the present disclosure provides a sparse matrix multiplication acceleration device 300 based on a double-layer NOC, including a processor 304 and a memory 301. Optionally, the device may further include a communication interface 302 and a bus 303. Among them, the processor 304, the communication interface 302, and the memory 301 can complete mutual communication through the bus 303. The communication interface 302 can be used for information transmission. The processor 304 can call the logical instructions in the memory 301 to execute the sparse matrix multiplication acceleration method based on the double-layer NOC in the above embodiment.
[0067] In addition, when the logical instructions in the above-mentioned memory 301 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium.
[0068] The memory 301, as a computer-readable storage medium, can be used to store software programs and computer-executable programs, such as the program instructions / modules corresponding to the methods in the embodiments of the present disclosure. The processor 304 executes functional applications and data processing by running the program instructions / modules stored in the memory 301, that is, implements the sparse matrix multiplication acceleration method based on the double-layer NOC in the above embodiment.
[0069] The memory 301 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the terminal device, etc. In addition, the memory 301 may include a high-speed random access memory and may also include a non-volatile memory.
[0070] Embodiments of the present disclosure provide a computer-readable storage medium storing computer-executable instructions, and the computer-executable instructions are configured to execute the above-mentioned sparse matrix multiplication acceleration method based on a double-layer NOC.
[0071] The above-mentioned computer-readable storage medium may be a transient computer-readable storage medium or a non-transient computer-readable storage medium.
[0072] The technical solution of the embodiments of the present disclosure may be embodied in the form of a software product. This computer software product is stored in a storage medium and includes one or more instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in the embodiments of the present disclosure. The foregoing storage medium may be a non-transient storage medium, including: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs, or may also be a transient storage medium.
[0073] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art may still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A sparse matrix multiplication acceleration method based on a double-layer NOC, characterized in that It includes the following steps: Preprocess and encode the input sparse matrix to generate a set of triples containing row coordinates, column coordinates, and non-zero data, and randomly store the set of triples in off-chip memory; Store the weight matrix in the on-chip memory bank in row-major order, and each row memory independently accesses the on-chip interconnect network and maintains the corresponding network coordinates; Independently transmit the preemption signal and the row memory activation signal through the control path, and simultaneously transmit the operation data in parallel through the data path; the control path initiates a multiplication unit preemption request according to the column-major order of the input matrix, and assigns adjacent multiplication units to each non-zero element of the input matrix based on the proximity priority principle by the routing controller; The data path sends the row data of the weight data to the target multiplication unit by multicast; Parallelly execute the product operation of the input element and the corresponding weight row data in the multiplication unit, generate the intermediate result and store it in the output buffer according to the column index; Route the intermediate result to the corresponding adder unit based on the row label of the input matrix, and perform the accumulation operation through the lazy listening mode until the sparse matrix multiplication result is obtained.
2. The sparse matrix multiplication acceleration method based on a double-layer NOC according to claim 1, wherein Determine the storage location through the row and column indexes of the weight matrix during the process of storing the weight matrix in the on-chip memory bank in row-major order; the memory bank performs data access operations with the row data as the finest granularity, and each row memory contains an independent data interface to access the on-chip interconnect network, accesses the network as an on-chip interconnect network node, and maintains the on-chip interconnect network coordinates of each row memory based on the physical layout.
3. The sparse matrix multiplication acceleration method based on a double-layer NOC according to claim 1, wherein The allocation and release of the multiplication unit are maintained by the routing controller through a lookup table. For the elements to be operated on in the input matrix, after the multiplication unit allocation is completed, the on-chip interconnect network node coordinates where the target multiplication unit is located will be encapsulated into the data packet and enter the on-chip interconnect network data path for transmission.
4. The sparse matrix multiplication acceleration method based on a double-layer NOC according to claim 1, wherein The preemption signal is transmitted prior to the row memory activation signal.
5. The sparse matrix multiplication acceleration method based on a double-layer NOC according to claim 1, characterized in that When sending the row weight data in the multicast mode, after receiving the data packet from the on-chip interconnect network, the multiplication unit compares the target row label and the data packet source label, and only loads the matching data packet.
6. The sparse matrix multiplication acceleration method based on a double-layer NOC according to claim 5, characterized in that, The multiplication unit immediately starts parallel multiplication operations as soon as the input matrix element and the corresponding weight row data are both ready, completes all operations in a single cycle, and stores the operation result in the output buffer of the multiplication unit.
7. The sparse matrix multiplication acceleration method based on a double-layer NOC according to claim 1, wherein The adder unit maintains a passive receiving state, activates the accumulation operation only when it detects that the row label in the input data packet matches, and the bit width of each adder unit is equal to the depth of the output buffer of the multiplication unit. The number of adder units is equal to the number of rows of the input matrix, and the bit width of each adder unit is consistent with the number of columns of the weight matrix.
8. The sparse matrix multiplication acceleration method based on a double-layer NOC according to claim 7, characterized in that The working process of the adder unit is as follows: The routing controller generates initialization signals for the corresponding number of adder units according to the row dimension of the input matrix. Each adder unit configures an accumulator with a bit width equal to the depth of the output buffer of the multiplication unit and binds the row label of the input matrix; After the multiplication unit completes the operation, send the operation result containing the row label to the adder unit corresponding to the row label for accumulation. After the accumulation is completed, send the row data of the final result matrix to off-chip storage and release the adder unit resources.
9. A sparse matrix multiplication acceleration device based on a double-layer NOC, comprising a processor and a memory storing program instructions, characterized in that, The processor is configured to execute the sparse matrix multiplication acceleration method based on a double-layer NOC according to any one of claims 1-8 when running the program instructions.
10. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and when the program is executed by a processor, it implements the sparse matrix multiplication acceleration method based on a double-layer NOC according to any one of the above claims 1-8.
Citation Information
Patent Citations
Transform model irregular sparse matrix multiplication method and hardware architecture
CN115357850A
Graph convolutional neural network acceleration method and apparatus, and electronic device
CN116451755A