Event-driven CSR resolver supporting multiple calculation modes
Through the event-driven CSR parser, it supports multiple computing modes and asynchronous circuit design, solving the problems of insufficient flexibility and high power consumption of existing CSR parsers, and achieving efficient and low-energy sparse matrix calculation.
Patent Information
- Application Number
- CN202510711926.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-05-29
AI Technical Summary
The existing CSR parser has a single computing mode, insufficient flexibility, low resource utilization, large power consumption, and difficult to adapt to the needs of multiple computing modes and complex systems.
The event-driven CSR parser adopts event-driven, including pointer counter, index counter, flag array generator and data matcher, combined with the coarse-grained asynchronous micropipeline structure, supports multiple computing modes, reduce power consumption through asynchronous circuit design and optimize memory access behavior.
It improves the adaptability of the computing mode and resource reuse rate, reduces dynamic energy consumption, improves system universality and execution efficiency, and avoids blocking and waiting problems in the synchronization model.
Smart Images

Figure CN120234515A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of CSR parsers, and particularly relates to an event-driven CSR parser that supports multiple computing modes. Background Art
[0002] Low-order tensors are basic data structures that represent data in the form of multi-dimensional arrays and support mathematical operations and hardware acceleration. Low-order tensor operations play an important role in scientific computing fields such as finite element analysis, computational fluid dynamics, and image processing. Matrices and vectors involved in scientific computing fields are usually irregular and sparse, posing the following challenges to the efficient calculation of low-order tensor operations: (1) efficiently accessing non-zero elements in sparse matrices and sparse vectors; (2) reducing operations on intermediate results.
[0003] Sparse matrices are usually stored in the CSR (Compressed Sparse Row) storage format to improve storage efficiency. Therefore, a CSR parser is required to parse sparse matrices and sparse vectors. The role of the CSR parser is to parse and schedule compressed-stored sparse matrices and sparse vectors in a matrix accelerator, extract the positions of non-zero elements according to the corresponding computing mode, generate executable computing tasks and distribute them to computing units, which can dynamically adapt the execution path, support data flow parallelism and asynchronous operations, and improve system throughput.
[0004] However, the existing CSR parsers have the following defects: 1. Single computing mode and lack of flexibility. Most existing CSR parsers use a certain specific operation mode, such as sparse matrix multiplication, sparse matrix-vector multiplication, etc. However, in the actual scientific computing field, a single computing mode is difficult to meet the system requirements, and often multiple computing modes need to cooperate with each other to calculate to obtain the final result. This single CSR parsing technology limits its applicability in general sparse computing tasks; 2. Low resource utilization rate and high power consumption. Most existing CSR parsers adopt a synchronous execution model, which needs to wait for the completion of the previous round of calculation or memory access before the next round of operation can be performed, easily causing computing units to be idle and memory access waiting, etc. In terms of timing, the internal of the synchronous chip needs to satisfy that the combinational logic delay is strictly less than the clock cycle; in terms of power consumption, the synchronous chip cannot meet the requirements of extremely low power consumption application scenarios. The sharp increase in chip scale makes the synchronous clock circuit design face substantial challenges brought by problems such as clock skew and clock jitter. Summary of the Invention
[0005] Aiming at the problems existing in the above background art, the purpose of the present invention is to provide an event-driven CSR parser that supports multiple computing modes.
[0006] To achieve the above purpose, the present invention adopts the following technical solutions: An event-driven CSR parser that supports multiple computing modes, including a pointer fetcher, an index fetcher, a flag bit array generator, and a data matcher, where: Pointer fetcher: Receives external start events or events sent by the data matcher, used to access the pointer data of the input matrix or vector, extracts the pointer group and selects the current start position pointer and end position pointer, outputs the matching line number currently being parsed, determines whether to update the start position pointer and end position pointer according to the relevant flag bits provided by the data matcher, and at the same time outputs a flag bit indicating whether the offset buffer is empty; when the input is a vector, the start position pointer and end position pointer are only updated once and then remain unchanged; Index fetcher: Used to extract index data from the compressed format; during the parsing process, if the current data counter is less than the end position pointer set by the pointer, it outputs the index of the current valid element and drives the generation of flag bits; otherwise, it sends a drive event to the flag bit array generator; Flag bit array generator: Used to generate the flag bit array of the compressed matrix or vector according to the index of the current valid element output by the index fetcher, and calculate the current minimum matching index; Data matcher: Used to determine whether the current non-zero element is successfully matched according to the matching situation in the flag bit array, calculate the accurate position of the current non-zero element in the compressed sparse matrix or sparse vector, and at the same time generate a control signal according to the current computing mode and the status of the matching counter to guide the pointer fetcher and the index fetcher to perform the next round of decoding. After successful matching, it sends the index of the non-zero element that needs to be calculated and is matched to the computing unit for calculation and terminates the CSR parsing stage.
[0007] Furthermore, the control circuit in the CSR parser is based on the "send-relay-receive" model and is implemented using a coarse-grained asynchronous micro-pipeline structure. The control circuit consists of three basic modules: a FIFO module, a Merge module, and a Split module.
[0008] Furthermore, the methods for the three basic modules to transmit events and data are as follows: FIFO module: After receiving the activation event, the FIFO module sequentially generates multiple pulse signals to control the gradual transmission of data. After the last pulse event is generated, after a certain delay, it sends an acknowledgment signal to the upper-level micro-pipeline and a request signal to the lower-level micro-pipeline, thereby realizing the complete relay of the event; Merge Module: The Merge module contains two control structures, ConfMerge and WaitMerge. In ConfMerge, the input events of each path are mutually exclusive. Only when an event arrives on one of the paths at a certain moment, ConfMerge will send the arriving event and data to the adjacent module. In WaitMerge, after all events of each path arrive, the events and data are sent to the next-level micro-pipeline. Split Module: The core module is the SelSplit module. The SelSplit module contains multiple valid signals of validity flags inside. When the Split module receives an activation event, only the branches with the valid signal being 1 will pass the event and data to the next-level micro-pipeline, realizing selective branching based on conditional judgment.
[0009] Furthermore, the pointer fetcher includes a first ConfMerge control structure, a second ConfMerge control structure, a first-level FIFO module, a second-level FIFO module, and a SelSplit module. After the pointer fetcher receives an external start event or an event sent by the data matcher, the event is passed through the first ConfMerge control structure to the SelSplit module for selection. If the internal offset buffer is not empty, an event is directly generated and driven to the second-level FIFO module after passing through the second ConfMerge control structure. If the internal offset buffer is empty, it drives the first-level FIFO module to generate a pulse signal, drives the generation of the address of the next set of pointers, and extracts the pointers from the storage. The event returned by the storage is passed through the second ConfMerge control structure to the second-level FIFO module, and a pulse signal is generated to drive the extracted pointers to be stored in the pointer memory. And it decides whether to update the start position pointer and the end position pointer according to the relevant flag bits provided by the data matcher, and at the same time outputs the flag bit indicating whether the buffer is empty. Finally, the event generated by the second-level FIFO module is sent to the index fetcher.
[0010] Further, the index fetcher includes a first ConfMerge control structure, a second ConfMerge control structure, a first-level FIFO module, a second-level FIFO module, and a SelSplit module. After the index fetcher receives an event sent by the pointer fetcher or the data matcher, the event is passed to the SelSplit module for selection after passing through the first ConfMerge control structure. Then, there are three cases for the transmission of the event: (1) If the indexes of the current row are not all traversed and the internal index buffer is not empty, an event is directly generated and, after passing through the first ConfMerge control structure, drives the second-level FIFO module; (2) If the indexes of the current row are not all traversed and the internal index buffer is empty, the first-level FIFO module is driven to trigger the index address generation module to generate the address of the next set of indexes and retrieve the indexes from the storage; the event returned by the storage is passed to the second-level FIFO module through the second ConfMerge control structure; the second-level FIFO module generates a pulse signal to write a set of indexes retrieved from the storage into the index buffer; during the parsing process, if the current data counter is less than the end position pointer set by the pointer, the index of the current valid element is output through the index selection module, and the flag bit array generator is driven; (3) If the elements of the current row have been traversed, the event is directly sent to the flag bit array generator.
[0011] Further, the flag bit array generator includes a FIFO module, a ConfMerge control structure, and a WaitMerge control structure. After the flag bit array generator receives an event sent by the index fetcher, the event drives the FIFO module through the ConfMerge control structure, and the events generated by the two FIFO modules are merged into one event through the WaitMerge control structure to drive the data matcher.
[0012] Further, the data matcher includes a ConfMerge control structure, a FIFO module, and a SelSplit module. After the data matcher receives an event sent by the flag bit array generator, the event drives the FIFO module through the ConfMerge control structure to generate a pulse signal to start the current round of matching, and the SelSplit module generates events to control different modules.
[0013] Further, the data matcher calculates the accurate position of the current non-zero element in the compressed sparse matrix or sparse vector according to the start position pointer of the current row and the offset of the non-zero element in the flag bit array through the element index calculation logic.
[0014] Further, the control signals generated in the data matcher include a line feed flag bit, a column change flag bit, and an element completion flag bit.
[0015] Furthermore, the information output by the CSR parser to the computing unit includes: (1) The index of the matching non-zero element in the data area of the storage space, which is used to calculate the starting position of the non-zero element; (2) The non-zero element flag bit, which is used to indicate whether the non-zero element is valid; (3) The column index of the non-zero element of the result matrix, or the row / column index of the non-zero element of the result vector; (4) The non-zero element completion flag bit, which is used to indicate whether the current non-zero element matching is completed; (5) The row completion flag bit, which is used to indicate whether a row of the result matrix is matched and completed; (6) The mode end flag bit, which is used to indicate whether all the matches in the current calculation mode are completed.
[0016] Compared with the disadvantages and deficiencies of the prior art, the present invention has the following beneficial effects: 1. The CSR parser provided by the present invention is compatible with a variety of common sparse matrix calculation modes, including matrix-vector multiplication, matrix-matrix multiplication, vector-vector multiplication, matrix-matrix addition, and vector-vector 5 calculation modes, and has good calculation mode adaptability and versatility; this CSR parser can not only extract the effective element positions and data contents in the CSR format, but also flexibly adjust the execution path of the accelerator according to different calculation modes to improve the execution efficiency; the built-in matching mechanism enables it to have memory access optimization capabilities, and can intelligently schedule access behaviors according to data requirements, only loading necessary data, thereby improving the bandwidth utilization rate; 2. The CSR parser provided by the present invention can deploy various calculations related to matrices and vectors on the same matrix computing unit at a relatively low cost, can dynamically parse and adapt to different calculation requests, and significantly improves the system versatility and resource reuse rate; 3. The CSR parser provided by the present invention adopts an asynchronous and clockless circuit design, and the modules are driven by events (pulse signals), effectively reducing the dynamic power consumption. Compared with traditional synchronous circuits, the asynchronous architecture effectively alleviates the clock skew and jitter problems, especially suitable for systems with complex structures or high uncertainties in the operating environment. At the same time, the asynchronous circuit binds data and events to achieve stable operation of modules and system performance while reducing power consumption. This CSR parser can work in parallel with the computing unit, avoiding the common blocking and waiting in the synchronous model, thereby significantly improving the overall operation efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 is a schematic diagram of the internal event flow of the event-driven CSR parser supporting multiple calculation modes provided by the embodiment of the present invention; Figure 2It is a schematic structural diagram of three basic modules that make up the control circuit provided by an embodiment of the present invention. In the figure, (a) represents the FIFO module, (b) represents the Merge module, and (c) represents the Split module; Figure 3 It is the microarchitecture of PtrFetcherA provided by an embodiment of the present invention; Figure 4 It is the microarchitecture of IdxFetcherA provided by an embodiment of the present invention; Figure 5 It is the microarchitecture of BitArray Generator provided by an embodiment of the present invention; Figure 6 It is the microarchitecture of Element Matcher provided by an embodiment of the present invention. Detailed implementation manners
[0018] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0019] The present invention provides an event-driven CSR parser that supports multiple computing modes. Based on the inner product algorithm process, it supports the multiplication operations of sparse matrices and sparse vectors. For addition operations, the CSR parser is implemented using a naive matching idea. The CSR parser aims to asynchronously decode the pointer and index information in the CSR format and drive the downstream matrix calculation unit to achieve efficient calculation.
[0020] 1. Overall architecture of the CSR parser The schematic diagram of the internal event flow of the CSR parser is as Figure 1 shown. The internal working process of the CSR parser is completed by the cooperation of four major modules: the pointer fetcher PtrFetcher, the index fetcher IdxFetcher, the flag bit array generator BitArrayGenerator, and the data matcher Element Matcher. After the CSR parser receives the start event, it drives the pointer fetcher PtrFetcher, and then the index fetcher IdxFetcher, the flag bit array generator BitArray Generator, and the data matcher Element Matcher are sequentially triggered. The data matcher Element Matcher determines how the CSR parser proceeds with the subsequent loop calculation process.
[0021] The core functions of the CSR parser are: (1) sending the indexes of the non-zero elements that are matched and need to be calculated to the row calculation of the calculation unit; (2) generating relevant flag bit information.
[0022] The CSR parser outputs the following important information to the computing unit: (1) Non-zero element A index and non-zero element B index: The indices of the matching non-zero elements in the data area of the storage space. The computing unit calculates the starting positions of the non-zero elements based on the indices; (2) Non-zero element A flag bit and non-zero element B flag bit: Used to indicate whether the non-zero elements are valid; Whether the non-zero elements are valid is one of the judgment conditions for whether the computing unit needs to fetch data; (3) Result block column index: The column index of the non-zero elements of the result matrix, or the row / column index of the non-zero elements of the result vector; (4) Non-zero element completion flag bit: Used to indicate whether the current non-zero element matching is completed; (5) Row completion flag bit: Used to indicate whether a row of the result matrix has been completely matched; When the calculation result is a matrix and the row completion flag is 1, a row of the result matrix has been completely matched; (6) Mode end flag bit: Used to indicate whether all the matching in the current calculation mode is completed.
[0023] 2. Control Circuit in the CSR Parser The control circuit in the CSR parser is implemented based on the "send-relay-receive" model and adopts a coarse-grained asynchronous micro-pipeline structure. The control circuit consists of three basic modules: the FIFO module, the Merge module, and the Split module. The schematic diagrams of the structures of the three basic modules are as Figure 2 shown. The methods for transmitting events and data are as follows: FIFO module: After receiving the activation event, it generates multiple pulse signals in sequence to control the gradual transmission of data. After the last pulse event is generated, after a certain delay, it sends an acknowledgment signal to the upper-level micro-pipeline and a request signal to the lower-level micro-pipeline, thereby realizing the complete relay of the event; The FIFO module used in the present invention for the CSR parser includes cFifo1 and cFifo2. cFifo1 generates 1 pulse signal, and cFifo2 generates 2 pulse signals; Merge module: It includes two control structures, ConfMerge and WaitMerge. In ConfMerge, the input channels are mutually exclusive. Only when one of the events arrives at the same time, ConfMerge sends the arriving event and data to the adjacent module; WaitMerge sends the event and data to the lower-level micro-pipeline after all the events have arrived; Split Module: The core module is the SelSplit module, which contains multiple validity flag signals such as valid0 and valid1. When the Split module receives an activation event, only the branches with the validity flag signal equal to 1 will pass the event and data to the next-level micro-pipeline, realizing conditional judgment-based selective branching.
[0024] 3. Working Process of CSR Parser After receiving the start event, the CSR parser drives the pointer fetcher PtrFetcher to work. The micro-architectures and working processes of PtrFetcherA and PtrFetcherB are the same, and they are responsible for accessing the pointer data of two input matrices or vectors respectively. Taking PtrFetcherA as an example, the micro-architecture of PtrFetcherA is as follows Figure 3 shown. The pointer fetcher receives an external start event or an event sent by the element matcher Element Matcher. After the event passes through the first ConfMerge2, it goes to SelSplit2 for selection. If the internal offset buffer PtrBufferA is not empty, it directly generates an event, which drives the subsequent cFifo2 after passing through the second ConfMerge2. If the internal offset buffer PtrBufferA is empty, it drives cFifo1 to generate a pulse signal, which drives the pointer address generation module GenPtrAddr to generate the address of the next set of pointers and extracts the pointers from the memory Memory. The event returned by the memory Memory is passed to the next-level cFifo2 through the second ConfMerge2, generating a pulse signal to drive the extracted pointers to be stored in the pointer storage module PtrBufferA, providing boundary conditions for subsequent index decoding. At the same time, it drives the pointer selection module PtrChooserA to decide whether to update the start position pointer and the end position pointer according to the relevant flag bits provided by the data matcher, and outputs the flag bit indicating whether the buffer is empty. When the input is a vector, PtrChooser only updates the start position pointer and the end position pointer once and then remains unchanged. Finally, the event generated by cFifo2 drives the index fetcher IdxFetcherA to start working.
[0025] The index fetcher IdxFetcher is used to extract index data from the compressed format. The micro-architectures and working processes of IdxFetcherA and IdxFetcherB are the same. Taking IdxFetcherA as an example, the micro-architecture of IdxFetcheA is as follows Figure 4As shown, after the index fetcher receives an event sent by the pointer fetcher PtrFetcherA or the data matcher ElementMatcher, the event is selected at SelSplit3 after passing through the first ConfMerge2, and there are three cases: (1) If the indexes of the current row have not been fully traversed and the internal index buffer IdxBufferA is not empty, an event is directly generated and drives the subsequent cFifo2 after passing through the second ConfMerge2; (2) If the indexes of the current row have not been fully traversed and the internal index buffer IdxBufferA is empty, cFifo1 is driven to trigger the index address generation module GenIdxAddrA to generate the address of the next set of indexes and retrieve the indexes from the memory Memory. The event returned by the memory Memory is passed to cFifo2 through the second ConfMerge2 to generate a pulse signal, and a set of indexes retrieved from the memory is written into IdxBufferA. During the parsing process, if the current data counter is less than the end position pointer set by the pointer, the index of the current valid element is output through the index selection module IdxChooserA and the flag bit array generator is driven; (3) If the elements of the current row have been fully traversed, an event is directly sent to the flag bit array generator BitArray Generator.
[0026] The microarchitecture of the flag bit array generator BitArray Generator is as Figure 5As shown, the flag bit array generator will generate the flag bit arrays of the matrix or vector according to the indexes of different current valid elements of the index fetcher respectively, and restore the positions of non-zero elements in the sparse matrix and sparse vector. Whether the indexes bound by the events sent by the IdxFetcher are valid or not, the events will drive cFifo1 through ConfMerge2. If there are indexes of valid elements in the output of the IdxFetcher, the corresponding flag bit arrays bitArrayA and bitMapB will be generated, and the current minimum matching index will be calculated at the same time. The events generated by the two cFifo1 are merged into one event by WaitMerge2 to drive the data matcher. Taking matrix × matrix as an example, the matrix on the right side of the operator will be parsed repeatedly, so a two-dimensional flag bit array bitMapB with a size of 32×32 is placed in the flag bit array generator. Therefore, for the sparse matrix and sparse vector on the right side of the operator, IdxFetcherB only needs to traverse and parse once. A flag bit array with a length of 32 is enough to record the position information of non-zero elements in the input on the left side of the operator. After the flag bit array generator generates a flag bit each time, it outputs bitArrayA and a certain row in bitMapB determined by the pointer counter. In the flag bit array generator, the farthest position that the data matcher can match forward is obtained by comparing through the minimum index judgment logic SetMinIdx. Taking matrix 𝐴 × matrix 𝐵 as an example, at the beginning of the calculation, bitArrayA and bitMapB are all 0. If the indexes of the valid elements obtained from IdxFetcherA and IdxFetcherB are 4 and 2 respectively for the first time, it means that the farthest position that can be matched in the element matcher is 2, because it is impossible to confirm whether there are non-zero elements with indexes 3 or 4 in the current column of matrix 𝐵. When the bitArrayA of the current row in matrix 𝐴 has been generated and a certain column of matrix 𝐵 is being parsed, the minimum matching index will be set to the index of the valid element obtained by IdxFetcherB.
[0027] The Element Matcher is the control center of the CSR parser and controls the overall work of the CSR parser. The microarchitecture of the Element Matcher is as Figure 6As shown in the figure. After the Element Matcher receives the event sent by the flag bit array generator, this event drives cFifo2 through ConfMerge2 to generate a pulse signal to start the current round of matching. The ElementMatcher determines whether the current two elements are successfully matched in sequence by the internal valid element matching logic ElementMatch according to the matching situation in the bitArrayA and bitMapB output by the BitArray Generator: in the multiplication mode, both the left and right inputs are required to be non-zero elements at the same position; in the addition mode, a match is successful if there is a non-zero element on either side. Subsequently, the element index calculation logic GetElementID calculates the accurate position of the current non-zero element in the compressed sparse matrix or sparse vector according to the start position pointer of the current row and the offset of the non-zero element in the flag bit array. At the same time, the parser control logic ParserControl generates control signals according to the current calculation mode and the current value of the match counter, including line break flag bits, column change flag bits, element completion flag bits, etc., to guide whether PtrFetcher and IdxFetcher enter the next round of decoding. Finally, an event for controlling different modules is generated through a SelSplit4. If the currently obtained flag bit array is not completely matched, an event is generated to drive ConfMerge2 to drive cFifo2 to generate a pulse signal, and then the next match is performed. If all the data sent this time is completely matched, but the overall two matrices are not completely matched, the data matcher will drive whether PtrFetcher or IdxFetcher enters the next round of decoding. And send the indexes of the non-zero elements that are matched and need to be calculated to the calculation unit for calculation; if the current round of matching is completed and all relevant information of the two matrices has been matched, the indexes of the non-zero elements that are matched and need to be calculated are sent to the calculation unit for calculation and the CSR parsing stage is terminated.
[0028] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.
Claims
1. An event-driven CSR parser that supports multiple computing modes, characterized in that It includes a pointer fetcher, an index fetcher, a flag bit array generator, and a data matcher, where: Pointer fetcher: Receives external start events or events sent by the data matcher, used to access the pointer data of the input matrix or vector, extracts the pointer group and selects the current start position pointer and end position pointer, outputs the currently parsed matching line number, determines whether to update the start position pointer and end position pointer according to the relevant flag bits provided by the data matcher, and at the same time outputs the flag bit indicating whether the offset buffer is empty; when the input is a vector, the start position pointer and end position pointer are only updated once and then remain unchanged; Index fetcher: Used to extract index data from the compressed format; during the parsing process, if the current data counter is less than the end position pointer set by the pointer, it outputs the index of the current valid element and drives the generation of flag bits; otherwise, it sends a drive event to the flag bit array generator; Flag bit array generator: Used to generate the flag bit array of the compressed matrix or vector according to the index of the current valid element output by the index fetcher, and calculate the current minimum matching index; Data matcher: Used to determine whether the current non-zero element is successfully matched according to the matching situation in the flag bit array, calculate the accurate position of the current non-zero element in the compressed sparse matrix or sparse vector, and generate a control signal according to the current calculation mode and the status of the matching counter to guide the pointer fetcher and the index fetcher to perform the next round of decoding. After successful matching, it sends the index of the non-zero element that needs to be calculated and is matched to the calculation unit for calculation and terminates the CSR parsing stage.
2. The event-driven CSR parser that supports multiple computing modes according to claim 1, wherein The control circuit in the CSR parser is implemented based on the "send-relay-receive" model and adopts a coarse-grained asynchronous micro-pipeline structure. The control circuit consists of three basic modules: a FIFO module, a Merge module, and a Split module.
3. The event-driven CSR parser supporting multiple computing modes according to claim 2, wherein The methods for the three basic modules to transmit events and data are as follows: FIFO module: After the FIFO module receives the activation event, it generates multiple pulse signals in sequence to control the gradual transmission of data. After the last pulse event is generated, after a certain delay, it sends an acknowledgment signal to the upper-level micro-pipeline and sends a request signal to the lower-level micro-pipeline, thus realizing the complete relay of the event; Merge module: The Merge module includes two control structures, ConfMerge and WaitMerge. In ConfMerge, the input paths are mutually exclusive. Only when one of the events arrives at the same time, ConfMerge sends the arriving event and data to the adjacent module; WaitMerge sends the event and data to the lower-level micro-pipeline after all the events arrive; Split module: The core module is the SelSplit module. The SelSplit module contains multiple valid signals of the validity flag. When the Split module receives the activation event, only the branch with the valid signal equal to 1 will transmit the event and data to the lower-level micro-pipeline, realizing the selective branch based on conditional judgment.
4. The event-driven CSR parser supporting multiple computing modes according to claim 3, characterized in that, The pointer fetcher includes a first ConfMerge control structure, a second ConfMerge control structure, a first-level FIFO module, a second-level FIFO module, and a SelSplit module. After receiving an external start event or an event sent by a data matcher, the event is passed through the first ConfMerge control structure to the SelSplit module for selection. If the internal offset buffer is not empty, an event is directly generated and driven to the second-level FIFO module after passing through the second ConfMerge control structure; if the internal offset buffer is empty, the first-level FIFO module is driven to generate a pulse signal, driving the generation of the address of the next set of pointers and fetching the pointers from the storage. The event returned by the storage is passed through the second ConfMerge control structure to the second-level FIFO module, generating a pulse signal to drive the fetched pointers to be stored in the pointer memory. And it decides whether to update the start position pointer and the end position pointer according to the relevant flag bits provided by the data matcher, and at the same time outputs the flag bit indicating whether the buffer is empty. Finally, the event generated by the second-level FIFO module is sent to the index fetcher.
5. The event-driven CSR parser that supports multiple computing modes according to claim 3, characterized in that The index fetcher includes a first ConfMerge control structure, a second ConfMerge control structure, a first-level FIFO module, a second-level FIFO module, and a SelSplit module. After receiving an event sent by the pointer fetcher or the data matcher, the event is passed through the first ConfMerge control structure and then to the SelSplit module for selection. After that, there are three cases for the event transfer: (1) If the indexes of the current row have not been fully traversed and the internal index buffer is not empty, an event is directly generated and driven to the second-level FIFO module after passing through the first ConfMerge control structure; (2) If the indexes of the current row have not been fully traversed and the internal index buffer is empty, the first-level FIFO module is driven to trigger the index address generation module to generate the address of the next set of indexes and retrieve the indexes from the storage; the event returned by the storage is passed through the second ConfMerge control structure to the second-level FIFO module; The second-level FIFO module generates a pulse signal to write a set of indexes fetched from the storage into the index buffer. During the parsing process, if the current data counter is less than the end position pointer set by the pointer, the index of the current valid element is output through the index selection module, and the flag bit array generator is driven; (3) If the elements of the current row have been fully traversed, an event is directly sent to the flag bit array generator.
6. The event-driven CSR parser supporting multiple computing modes according to claim 3, characterized in that, The flag bit array generator includes a FIFO module, a ConfMerge control structure, and a WaitMerge control structure. After receiving an event sent by the index fetcher, the event drives the FIFO module through the ConfMerge control structure, and the events generated by the two FIFO modules are merged into one event through the WaitMerge control structure to drive the data matcher.
7. The event-driven CSR parser supporting multiple computing modes according to claim 3, characterized in that, The data matcher includes a ConfMerge control structure, a FIFO module, and a SelSplit module. After receiving an event sent by the flag bit array generator, the event drives the FIFO module to generate a pulse signal through the ConfMerge control structure to start the current round of matching, and the SelSplit module generates events to control different modules.
8. The event-driven CSR parser supporting multiple computing modes according to claim 1, characterized in that, The data matcher calculates the accurate position of the current non-zero element in the compressed sparse matrix or sparse vector according to the starting position pointer of the current row and the offset of the non-zero element in the flag bit array through the element index calculation logic.
9. The event-driven CSR parser supporting multiple computing modes according to claim 1, characterized in that The control signals generated in the data matcher include a line break flag bit, a column change flag bit, and an element completion flag bit.
10. The event-driven CSR parser supporting multiple computing modes according to claim 1, characterized in that The information output by the CSR parser to the calculation unit includes: (1) The index of the matched non-zero element in the data area of the storage space, which is used to calculate the starting position of the non-zero element; (2) The non-zero element flag bit, which is used to indicate whether the non-zero element is valid; (3) The column index of the non-zero element of the result matrix, or the row / column index of the non-zero element of the result vector; (4) The non-zero element completion flag bit, which is used to indicate whether the current non-zero element has completed matching; (5) The row completion flag bit, which is used to indicate whether a row of the result matrix has completed matching; (6) The mode end flag bit, which is used to indicate whether all the matches in the current calculation mode have been completed.
Citation Information
Patent Citations
Asynchronous sparse matrix outer product multiplier based on event-driven circuit design
CN115617305A
Sparse matrix vector multiplication acceleration method and device based on non-null column storage
CN117539546A
Matrix processing method and device, electronic equipment and storage medium
CN118051264A
Data processing acceleration method and device for sparse matrix multiplication operator
CN119474631A
Sparse spiking neural network accelerator based on ping-pong architecture
WO2024216857A1