Implementing N: M sparsity in compute accelerators in digital memory
By using the FlexCiM design, the DCiM macro is divided into sub-macros and buffers and multiplexers are introduced, which solves the area and efficiency problems of the DCiM accelerator when adapting to structured sparsity patterns, and achieves high computational throughput and reduced energy consumption.
Patent Information
- Application Number
- CN202511428247.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-08-22
- Filing Date
- 2025-09-30
- Publication Date
- 2026-05-15
AI Technical Summary
Existing digital in-memory computing (DCiM) accelerators struggle to flexibly adapt to different structured sparsity patterns, leading to increased area overhead and reduced computational efficiency, particularly when implementing N:M sparsity patterns.
The FlexCiM design divides the DCiM macro into multiple sub-macros and introduces input activation buffers, distribution networks, and merging networks. Flexible N:M sparsity modes are achieved through 2-1 multiplexers and P-1 multiplexers, supporting multiple sparsity ratios and reducing area overhead.
It achieves higher computational throughput, throughput per watt, and area efficiency, significantly reduces inference latency and energy consumption, while maintaining model accuracy and improving computational efficiency.
Smart Images

Figure CN122047327A_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This application claims priority to and / or benefits from U.S. Provisional Patent Application No. 63 / 720,263, filed November 14, 2024, entitled "In-Memory Computing Architecture for Accelerating Neural Network Operations," which is incorporated herein by reference in its entirety. Background Technology
[0003] Digital in-memory computing (DCiM) performs data processing directly within or near memory cells in the analog domain, rather than transferring data back and forth between separate memory and processing units. DCiM implementations can employ non-volatile memory technologies and digital circuits that enable digital computation to be performed directly within memory arrays. Attached Figure Description
[0004] The embodiments will be readily understood from the following detailed description taken in conjunction with the accompanying drawings. For ease of description, similar reference numerals designate similar structural elements. In the accompanying drawings, embodiments are illustrated by way of example rather than limitation.
[0005] Figure 1A It is a digital accelerator that implements the von Neumann architecture according to some embodiments of this disclosure.
[0006] Figure 1B It is a digital accelerator that implements the DCiM architecture according to some embodiments of the present disclosure.
[0007] Figure 2 The illustration shows a digital accelerator with a DCiM macro partitioned into P sub-macros according to some embodiments of the present disclosure.
[0008] Figure 3 The illustration shows a column of submacros according to some embodiments of the present disclosure.
[0009] Figure 4 The illustration shows an in-memory computed cell with a 2-1 (2:1) multiplexer for a submacro according to some embodiments of the present disclosure.
[0010] Figure 5 The illustration shows an implementation of an allocation network including a P number of P-1 (P:1) multiplexers according to some embodiments of the present disclosure.
[0011] Figure 6 The illustration shows an implementation of a merging network including an adder tree according to some embodiments of the present disclosure.
[0012] Figures 7A-7B The illustration shows a demonstrative example of implementing a 1:4 sparsity pattern according to some embodiments of the present disclosure.
[0013] Figures 8A-8B The illustration shows a demonstrative example of implementing the 4:8 sparsity pattern according to some embodiments of the present disclosure.
[0014] Figures 9A-9C The illustrations depict different optimal sparsity patterns according to some embodiments of the present disclosure for different outlier presences and outlier distributions.
[0015] Figure 10 The illustration shows a system for optimizing N:M sparsity patterns and generating sparse neural networks according to some embodiments of the present disclosure.
[0016] Figure 11 This is a flowchart illustrating a method for determining the count of outliers and measuring the locality of outliers according to some embodiments of the present disclosure.
[0017] Figure 12 This is a flowchart illustrating a method for determining an optimal N:M sparsity pattern according to some embodiments of the present disclosure.
[0018] Figure 13 This is a flowchart illustrating a method for determining an optimal N:M sparsity pattern according to some embodiments of the present disclosure.
[0019] Figure 14 This is a flowchart illustrating a method for accelerating the multiplication and accumulation operations of input activations and weights by utilizing an N:M sparsity pattern, according to some embodiments of the present disclosure.
[0020] Figure 15 This is a block diagram of an exemplary computing device according to some embodiments of the present disclosure. Detailed Implementation
[0021] Overview
[0022] Recent advances in deep neural network (DNN) research have shown that many DNNs are highly overparameterized, allowing for important parameter pruning. Two types of sparsity exist: unstructured sparsity and structured sparsity. Unstructured sparsity refers to the random removal of weights in a neural network, resulting in an irregular distribution of zero elements within the model. While this method achieves high sparsity and minimal accuracy loss, its unpredictability poses a challenge to effective hardware acceleration, often slowing down computation. In contrast, structured sparsity removes weights using regular, predefined sparsity patterns, such as eliminating the entire channel or following a specific N:M ratio. Structured sparsity allows for more efficient hardware accelerators and simplifies processing logic, thus avoiding the bottlenecks inherent in unstructured methods. Specifically, an N:M sparsity pattern means that only N elements in every M groups are non-zero. In other words, N:M refers to a pattern where only N parameters out of every M consecutive parameters in the network are non-zero. While unstructured sparsity leads to irregular computations and hardware inefficiencies, structured sparsity methods have gained attention for their balance between accuracy and speedup efficiency, as seen in some accelerators that support specific sparsity patterns such as 1:2, 2:4, and 4:8. Specifically, structured sparsity ensures a fixed sparsity ratio or number of zeros (or the number of non-zero parameters) within a block or window of parameters. The constrained selection window can be implemented with low-overhead multiplexing logic, resulting in significant benefits.
[0023] While some DCiM implementations can support a fixed structured sparsity ratio, adapting to flexible structured sparsity presents significant challenges. Specifically, for DCiM-based accelerators, implementing N:M sparsity patterns is not straightforward, as integrating numerous multiplexers within each memory cell to accommodate different sparsity patterns would significantly increase area overhead and compromise the regular architecture of DCiM macros.
[0024] To address this issue, the FlexCiM design (Flexible DCiM Architecture) can be implemented in DCiM-based accelerators. The FlexCiM design provides a low-overhead, flexible, structured sparsity DCiM accelerator that achieves significantly higher computational throughput, throughput per watt, and area efficiency compared to dense or unstructured sparsity-based digital accelerators and other DCiM-based accelerators. In some embodiments, FlexCiM supports multiple N:M sparsity ratios in INT8 mode of operation.
[0025] By supporting a flexible N:M sparsity pattern, FlexCiM accelerates multiplication and accumulation (MAC) operations in hardware accelerators used for neural networks. FlexCiM divides a DCiM macro into multiple sub-macros based on a partitioning factor P. A complete DCiM macro can have a grid or array of in-memory compute cells with dimensions X multiplied by Y (e.g., an in-memory compute cell with X rows and Y columns). In other words, a DCiM macro is subdivided into P number of sub-macros. Each sub-macro has a grid of in-memory compute cells with dimensions X divided by P (X / P) multiplied by Y (e.g., an in-memory compute cell with X / P rows and Y columns). An in-memory compute cell can have B number of memory elements to store B number of bits, such as B bit weights. The in-memory compute cells are equipped with 2-1 multiplexers. The in-memory compute cells enable bit-sequential multiplication of weights and input activations. The 2-1 multiplexers allow the in-memory compute cells to select between input activation pairs to achieve 1:2 sparsity. The multiplication results generated by the columns of the in-memory computed cell in the submacro can be summed together to form a partial sum. Taking full advantage of the partition design, submacros that achieve a baseline 1:2 sparsity ratio can be combined to support different N:M sparsity modes, and can also operate in a fully dense M:M sparsity mode.
[0026] To support a flexible N:M sparsity pattern, an input activation buffer is included to buffer input activations according to N and M of the sparsity pattern, and an allocation network is implemented to direct the pairs of input activations received from the input activation buffer to the appropriate submacros. The allocation network can include P number of P-1 multiplexers, each having P outputs to a number of P submacros. Specifically, the M value in the N:M sparsity pattern indicates the number of input activations to be allocated by each P-1 multiplexer. The N value in the N:M sparsity pattern indicates the number of submacros that are aggregated together to process the same block M. If N > 1, multiple P-1 multiplexers will receive the same set of input activations. In other words, P number of P-1 multiplexers can receive N pairs of the same set of input activations.
[0027] A P-1 multiplexer can have P inputs, each of which can receive two input activation words or a pair of input activations. The P-1 multiplexer can selectively route a pair of input activations to the corresponding submacro. The P-1 multiplexer can select a pair of input activations at one of the P inputs. The selection signal to the P-1 multiplexer can be generated based on metadata encoded with coordinates of non-zero weights. A 2-1 multiplexer for in-memory computed cells can selectively select one of the input activations received from a pair of input activations from the P-1 multiplexer and perform a multiplication of the selected input activation with weights stored in the in-memory computed cell. The selection signal to the 2-1 multiplexer can be generated based on metadata encoded with coordinates of non-zero weights.
[0028] Compared to dense DCiM accelerators, FlexCiM accelerators achieve significant improvements in computational throughput, throughput per watt, and area efficiency for the LLaMA3-8B model. Compared to other sparse accelerators, FlexCiM offers up to 1.75x lower inference latency and 1.5x lower power consumption.
[0029] Algorithms can prune weights with minimal performance degradation. However, these algorithms have several drawbacks. Some algorithms assign a consistent sparsity pattern to all layers of the neural network, which may be suboptimal for some neural networks because outlier features can change from one layer to another. Another algorithm examines the number of outliers in the layer and the different changes in the value of N when assigning an N:M sparsity pattern. This algorithm has the constraint that the accuracy of the model implemented at high sparsity ratios is limited.
[0030] To address this issue, the FLOW framework (a flexible, layer-by-layer out-of-point density-aware algorithm) can be implemented to determine the optimal N:M sparsity pattern for a given layer. The value A in the sparsity ratio A / B is determined based on the number of out-of-points in the layer, and the value B in the sparsity ratio A / B is determined based on a locality measurement of the out-of-points representing their spatial distribution. This locality measurement indicates how close or far apart the identified out-of-points are distributed within the layer. Furthermore, the optimal N:M sparsity pattern, consistent with the determined sparsity ratio A / B, can be selected based on whether latency or accuracy is prioritized, or a balance is struck between the two.
[0031] Compared to other frameworks, FLOW, while considering model performance on the FlexCiM architecture, results in significantly better pruning models with minimal accuracy reduction. FLOW outperforms other frameworks by 36% in accuracy.
[0032] In this paper, the sparsity ratio A / B refers to the presence of A non-zero elements out of B elements. The structured sparsity pattern N:M refers to the presence of N non-zero elements in a block of M consecutive elements.
[0033] Comparison of von Neumann architecture and DCiM architecture
[0034] Figure 1A This is a digital accelerator 102 implementing a von Neumann architecture according to some embodiments of the present disclosure. The digital accelerator 102 (shown as an "on-chip" region) is connectable to external memory 104 via memory interface 106. The digital accelerator 102 may include a grid 120 of processing elements (PEs), such as an array of PEs.
[0035] Memory interface 106 manages the data flow between components in digital accelerator 102 and external memory 104. External memory 104 serves as main memory for model parameters and input data, and provides weights and input activations to digital accelerator 102. Memory interface 106 bridges external memory 104 with on-chip memories / buffers, such as input activation buffer 108, weight buffer 110, and output activation buffer 112. Memory interface 106 handles data acquisition and synchronization, and ensures that weights and input activations are delivered to the appropriate on-chip memories / buffers.
[0036] Weight buffer 110 stores weights retrieved from external memory 104. Input activation buffer 108 stores input activations retrieved from external memory 104. Weight buffer 110 provides weights to the grid 120 of the PE, and input activation buffer 108 provides input activations to the grid 120 of the PE. The grid 120 of the PE can perform multiplication of input activations and weights independently, and multiple PEs can be interconnected to allow the multiplication results to be accumulated. Output activation buffer 112 collects and stores output activations generated by the grid 120 of the PE. Output activation buffer 112 can write output activations to external memory 104 via memory interface 106.
[0037] The decoding phase of large language model inference is memory-constrained, with loading weights from the weight buffer 110 to the grid 120 of the PE being a significant bottleneck. The back-and-forth transfer of weights from the weight buffer 110 to the grid 120 consumes a large amount of energy and time.
[0038] Figure 1BThis is a digital accelerator implementing a DCiM architecture according to some embodiments of the present disclosure. Digital accelerator 132 (shown as an "on-chip" region) is capable of being connected to external memory 130 via memory interface 136. Digital accelerator 132 may include a DCiM macro 134, which includes a grid of in-memory compute cells (shown as "SRAM cells") and an adder tree summing the outputs from columns of in-memory compute cells. SRAM stands for Static Random Access Memory.
[0039] Memory interface 136 manages the data flow between components in digital accelerator 132 and external memory 104. External memory 130 serves as main memory for model parameters and input data, and provides weights and input activations to digital accelerator 132. Memory interface 136 bridges external memory 130 with on-chip memories / buffers, such as input activation buffer 138 and output activation buffer 140. Memory interface 136 handles data acquisition and synchronization, and input activations are transferred to input activation buffer 138. Memory interface 136 bridges external memory 130 with in-memory compute cells in DCiM macro 134. Memory interface 136 can directly load weights into memory / memory elements within the in-memory compute cells.
[0040] Input activation buffer 138 can store input activations obtained from external memory 130. Input activation buffer 138 can provide input activations to the in-memory compute cells of DCiM macro 134. Input activation buffer 138 can be positioned close to digital accelerator 132 to reduce latency. The in-memory compute cells can perform multiplication of input activations and weights independently. Multiplication can be performed via bit-serial multiplication or bit-parallel multiplication. Advantageously, computation occurs directly within the in-memory compute cells, and there is no need to transfer weights back and forth from the on-chip weight buffer to the PE. An adder tree of columns for each in-memory compute cell can accumulate the multiplication results of the in-memory compute cells. Output activation buffer 140 can collect and store output activations generated by DCiM macro 134. Output activation buffer 140 can write output activations to external memory 130 via memory interface 136. Output activation buffer 140 can include a register file for storing output activations. Using the register file can reduce the output activation throughput to memory interface 136.
[0041] FlexCiM architecture
[0042] Supporting flexible N:M sparsity in the DCiM architecture is not easy. The FlexCiM architecture can support a variety of N:M sparsity patterns (e.g., 1:1, 1:2, 1:4, 1:8, 1:16, 2:4, 2:8, 2:16, 4:8, 4:16, 8:16, etc.) while retaining the energy efficiency and low data movement costs of the CiM architecture, and adding only a small amount of area overhead. In some embodiments, N and M are powers of 2. Supporting flexible N:M sparsity allows the sparsity ratio and pattern to be adjusted for a given neural network and for a given layer, thereby balancing model accuracy and latency.
[0043] Figure 2 The illustration shows a digital accelerator 280 with a DCiM macro 202 partitioned into P sub-macros 204, according to some embodiments of the present disclosure. The digital accelerator 280 can include an integrated circuit that accelerates MAC operations of input activation and weights in an N:M sparse pattern. The N:M sparse pattern can be flexible, meaning that the digital accelerator 280 can support multiple N:M sparse patterns.
[0044] Digital accelerator 280 may include one or more instances of DCiM macro 202. To utilize DCiM macro 202, digital accelerator 280 may include an input activation buffer 208, an allocation network 210, and a merging network 216. Digital accelerator 280 may also include a memory interface 136 and an output activation buffer 140. Memory interface 136 may be connected to external memory 130.
[0045] The DCiM macro 202 can have in-memory compute cells with X rows and Y columns. For example, the DCiM macro 202 can have dimensions X×Y×B, where X and Y represent the crossbar switch array dimensions and B represents the memory word length. Exemplary values for B can be 1, 2, 4, 6, 8, 12, or 16. In one example, the DCiM macro 202 can have in-memory compute cells with X = 128 rows and Y = 32 columns, where each in-memory compute cell stores B = 8-bit memory words (e.g., 8-bit weights). The DCiM macro 202 can be divided into P number of submacros 204 or arranged in P number of submacros 204.
[0046] P refers to the partitioning factor. In some examples in this paper, P = 4, and therefore DCiM macro 202 comprises P = 4 submacros. DCiM macro 202 is subdivided along the row dimension, and therefore DCiM macro 202 is divided into P number of X / P×Y×B submacros. Submacros in the P number of submacros 204 can have an in-memory compute cell of X divided by P (X / P) rows and Y columns, for example, the submacro can have a dimension X / P×Y×B. In one example, submacros in the P number of submacros 204 can have an in-memory compute cell of X / P = 128 / 4 = 32 rows and Y = 32 columns, and wherein each in-memory compute cell stores B = 8-bit memory words (e.g., 8-bit weights).
[0047] In the example in this article, P divides DCiM macro 202 into submacros along the row dimension because DCiM macro 202 is summed along the column dimension. Imagine that if DCiM macro 202 is summed along the row dimension, then P can equivalently divide DCiM macro 202 into submacros along the column dimension.
[0048] exist Figure 4 The in-memory compute cell, described in more detail, can include a 2-1 (2:1) multiplexer that selects one of two input activations and multiplies it with a weight stored in the in-memory compute cell.
[0049] Naturally, in-memory computation cells with 2-1 multiplexers support 1:2 structured sparsity. 2-1 multiplexers utilize a small transistor count and use only 1 bit of metadata storage for controlling the 2-1 multiplexer. To support flexible N:M sparsity, one or more components (e.g., input activation buffer 208, allocation network 210, merging network 216, and controller 206) are introduced to orchestrate the computation of P number of submacros 204. Specifically, P number of submacros 204 can be grouped and activated together according to an N:M sparsity pattern. Appropriate input activations can be buffered by input activation buffer 208 to allocation network 210 according to the N:M sparsity pattern. Appropriate input activations can be selected by allocation network 210 and sent to the appropriate submacro according to the N:M sparsity pattern.
[0050] P-number submacros 204 can compute partial sums along the columns of submacro 204. P-number submacros 204 can include partial sum buffers 212 for storing the partial sums computed by the individual columns of the submacros. The partial sums generated by P-number submacros 204 (e.g., stored in the corresponding partial sum buffers 212) can be appropriately added by merging network 216 according to the N:M sparsity pattern. Merging network 216 in Figure 6 It is described in more detail in the middle.
[0051] Each of the P-number submacros 204 is functionally independent of each other to support 1:2 sparsity, while the input activation buffer 208 and the allocation network 210 are able to orchestrate multiple neighboring submacros in the P-number submacros 204 to operate together, thereby supporting different N:M sparsity patterns.
[0052] Input activation buffer 208 can buffer input activations according to N and M. The input activations buffered by input activation buffer 208 can be arranged in different ways according to N and M. Input activation buffer 208 is responsible for feeding or streaming input activations to distribution network 210 according to N and M. Distribution network 210 can receive input activations from input activation buffer 208. Distribution network 210 can select appropriate input activations to be fed to P number of submacros 204. Merging network 216 can sum P partial sums calculated from P number of submacros 204. The summation that can be performed by merging network 216 can be based on which submacros in P number of submacros 204 are activated to support the N:M sparsity pattern.
[0053] The distribution network 210 can include a P number of P-1 multiplexers, each having P outputs respectively to a P number of submacros 204. Figure 5 The P-number P-1 multiplexers described in more detail can be provided for rows at the same row position in the P-number submacros 204 (e.g., rows that are spatially identical to each other within the P-number submacros 204). For example, rows at position 0 in the P-number submacros 204 can share the P-number P-1 multiplexers. In other words, the allocation network 210 can have X / P sets of P-number P-1 multiplexers, for example, each set of P-number P-1 multiplexers for each spatially identical row in the submacros 204. Figure 5 As shown in the diagram, the P-1 multiplexer has P inputs and an output for outputting two selected input activation words (or a selection pair of input activation words). The inputs receive two input activation words (or a selection pair of input activation words).
[0054] Memory interface 136 can be connected to external memory 130 to write input activations to input activation buffer 208. Memory interface 136 can be connected to DCiM macro 202 to write weights to the in-memory computed cell of DCiM macro 202. Memory interface 136 can be connected to output activation buffer 140 to write output activations to external memory 130.
[0055] The digital accelerator 280 may include a controller 206. The controller 206 may receive metadata encoded with coordinates of non-zero or dense weights from external memory 130 via memory interface 136, and generate selection signals for allocating network 210 based on the metadata. In some cases, the selection signals for allocating network 210 are generated based on the sparsity patterns N and M.
[0056] In some embodiments, controller 206 may generate selection signals for 2-1 multiplexers in a number of P submacros based on metadata. In some cases, the selection signals for 2-1 multiplexers in a number of P submacros are generated based on the sparsity patterns N and M.
[0057] A submacro within a submacro of number P can contain Y columns. The columns of the submacro (e.g., column 214) are in... Figure 3 It is described in more detail in the middle.
[0058] The FlexCiM design is modular, and any value of P can be used. The higher the value of P, the higher the value of M in the N:M sparsity pattern can be supported.
[0059] In one implementation, P = 4. The DCiM macro 202 can include four sub-macros 204. The allocation network 210 can include four 4-1 (4:1) multiplexers, each multiplexer routing input activations to the corresponding sub-macro. Each sub-macro supports 1:2 sparsity. In the 1:2 sparsity base case, each 4-1 multiplexer in the allocation network acquires a pair of input activations at one of a plurality of inputs, and this pair of input activations is directly routed to the corresponding sub-macro. Once two pairs of input activations arrive at the sub-macro, a 2-1 multiplexer for the in-memory computation cell selects one of the two input activations based on stored 1-bit metadata. Once the sub-macro has completed computation, a merging network 216 (including a four-input adder tree for each column of the DCiM macro 202, or a Y number of four-input adder trees corresponding to the Y columns of the DCiM macro 202) adds together the partial sums or contributions of each sub-macro.
[0060] To support different N:M sparsity modes, the grouping or orchestration logic followed by the input activation buffer 208, the allocation network 210, and the merging network 216 can be summarized as follows. In the N:M sparsity mode, N represents the number of submacros that will cooperate to process a block of M. Since the submacros 204 are identical, the same orientation in each space of the submacros 204 can process the same block of M. In other words, N specifies the number of the same set of input activations (paired) received by P number of P-1 multiplexers. N identifies the number of submacros in P number of submacros 204 that are grouped or aggregated together to process the same block M. If N is greater than 1, multiple P-1 multiplexers will receive the same set of input activations. In the N:M sparsity mode, M represents the number of activation signals buffered to the P inputs of each P:1 multiplexer in the allocation network 210. In other words, M specifies the number of input activations (paired) processed or received by each P-1 multiplexer in the allocation network 210.
[0061] For example, in a basic 1:2 sparsity pattern scenario, since M=2, each P-1 multiplexer can receive two input activations or a pair of input activations at its input. Since N=1, two input activations can be buffered into a single P-1 multiplexer. Two further input activations can be buffered into a further P-1 multiplexer, and so on. In other words, a submacro is activated to process blocks of two input activations. By default, the P-1 multiplexer passes the two input activations received at its input to the corresponding submacro, where the P-1 multiplexer does not perform selection based on metadata encoded for coordinates of non-zero or dense weights. Based on metadata encoded for coordinates of dense or non-zero weights, the submacro's in-memory computation cell can use a 2-1 multiplexer via its 1-bit select signal to select the appropriate input activation to perform computation with the loaded weights.
[0062] For example, in a 2:4 sparsity pattern scenario, since M=4, each P-1 multiplexer can receive four input activations or two pairs of input activations at its two inputs respectively. Since N=2, four input activations from the same set can be buffered to two P-1 multiplexers. In other words, groups of two submacros are activated to process the same block of four input activations together. Spatially identical in-memory computation cells within the two grouped submacros will cooperate to process the same block of four input activations. Spatially identical in-memory computation cells in the two submacros that process the block of four input activations receive two corresponding non-zero or dense weights. Since M=4, each P-1 multiplexer receives four input activations. Based on metadata encoded with the coordinates of the dense or non-zero weights, each P-1 multiplexer selects a single pair of input activations and directs that pair to the corresponding submacro. The selection performed by the P-1 multiplexer can be based on a selection signal with one or more bits. If P=4, the selection signal has two bits to indicate which of the P inputs of the P-1 multiplexer is selected and output to the corresponding submacro. Based on metadata encoding the coordinates of dense or non-zero weights, the in-memory computation cell of the submacro can use a 2-1 multiplexer via its 1-bit selection signal to select the appropriate input activation to perform the computation with the loaded weights.
[0063] In addition to flexible N:M sparsity, dense operations (e.g., no sparsity) can also be supported or maintained.
[0064] In some embodiments, dense operations can be supported, as if the sparsity pattern were 2:2. Since M=2, each P-1 multiplexer can receive two input activations or a pair of input activations at its input. Since N=2, the same two input activations can be buffered to two P-1 multiplexers, thus mapping the same two input activations to two submacros. In other words, two submacros can be activated to process blocks of two input activations. By default, the P-1 multiplexer passes the two input activations received at its input to the corresponding submacro, where the P-1 multiplexer does not perform selection based on metadata encoding the coordinates of non-zero or dense weights. A 2-1 multiplexer in a submacro can select one of a pair of input activations for processing, and another 2-1 multiplexer in the submacro can select the other of the pair of input activations for processing.
[0065] In some embodiments, dense operations can be supported, as if the sparsity mode were 1:1. Since M=1, each P-1 multiplexer can receive two instances or a pair of identical input activations at its input. Since N=1, two instances of identical input activations can be buffered into a single P-1 multiplexer, mapping one input activation word to a submacro. Further instances of identical input activations can be buffered into further P-1 multiplexers. In other words, a submacro can be activated to process one input activation. By default, the P-1 multiplexer passes two instances of identical input activations received at its input to the corresponding submacro, where the P-1 multiplexer does not perform selection based on metadata encoding the coordinates of non-zero or dense weights. A 2-1 multiplexer within the submacro can select one of the two instances of identical input activations for processing based on an "irrelevant" selection signal.
[0066] It is evident that in both the 1:2 sparse mode base case and the dense inference case, the P-1 multiplexer in the distribution network does not perform a selection operation. In the dense inference case, the same input activation is streamed along both bit lines, and the selection signal of the 2-1 multiplexer is an "irrelevant" selection signal.
[0067] When M=8, three bits of information can be used to encode the sparsity selection information. When M=4, two bits of information can be used to encode the sparsity selection information. When M=2, one bit of information can be used to encode the sparsity selection information. The 2-to-1 multiplexer in the in-memory computing unit uses one bit of the selection signal or one bit of the sparsity selection information to select between two input activations, while supporting all N:M modes. The remaining bits(s) of the sparsity selection information are used in the allocation network 210 to support N:M modes.
[0068] Figure 3 Column 214 of submacros is illustrated according to some embodiments of the present disclosure. Column 214 illustrates... Figure 2 The column 214 is a submacro in memory compute cell 308 of the P-number submacros. Column 214 can include a column controller 302, an X / P-number of in memory compute cells 308, an input activation serializer, and a word line (WL) driver 310. In the example shown, each in memory compute cell in the in memory compute cells 308 stores an 8-bit memory word, for example, an 8-bit weight.
[0069] The in-memory computation cell 308, representing the X / P number in column 214, can perform the multiplication of the X / P number of input activations and weights in parallel. Memory cell 304 in Figure 4The details are described in more detail below. Each in-memory compute cell in in-memory compute cell 308 is capable of performing bit-by-bit multiplication on the stream input activation bits from the input activation serializer and WL driver 310.
[0070] The result of the multiplication generated by cell 308 in the memory of column 214 is summed (or accumulated) using adder tree 306 to generate a partial sum. Adder tree 306 can be a column-wise adder tree. Adder tree 306 can be an X / P input adder tree, such as a 32-input adder tree. The partial sum can be stored... Figure 2 In the part and buffer 212.
[0071] Subdividing the DCiM macro 202 into P submacros 204 has an additional benefit: the adder tree 306 used to sum X / P multiplications is much smaller than the adder tree used to sum X multiplications in the unpartitioned DCiM macro. The complexity of the adder tree for the unpartitioned DCiM macro is X*log(X), while the complexity of the adder tree 306 for the P submacros 204 is Xlog(X / P). Compared to the unpartitioned DCiM macro design, the reduced size of the adder tree in the partitioned design significantly reduces area and power consumption.
[0072] The input activation serializer and WL driver 310 can output signals to the word lines of the in-memory computing cell 308 to switch the state of the memory cell, thereby enabling operations such as writing, reading and erasing of memory elements, and controlling the digital circuits in the in-memory computing cell 308 to perform in-memory calculations.
[0073] Column controller 302 is capable of outputting selection signals for the 2-to-1 multiplexers in column 214. The selection signals for the 2-to-1 multiplexers in a given column's memory cells can be generated by column controller 302 based on metadata encoded with the coordinates of dense weights. Column controller 302 can be dedicated to generating selection signals for the 2-to-1 multiplexers in a given column 214.
[0074] In some embodiments, column controller 302 can output control signals to enable the in-memory computing cell 308 of column 214 to be in memory mode or computing mode. The control signals can be generated based on a column enable signal (e.g., EN_COL). The in-memory computing cells 308 of column 214 can share the same column enable signal.
[0075] Figure 4The illustration shows an in-memory compute cell 304 with a 2-1 multiplexer 404 for a submacro according to some embodiments of the present disclosure. The in-memory compute cell 304 can include B-number of memory elements 402 to store B-bit memory words. The B-number of memory elements 402 can include B-number of bit cells. In some embodiments, the B-number of memory elements 402 are implemented using a memory cell based on 28-nanometer latches (structurally similar to a 6T SRAM cell).
[0076] In-memory computing cell 304 can have two bit lines BL and The two bit lines are shared by a number of memory elements 402 of size B. The in-memory compute cell 304 can have two word lines WL and The two word lines are shared by memory elements that calculate the number of X / P cells in the column of memory.
[0077] The in-memory compute cell 304 can have a 2-to-1 multiplexer 404 to receive two input activations. The input activations are bit-serialized, therefore the 2-to-1 multiplexer 404 receives two input activation bits or iAct[1:0] at a time and outputs the selected input activation bit. The 2-to-1 multiplexer 404 can be configured by a selection signal or selection bit i... sel Control. A number of memory elements 402 share the same 2-1 multiplexer 404. The 2-1 multiplexer 404 selects one of the two input active bits to stream to a serial multiplier circuit 406. The 2-1 multiplexer 404 implements 1:2 structured sparsity as a baseline case.
[0078] The in-memory computing cell 304 may also include a bit-serial multiplier circuit 406 to multiply a selected one of the two input activations by a weight. The bit-serial multiplier circuit 406 may include a bitwise AND operator to perform a bitwise AND operation between bits of the memory word and a stream of input activation bits. The bitwise AND operation result is weighted by the bit positions and accumulated to form the multiplication result of the input activations and the weight. In some embodiments, the bit-serial multiplier circuit 406 may be implemented using NOR gates to perform the multiplication between the selected input activation bit and the weight bit.
[0079] The in-memory computing cell 304 can have two operating modes: memory mode and computing mode. During memory mode, read and write operations are performed on B-number of memory elements 402. Write lines (e.g., WL and...) ) can be activated to access all X / P number of memory elements in the column of the computed cell within memory, while bit lines (e.g., BL and This is used to read / write data to B-number of memory elements 402. During computation mode, the bit serial multiplier circuit 406 performs computation operations. When WL=0, computation mode can be enabled for all X / P-number of memory elements in the column of the computation cell within the memory, and the corresponding column enable signal (e.g., EN_COL=1) enables the column to perform computation operations.
[0080] Figure 5 The illustration shows an implementation of a distribution network 210 comprising a P number of P-1 multiplexers 502 according to some embodiments of the present disclosure. The distribution network 210 is illustrated as being capable of being used for... Figure 2 A collection of P-1 multiplexers 502 that are spatially identical to each other within a P-number submacro 204. Each P-number P-1 multiplexer 502 serves a P-number submacro individually.
[0081] The P-1 multiplexer 502 can have P inputs and one output. A pair of input activations can be fed to the inputs, and a pair of selected input activations can be directed to the corresponding submacro row R. The selection signal for the P-1 multiplexer can be controlled by a controller (such as...). Figure 2 The controller 206) is generated.
[0082] The allocation network 210 is responsible for feeding appropriate pairs of input activations to the submacro based on N and M in the N:M sparsity pattern. The P inputs of the P-1 multiplexer are loaded with one (or more) pairs of input activations based on M. The N sets of P-1 multiplexers in the allocation network 210 are loaded with the same M number of input activations.
[0083] Because the in-memory compute cells support a 1:2 sparsity mode, the output of the P-1 multiplexer has selected pairs of input activations. In an embodiment where the input activations are represented by 8-bit words, the P inputs and outputs of the P-1 multiplexer have a 16-bit width to pack two 8-bit input activations together.
[0084] The selection of appropriate pairs of input activations performed by P-1 multiplexers 502 is controlled or determined by metadata encoding coordinates with dense or non-zero weights. The allocation network 210 abstracts the complexity of supporting large multiplexers in in-memory compute cells by performing the selection of M-1 input activations at a global level, thereby efficiently selecting the correct / appropriate pairs of input activations to be fed into the in-memory compute cell.
[0085] Figure 6The illustration shows an implementation of a merging network 216 including an adder tree according to some embodiments of the present disclosure. Due to FlexCiM's partitioning scheme, the partial sums calculated for each column of a submacro can be accumulated or added together by the merging network 216. The merging network 216 can include a P-input adder tree that reads partial sums from the corresponding portions and buffers of the submacro and outputs the final partial sum.
[0086] Row and column pipelines in FlexCiM
[0087] Briefly return to reference Figure 2 The input activation buffer 208 has a limited bandwidth. In practice, the input activation buffer 208 can buffer a certain number of input activations for a given clock cycle (e.g., 1024 bits / cycle), or 128 8-bit input activations. For dense LLM inference, X (e.g., X = 128) input activations can be streamed or buffered to the column, and X parallel MAC operations can be performed by the in-memory computed cell of the column. However, for sparse inference with the highest sparsity ratio (1:8 sparsity mode in some implementations), the P-1 multiplexer for the row is selected based on the 8 input activations converted to the 1024 input activations to be streamed to the column. Considering the upper limit of the bandwidth constraint of the input activation buffer 208, only 4 rows of the column can perform parallel MAC operations during 1:8 sparse inference, only 8 rows of the column can perform parallel MAC operations for 1:4 sparse inference, and only 16 rows of the column can perform parallel MAC operations for 1:2 sparse inference.
[0088] Because in-memory computation cells perform bit-sequential multiplication, the MAC operation itself can occupy multiple cycles, for example, eight cycles. This creates an opportunity to perform row pipelined operations by overlapping computation cycles with memory access cycles.
[0089] Row pipelines are techniques used to improve the throughput of computing systems within bit-serial memories. In architectures like FlexCiM, due to bandwidth constraints of the input activation buffer 208, especially in highly sparsity modes like 1:8, not all rows can be fed input activation simultaneously or within a single cycle. Row pipelines address this problem by dividing rows into smaller groups and buffering input activations in stages (e.g., over multiple clock cycles) to these smaller groups. The number of row pipeline stages (#stages) can be equal to X divided by the number of rows grouped together, e.g., #stages = X / #grouped rows. The input activation buffer 208 buffers input activations in pairs to the allocation network 210 over multiple clock cycles. During a clock cycle, input activations are buffered by the input activation buffer 208 to a subset of the allocation network, which corresponds to a group of P number of submacros. A subset of the allocation network means a set of P number of P-1 multiplexers serving the group of P number of rows.
[0090] Considering the implementation of the DCiM macro 202 with X = 32 rows and a 1:8 sparsity pattern, each row requires eight input activations, but the input activation buffer 208 only has bandwidth to feed four rows per cycle. The row pipeline groups the 32 rows into eight sets of four rows, for example: #group of rows = 4, #stage = X / #group of rows = 32 / 4 = 8. In each cycle, one group of rows in the activation column (e.g., EN_COL = 1) (i.e., the set of four rows) is fed input activation. In the first clock cycle, the group of four rows (e.g., the corresponding P sets serving the P-1 multiplexer of four rows) is fed input activation, and in the second clock cycle, further groups of four rows (e.g., the corresponding P sets serving the P-1 multiplexer of four rows) are fed further input activation, and so on. The 32 rows are fed into the input in groups of four within eight clock cycles or in an eight-row pipeline stage.
[0091] The eight clock cycles of buffered input activation can overlap with computation cycles. This overlap between memory accesses and computations ensures that the pipeline remains fully loaded and no cycles are wasted.
[0092] During the # phase of a clock cycle, input activations are fed in groups to rows of specific columns. A column can be activated or enabled by setting the column enable signal EN_COL = 1 for the # phase of the clock cycle. At the end of the # phase of the clock cycle, the next column / adjacent column is activated using the column enable signal EN_COL = 1 for the next # phase of the clock cycle. Once all columns have been fed input activations (one by one), the first column is deactivated using the column enable signal EN_COL = 1, and the process is repeated. This process implements column pipelined operation, where input activation is fed to column one at a time.
[0093] Column pipelines are a technique used to reduce the hardware overhead associated with the allocation network 210 and improve the throughput of DCiM systems. Without column pipelines in FlexCiM, each column would utilize a dedicated allocation network 210, which would be extremely expensive for memory-centric designs. In other words, simultaneously feeding input activations to all columns would require a dedicated allocation network 210 for each column, which is costly in terms of area. Column pipelines address this problem by using a shared allocation network 210 across columns in a time-division multiplexing manner. Input activation buffers 208 buffer input activations in pairs to the allocation network 210 over multiple time periods, where, during a time period, input activations are buffered to columns of P number of sub-macros (e.g., active columns enabled by the column enable signal EN_COL = 1). Input activations are buffered by input activation buffers 208 to P sets of P-1 multiplexers serving all rows. Therefore, input activation buffers 208 buffer input activations to one column at a time. A time period can include the amount of time that input activation buffers 208 spend buffering input activations to columns. The time period can be the # phase of a clock cycle, where the input activation buffer 208 can perform row pipelines within the active column.
[0094] In the FlexCiM architecture, columns are activated sequentially. During a given time period, one column is activated by input fed into input activation buffer 208, and during the next time period, that same column can perform computations and adjacent columns are activated by input fed into input activation buffer 208. The time period can comprise a fixed number of cycles, determined by the bit-serialized MAC operation delay. Essentially, the allocation network 210 can serve multiple columns over time, rather than replicating hardware for each column.
[0095] Row and column pipelines allow FlexCiM to maintain high throughput and energy efficiency while supporting flexible N:M sparsity in compact in-memory computing designs.
[0096] Figure 2 Controller 206 and Figure 3 The column controller 302 is capable of generating control signals for the row and column pipelines of P number of sub-macros 204. The controller 206 is capable of generating selection signals for the P-1 multiplexers in the distribution network 210. The column controller 302 is capable of generating selection signals for the 2-1 multiplexers in the memory-based computed cells of P number of sub-macros 204.
[0097] Sparse storage formats and demonstration examples: 1:4 sparse mode and 4:8 sparse mode
[0098] The FlexCiM architecture utilizes the Compressed Sparse Columns (CSC) format as the storage format for sparse weights with an N:M sparsity pattern. Each column of the weights can be stored as a list of all dense or non-zero weight values, and the orientation, position, or coordinates of the N dense or non-zero weight values within a block of size M can be stored as corresponding metadata. This metadata informs... Figure 2 The DCiM macro 202 identifies which weights are dense or non-zero and where they are located, thus allowing selective or sparse computation without performing calculations for zero-value weights. Specifically, the metadata can be used to generate selection signals for the P-1 and 2-1 multiplexers of the DCiM macro 202. Furthermore, the metadata can be used to instruct the merging network 216 to merge partial sums of sub-macros from packets.
[0099] Figures 7A-7B The illustration shows a demonstrative example of implementing a 1:4 sparsity pattern according to some embodiments of the present disclosure. Figures 8A-8B The illustration shows a demonstrative example of implementing a 4:8 sparsity pattern according to some embodiments of the present disclosure. Operations are marked with circled numbers. and The demo examples demonstrate how FlexCiM can be configured to support flexible N:M sparsity patterns.
[0100] For illustration, when the partition factor P = 4, and the DCiM macro has P = 4 submacros, for example, submacro 0, submacro 1, submacro 2, and submacro 4. For simplicity, a P-1 (e.g., 4-1) multiplexer of P = 4 number is depicted for row 0. It is envisioned that other sets of four 4-1 multiplexers are provided as parts of the distribution network for the other rows. For illustration, the submacro has two rows and two columns. It is envisioned that the submacro can be generalized to have X / P rows and Y columns.
[0101] For in Figures 7A-7B The 1:4 sparsity pattern illustrated in the diagram involves each submacro independently operating on blocks of four weights, where each block retains a dense or non-zero weight. Each 4-1 multiplexer for row 0, corresponding to the four submacros, receives four different input activations (e.g., two pairs of input activations). Metadata related to the sparse weight tensor, encoded in CSC format (e.g., in...) Figure 7A (As depicted in the text) specifies the location of non-zero weights within each block of the four weights.
[0102] For in Figures 8A-8BThe 4:8 sparsity pattern illustrated in the diagram shows four submacros aggregating to process a single block of eight weights, where this single block retains four non-zero weights. Input buffers buffer eight input activations (e.g., four pairs of input activations) from the same set into the same spatially identical rows or orientations in the four submacros. Each 4-1 multiplexer for row 0 in the four submacros receives the same eight input activations (e.g., four pairs of input activations). Metadata related to the sparse weight tensor, encoded in CSC format (e.g., in...) Figure 8A (As depicted in the text) specifies the location of non-zero weights within each of the eight weight blocks.
[0103] In operation In this process, controller 206 receives metadata encoded with densely packed or non-zero weighted coordinates and thus generates signals for allocating 4-to-1 multiplexers in the network and 2-to-1 multiplexers in the in-memory computed cell. The least significant bit (LSB) of the metadata can be used to generate selection signals for 2-to-1 multiplexers, and the remaining bits of the metadata can be used to generate selection signals for 4-to-1 multiplexers. For M=4, the metadata has two bits: one LSB for generating selection signals for 2-to-1 multiplexers and one most significant bit (MSB) for generating selection signals for 4-to-1 multiplexers. For M=8, the metadata has three bits: one LSB for generating selection signals for 2-to-1 multiplexers and two MSBs for generating selection signals for 4-to-1 multiplexers.
[0104] In operation In the distribution network, the 4-1 multiplexer for rows (e.g., row 0) selects an appropriate pair of input activations at one of the four inputs of the 4-1 multiplexer based on a selection signal generated by the controller 206, and directs the pair of input activations to the corresponding sub-macro of the column (e.g., col0).
[0105] In the steps In this configuration, a column (e.g., col0) is activated to receive a pair of input activations selected by a 4-1 multiplexer (e.g., EN_COL = 1). A pair of input activations can be serialized and fed to the bit lines of the computed cell in memory. Row pipelines can be executed to feed one group of input activations to a row at a time.
[0106] In operation In the column, the 2-1 multiplexer calculates the input activation received at the cell from the memory on the bit line based on the selection signal generated by the controller 206, and selects the appropriate input activation. MAC operations are performed in the column.
[0107] In operation In this process, during the # phase of the cycle in which the MAC operation is performed in a column, adjacent columns can be activated to receive input activation pairs (e.g., EN_COL = 1). Row pipelines can be executed to feed input activations to a group of rows each time.
[0108] After the MAC operation is performed in the column, as the partial sums of the column are calculated and generated, the merge network (not shown) is able to perform the final partial sum accumulation by reading from the partial sum buffer (not shown) of each column.
[0109] In some implementations, DCiM macros supporting a 1:4 sparsity pattern exhibit 1.63x lower latency than dense inference. In other implementations, DCiM macros supporting a 2:8 sparsity pattern exhibit 1.42x lower latency than dense inference.
[0110] Implementing the DCiM architecture in a neural processing unit
[0111] In some embodiments, the DCiM macro 202 can be integrated as an accelerator within a neural processing unit to perform operations for layers of a neural network (which can be represented as GEMM or MAC operations) and supports flexible N:M sparsity ratios. The DCiM macro 202 can be implemented with or in parallel with a von Neumann-based digital accelerator with processing elements, wherein the DCiM macro 202 is capable of performing 8-bit integer (INT8) GEMM or MAC operations. The compiler for the neural processing unit generates machine-readable instructions capable of offloading layer GEMM or MAC operations to the DCiM macro 202 for execution with respect to a specific N:M sparsity ratio determined for the layer. The compiler can also assign other types of operations, such as element size operations, depth orientation operations, and layer normalization operations, to be performed by the von Neumann-based digital accelerator with processing elements.
[0112] FLOW framework
[0113] To further improve other weight pruning algorithms, the FLOW framework employs a unique approach to identify optimal layer-by-layer N:M structured sparse patterns that balance accuracy reduction with efficient inference performance. Neural networks pruned using the FLOW framework can be readily executed on digital accelerators that support N:M sparse patterns. The FLOW framework identifies layer-by-layer N:M patterns based on one or more factors, including outlier presence, outlier distribution (or outlier density or outlier clustering), and latency / performance importance versus accuracy importance. Appropriate algorithms can be used to determine outliers, identifying weights that are outliers and may not be suitable for pruning. In particular, the FLOW framework can be executed on FlexCiM-based digital accelerators as described in this paper.
[0114] In some experiments, compared to other layer-by-layer N:M pruning techniques, FLOW was able to improve the performance of pruning models by 25%-36% at high sparsity ratios when tested on some DNN models.
[0115] Figures 9A-9C The illustrations depict different optimal sparsity patterns according to some embodiments of the present disclosure for different outlier presences and outlier distributions.
[0116] exist Figure 9A In the example shown in the diagram, the presence of outliers is low, and the outlier distribution is sparse. The optimal N:M sparsity pattern can be 1:4.
[0117] exist Figure 9B In the example illustrated in the diagram, the presence of outliers is high, and the outlier distribution is sparse. The optimal N:M sparsity pattern can be 2:4. Specifically, for... Figure 9B The value N in the example is relative to Figure 9A Higher, to indicate the existence of higher exterior points.
[0118] exist Figure 9C In the example illustrated in the diagram, outlier presence is high, and the outlier distribution is clustered / nearby. The optimal N:M sparsity pattern can be 4:8. Specifically, for Figure 9C The value M in the example is relative to Figure 9A and Figure 9B Higher, to indicate a more clustered distribution of outliers.
[0119] Figures 9A-9C The illustrated example reveals that outlier distribution plays a crucial role in determining the optimal N:M sparsity pattern (specifically, the value of M). FLOW tends to favor larger M values when outlier distributions are dense or clustered, minimizing the likelihood of pruning outliers within blocks of M and maintaining a high degree of flexibility in deciding which weights to prune. Considering both N and M values when assigning N and M values can lead to significantly higher confusion scores.
[0120] Figure 10 The illustration shows a system 1000 for optimizing N:M sparsity patterns and generating sparse neural networks according to some embodiments of the present disclosure. The system 1000 may include an N:M sparsity optimizer 1002, a weight pruner 1004, and a model compiler 1006.
[0121] The N:M sparsity optimizer 1002 accepts a neural network model, such as the definition of a neural network with multiple layers. This definition specifies the operations performed on the layers and their inputs and outputs. The definition may include layer parameters, such as weight tensors. The N:M sparsity optimizer 1002 can accept a target delay or a delay target. The N:M sparsity optimizer 1002 can accept a target sparsity or sparsity ratio target.
[0122] The N:M sparsity optimizer 1002 may include outlier identification 1090 to identify outliers for layers in a neural network. For illustration, the N:M sparsity optimizer 1002 focuses on determining the N:M sparsity pattern used for weights. Outlier identification 1090 identifies weights in the neural network model that are outliers. Suitable techniques can be used to identify outliers. Outlier identification 1090 can calculate a statistical or functional score that reflects the atypicality of individual weight parameters relative to the overall distribution of weights within a layer or model.
[0123] Assume the layer has a weight tensor And receive input activation T represents the number of tokens, and K represents the sequence length. Identifying outliers 1090 allows for the assignment of importance scores I to individual weights. W =|W ij |·‖X j ||2, where ||X j ||2 is the L2 norm of the input activation connected to the weight elements. W This can serve as a metric for identifying outliers based on both weight values and associated activation values. Based on the weight scores for the weights in the weight tensor used for the layer, the mean μ and / or standard deviation σ of the importance scores of the weights can be calculated for each layer. The mean μ and / or standard deviation σ can be used to set a threshold to determine whether a weight element is an outlier. For example, the threshold can be set as a multiple of the standard deviation, such as τ·σ, where τ can be a hyperparameter. In some implementations, τ = 3. In some implementations, τ = 5. If the difference between the importance score of a weight element and the mean μ exceeds or surpasses the threshold, the weight element is classified or identified as an outlier.
[0124] In some embodiments, a weight element is classified as an outlier if the absolute value of its score exceeds a calculated threshold (such as a percentile-based cutoff or a multiple of the standard deviation of the mean score). The threshold used for outlier classification can be layer-specific and sign-sensitive, allowing for fine-grained control and robustness to noise within sparse patterns.
[0125] In some embodiments, the score can be computed as a function of the weight values. In other embodiments, the score can be computed as a joint function of the weights and their associated input activations, for example, by evaluating the product of the weight metric and the norm of the input activation vector. This joint score captures both the importance of static parameters and the relevance of dynamic features, thereby enabling selective pruning of weights that contribute least to the model output while preserving weights linked to high activation paths.
[0126] The N:M sparsity optimizer 1002 can include a layer-by-layer outlier counter 1010. The layer-by-layer outlier counter 1010 can determine the count of identified outliers. The layer-by-layer outlier counter 1010 can quantize the presence of outliers for a given layer. The presence of outliers can differ from one layer to another. The presence of outliers can be quantized based on the number of outliers in each layer.
[0127] The N:M sparsity optimizer 1002 can include a layer-by-layer out-point locality index calculator 1020. The layer-by-layer out-point locality index calculator 1020 can determine the locality measurement of the identified out-points. The layer-by-layer out-point counter 1010 can quantize the out-point distribution (e.g., density and / or clustering) for a given layer. The out-point distribution can differ from one layer to another. The out-point distribution can be quantized based on out-points that are closely clustered together and / or based on whether out-points are sparsely distributed or spaced far apart from each other. An exemplary algorithm is described in... Figure 11 The details are as follows.
[0128] The N:M sparsity optimizer 1002 can include estimating the A / B sparsity ratio 1030. The A / B sparsity ratio estimation 1030 can determine the value A in the A / B sparsity ratio based on counts determined by the layer-by-layer outlier counter 1010. Based on outlier presence (such as counts), the A / B sparsity ratio estimation 1030 can determine the value A in the A / B sparsity ratio that causes the minimum reduction in accuracy. A higher outlier presence can lead to a higher value A. The A / B sparsity ratio estimation 1030 can also determine the value B in the A / B sparsity ratio based on locality measurements. Based on outlier distribution (such as locality measurements), the A / B sparsity ratio estimation 1030 can determine the value B in the A / B sparsity ratio that causes the minimum reduction in accuracy. A denser or more clustered outlier distribution can lead to a higher value B. The values A and B in the A / B sparsity ratio are determined in... Figure 12 The details are as follows.
[0129] The N:M sparsity optimizer 1002 can include a final optimal N:M sparsity pattern 1040. The final optimal N:M sparsity pattern 1040 can assign N:M sparsity patterns to layers based on a sparsity ratio A / B. Based on the identified A / B sparsity ratio, the final optimal N:M sparsity pattern 1040 can determine one or more N:M sparsity patterns that retain the identified A / B sparsity ratio. The final optimal N:M sparsity pattern 1040 can consider the importance of latency / performance to accuracy to select or choose the optimal N:M sparsity pattern from candidate N:M sparsity patterns that satisfy the A / B sparsity ratio. Specifically, the final optimal N:M sparsity pattern 1040 can quantify the accuracy of a given sparsity pattern based on a value M. The higher the value M, the larger the block of weights the pruning algorithm must select for pruning, which reduces the risk of pruning outliers. Furthermore, the final optimal N:M sparse pattern 1040 can quantify the latency of a given sparse pattern based on the value M, or use a latency calculator 1080 to resolvely determine the latency of a given sparse pattern when implemented on a specific digital accelerator that supports sparsity. When using the FlexCiM architecture, a higher value of M can lead to a longer execution time because more input activations are buffered into rows for selection during a given clock cycle. The accuracy and / or latency of a given candidate sparse pattern can be used to guide the selection of the optimal N:M sparse pattern from the candidate N:M sparse patterns by weighting and balancing the importance of latency / performance against the importance of accuracy. The final optimal N:M sparse pattern is determined based on the candidate sparse patterns. Figure 12 The details are as follows.
[0130] Based on the optimal N:M sparsity pattern determined by the N:M sparsity optimizer 1002, the weight pruner 1004k can prune the layer weights according to the N:M sparsity pattern and determine and encode the sparse weight tensor according to the N:M structured sparsity. The weight pruner 1004k can encode the sparse weight tensor using the CSC format.
[0131] Based on the definition of sparse weight tensors and neural network models, model compiler 1006 can generate machine-readable instructions or configurations for neural networks according to the N:M sparsity pattern. Model compiler 1006 can package the pruned weights determined by weight pruner 1004 together with the machine-readable instructions or configurations for the neural network. The instructions or configurations can be loaded onto a processor (such as a neural processing unit with a FlexCIM-based digital accelerator) to enable the processor to execute operations on the neural network model.
[0132] Figure 11 This is a flowchart illustrating a method 1100 for determining the count of outliers and measuring the locality of outliers according to some embodiments of the present disclosure. Method 1100 can be performed by... Figure 10 The N:M sparsity optimizer 1002 is executed. Method 1100 illustrates an exemplary implementation of a process for counting outliers and calculating locality measures (e.g., locality indices). The outlier count is an indication of the existence of outliers. The locality measure is an indication of the distribution of outliers (e.g., density or clustering). The weight tensor W for the layer can be loaded. Optionally, representative input activations for the layer can be loaded.
[0133] In 1102, the importance scores of the weight elements used in the weight tensor W can be determined. The importance score can be a function of the weight values. Alternatively, it can be a function of the weight values and the associated input activations.
[0134] In 1104, the mean and standard deviation can be determined based on importance scores.
[0135] In 1106, the outlier threshold can be determined using the standard deviation and optional hyperparameters. The outlier threshold can be a multiple of the standard deviation. In some embodiments, the hyperparameters can be user-provided. The outlier threshold can be used to identify the range of importance scores that classify a weighted element as an outlier and the range of importance scores that classify a weighted element as an inlier.
[0136] In 1108, outliers can be identified using outlier thresholds and means.
[0137] Operations 1110, 1112, and 1114 can be performed on individual neighborhoods of the weight tensor (such as a 128x128 neighborhood of an element). The weight tensor can be divided into the number of neighborhoods of an element or the number of non-overlapping blocks of an element.
[0138] In 1110, the outlier count can be determined. The count can be determined separately for individual neighborhoods and aggregated / averaged across neighborhoods to represent the outlier count or presence of the layer. The count can also be determined across the entire weight tensor to represent the outlier count or presence of the layer.
[0139] In 1112, pairwise distances between outgoing points (e.g., Euclidean distances) can be computed in the neighborhood to form a distance matrix.
[0140] In 1114, the upper triangular portion of the distance matrix can be summed, and this sum can be partitioned by the number of outliers to obtain an average distance measurement for the neighborhood. The average distance measurements for individual neighborhoods can be aggregated to represent the outlier distribution of the layer or for use in layer locality measurements.
[0141] In some embodiments, for a weight tensor l, the exterior point distribution D l It can be computed to represent a surrogate metric that measures the average of the sum of pairwise distances between outliers and the number of outliers:
[0142]
[0143] Here, dist(.) measures the pairwise L1 distance between two exterior points, and n C2 is the total number of outlier pairs in a layer with n outliers. The normalized outlier distribution (ND) can be calculated using min-max normalization across all layers.
[0144]
[0145] D represents the distribution of all outpoints in the entire layer. l A set of.
[0146] In some embodiments, Formula 1 is used to calculate the neighborhood b within the weight tensor, and the outlier distribution is obtained by applying the distribution across the neighborhood. The summation is used to obtain the normalized outlier distribution ND. l It can be obtained using Equation 2.
[0147] Higher ND l It can indicate the sparse distribution of outgoing points, or the large distances between outgoing points within a layer. Smaller ND l It can indicate a large number of closely located or clustered out points within a layer.
[0148] Figure 12 This is a flowchart illustrating a method 1200 for determining an optimal N:M sparsity pattern according to some embodiments of the present disclosure. Method 1200 can be... Figure 10 The N:M sparsity optimizer 1002 is executed.
[0149] In 1202, candidate N and candidate M can be determined. For example, candidate N and candidate M can be powers of 2. Candidate N can include {1, 2, 4, 8}, and candidate M can include {2, 4, 8}.
[0150] Subsequent operations can be performed on a per-layer basis.
[0151] In 1204, the number of out-of-field points (such as out-of-field percentage or proportion) can be determined for a layer.
[0152] In 1206, the value A in the sparsity ratio A / B can be estimated based on the outlier count. A higher outlier count tends to result in a higher value A. The value A can be selected from the candidate N. More outliers in the layer indicate that more non-zero elements (larger values A) should be retained for the layer.
[0153] In 1208, the locality of external points can be determined for the layer.
[0154] In 1210, the value B in the sparsity ratio A / B can be estimated based on the out-point locality measurement. A higher locality measurement (indicating dense / clustered out-points) tends to result in a higher value B. The value B can be selected from candidate M. The denser the out-points, the larger the block (larger value B) should be used during pruning to avoid pruning the out-points.
[0155] Users can provide or set a latency importance category, which indicates whether latency / performance has high, medium / balanced, or low priority relative to model accuracy. The latency importance category indicates the trade-off between latency / performance and model accuracy.
[0156] In step 1212, method 1200 checks whether latency / performance has a low priority relative to accuracy (e.g., latency importance = low or 0). When latency / performance has a low priority, model accuracy has a high priority. If yes, method 1200 proceeds to step 1214. If no, method 1200 proceeds to step 1216.
[0157] In 1214, values A and B can be set to values N and M in an N:M structured sparsity pattern. In some cases, the N:M sparsity pattern with the highest value M that preserves the sparsity ratio of A / B is selected as the optimal N:M sparsity pattern. The N:M sparsity pattern with the highest value M prioritizes model accuracy over latency.
[0158] In step 1216, method 1200 checks whether latency / performance and accuracy have a balanced priority (e.g., latency importance = medium or 1). In other words, the model's latency / performance and accuracy have the same or equal priority. If yes, method 1200 proceeds to step 1218. If no, the method proceeds to step 1222.
[0159] In step 1218, the N:M sparse pattern with the highest latency (e.g., the highest value M) that retains the sparsity ratio of A / B is removed from the candidate sparse patterns. In step 1220, the optimal N:M sparse pattern is determined from one or more remaining candidate sparse patterns (i.e., the patterns that result in the minimum amount of accuracy reduction) that have the highest value M that retains the sparsity ratio of A / B. The removal operation in step 1218 ensures that latency / performance priorities and accuracy priorities are given balanced or equal weight.
[0160] Starting from 1216, no path following means that latency / performance has a high priority relative to accuracy (e.g., latency importance = high or 0). When latency / performance has a high priority, the model's accuracy has a low priority.
[0161] In 1222, the N:M sparsity pattern with the lowest value M that preserves the sparsity ratio of A / B (e.g., the pattern that results in the lowest latency) is selected as the optimal N:M sparsity pattern. The N:M sparsity pattern with the lowest value M prioritizes latency over model accuracy.
[0162] In some embodiments, the optimal N:M sparsity pattern for different layers of the model can be formulated as an integer linear programming problem. Weights can be assigned for latency / performance and accuracy. The integer linear programming problem searches for a domain of possible N:M sparsity patterns for each layer of the model that will satisfy a target sparsity ratio for the overall model, ensuring that for a given sparsity pattern: 1 ≤ N ≤ M, while aligning N with outlier counts and M with outlier locality measurements.
[0163] Figure 13 This is a flowchart illustrating a method 1300 for determining an optimal N:M sparsity pattern according to some embodiments of the present disclosure. Method 1300 can be executed by one or more components of system 1000.
[0164] In 1302, the outgoing points of the layers used in the neural network can be identified.
[0165] In 1304, the count of the outer points can be determined.
[0166] In 1308, a locality measurement of the outgoing points can be determined. In some embodiments, the locality measurement indicates how close or far the outgoing points are. In some embodiments, determining the locality measurement includes calculating the locality measurement based on one or more pairwise distances between outgoing points within a neighborhood. In some embodiments, determining the locality measurement includes calculating the locality measurement based on the average of one or more pairwise distances between outgoing points within a neighborhood.
[0167] In 1306, the value A in the sparsity ratio of A / B can be determined based on counts. In some embodiments, determining the value A in the sparsity ratio A / B includes setting the value A higher when the count is higher.
[0168] In 1310, the value B in the sparsity ratio of A / B can be determined based on locality measurements. In some embodiments, determining the value B in the sparsity ratio of A / B includes setting the value B higher when the locality measurement indicates that outliers are more clustered rather than widely spaced.
[0169] In 1312, the sparsity pattern for N:M of a layer can be assigned based on the sparsity ratio of A / B.
[0170] In some embodiments, assigning a sparsity pattern includes: determining an N:M sparsity pattern with the lowest value M that preserves the sparsity ratio of A / B.
[0171] In some embodiments, assigning a sparsity pattern includes: determining an N:M sparsity pattern with the highest value M that preserves the sparsity ratio of A / B.
[0172] In some embodiments, assigning a sparsity pattern includes: removing the candidate sparsity pattern with the highest latency, and determining the N:M sparsity pattern from one or more remaining candidate sparsity patterns with the highest value M of the sparsity ratio that retains A / B.
[0173] In some embodiments, assigning a sparsity pattern includes assigning a sparsity pattern based on a delay importance category.
[0174] Methods for accelerating neural network operations using the FlexCiM architecture
[0175] Figure 14 This is a flowchart illustrating a method 1400 for accelerating MAC operations on input activation and weights using an N:M sparsity pattern, according to some embodiments of the present disclosure. The method can be performed by, as in... Figure 2 The diagram shows the execution of the digital accelerator 280, thereby realizing the FlexCiM architecture.
[0176] In 1402, input activations are buffered in pairs to the allocation network according to N and M. In some embodiments, buffering input activations includes buffering M pairs of input activations to a P-1 multiplexer of the allocation network. In some embodiments, buffering input activations includes buffering N pairs of the same set of input activations to the allocation network.
[0177] In some embodiments, buffered input activation in 1402 includes: buffering input activations in pairs to an allocation network over multiple clock cycles to perform row pipelined operations. During a clock cycle, input activations are buffered to a subset of the allocation network corresponding to a group of rows(multiple) of P number of submacros.
[0178] In some embodiments, buffered input activation in 1402 includes: buffering input activations in pairs to an allocation network over multiple clock segments to perform column pipelined operations. During a clock segment, the input activations are buffered to columns of P number of sub-macros.
[0179] In 1404, a pair of inputs is selected to activate at the input of the P-1 multiplexer of the distribution network. P is the number of sub-macros in the DCiM macro.
[0180] In 1406, the P-1 multiplexer selects the right input to activate the output to the submacro.
[0181] In 1408, the submacro's in-memory computed cell's 2-1 multiplexer selects input activation from the input activation of the selection pair.
[0182] In some embodiments, method 1400 further includes: performing a multiplication of the selected input activation with weights stored in memory in the computed cell, and summing the P partial sums of the P number of sub-macro outputs by merging the network.
[0183] In some embodiments, method 1400 further includes: generating a selection signal for a P-1 multiplexer based on metadata encoding. In some embodiments, method 1400 includes: generating a selection signal for a 2-1 multiplexer based on metadata encoded with the coordinates of dense weights.
[0184] Exemplary computing device
[0185] Figure 15 This is a block diagram of an apparatus or system (e.g., an exemplary computing device 1500) according to some embodiments of the present disclosure. One or more computing devices 1500 may be used to implement the functions described with reference to the drawings and herein. Figure 15 The various components illustrated in the figure can be included in computing device 1500, but one or more of these components may be omitted or duplicated depending on the application. In some embodiments, some or all of the components included in computing device 1500 may be attached to one or more motherboards. In some embodiments, some or all of these components are fabricated on a single system-on-a-chip (SoC) die. Additionally, in various embodiments, computing device 1500 may not be included in... Figure 15 One or more of the components illustrated may be included, and computing device 1500 may include interface circuitry for coupling to one or more components. For example, computing device 1500 may not include display device 1506, but may include display device interface circuitry (e.g., connector and driver circuitry) to which display device 1506 may be coupled. In another set of examples, computing device 1500 may not include audio input device 1518 or audio output device 1508, but may include audio input or output device interface circuitry (e.g., connector and support circuitry) to which audio input device 1518 or audio output device 1508 may be coupled.
[0186] Computing device 1500 may include processing device 1502 (e.g., one or more processing devices, one or more of the same type of processing devices, or one or more of different types of processing devices). Processing device 1502 may include electronic circuitry that processes electronic data from data storage elements (e.g., registers, memories, resistors, capacitors, qubit cells) to convert the electronic data into other electronic data that can be stored in registers and / or memories. Examples of processing device 1502 may include: CPU, GPU, quantum processor, machine learning processor, artificial intelligence processor, neural network processor, artificial intelligence accelerator, application-specific integrated circuit (ASIC), analog signal processor, analog computer, microprocessor, digital signal processor, field-programmable gate array (FPGA), tensor processing unit (TPU), neural network hardware accelerator, DNN hardware accelerator (e.g., having DCiM macros as described herein), etc.
[0187] In some embodiments, the processing device 1502 may include a digital accelerator as described and illustrated herein (including a digital accelerator implementing the FlexCiM architecture).
[0188] Computing device 1500 may include memory 1504, which itself may include one or more memory devices, such as volatile memory (e.g., DRAM), non-volatile memory (e.g., read-only memory (ROM)), high-bandwidth memory (HBM), flash memory, solid-state memory, and / or hard disk drives. Memory 1504 includes one or more non-transitory computer-readable storage media. In some embodiments, memory 1504 may include memory that shares a die with processing device 1502.
[0189] In some embodiments, memory 1504 includes one or more non-transitory computer-readable storage media that store executable instructions to perform the operations described herein with respect to the accompanying drawings. Exemplary components are depicted, such as an N:M sparsity optimizer 1002, a weight pruner 1004, and a model compiler 1006, which may be encoded as instructions and stored in memory 1504. Memory 1504 may store instructions encoding one or more exemplary components, such as one or more components of the N:M sparsity optimizer 1002. Instructions stored in one or more non-transitory computer-readable storage media may be executed by processing device 1502. Memory 1504 may store instructions that cause processing device 1502 to perform one or more of the following: method 1100, method 1200, and method 1300.
[0190] In some embodiments, memory 1504 may store data as described with respect to the accompanying drawings and herein, such as data structures, binary data, bits, metadata, files, binary large objects, etc. Memory 1504 is capable of storing input data, intermediate data, and output data for various components, such as the N:M sparsity optimizer 1002, the weight pruner 1004, and the model compiler 1006.
[0191] In some embodiments, memory 1504 may store one or more DNNs (and / or components thereof). Memory 1504 may store training data used to train (trained) DNNs. Memory 1504 may store instructions for performing operations related to training the DNNs. Memory 1504 may store input data, output data, intermediate outputs, and intermediate inputs of one or more DNNs. Memory 1504 may store one or more parameters used by one or more DNNs. Memory 1504 may store information encoding how the nodes of one or more DNNs are connected to each other. Memory 1504 may store instructions to perform one or more operations of one or more DNNs. Memory 1504 may store model definitions specifying one or more operations of the DNNs. Memory 1504 may store machine-readable instructions, such as configuration descriptors, generated by a compiler based on the model definitions.
[0192] In some embodiments, computing device 1500 may include communication device 1512 (e.g., one or more communication devices). For example, communication device 1512 may be configured to manage wired and / or wireless communications to transmit data to and from computing device 1500. The term “wireless” and its derivatives may be used to describe circuits, devices, systems, methods, techniques, communication channels, etc., that can transmit data over a non-solid medium using modulated electromagnetic radiation. The term does not imply that the associated device does not contain any wires, although in some embodiments the associated device may not contain any wires. Communication device 1512 may implement any of a number of wireless standards or protocols, including but not limited to Institute of Electrical and Electronics Engineers (IEEE) standards (including Wi-Fi (IEEE 802.10 series), IEEE 802.16 standards (e.g., IEEE 802.16-2005 revision)), Long Term Evolution (LTE) projects, and any revisions, updates, and / or modifications (e.g., Advanced LTE project, Ultra Mobile Broadband (UMB) project (also known as “3GPP2”), etc.). IEEE 802.16 compliant Broadband Wireless Access (BWA) networks are commonly referred to as WiMAX networks. WiMAX stands for Global Microwave Access Interoperability and is a certification mark for products that have passed the IEEE 802.16 standard's compliance and interoperability testing. Communication equipment 1512 can operate according to Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Universal Mobile Telecommunications System (UMTS), High-Speed Packet Access (HSPA), Evolved HSPA (E-HSPA), or LTE networks. Communication equipment 1512 can operate according to Enhanced GSM Evolved Data (EDGE), GSM EDGE Radio Access Network (GERAN), Universal Terrestrial Radio Access Network (UTRAN), or Evolved UTRAN (E-UTRAN). Communication equipment 1512 can operate according to Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Digital Enhanced Cordless Telecommunications (DECT), Evolved Data Optimized (EV-DO), and their derivatives, as well as any other wireless protocols designated as 3G, 4G, 5G, and higher. Communication device 1512 may operate according to other wireless protocols in other embodiments. Computing device 1500 may include antenna 1522 to facilitate wireless communication and / or receive other wireless communications (such as radio frequency transmissions). Computing device 1500 may include receiver circuitry and / or transmitter circuitry. In some embodiments, communication device 1512 may manage wired communication, such as electrical, fiber optic, or other suitable communication protocols (e.g., Ethernet). As described above, communication device 1512 may include multiple communication chips.For example, the first communication device 1512 may be dedicated to short-range wireless communication such as Wi-Fi or Bluetooth, and the second communication device 1512 may be dedicated to long-range wireless communication such as Global Positioning System (GPS), EDGE, GPRS, CDMA, WiMAX, LTE, EV-DO, etc. In some embodiments, the first communication device 1512 may be dedicated to wireless communication, and the second communication device 1512 may be dedicated to wired communication.
[0193] The computing device 1500 may include a power supply / power circuit 1514. The power supply / power circuit 1514 may include one or more energy storage devices (e.g., batteries or capacitors) and / or circuitry for coupling components of the computing device 1500 to an energy source (e.g., a DC power supply, an AC power supply, etc.) that is separate from the computing device 1500.
[0194] The computing device 1500 may include a display device 1506 (or a corresponding interface circuit as discussed above). For example, the display device 1506 may include any visual indicator, such as a head-up display, computer monitor, projector, touch screen display, liquid crystal display (LCD), light-emitting diode display, or flat panel display.
[0195] The computing device 1500 may include an audio output device 1508 (or a corresponding interface circuit as discussed above). For example, the audio output device 1508 may include any device that generates an audible indicator, such as a speaker, headphones, or earphones.
[0196] The computing device 1500 may include an audio input device 1518 (or a corresponding interface circuit as discussed above). The audio input device 1518 may include any device that generates a signal representing sound, such as a microphone, microphone array, or digital instrument (e.g., an instrument with a Music Instrument Digital Interface (MIDI) output).
[0197] The computing device 1500 may include a GPS device 1516 (or a corresponding interface circuit as discussed above). The GPS device 1516 may communicate with a satellite-based system and may receive the location of the computing device 1500, as known in the art.
[0198] The computing device 1500 may include a sensor 1530 (or one or more sensors). The computing device 1500 may include the corresponding interface circuitry discussed above. The sensor 1530 can sense physical phenomena and convert them into electrical signals that can be processed by, for example, the processing device 1502. Examples of the sensor 1530 may include: capacitive sensors, inductive sensors, resistive sensors, electromagnetic field sensors, light sensors, cameras, imagers, microphones, pressure sensors, temperature sensors, vibration sensors, accelerometers, gyroscopes, strain sensors, moisture sensors, humidity sensors, distance sensors, ranging sensors, time-of-flight sensors, pH sensors, particle sensors, air quality sensors, chemical sensors, gas sensors, biosensors, ultrasonic sensors, scanners, etc.
[0199] The computing device 1500 may include another output device 1510 (or a corresponding interface circuit as discussed above). Examples of the other output device 1510 may include: an audio codec, a video codec, a printer, a wired / wireless transmitter for providing information to other devices, a haptic output device, a gas output device, a vibration output device, a lighting output device, a home automation controller, or an additional storage device.
[0200] The computing device 1500 may include another input device 1520 (or a corresponding interface circuit as discussed above). Examples of the other input device 1520 may include: an accelerometer, a gyroscope, a compass, an image acquisition device, a keyboard, a cursor control device such as a mouse, a stylus, a touchpad, a barcode reader, a quick-response (QR) code reader, any sensor, or a radio frequency identification (RFID) reader.
[0201] The computing device 1500 may take any desired form, such as a handheld or mobile computer system (e.g., a mobile phone, smartphone, mobile internet device, music player, tablet computer, laptop computer, netbook computer, personal digital assistant (PDA), personal computer, remote control, wearable device, headgear, glasses, shoes, electronic clothing, etc.), desktop computer system, server or other networked computing component, printer, scanner, monitor, set-top box, entertainment control unit, vehicle control unit, digital camera, digital video recorder, Internet of Things device, or wearable computer system. In some embodiments, the computing device 1500 may be any other electronic device that processes data.
[0202] Select Example
[0203] Example 1 provides an integrated circuit for accelerating the multiplication and accumulation operations of activations and weights in an N:M sparsity pattern, comprising: a digital in-memory computation macro having in-memory computation cells of X rows and Y columns, the digital in-memory computation macro being arranged into P number of submacros, wherein the submacros of the P number of submacros have X divided by P rows and Y columns of in-memory computation cells, and the in-memory computation cells have 2-1 multiplexers to select one of two activations to be multiplied by weights stored in the in-memory computation cells; an activation buffer that buffers activations according to N and M; an allocation network that receives activations from the activation buffer, the allocation network including P number of P-1 multiplexers having P outputs respectively to the P number of submacros; and a merging network that sums the P partial sums computed by the P number of submacros.
[0204] Example 2 provides the integrated circuit of Example 1, wherein the P-1 multiplexer of P number has: P inputs, each of which receives two activation words from an activation buffer; and an output that outputs two selected activation words.
[0205] Example 3 provides an integrated circuit of Example 1 or Example 2, wherein a P-1 multiplexer of a P-number receives an M-number of activations in pairs from an activation buffer.
[0206] Example 4 provides an integrated circuit from any of Examples 1-3, wherein P-number P-1 multiplexers receive activations from N identical sets arranged in pairs.
[0207] Example 5 provides an integrated circuit from any of Examples 1-4, and further includes a controller that outputs selection signals for a P-number of P-1 multiplexers.
[0208] Example 6 provides the integrated circuit of Example 5, wherein the selection signal for the P-number of P-1 multiplexers is generated by the controller based on metadata encoded with the coordinates of the dense weights.
[0209] Example 7 provides an integrated circuit from any of Examples 1-6, wherein the submacros, including a number of P submacros, further include: a column controller that outputs a selection signal for a 2-1 multiplexer for in-memory computation of a cell for a given column.
[0210] Example 8 provides an integrated circuit from Example 7, wherein the selection signal of a 2-1 multiplexer for in-memory computation of a cell for a given column is generated by the column controller based on metadata encoded with coordinates of dense weights.
[0211] Example 9 provides an integrated circuit from any of Examples 1-8, wherein rows at the same row position in a number of P submacros share a number of P-1 multiplexers.
[0212] Example 10 provides an integrated circuit of any of Examples 1-9, wherein the in-memory computation cell includes a bit-serial multiplier circuit that multiplies one of the two active cells by a weight.
[0213] Example 11 provides an integrated circuit of any of Examples 1-10, wherein the in-memory compute cell stores an 8-bit memory word.
[0214] Example 12 provides an integrated circuit from any of Examples 1-11, where activation is represented by an 8-bit word.
[0215] Example 13 provides an integrated circuit of any one of Examples 1-12, wherein the in-memory computing cell has a memory mode and a computing mode.
[0216] Example 14 provides a method for accelerating the multiplication and accumulation operations of activations and weights using a digital in-memory computational macro with a number of P submacros in an N:M sparsity pattern. The method includes: buffering activations in pairs to an allocation network according to N and M; selecting a pair of activations at the input of a P-1 multiplexer of the allocation network; outputting the selected pair of activations to the submacro via the P-1 multiplexer; and selecting activations from the selected pair of activations via a 2-1 multiplexer of the in-memory computational cell of the submacro.
[0217] Example 15 provides the method of Example 14, and further includes: performing a multiplication of the selected activation from a pair of selected activations with the weights stored in memory in the computed cell; and adding the P partial sums of the outputs of the P number of submacros.
[0218] Example 16 provides a method of Example 14 or Example 15, wherein pairwise buffered activation includes: buffering M pairs of activations into a P-1 multiplexer of the distribution network.
[0219] Example 17 provides a method from any of Examples 14-16, wherein pairwise buffered activations include: buffering the activations of N identical sets arranged in pairs into the allocation network.
[0220] Example 18 provides a method from any of Examples 14-17, wherein pairwise buffered activation comprises: buffering activations in pairs to the allocation network over multiple clock cycles, wherein during a clock cycle, a subset of activations is buffered to the allocation network, the subset corresponding to a group of rows(multiple) of P number of submacros.
[0221] Example 19 provides a method from any of Examples 14-18, wherein pairwise buffered activation includes: activating pairs of buffers to the allocation network over multiple time periods, wherein, during the time periods, the activation is buffered to columns of P number of sub-macros.
[0222] Example 20 provides a method from any of Examples 14-19, and further includes generating a selection signal for a P-1 multiplexer based on metadata encoded with the coordinates of the dense weights.
[0223] Example 21 provides a method from any of Examples 14-20, and further includes generating a selection signal for a 2-1 multiplexer based on metadata encoded with the coordinates of the dense weights.
[0224] Example 22 provides a method comprising: identifying multiple outliers for a layer of a neural network, the multiple outliers corresponding to multiple weights of the layer of the neural network; determining a count of the outliers; determining a locality measure of the outliers; determining a value A in a sparsity ratio of A / B based on the count; determining a value B in a sparsity ratio of A / B based on the locality measure; and assigning a sparsity pattern of N:M to the layer based on the sparsity ratio of A / B.
[0225] Example 23 provides the method of Example 22, where locality measurement indicates how close or far apart multiple outlying points are.
[0226] Example 24 provides a method from Example 22 or Example 23, wherein determining the locality measurement includes: calculating the locality measurement based on one or more pairwise distances between multiple outliers in the neighborhood.
[0227] Example 25 provides a method from any of Examples 22-24, wherein determining the locality measure includes: calculating the locality measure based on the average of one or more pairwise distances from multiple outliers within the neighborhood.
[0228] Example 26 provides a method from any of Examples 22-25, wherein assigning a sparsity pattern includes: determining an N:M sparsity pattern with the lowest value M that preserves the sparsity ratio of A / B.
[0229] Example 27 provides a method from any of Examples 22-26, wherein assigning a sparsity pattern includes: determining an N:M sparsity pattern with the highest value M that preserves the sparsity ratio of A / B.
[0230] Example 28 provides a method for any of Examples 22-27, wherein assigning a sparsity pattern includes: removing the candidate sparsity pattern with the highest latency; and determining the N:M sparsity pattern from one or more remaining candidate sparsity patterns with the highest value M of the sparsity ratio that retains A / B.
[0231] Example 29 provides a method from any of Examples 22-28, where assigning a sparsity pattern includes assigning a sparsity pattern based on a delay importance category.
[0232] Example 30 provides a method from any of Examples 22-29, wherein determining the value A in the sparsity ratio A / B involves setting the value A higher when the count is higher.
[0233] Example 31 provides a method for any of Examples 22-30, wherein determining the value B in the sparsity ratio A / B includes setting the value B higher when the locality measurement indicates that multiple outliers are more clustered rather than widely spaced.
[0234] Example 32 provides a method from any of Examples 22-31, and also includes: pruning multiple weights of the layer according to the N:M sparsity pattern.
[0235] Example 33 provides a method from any of Examples 22-32, and also includes: using an N:M sparsity pattern for layers to generate machine-readable instructions or configurations for neural networks.
[0236] Example 34 provides one or more non-transitory computer-readable media that store instructions that, when executed by one or more processors, cause one or more processors to perform a method according to any one of Examples 22-33.
[0237] Example 35 provides an apparatus including means for performing a method according to any one of Examples 14-33.
[0238] Example 36 provides an apparatus for accelerating the multiplication and accumulation operations of activations and weights in an N:M sparsity pattern, comprising: an in-memory computation macro having an array of in-memory computation cells and P number of submacros arranged in a subdivided array along a dimension, wherein the in-memory computation cells have 2-1 multiplexers to select one of two activations to be multiplied with weights stored in the in-memory computation cells; a memory interface for retrieving activations from memory; an activation buffer for buffering activations from the memory interface according to N and M; an allocation network for receiving activations from the activation buffer, the allocation network including P number of P-1 multiplexers having P outputs respectively to the P number of submacros; and a merging network for summing P partial sums computed by the P number of submacros.
[0239] Example 37 provides the apparatus of Example 36, wherein a P-1 multiplexer of a number of P multiplexers has: P inputs, each receiving two activation words from an activation buffer; and an output that outputs two selected activation words.
[0240] Example 38 provides the apparatus of Example 36 or Example 37, wherein a P-1 multiplexer of a P-number of P-1 multiplexers receives a pairwise arrangement of M-number of activations from an activation buffer.
[0241] Example 39 provides an apparatus of any of Examples 36-38, wherein P number of P-1 multiplexers receive activations of N number of identical sets arranged in pairs.
[0242] Example 40 provides a digital in-memory computing device for accelerating structured sparse deep neural networks, comprising: a plurality of digital in-memory computation submacros, each submacro including a memory cell configured to store weight data and perform a 1:2 structured sparsity operation; an activation buffer configured to store activation data for processing by the plurality of submacros; an allocation network including a plurality of multiplexers, each multiplexer configured to selectively route activation data from the activation buffer to the corresponding submacro based on sparsity metadata to support multiple N:M structured sparsity patterns; and a merging network configured to combine partial sums of the outputs from the plurality of submacros to generate a final computation result.
[0243] Example 41 provides a computing device in digital memory of Example 40, and also includes one or more aspects of Examples 1-13.
[0244] Variations and other notes
[0245] As used herein, the terms “coupled to” or “coupled with” refer to a relationship between electronic components or circuit elements, wherein the components communicate electrically with each other and are capable of transmitting and / or receiving electrical signals therebetween. The term “coupled to” does not require a direct physical or electrical connection between the coupled components. Rather, “coupled to” can encompass an arrangement in which components are connected via one or more intermediary elements, components, circuits, or transmission paths. For example, a first component may be “coupled to” a second component via an intermediate component (such as a resistor, capacitor, inductor, transistor, logic gate, bus, transformer, or other electronic component) or via an intermediate transmission path, while still maintaining the ability to communicate electrically between the first and second components.
[0246] The foregoing description of implementations of this disclosure (including those described in the abstract) is not intended to be exhaustive or to limit this disclosure to its precise form. While specific implementations of this disclosure and examples used in this disclosure have been described herein for illustrative purposes, various equivalent modifications may be made within the scope of this disclosure, as will be recognized by those skilled in the art. These modifications may be made to this disclosure based on the detailed description above.
[0247] For illustrative purposes, specific figures, materials, and configurations have been set forth to provide a full understanding of the illustrative implementation. However, it will be apparent to those skilled in the art that this disclosure may be practiced without specific details and / or only some of the aspects that may be described in this disclosure. In other instances, well-known features have been omitted or simplified so as not to obscure the illustrative implementation.
[0248] Furthermore, reference is made to the accompanying drawings, which form part of this document and illustrate, by way of illustration, embodiments that can be practiced. It should be understood that other embodiments may be utilized, and structural or logical changes may be made, without departing from the scope of this disclosure. Therefore, the following detailed description is not to be considered limiting.
[0249] Various operations can be described in a manner most conducive to understanding the disclosed subject matter as a series of discrete actions or operations in sequence. However, the order of description should not be construed as implying that these operations must be sequentially dependent. In particular, these operations may not be performed in the order presented. The described operations may be performed in a different order than in the described embodiments. Various additional operations may be performed or the described operations may be omitted in additional embodiments.
[0250] For the purposes of this disclosure, the phrase "A or B" or the phrase "A and / or B" means (A), (B), or (A and B). For the purposes of this disclosure, the phrase "A, B, or C" or the phrase "A, B, and / or C" means (A), (B), (C), (A and B), (A and C), (B and C), or (A, B, and C). The term "between" when used with respect to a measurement range includes the endpoints of the measurement range.
[0251] The description uses the phrases "in one embodiment" or "in an embodiment," which may refer to one or more of the same or different embodiments. The terms "comprising," "including," "having," etc., are synonymous when used with respect to embodiments of this disclosure. This disclosure may use perspective-based descriptions (such as "above," "below," "top," "bottom," and "side") to explain various features of the drawings, but these terms are for convenience of discussion only and do not imply a desired or required orientation. The drawings are not necessarily drawn to scale. Unless otherwise specified, the use of ordinal adjectives such as "first," "second," "third," etc., to describe common objects merely indicates that different instances of similar objects are referenced and is not intended to imply that objects so described must be in a given order in time, space, hierarchy, or any other manner.
[0252] In the following detailed description, various aspects of the illustrative implementations will be described using terminology commonly employed by those skilled in the art, in order to convey the substance of their work to those skilled in the art.
[0253] The terms “substantially,” “near,” “approximately,” “near,” and “about” generally refer to within ±20% of the target value as described herein or as known in the art. Similarly, terms indicating the orientation of individual elements (e.g., “coplanar,” “perpendicular,” “orthogonal,” “parallel,” or any other angle between elements) generally refer to within ±5%–20% of the target value as described herein or as known in the art.
[0254] Furthermore, the terms “comprising,” “including,” “containing,” “having,” “with,” or any other variations thereof are intended to cover non-exclusive inclusion. For example, a method, process, or apparatus that includes a list of elements is not necessarily limited to only those elements, but may include other elements not expressly listed or inherent to the method, process, or apparatus. Additionally, the term “or” refers to an inclusive “or,” not an exclusive “or.”
[0255] The systems, methods, and apparatus of this disclosure have several innovative aspects, and no single one of these innovative aspects can individually account for all the expected properties disclosed herein. Details of one or more implementations of the subject matter described in this specification are set forth in the description and accompanying drawings.
Claims
1. An integrated circuit, said integrated circuit using an N:M sparsity mode to accelerate activation and weight multiplication and accumulation operations, comprising: A digital memory-in-computation macro has X rows and Y columns of in-memory-in-computation cells. The digital memory-in-computation macro is arranged into P sub-macros, wherein each of the P sub-macros has X divided by P rows and Y columns of in-memory-in-computation cells, and each in-memory-in-computation cell has a 2-1 multiplexer to select one of two activations to be multiplied by a weight stored in the in-memory-in-computation cell. An activation buffer is provided, which buffers the activation based on N and M; An allocation network that receives the activation from the activation buffer, the allocation network comprising P P-1 multiplexers, each P-1 multiplexer having P outputs respectively to the P-1 sub-macros; and A merging network that sums P partial sums calculated from the P number of submacros.
2. The integrated circuit according to claim 1, wherein, The P-1 multiplexers of the P number have the following characteristics: P input terminals, wherein each input terminal receives two activation words from the activation buffer; and The output terminal outputs two selected activation words.
3. The integrated circuit according to claim 1 or 2, wherein, The P-1 multiplexers of the P number receive M number of activations arranged in pairs from the activation buffer.
4. The integrated circuit according to claim 1 or 2, wherein, The P-number of P-1 multiplexers receive activations from the same set of N pairs.
5. The integrated circuit according to claim 1 or 2, further comprising: A controller that outputs selection signals for the P number of P-1 multiplexers.
6. The integrated circuit according to claim 5, wherein, The selection signal for the P-number of P-1 multiplexers is generated by the controller based on metadata encoded from the coordinates of dense weights.
7. The integrated circuit according to claim 1 or 2, wherein the submacros in the P-number of submacros further include: A column controller that outputs a selection signal for a 2-to-1 multiplexer for in-memory computation of a given column.
8. The integrated circuit according to claim 7, wherein, The selection signal of the 2-1 multiplexer used for in-memory computation of the given column is generated by the column controller based on metadata encoded with coordinates of dense weights.
9. The integrated circuit according to claim 1 or 2, wherein, Rows at the same row position in the P number of submacros share the P number of P-1 multiplexers.
10. The integrated circuit according to claim 1 or 2, wherein, The in-memory computing cell includes a bit-serial multiplier circuit that multiplies one of the two activations with the weight.
11. The integrated circuit according to claim 1 or 2, wherein, The memory cell stores 8-bit memory words.
12. The integrated circuit according to claim 1 or 2, wherein, The activation is represented by an 8-bit word.
13. The integrated circuit according to claim 1 or 2, wherein, The in-memory computing cell has a memory mode and a computing mode.
14. A method that uses in-digital memory computation of macros having a number of P sub-macros in an N:M sparsity pattern to accelerate activation and weight multiplication and accumulation operations, the method comprising: Based on N and M, the activation pairs are buffered into the allocation network; At the input of the P-1 multiplexer in the distribution network, select a pair to activate; The selected pair of activations is output to the submacro via the P-1 multiplexer; as well as The 2-1 multiplexer of the cell is computed within the memory of the sub-macro, and an activation is selected from the selected pair of activations.
15. The method of claim 14, further comprising: Perform a multiplication of the selected activation from the chosen pair of activations with the weights stored in the memory for calculating the cell; as well as The P parts output by the P number of submacros are summed together.
16. The method according to claim 14 or 15, wherein, The activation of the paired buffer includes: The M pairs of activation buffers are fed into the P-1 multiplexer of the distribution network.
17. The method according to claim 14 or 15, wherein, The activation of the paired buffer includes: The activation buffers of N identical sets arranged in pairs are fed into the allocation network.
18. The method according to claim 14 or 15, wherein, The activation of the paired buffer includes: Over multiple clock cycles, the activations are buffered in pairs into the allocation network, wherein during a clock cycle, the activations are buffered into a subset of the allocation network, the subset corresponding to a group of rows of the P number of sub-macros.
19. The method according to claim 14 or 15, wherein, The activation of the paired buffer includes: The activations are buffered in pairs to the allocation network over multiple time periods, wherein during the time period, the activations are buffered to columns of the P number of sub-macros.
20. The method according to claim 14 or 15, further comprising: Based on metadata encoded with the coordinates of the dense weights, a selection signal for the P-1 multiplexer is generated.
21. The method according to claim 14 or 15, further comprising: Based on metadata encoded with the coordinates of the dense weights, a selection signal for the 2-1 multiplexer is generated.
22. An apparatus comprising means for performing the method according to any one of claims 14-21.
23. An apparatus for accelerating activation and weight multiplication and accumulation operations in an N:M sparsity pattern, comprising: A digital memory-in-computation macro having an array of in-memory computing cells and P number of sub-macros arranged to subdivide the array along a dimension, wherein the in-memory computing cells have 2-1 multiplexers to select one of two activations to be multiplied with a weight stored in the in-memory computing cell; A memory interface that retrieves the activation from memory; An activation buffer, which buffers the activation from the memory interface according to N and M; An allocation network that receives the activation from the activation buffer, the allocation network comprising P P-1 multiplexers, each P-1 multiplexer having P outputs respectively to the P-1 sub-macros; and A merging network that sums P partial sums calculated from the P number of submacros.
24. The apparatus according to claim 23, wherein, The P-1 multiplexers of the P number have the following characteristics: P input terminals, wherein each input terminal receives two activation words from the activation buffer; and The output terminal outputs two selected activation words.
25. The apparatus according to claim 23 or 24, wherein, The P-1 multiplexers of the P number receive M pairs of activations from the activation buffer, and The P-number of P-1 multiplexers receive activations from the same set of N pairs.