Sparsity-Aware Neural Processing Unit and Method with Constant Probability Matching Metrics

By introducing index matching units and priority encoder into sparse convolutional neural networks, and arranging and matching input activation and weighting matrices are solved, the problem that matrix multiplication calculation efficiency is affected by density changes is achieved, and the stability of multiplier utilization and performance guarantee of the sparseness model is achieved.

CN113469323BActive Publication Date: 2025-06-24POSTECH ACADEMY INDUSTRY FOUNDATION
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011353674.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-03-30
Filing Date
2020-11-27
Publication Date
2025-06-24
Estimated Expiration
2040-11-27

AI Technical Summary

Technical Problem

In prior art In sparse convolutional neural networks, the efficiency of matrix multiplication calculation is affected by input activation and weighted matrix density changes, resulting in unstable utilization of the multiplier and the performance of the sparse model cannot be maintained constantly.

Method used

By arranging and matching input activation and weighting matrices by constant probability in the sparse perceptual neural processing unit, the utilization of the multiplier is always stable.

Benefits of technology

It realizes that the utilization rate of the multiplier is maintained constantly under different input activation and weighted matrix density conditions, avoids fluctuations in computing efficiency, and ensures the stable performance of the sparse model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113469323B_ABST
    Figure CN113469323B_ABST
Patent Text Reader

Abstract

The present invention relates to a sparsity-aware neural processing unit that performs index matching of a constant probability regardless of IA and weighted density, and a processing method thereof. According to an embodiment of the present invention, the sparsity-aware neural processing unit processing method may include: receiving a plurality of input activations (IA); obtaining weights with non-zero values from each weighted output channel; obtaining an input channel index, wherein the input channel index stores the weights and IA in a memory and includes a memory address location where the weights and IA are stored; and in an index matching unit (IMU) including a buffer memory storing the input channel index, arranging the weights with non-zero values of each weighted output channel according to the row size of the IMU and matching the weights and IA.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a sparsity-aware neural processing unit that performs index matching with a constant probability regardless of input activation (IA) and weighting, and a processing method thereof. Background Art

[0002] In Figure 1a and Figure 1b the existing literature shown, a Convolution Neural Network uses the input activation output by applying a weighted filter as the input to the output layer, and requires multiple matrix multiplications in a large-dimensional space. Among them, when IA is a zero value, a Sparse Convolution Neural Network performs pruning and uses a calculation method of performing multiple array multiplications only with non-zero values.

[0003] That is, in IA and weighting, the sparse convolution neural network approximates values close to 0 as 0, removes zero values, retains only non-zero values, finds indexes having non-zero values, and after importing the values at the corresponding indexes into a processing element, uses a calculation multiplication superposition calculation method. Through this, the number of memory communications and the number of multiplication calculations for a neural network algorithm can be considerably reduced, and matrix multiplication between input activation and weighting can be effectively performed.

[0004] However, according to the existing literature, the density of the IA and weighting matrices is shown differently for each layer, and in order to effectively perform matrix multiplication between IA and weighting using these sparse matrices, it is necessary to maintain a constant probability without being restricted by the density changes of the IA and weighting arrays of each layer for the multiplication calculation.

[0005] That is, there is a need for an invention for constantly maintaining the performance of a model having sparse variables.

[0006] Prior Art Documents

[0007] Non-Patent Documents

[0008] Zhang Jiefang et al., "SNAP: A 1.67 - 21.55 TOPS / W Sparse Neural Accelerator Processor for Unstructured Sparse Deep Neural Network Inference in 16nm CMOS", 2019 Symposium on VLSI Circuits, IEEE, 2019. Summary of the Invention

[0009] The utilization of a multiplier according to the density of an IA and a weight matrix changes sharply. It relates to a sparsity-aware neural processing unit and a processing method that maintain the utilization of the multiplier with a constant probability regardless of the change in the density of the IA and the weight matrix, and arrange weighted input and output channels in a matrix.

[0010] According to an embodiment of the present invention, there is provided a sparsity-aware neural processing unit processing method, which includes: a step of receiving a plurality of input activations (IAs); a step of obtaining weights with non-zero values from each weighted output channel; a step of obtaining an input channel index, where the input channel index stores the weights and IAs in a memory and includes a memory address location where the weights and IAs are stored; and a step of arranging the non-zero value weights of each weighted output channel according to the row size of an index matching unit (IMU) including a buffer memory storing the input channel index and matching the weights and IAs in the IMU.

[0011] Furthermore, the IMU includes a comparator array, a weight buffer memory, and an IA buffer memory. The IA buffer memory storing the IA input channel index is arranged in the column direction at the boundary of the comparator array, and the weight buffer memory storing the non-zero value weight input channel index is arranged in the row direction so that the weights and IAs can be matched.

[0012] At this time, in the weight buffer memory, the non-zero value weights of each weighted output channel are arranged in ascending order. When exceeding at least one of the input channel size of the output channel of the weight and the number of non-zero value weights, they are arranged to move to the next output channel of the weight, or the IA buffer memory matches each input channel index of the IA one-to-one with a pixel dimension and arranges each pixel dimension in ascending order.

[0013] According to an embodiment of the present invention, it further includes: in the weighted buffer memory, a step of determining an average number n of non-zero input channel indicators for each weighted output channel according to a density d between the weighting and the IA; a step of determining an average number m of weighted output channels of the weighted buffer memory according to the average number n of non-zero input channel indicators, where n and m are determined by the following formula, n = s * d, where s may be the input channel size of the weighted output channel, and m = s' / n, where s' is the column size of the IMU. For each CLK cycle, an average number Pm of weighted input channel indicators matching the IA input channel indicators is maintained by the following formula regardless of the density d between the weighting and the IA, Pm = d * m = d * (s' / n) = d * (s' / s / d) = s' / s, and when s = s' is set, Pm = 1.

[0014] According to another embodiment of the present invention, it may further include: a step of displaying a flag signal of matching indicators between the IA and the weighting, and transmitting the indicators with the most p matches to a p-way priority encoder; and a step of receiving the p-way priority encoder for the IA and weighting pairs with the most p matches, and sequentially transmitting the matched IA and weighting pairs to a First In First Out (FIFO) queue one by one.

[0015] According to another embodiment of the present invention, it may further include: after sending the matched IA and weighting pairs stored in the FIFO to a multiplier, deleting the weighting values sent from the priority encoder to the FIFO in the weightings stored in the weighted buffer memory, and then rearranging the weightings stored in the weighted buffer memory that are not sent to the FIFO.

[0016] As another form of the present invention, in a Processing Element, a sparsity-aware neural processing unit may include: an IMU including a buffer memory and a comparator array, where the buffer memory includes a weighted buffer memory and an IA buffer memory for storing non-zero values of the weighting and the IA, and the comparator array matches the indicators of the weighting and the IA; a p-way priority encoder receiving the weighting and IA pairs with the most p matches from the comparator array; and a FIFO receiving the weighting and IA pairs matched from the p-way priority encoder one by one and transmitting them to a multiplier one by one.

[0017] Through the weighted output channel arrangement according to the present invention, when the matrix density between the IA and the weighting is low, the matched IA and weighting pairs cannot be sufficiently supplied to the IMU, resulting in a reduction in the utilization rate of the multiplier. When the density is high, all calculations of the IA and weighting pairs for the matched multiplier cannot be performed, thereby preventing a bottleneck phenomenon in the memory and maintaining the utilization rate of the multiplier constantly. Description of the Drawings

[0018] Figure 1a and Figure 1b show a graph of the multiplier utilization variation according to an existing sparse sensing neural processing unit and IA / weight density.

[0019] Figure 2 show an IMU (Index Matching Unit) and a p-way priority encoder according to an embodiment of the present invention.

[0020] Figure 3 show the processing process of the IMU outputting IA and weight pairs matched according to IA and weight density according to an embodiment of the present invention.

[0021] Figure 4 show the processing process of the IMU outputting IA and weight pairs matched according to IA and weight density according to another embodiment of the present invention.

[0022] Figure 5 show the processing process of the IMU outputting IA and weight pairs matched according to IA and weight density according to another embodiment of the present invention.

[0023] Figure 6 show the processing process of a processing element (Processing Element) including an IMU and matching IA and weight pairs of the processing element according to an embodiment of the present invention.

[0024] Figure 7 show a block diagram of a processing element performing a processing process according to an embodiment of the present invention.

[0025] Symbol Description

[0026] 100: IMU 101: Comparator array

[0027] 102: p-way priority encoder 103: Weight buffer memory

[0028] 104: IA buffer memory 105: FIFO

[0029] 106: Multiplier Detailed Description of the Invention

[0030] Refer to Figure 1a, in the existing Index Matching Unit (IMU) 10 for input activation (IA) and weighting, for IA and weighting, the IA is arranged in a (1xN) matrix, the weighting W is arranged in an (Nx1) matrix, and the indices of IA and weighting are matched. For the (1xN) matrix of IA, the matching is performed in the order of (W1), 2(W2), 12(W3), …, WN, and at most one matching index item of the matched IA and weighting pair is transmitted to the priority encoder. The multiplier successively receives the matched indices of the IA and weighting pairs matched by the priority encoder and performs calculations.

[0031] In the weighting arranged in the (Nx1) matrix, depending on the number of non-zero weightings according to the weighting density, the lower the weighting density, the significantly smaller the number of non-zero weightings compared to N.

[0032] If zero values are deleted from the weighting and the IA and index matching are performed in a state where only non-zero values remain, when the matrices of IA and weighting have a low density, the number of matched IA and weighting pairs is small, so the IA and weighting pairs supplied from the IMU to the multiplier also decrease, and the efficiency of the multiplier decreases. On the contrary, when the matrices of IA and weighting have a high density, the number of matched IA and weighting pairs exceeds the capacity of the multiplier, and the matched IA and weighting pairs cannot be transmitted to the multiplier per CLOCK cycle and accumulate in the IMU, and the multiplier cannot process all of them at once, resulting in a bottleneck phenomenon.

[0033] Observation Figure 1b shows that when the IA / weighting density is 1.0 / 1.0, the multiplier can be utilized nearly 100% when the size of the comparator array is 64x64, 32x32, or 16x16. However, when the weighting density is higher than this, the matched IA and weighting pairs that have already used up all the multiplier performance continue to be superimposed, so the matched pairs transmitted from the IMU to the priority encoder continue to accumulate, causing the IMU to stop.

[0034] On the contrary, when the IA / weighting density is 0.1 / 0.1, the smaller the size of the comparator array, the less than 10% of the multiplier performance can be utilized, resulting in a sharp decrease in the calculation efficiency.

[0035] In addition, according to the matrix density of IA and weighting, multiple matches between weightings can occur for one IA. However, when using the priority encoder, the IMU only receives one matched IA and weighting pair in each cycle. Therefore, until all the multiple matched IA and weighting pairs are transmitted to the multiplier, the unmatched IA and weighting indices wait in the IMU for multiple cycles, resulting in a decrease in the efficiency of the entire processing unit.

[0036] Hereinafter, a detailed description will be given according to an embodiment of the present invention. According to an embodiment of the present invention, the matching probability between IA and weighting can be constantly maintained so that the change in the efficiency of the IA and the weighting matrix density multiplier is not significant and a constant efficiency can be maintained.

[0037] Figure 2 As an illustration of the IMU 100 according to an embodiment of the present invention, an IA and a weighting matrix arrangement are shown.

[0038] As Figure 2 shown, the IMU 100 includes an IA buffer memory 104 for storing IA and weighting, and a weighting buffer memory 103, and may include a comparator array 101 for comparing indices between IA and weighting and performing matching.

[0039] The size of the comparator array 101 can be changed according to the user's settings. In the following embodiments of the present invention, it is assumed to be 32x32 for illustration. This is only an exemplary embodiment, and the scope of the claims is not limited to the embodiment.

[0040] As Figure 2 shown, the data elements of 1, 12, 3, 4, 10,..., 12 arranged in a (32x1) matrix in the IA buffer memory 104 represent the input channel index values of data activation (IA). Each row corresponds to the depth of the input of the neural network layer where the IA is located. In IA#n, n represents the pixel-dimension position of each IA, and n can be a value from 0 to 31.

[0041] Furthermore, the data elements of 0, 2, 12, 1, 3, 10, 14,..., 4, 6 arranged in a (1x32) matrix in the weighting buffer memory 103 are the input channel index values of the weighting obtained by approximating zero values as 0 and arranging the remaining non-zero values. The non-zero value weightings are represented in a matrix structure. In W#m, m represents the index of the weighting output channel, and m can be determined according to the density of the weighting. The process of determining the value of m according to the weighting density will be described Figures 3 to 5 together.

[0042] The weighting matrix arranges non-zero values from low to high in each output channel of the weighting, compares the indices of the input channels of the IA and the weighting matrix, and can match the consistent values.

[0043] When the IA and the weights are matched, the matched IA and weights are transmitted to the p-way priority encoder 102, and the matched IA and weights can be one or more. When there is one pair of matched IA and weights, one pair is transmitted to the p-way priority encoder 102. When there are multiple pairs of matched IA and weights, at most p pairs of the multiple pairs can be transmitted to the p-way priority encoder 102 at one time.

[0044] Therefore, according to an embodiment of the present invention, it may include the steps of receiving a plurality of input activations (IA); obtaining weights with non-zero values from each weighted output channel; obtaining an input channel index, wherein the input channel index stores the weights and IA in a memory and includes a memory address location storing the weights and IA; in an IMU (Index Matching Unit) including a buffer memory storing the input channel index, arranging the non-zero value weights of each weighted output channel according to the row size of the IMU and matching the weights and IA.

[0045] In addition, according to an embodiment of the present invention, the IMU includes a comparator array, a weight buffer memory 103, and an IA buffer memory 104. The IA buffer memory 104 storing the IA input channel indexes is arranged in the column direction at the boundary of the comparator array, and the weight buffer memory 103 storing the non-zero value weight input channel indexes is arranged in the row direction so as to match the weights and IA.

[0046] Specifically, the non-zero value weights of each weighted output channel in the weight buffer memory 103 are configured in ascending order. When at least one of the input channel size of the output channel exceeding the weight and the number of non-zero value weights is satisfied, that is, when the number of non-zero value weights exceeds the input channel size of the output channel of the weight, it can be arranged to move to the next output channel of the weight. The IA buffer memory matches each input channel index of the IA one-to-one with the pixel dimension and arranges each pixel dimension in ascending order.

[0047] Figures 3 to 5 Shows according to Figure 2 When arranging IA and weights in a matrix, the probability of outputting pairs of IA and weights matched with the IMU100 according to the IA and weight density. Also, Figures 3 to 5 For ease of explanation, the IA is represented by only one input index channel out of 0 to 31 input index channels.

[0048] Figure 3The weighted density is 0.1 (d = 0.1). The weighted density of 0.1 means that among the weighted input channel indices 0 - 31, the probability of having a non-zero value is 1 out of 10. That is, when the density of non-zero values in the weighted input channels of size 32 is 0.1, the number of weighted non-zero values can be 32 * 0.1 = 3.2. However, this is an average probability, so it is not necessarily the case that there must be 3.2. As Figure 3 shown, there are three non-zero values 0, 2, 12 in W#0, four input channel indices 1, 3, 10, 14 in W#1, and two channel input indices 4, 6 in the case of W#9.

[0049] In addition, according to an embodiment of the present invention, in the 32x32 matrix between IA and weighting, the columns of the matrix can be filled with 32 weighted non-zero values. In the weighting of size [0:31], only the remaining non-zero values are arranged by an approximation method, and the columns of the matrices W#0, W#1, …, W#9 are filled in the order of the lower output channel indices, and the weighted channel input indices of size 32 of the matrix can be filled.

[0050] As a result, the non-zero input channel index values of the weighted W#m with an average of 3.2 non-zero values are arranged according to the matrix size, so that a total of 10 weighted output channels from 0 - 9 can be arranged.

[0051] However, as described above, this is a probability and an average of 10 weighted output channels can be arranged, and it is not necessarily the case that 10 weighted output channels must be arranged. Therefore, the scope of the claims is not limited thereto.

[0052] Therefore, as Figure 3 shown, in the state where a total of 10 weighted output channels are arranged from W#0 - W#9, there are 10 weighted output channels W#m with a matching probability of 0.1 for the input channel index 12 of IA#1, and an average of one match between the IA and the weighted matrix (1x32) of 0.1 * 10 = 1 can be performed.

[0053] To prevent the matching probability between IA and weighting from decreasing due to a low weighted density and the efficiency of the multiplier from sharply decreasing, the weighted non-zero values are filled in the weighted output channels of the matrix size, so as to constantly maintain the matching probability of the entire IMU matrix, and through this, the efficiency of the multiplier can be constantly maintained.

[0054] Figure 4 The weighted density is 0.5 (d = 0.5). The weighted density of 0.5 means that among the weighted input channel indices 0 - 31, the probability of having a non-zero value is 5 out of 10. That is, when the density of non-zero values in the weighted input channels of size 32 is 0.5, the number of weighted non-zero values can be 32 * 0.5 = 16. However, this is an average probability, so it is not necessarily the case that there must be 16. As Figure 4As shown, there are 16 non-zero values of W#0 existing as 0, 1, 3, …, 30, and 16 input channel indicators for W#1 existing as 2, 3, …, 28, 29. From this, it can be seen that two weighted output channels are used to fill the columns of the 1x32 matrix.

[0055] That is, according to an embodiment of the present invention, in the 32x32 matrix between IA and weighting, the columns of the matrix can be filled with 32 weighted non-zero values. Among the weightings of size [0:31], only the remaining non-zero values are arranged by an approximation method, and the columns of the W#0 and W#1 matrices are filled in the order of the low output channel indicators, and the size of the matrix that can be filled is 32 weighted channel input indicators.

[0056] As a result, the non-zero input channel indicator values of the weighting W#m with an average of 16 non-zero values are arranged according to the matrix size, so that a total of 2 weighted output channels of 0-1 can be arranged.

[0057] However, as described above, this is a probability and an average of 2 weighted output channels can be arranged, and it is not necessarily required to arrange 2 weighted output channels. Therefore, the scope of the claims is not limited thereto.

[0058] Therefore, as Figure 4 shown, in the state where a total of 2 weighted output channels are arranged from W#0 - W#1, there are 2 weighted output channels W#m with a matching probability of 0.5 for the input channel indicator 1 of IA#0, and an average of one match between the IA and the weighting matrix (1x32) of 0.5 * 2 = 1 can be performed.

[0059] In order to prevent the matching probability between IA and weighting from decreasing due to low weighting density or too many matches between IA and weighting due to high weighting density, resulting in a bottleneck phenomenon and a sharp reduction in the efficiency of the multiplier, the non-zero value weightings are filled into the weighted output channels of the matrix size, so as to constantly maintain the matching probability of the entire IMU matrix, and through this, the efficiency of the multiplier can be constantly maintained.

[0060] Figure 5 The weighting density of Figure 5 is 1.0 (d = 1.0). A weighting density of 1.0 means that among the input channel indicators 0 - 31 of the weighting, the probability of having non-zero values is 10 out of 10. That is, when the density of non-zero values in the weighted input channels of size 32 is 1.0, 32 * 1.0 = 32 weighted non-zero values can exist. However, this is an average probability, so it is not necessarily required to have 32. As

[0061] ​That is, according to an embodiment of the present invention, in the 32x32 matrix between IA and weighting, the columns of the matrix can be filled with 32 weighted non-zero values. Among the weightings of size [0:31], only the remaining non-zero values are arranged by an approximation method, and the columns of the W#0 matrix are filled in the order of the low output channel index. When the weighting density is 1, the input channel index can be all non-zero values. Therefore, the size of the matrix that can be filled by only one weighted output channel is 32 weighted channel input indices.

[0062] As a result, the non-zero input channel index values of the weighting W#m with an average of 32 non-zero values are arranged according to the matrix size, so that a total of 1 weighted output channel can be arranged only by the weighted output channel 0.

[0063] However, as described above, this is a probability and an average of 1 weighted output channel can be arranged, and it is not necessarily required to arrange 1 weighted output channel. Therefore, the scope of the claims is not limited thereto.

[0064] Therefore, as Figure 5 shown, in the state where a total of 1 weighted output channel is arranged by W#0, there is 1 weighted output channel W#m whose matching probability with the input channel index 5 of IA is 1.0. An average of one match between IA and weighting can be performed by the IA and weighting matrix (1x32) of 1.0 * 1 = 1.

[0065] To prevent the matching between IA and weighting from increasing due to excessive weighting density, resulting in a bottleneck phenomenon in the transmission of the pair during IMU matching and a sharp reduction in the efficiency of the multiplier, the non-zero weightings are filled in the weighted output channels of the matrix size, thereby constantly maintaining the entire IMU matrix matching probability, and the efficiency of the multiplier can be constantly maintained through this.

[0066] That is, according to an embodiment of the present invention, in the weighting buffer memory 103, the average number n of non-zero input channel indices of each weighted output channel is determined according to the density d between the weighting and IA. According to the average number n of non-zero input channel indices, the average number m of weighted output channels of the weighting buffer memory 103 can be determined, and the n and m can be determined by the following formula.

[0067]

Equation 1

[0068] n = s * d, m = s' / n

[0069] (At this time, s is the input channel size of the weighted output channel, and s' is the column size of the IMU)

[0070] In addition, as a result, in each CLK cycle, the average number Pm of weighted input channel indicators that match each IA input channel indicator is independent of the weight and the density d between the IAs and is maintained by d*m=d*(s' / n)=d*(s' / s / d)=s' / s. At this time, if s=s' is set, Pm=1 can be maintained constantly.

[0071] In other words, Figures 3 to 5 As shown, even if the weight density changes, the average number Pm of IA weights matched can be maintained constant by 1 (=d*m=d*s / n=d*s / (s*d)), so that the multiplier efficiency when the IA and weighted pairs are transmitted from the IMU to the next stage of matching can be maintained constant. Figure 6 As described later, each IA is configured with a p-way priority encoder, FIFO, and multiplier accordingly. When a match is formed between an IA and the weighting by a constant ratio, the overall efficiency of the multiplier is maintained constant, so that there is no difference in computing performance depending on the weighting density.

[0072] Figure 6 The figure shows a processing process including an IA and a weighted pair of a processing element (ProcessingElement) of the IMU 100 and a matching processing element according to an embodiment of the present invention.

[0073] like Figure 6 As shown in (a), a processing element according to an embodiment of the present invention may include an IMU 100 including a matrix arrangement IA and a weighted comparator array 101, a p-way priority encoder 102 allocated to each row of the IMU 100, a FIFO 105, and a multiplier 106.

[0074] Specifically, the IA buffer memory 104 and the weighted buffer memory 103 including IA and weights are arranged in rows and columns, and the IMU 101 compares and matches the IA and weighted indicators through the comparator array 101, and the matched IA and weighted pairs can be transmitted to the p-way priority encoder 102.

[0075] like Figures 3 to 5 As shown, regardless of density, a matching IA and weighted pair may appear in each row, and there may be multiple matching IA and weighted pairs. Each time a matching IA and weighted pair is transmitted to the priority encoder, one cycle is consumed. Therefore, when multiple matching IA and weighted pairs are transmitted one by one, more time is consumed and the efficiency of the processing element is reduced. Therefore, using the p-way priority encoder 102, a maximum of p multiple matching IA and weighted pairs are transmitted to the p-way priority encoder 102 at one time, thereby reducing the waste of consumed cycles.

[0076] The p-way priority encoder 102 transmits the received up to p matching IA and weighted pairs one by one to the FIFO 105, and the FIFO 105 can store them with different depths in the received order. The FIFO 105 transmits the received matching IA and weighted pairs one by one to the multiplier 106, enabling the multiplier 106 to perform calculations.

[0077] At this time, as shown in (b) of Figure 6 During the period of transmitting the matching IA and weighted pairs to the p-way priority encoder 102, the weights used for metric matching are deleted in the comparator array 101, and the unused weights are rearranged in order.

[0078] The maximum number of matching IA and weighted pairs that can be transmitted to the p-way priority encoder 102 is p. If there are more than p matching IA and weighted pairs, p of them are transmitted to the p-way priority encoder 102, and the remaining matching IA and weighted pairs are transmitted to the p-way priority encoder 102, waiting during several cycles, which reduces the efficiency of the processing element.

[0079] Therefore, according to an embodiment of the present invention, when p matching IA and weighted pairs that can be processed by the p-way priority encoder 102 occur, first, the matching IA and weighted pairs are transmitted to the p-way priority encoder 102, the weights transmitted from the p-way priority encoder to the FIFO are deleted, and for the remaining weights not transmitted from the p-way priority encoder to the FIFO, the weight column is rearranged so that matching can be performed again.

[0080] With the processing element configured as described above, although the matching probability between the IA and the weights is constantly maintained in the IMU 100, it is also possible to prevent subsequent processing delays from degrading the performance of the multiplier.

[0081] Figure 7 Shown in Figure 6 from (a) to Figure 6 is a block diagram illustrating the configuration of the processing element described in (b).

[0082] The processing element may include an IMU 100, a p-way priority encoder 102, a FIFO 105, a multiplier, and a psum buffer 106. The IMU 100 may include a weight buffer memory 103 that stores weight input channel metrics, an IA buffer memory 104 that stores IA input channel metrics, and a comparator array 101 that compares and matches the IA and weights stored in the weight buffer memory 103 and the IA buffer memory 104.

[0083] The p-way priority encoder 102 transmits up to p matching IA and weighted pairs received from the IMU 100 to the FIFO 105 one by one. The FIFO 105 stores the received matching IA and weighted pairs according to the depth in order and transmits them to the multiplier 106 one by one.

[0084] The multiplier 106 performs calculations using the received matching IA and weighted pairs. The calculation results of the multiplier array corresponding to each row are sent to the psum buffer to store the combined results.

[0085] The processing element may include more than one processor (e.g., a microprocessor or a central processing unit (CPU)), a graphics processing unit (GPU), an application specific integrated circuit (ASIC), or a combination of other processors. As an example, the processing element may include other storage or computing resources / devices (e.g., buffers, registers, control circuits, etc.) for performing one or more of the decisions and calculations described in the description of the present invention to provide additional processing options.

[0086] In some embodiments, the processing element executes the programmed instructions stored in the memory, enabling the controller and the computing system to perform one or more of the functions described in the description of the present invention. The memory may include more than one non-transitory machine-readable storage medium. The non-transitory machine-readable storage medium may include solid-state memory, magnetic disks and optical disks, portable computer floppy disks, random access memory (RAM), read-only memory (ROM), erasable programmable ROM (e.g., EPROM, EEPROM, or flash memory), or any other medium capable of storing information.

[0087] Generally, the processing element is an exemplary computing unit or type and may include additional hardware structures for performing calculations related to multi-dimensional data structures such as matrices and / or data arrays. In some embodiments, in order to activate the structure, the input activation values may be pre-loaded in the memory, and the weight values may be pre-loaded using data values related to the neural network hardware from an external or upper-level control device.

[0088] In the specification and the figures, the same or similar reference numerals denote the same or similar structures.

[0089] One embodiment of the present invention is merely an exemplary embodiment and is not limited to the above values.

Claims

1. A processing method for a sparsity-aware neural processing unit, wherein, The processing method of the sparsity-aware neural processing unit includes: Receiving a plurality of input activations; Obtaining non-zero weighted values from each weighted output channel; Storing the weighted values and input activations in a memory and obtaining input channel metrics, where the input channel metrics include the memory address locations storing the weighted values and input activations; and Arranging the non-zero weighted values of each weighted output channel according to the row size of the metric matching unit and matching the weighted values and input activations in the metric matching unit, where the metric matching unit includes a buffer memory storing the input channel metrics; Wherein, the metric matching unit includes a comparator array, a weighted buffer memory, and an input activation buffer memory; At the boundary of the comparator array, the input activation buffer memory storing the input activation input channel metrics is arranged in the column direction, and the weighted buffer memory storing the non-zero weighted input channel metrics is arranged in the row direction to match the weighted values and input activations; Wherein, in the weighted buffer memory, the non-zero weighted values of each weighted output channel are arranged in ascending order. When the number of non-zero weighted values exceeds the input channel size of the weighted output channel, the non-zero weighted values are arranged in the next output channel.

2. The processing method of the sparsity-aware neural processing unit according to claim 1, wherein, The input activation buffer memory matches the input channel metrics of each input activation with the pixel size one-to-one and arranges each pixel size in ascending order.

3. The processing method of the sparsity-aware neural processing unit according to claim 1 or 2 further includes: In the weighted buffer memory, determining the average number n of non-zero input channel metrics of each weighted output channel according to the density (d) between the weighted values and input activations; And Determining the average number m of weighted output channels of the weighted buffer memory according to the average number n of non-zero input channel metrics; The n and m are determined by the following formula: n = s * d, where s is the input channel size of the weighted output channel, m = s' / n, where s' is the column size of the metric matching unit.

4. The processing method of the sparsity-aware neural processing unit according to claim 3, wherein, For each clock cycle, the average number (Pm) of weighted input channel metrics matched with the input activation input channel metrics is independent of the density (d) between the weighted values and input activations and is maintained by the following formula: Pm = d * m = d * (s' / n) = d * (s' / s / d) = s' / s. At this time, when s = s' is set, Pm = 1.

5. The processing method of the sparsity-aware neural processing unit according to claim 3 further includes: Displaying a flag signal on the matching metrics of the input activation and the weighted values and transmitting at most p matching metrics to a p-way priority encoder; And The p-way priority encoder receiving at most p matched input activation and weighted value pairs transmits the matched input activation and weighted value pairs to a first-in-first-out queue one by one continuously, Wherein, the p is the maximum number of input activation and weighted value pairs that can be transmitted to the p-way priority encoder all at once in the case of a plurality of input activation and weighted value pairs.

6. The processing method of the sparsity-aware neural processing unit according to claim 5 further includes: Deleting the weights sent by the p-way priority encoder to the first-in-first-out queue from the weights stored in the weight buffer memory, and rearranging the weights stored in the weight buffer memory that are not sent by the p-way priority encoder to the first-in-first-out queue.

7. A sparsity-aware neural processing unit, wherein, The sparsity-aware neural processing unit includes: An index matching unit, which includes a buffer memory and a comparator array. The buffer memory includes a weight buffer memory for storing non-zero weights and an input activation buffer memory for storing non-zero input activations. The comparator array matches the indexes of the non-zero weights and the indexes of the non-zero input activations; A p-way priority encoder, which receives at most p pairs of matching input activations and weights from the comparator array; and A first-in-first-out queue, which receives the matching input activations and weight pairs one by one from the p-way priority encoder and transmits the matching input activations and weight pairs one by one to the multiplier; Wherein, when transmitting at most p pairs of matching input activations and weights to the p-way priority encoder, the index matching unit deletes the weights sent by the p-way priority encoder to the first-in-first-out queue from the weight buffer memory and rearranges the weights that are not sent by the p-way priority encoder to the first-in-first-out queue.

Citation Information

Patent Citations

  • Neural network accelerator based on structured pruning and low-bit quantization

    CN110378468A

  • Deep neural network accelerator with fine-grained parallelism discovery

    US20190370645A1