Hardware Circuit of Neural Network Fully Connected Layer, Its Design Method and Usage Method
By clustering the weights of the fully connected layer of the neural network and designing a parallel binary matrix vector multiplier circuit, the problem of slowing calculation speed and excessive hardware resource usage caused by the increase in the number of weight values is solved, and the effect of improving inference speed and reducing energy consumption is achieved.
Patent Information
- Application Number
- CN202210445789.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-26
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-04-26
AI Technical Summary
In the fully connected layer of a neural network, as the number of weight values increases, serial processing slows down the calculation speed, and parallel processing requires a large amount of hardware resources and increases energy consumption.
By clustering the weights of the fully connected layer, the number of weights is reduced, and the circuit structure of the binary matrix vector multiplier is designed based on the clustering results, and parallel processing is adopted to improve the computing efficiency.
While ensuring the accuracy of neural network inference, it improves inference speed, saves hardware resources, and reduces energy consumption.
Smart Images

Figure CN114997381B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer technology, and particularly relates to a hardware circuit for a fully connected layer of a neural network, and a design method and a usage method thereof. Background Art
[0002] Stochastic computing is a new computing paradigm that exhibits strong fault tolerance for bit reversal. Numbers can be represented by bit streams processed by very simple circuits, and the numbers themselves are interpreted as probabilities under normal and faulty conditions. Therefore, it can implement arithmetic operations at very low cost using standard logic elements. Although stochastic computing has a low implementation cost, as the length of the stochastic sequence increases, the computing time also increases. Therefore, if all the multiplication calculations of a neural network are replaced with stochastic computing methods, it will lead to the problem of excessive computing time.
[0003] Sim, Hyeonuk. "Low-Cost Deep Convolutional Neural Network Acceleration with Stochastic Computing and Quantization." (2021). proposed a binary interfaced matrix-vector multiplier (BISC-MVM). This method adopts a deterministic conversion mode, which can avoid the influence of stochastic fluctuations on stochastic computing. As Figure 1 shown, the eigenvalue x is input into a matrix-vector multiplier, and then a finite state machine and a data selector are used to generate a stochastic sequence for calculation. At the same time, after the weight value w is input into the binary matrix vector, the down-counter is initialized according to the value of the weight value w and starts down-counting. When the value in the down-counter is 0, the up-counter is controlled to stop calculating the number of 1s in the stochastic sequence. At this time, the up-counter outputs the calculation result. This method can terminate the generation of the stochastic sequence in advance, thereby accelerating the stochastic computing process. However, in the above method, the processing of the weight value is a serial processing method. In the case where the number of weight values in the fully connected layer of the current neural network is increasing, it will affect the computing speed and reduce the computing efficiency. Summary of the Invention
[0004] Object of the Invention: The object of the present invention is to propose a design method for a hardware circuit of a fully connected layer of a neural network, reduce the number of weight values through clustering, and optimize the hardware circuit according to the clustering result, so as to improve the inference speed while maintaining the inference accuracy of the neural network.
[0005] Another object of the present invention is to provide a hardware circuit for the fully connected layer of a neural network designed by the above method and its usage method, which can perform different hardware design methods according to actual requirements to achieve an optimization and balance between the use of hardware resources and computational throughput.
[0006] Technical solution: The method for designing a hardware circuit for the fully connected layer of a neural network according to the present invention includes the following steps:
[0007] S1: Cluster all the weights in a fully connected layer of a neural network according to a set clustering interval threshold;
[0008] S2: Use the mean value of each class of weights after clustering as the weight after clustering, and use the set of eigenvalues corresponding to each class of weights after clustering as the corresponding eigenvalue set;
[0009] S3: Design the circuit structure of a binary matrix vector multiplier according to the weights after clustering and the corresponding eigenvalue set.
[0010] Further, the step S1 includes the following steps:
[0011] S1.1: Construct a table H = [(w0, (x0)), (w1, (x1)),...(w n-1 , (x n-1 ))] with the weights W and their corresponding eigenvalues X of a fully connected layer of a neural network;
[0012] S1.2: Sort the table H in ascending order according to the weight size to obtain the table
[0013] S1.3: Initialize the access subscript i = 0, initialize a table A for storing the clustering results, and set the clustering interval threshold τ;
[0014] S1.4: Take out the element from the table in order as the clustering starting point, denoted as α and add it to the table A, and initialize the access subscript m = 1 and the temporary variable t = w i , and define a clustering interval as
[0015] S1.5: If i + m < n, then access the element in the table If w i+m < w i + τ, then go to step S2.2, if not satisfied, jump to step S2.3;
[0016] The step S2 includes the following steps:
[0017] S2.1: If \(i + m\geq n\), then update the weights in \(\alpha\) to \(t / m\), and jump to step S2.4;
[0018] S2.2: Update the temporary variable Add the set \((x i+m )\) to the set of \(\alpha\), update the access index \(m = m + 1\), and jump to step S1.5;
[0019] S2.3: Update \(i = i + m\), update the weight of \(\alpha\) to \(t / m\). If \(i < n\), then jump to step S1.4; if \(i\geq n\), then enter step S2.4;
[0020] S2.4: Output table A, where table A is composed of the clustered weights and their corresponding eigenvalue sets.
[0021] Furthermore, the step S3 includes:
[0022] S3.1: Divide the weights in the clustered table into several groups;
[0023] S3.2: Design a parallel binary matrix - vector multiplier for each group of weights. The parallelism of the binary matrix - vector multiplier corresponding to each group of weights is the maximum length of the eigenvalue set corresponding to this group of weights.
[0024] The hardware circuit designed by the hardware - circuit design method for the fully - connected layer of the neural network according to the present invention includes at least one binary matrix - vector multiplier. The binary matrix - vector multiplier includes a data selector, a finite - state machine, multiple up - counters, a down - counter, and an accumulator. The data selector generates a random number sequence to the up - counters under the control of the finite - state machine. The up - counters output count values to the accumulator. The down - counter is initialized according to the weights and starts counting. The down - counter controls the up - counters to stop counting.
[0025] Furthermore, the number of bits of the data selector and the number of up - counters are both equal to the length of the longest eigenvalue set in the corresponding re - grouped weight groups after clustering.
[0026] The usage method of the hardware circuit for the fully - connected layer of the neural network according to the present invention includes the following steps:
[0027] Step 1: Pad the eigenvalue sets with 0s whose lengths are less than the parallelism of the corresponding binary matrix - vector multiplier so that the lengths of all eigenvalue sets are equal to the parallelism of the corresponding binary matrix - vector multiplier;
[0028] Step 2: Send the weights and their corresponding eigenvalues to the corresponding binary matrix - vector multipliers for calculation in sequence.
[0029] Beneficial effects: Compared with the prior art, the present invention has the following advantages: 1. By clustering the weights of the fully connected layer, the number of weights is reduced, and the hardware circuit is optimized according to the clustering results, improving the inference speed while ensuring the accuracy of neural network inference. 2. The hardware circuit is reused, saving hardware resources and reducing energy consumption at the same time. Description of the Drawings
[0030] Figure 1 is a schematic diagram of an existing binary matrix vector multiplier;
[0031] Figure 2 is a flowchart of the weight clustering process according to an embodiment of the present invention;
[0032] Figure 3 is a schematic diagram of the hardware circuit of the fully connected layer of the neural network according to an embodiment of the present invention;
[0033] Figure 4 is a schematic diagram of the binary matrix vector multiplier according to the first embodiment of the present invention;
[0034] Figure 5 is a schematic diagram of the binary matrix vector multiplier according to the second embodiment of the present invention;
[0035] Figure 6 is a schematic diagram of the binary matrix vector multiplier according to the third embodiment of the present invention. Detailed Embodiments
[0036] The technical solution of the present invention will be further described below with reference to the accompanying drawings.
[0037] Referring to Figure 1 , the existing binary matrix vector multiplier is shown in the figure. To show its calculation process, this schematic diagram takes the input feature value x with 4-bit fixed points as an example. Its calculation process is as follows: After the feature value x is input, a random sequence for calculation is generated through a finite state machine and a data selector. After the weight w is input, the down counter is initialized according to the value of the weight w and starts down counting. When the value in the down counter is 0, the up counter is controlled to stop calculating the number of 1s in the random sequence. At this time, the up counter outputs the calculation result. Using the above binary matrix vector multiplier as the multiplier of the fully connected layer hardware circuit, when the number of weights is large, serial processing will lead to slow operation speed, and if parallel processing is used, it will require more hardware resources and increase energy consumption.
[0038] According to the design method of the hardware circuit of the fully connected layer of the neural network according to the embodiment of the present invention, the following steps are included:
[0039] S1: Cluster all the weights in a fully connected layer of the neural network according to the set clustering interval threshold;
[0040] S2: Use the mean of the weights of each cluster as the weight after clustering, and use the set composed of the eigenvalues corresponding to the weights of each cluster after clustering as the corresponding eigenvalue set;
[0041] S3: Design the circuit structure of the binary matrix vector multiplier according to the weights after clustering and the corresponding eigenvalue set.
[0042] Through the above design method, by clustering the weights of the fully connected layer of the neural network, the number of weights is reduced. While retaining the accuracy of the existing binary matrix vector multiplier, the number of weights is reduced, thereby reducing the number of calculations and improving the inference speed.
[0043] Refer to Figure 2 , the specific clustering and merging process of the weights is as follows:
[0044] S1.1: Construct a table H = [(w0, (x0)), (w1, (x1)),...(w n-1 , (x n-1 ))] with the weights W of a fully connected layer of the neural network and their corresponding eigenvalues X;
[0045] S1.2: Sort the table H in ascending order by the weight size to obtain the table
[0046] S1.3: Initialize the access index i = 0, initialize a table A for storing the clustering results, and set the clustering interval threshold τ;
[0047] S1.4: Take out the element from the table in order as the clustering starting point, denoted as α and add it to the table A, and initialize the access index m = 1, initialize the temporary variable t = w i , and define a clustering interval as
[0048] S1.5: If i + m < n, then access the element in the table If w i+m < w i + τ, then enter step S2.2, if not satisfied, jump to step S2.3;
[0049] The said step S2 includes the following steps:
[0050] S2.1: If i + m ≥ n, then update the weight in α to t / m, and jump to step S2.4;
[0051] S2.2: Update the temporary variable Add the set (x i+m)Add it to the set of α, update the access subscript m = m + 1, and jump to step S1.5;
[0052] S2.3: Update i = i + m, update the weight value of α to t / m. If i < n, jump to step S1.4; if i ≥ n, enter step S2.4;
[0053] S2.4: Output table A, where table A is composed of the clustered weight values and their corresponding eigenvalue sets.
[0054] For ease of explanation, the structure of table A can be denoted as:
[0055]
[0056] where k1, k2, …, ks respectively represent the lengths of the eigenvalue sets corresponding to the respective weight values in table A.
[0057] It can be understood that the selection of the τ value is restricted by three factors: hardware area, processing time, and inference accuracy requirements. A larger τ value may reduce the required hardware area and can reduce the calculation time, but there is a large loss in the inference accuracy of the neural network. A smaller τ value can reduce the loss of the inference accuracy of the neural network, but it will affect the processing time and hardware area. Adopting a parallel operation mode can improve the calculation speed, but it may increase the hardware area, while the serial operation mode will reduce the calculation speed. Therefore, the selection of the τ value needs to be evaluated based on historical data to obtain a reasonable range. According to the existing experimental results, it can be inferred that the same τ value can be taken for neural networks with the same structure.
[0058] Refer to Figure 3 and Figure 4 According to the hardware circuit design method of the fully connected layer of the neural network according to the embodiment of the present invention, the designed hardware circuit includes at least one binary matrix vector multiplier, a controller, and a memory. The controller extracts the weight value w i and its corresponding eigenvalue [x i,0 , x i,1 , …, x i,m from the memory in sequence according to the address of the weight value and the corresponding eigenvalue, and sends them to the binary matrix vector multiplier. The binary matrix vector multiplier includes a data selector, a finite state machine, an up counter, a down counter, and an accumulator. The data selector and the finite state machine cooperate to receive the eigenvalue array input from the memory and generate a random number sequence for the up counter. The up counter counts the number of 1s. The down counter is initialized according to the weight value and starts counting. When the count of the down counter is 0, it controls the up counter to stop counting in advance. The accumulator accumulates the counting result of the up counter to obtain the calculation result.
[0059] Refer to Figure 4 ,Figure 5 and Figure 6 For the calculation of weights and eigenvalues, serial processing or parallel processing can be adopted. If parallel processing is adopted, a multiplier needs to be designed for each weight after clustering. The parallelism of the multiplier is the length of the eigenvalue set corresponding to the weight in Table A. By adopting the parallel processing method and clustering the weights, the occupation of hardware resources can be reduced and the energy consumption can be lowered, as shown in Figure 4 . If serial processing is adopted, the maximum length of the eigenvalue set in Table A after clustering needs to be used as the parallelism of the multiplier, and when calculating, the eigenvalue sets with lengths less than the parallelism are padded with zeros to make the lengths of all eigenvalue sets equal to the maximum value, as shown in Figure 5 . By adopting the serial processing method and clustering the weights, the number of calculations can be reduced and the inference speed of the neural network can be improved.
[0060] The calculation of weights and eigenvalues can also adopt a serial-parallel hybrid method, which can not only reduce the occupied hardware resources and energy consumption, but also ensure a certain inference speed. That is, the elements in Table A after clustering are grouped again. Each group shares a multiplier circuit, and the parallelism of each group of multipliers is the maximum length of the eigenvalue set in the group. When calculating, the eigenvalue sets with lengths less than the parallelism of the corresponding multiplier in each group are padded with zeros, as shown in Figure 6 . The grouping of weights can be based on the mean or median of the lengths of all eigenvalue sets. The weights corresponding to the eigenvalue sets with lengths less than the median or mean are grouped into one group, so that the number of eigenvalues of all weights in each group is close to the mean or median. In addition, various combination schemes can be evaluated based on the resource occupation and number of calculations of the designed multiplier, and the optimal grouping scheme can be selected to design the hardware circuit of the multiplier.
[0061] Next, taking the fully connected layer of a specific neural network as an example, the above method is used to design the hardware circuit of this fully connected layer. Suppose the weights W = [w0, w1, w2, w3, w4, w5, w6, w7, w8] of this fully connected layer are [-0.40, 0.3, -0.90, -0.85, 0.40, 0.35, -0.20, -0.15, 0.45], and the corresponding eigenvalues X = [x0, x1, x2, x3, x4, x5, x6, x7, x8] are [2, 4, 3, 1, 6, 3, 5, 1, 7]. The table H = [(-0.40, (2)), (0.30, (4)), (-0.90, (3)), (-0.85, (1)), (0.40, (6)), (0.35, (3)), (-0.20, (5)), (-0.15, (1)), (0.45, (7))] is obtained by combining the weights and their corresponding eigenvalues.
[0062] Sort Table H in ascending order according to the weights to obtain the sorted table Set the clustering interval threshold τ = 0.3. Taking the element (-0.90, (3)) in the table as an example of the clustering starting point, first add the clustering starting point element to Table A for storing the clustering results, getting A = [(-0.90, (3))], and denote it as α. According to the clustering interval threshold τ = 0.3, the clustering interval is [-0.90, -0.60], and initialize the temporary variable t to -0.90.
[0063] Since (i + m) = 1 < 9, the clustering condition is satisfied, and continue to execute the clustering process. Access the subsequent element (-0.85, (1)). Since -0.85 ≤ -0.60, according to the formula update the temporary variable t to -1.75, and add the set of eigenvalue corresponding to the weight value -0.85 in the table to the set of eigenvalues of α, getting Table A = [(-0.90, (3, 1))]. Update the value of m to 2 according to the formula m = m + 1. Since (i + m) = 3 < 9, the clustering condition is satisfied, and continue to execute the clustering process. Continue to access (-0.40, (2)). Since -0.40 ≥ -0.60, the clustering condition is not satisfied, and this round of clustering process ends. At this time, t = -1.75, m = 2. Update the weight value -0.90 in Table A to -0.875 according to the formula t / m. At this time, A = [(-0.875, (3, 1))].
[0064] Update i = 2. Since i < 9, continue with the clustering operation.
[0065] Taking the element (-0.40, (2)) in the table as an example of the clustering starting point. At this time, i = 2. First add the clustering starting point element to Table A for storing the clustering results, getting A = [(-0.875, (3, 1)), (-0.40, (2))]. Initialize the access subscript m to 1, initialize the temporary variable t to -0.40. According to the clustering interval threshold τ = 0.3, the clustering interval is [-0.40, -0.10], and update α to (-0.40, (2)).
[0066] Since (i + m) = 3 < 9, the clustering condition is satisfied, and continue to execute the clustering process.
[0067] Access the subsequent element (-0.20, (5)) in the table Since -0.20 ≤ -0.10, the clustering condition is satisfied, so continue with the clustering process
[0068] According to the formula update the temporary variable t to -0.60, and add the The eigenvalue set corresponding to the medium weight value -0.20 is added to the eigenvalue set of α, and α is updated to (-0.40, (2, 5)).
[0069] Update the access subscript m to 2 and continue the current clustering process.
[0070] Since (i + m) = 4 < 9, the clustering condition is satisfied, and the clustering process continues.
[0071] Continue to access the element (-0.15, (1)). Since -0.15 ≤ -0.10, the clustering condition is satisfied, so the clustering process continues.
[0072] According to the formula Update the temporary variable t to -0.75, and add the eigenvalue set corresponding to the weight value -0.15 in the table to the eigenvalue set of α, and α is updated to (-0.40, (2, 5, 1)).
[0073] Update the access subscript m to 3 and continue the current clustering process.
[0074] Since (i + m) = 5 < 9, the clustering condition is satisfied, and the clustering process continues.
[0075] Continue to access the element (0.30, (4)). Since 0.30 ≤ -0.10, the clustering condition is not satisfied, so the current clustering process ends.
[0076] Update i = 5. At this time, t = -0.75, m = 3. According to the formula Update the weight value -0.40 in α to -0.25. At this time, A = [(-0.875, (3, 1)), (-0.25, (2, 5, 1))]. Since i < 9, continue the clustering operation.
[0077] Take the element (0.30, (4)) in the table as an example of the clustering starting point. At this time, i = 5. First, add the clustering starting point element to the table A storing the clustering results, and get A = [(-0.875, (3, 1)), (-0.25, (2, 5, 1)), (0.30, (4))]. Initialize the access subscript m to 1, initialize the temporary variable t to 0.30. According to the clustering interval threshold τ = 0.3, the clustering interval is [0.30, 0.60], and update α to (0.30, (4)).
[0078] Since (i + m) = 6 < 9, the clustering condition is satisfied, and the clustering process continues.
[0079] Access the table The subsequent element (0.35, (3)), since 0.35 ≤ 0.60, satisfies the clustering condition, so the clustering process continues.
[0080] According to the formula update the temporary variable t to 0.65, and add the set of eigenvalue corresponding to the weight value 0.35 in the table to the set of eigenvalues of α. α is updated to (0.30, (4, 3)).
[0081] Update the access subscript m to 2, and continue this round of clustering process.
[0082] Since (i + m) = 7 < 9, it satisfies the clustering condition, and continue to execute the clustering process.
[0083] Continue to access the element (0.40, (6)). Since 0.4 ≤ 0.60, it satisfies the clustering condition, so the clustering process continues.
[0084] According to the formula update the temporary variable t to 1.05, and add the set of eigenvalue corresponding to the weight value 0.40 in the table to the set of eigenvalues of α. α is updated to (0.30, (4, 3, 6)).
[0085] Update the access subscript m to 3, and continue this round of clustering process.
[0086] Since (i + m) = 8 < 9, it satisfies the clustering condition, and continue to execute the clustering process.
[0087] Continue to access the element (0.45, (7)). Since 0.45 ≤ 0.60, it satisfies the clustering condition, so the clustering process continues.
[0088] Access the table The subsequent element (0.45, (7)). Since 0.45 ≤ 0.60, it satisfies the clustering condition, so the clustering process continues.
[0089] According to the formula update the temporary variable t to 1.50, and add the set of eigenvalue corresponding to the weight value 0.45 in the table to the set of eigenvalues of α. α is updated to (0.30, (4, 3, 6, 7)).
[0090] Update the access subscript m to 4. Continue to execute the clustering process.
[0091] Since (i + m) = 9 ≥ 9, it does not satisfy the clustering condition. At this time, t = 1.50, m = 4. According to the formula Update the weight value 0.30 in α to 0.375. At this time, A = [(-0.875, (3, 1)), (-0.25, (2, 5, 1)), (0.375, (4, 3, 6, 7))], and this round of clustering ends.
[0092] The finally obtained table A storing the clustering results is A = [(-0.875, (3, 1)), (-0.25, (2, 5, 1)), (0.375, (4, 3, 6, 7))].
[0093] According to the clustering results, design a multiplier circuit in the way of parallel operation of all weight values, then the embodiment 1 as shown in Figure 4 can be obtained. For the three weight values obtained after clustering, design binary matrix vector multipliers respectively, and the parallelism degrees of the corresponding multipliers are the lengths 2, 3, and 4 of the eigenvalue sets corresponding to each weight value.
[0094] Design a multiplier circuit in the way of serial processing of all weight values, that is, the calculation of all weight values uses the same multiplier circuit, then the embodiment 2 as shown in Figure 5 can be obtained. Take the maximum length 4 of the three eigenvalue sets as the parallelism degree of the multiplier, and fill in 0 in the eigenvalue sets of the weight values -0.875 and -0.25. The completed table A = [(-0.875, (3, 1, 0, 0)), (-0.25, (2, 5, 1, 0)), (0.375, (4, 3, 6, 7))].
[0095] Design a multiplier circuit in a hybrid way of serial and parallel, then the embodiment 3 as shown in Figure 6 can be obtained. In this embodiment, since the lengths of the three eigenvalue sets are 2, 3, and 4 respectively, and the length of the eigenvalue set of the weight value 0.375 is the longest, so divide the weight values -0.875 and -0.25 into one group, and divide the weight value 0.375 into one group, that is, [(-0.875, (3, 1)), (-0.25, (2, 5, 1))] and [(0.375, (4, 3, 6, 7)] two groups, and the weight values within each group reuse the same multiplier circuit. The parallelism degree of the first group of multipliers is the maximum length 3 of the eigenvalue set within the group, and the parallelism degree of the second group of multipliers is the length 4 of the eigenvalue set of the only weight value within the group. Compared with embodiment 1, embodiment 3 reduces one two-way multiplier circuit, saves hardware resources, and the maximum number of calculations is 2 times, and the number of calculations is less than 3 times of embodiment 2, and the inference speed is faster than that of embodiment 2.
Claims
1. A method for designing a hardware circuit of a fully connected layer of a neural network, characterized in that, It includes the following steps: S1: Cluster all the weights in a fully connected layer of the neural network according to the set clustering interval threshold; S2: Use the mean value of the weights in each cluster after clustering as the weight after clustering, and use the set composed of the eigenvalue corresponding to the weights in each cluster after clustering as the corresponding eigenvalue set; S3: Design the circuit structure of the binary matrix vector multiplier according to the weights after clustering and the corresponding eigenvalue set; The step S1 includes the following steps: S1.1: Construct a table H = [(w0, (x0)), (w1, (x1)),...(w n-1 , (x n-1 ))] with the weights W of a fully connected layer of a neural network and their corresponding feature values X; S1.2: Sort Table H in ascending order according to the weight value to obtain Table S1.3: Initialize the access subscript i = 0, initialize a table A for storing the clustering results, and set the clustering interval threshold τ; S1.4: Take out elements from the table in order As the clustering starting point, denote it as α and add it to table A, initialize the access subscript m = 1, and initialize the temporary variable t = w i , define a clustering interval according to the input value τ as S1.5: If i + m < n, then access the elements in the table in the If w i+m < w i + τ, then go to step S2.2; otherwise, jump to step S2.3 The step S2 includes the following steps: S2.1: If i + m ≥ n, update the weight in α to t / m, and jump to step S2.4; S2.2: Update the temporary variable Add the set (x i+m ) to the set of α, update the access index m = m + 1, and jump to step S1.5; S2.3: Update \(i = i + m\), and update the weight value of \(\alpha\) to \(t\). / For \(m\), if \(i < w\), then jump to step S1.4; If i ≥ n, enter step S2.4; S2.4: Output the table A, where the table A is composed of the clustered weights and their corresponding eigenvalue sets.
2. The method for designing a hardware circuit of a fully connected layer of a neural network according to claim 1, characterized in that, The S3 includes: S3.1: Divide the weights in the clustered table into several groups; S3.2: Design a parallel binary matrix vector multiplier for each group of weights, and the parallelism of the binary matrix vector multiplier corresponding to each group of weights is the maximum length of the eigenvalue set corresponding to this group of weights.
3. A hardware circuit of a fully connected layer of a neural network designed according to the method for designing a hardware circuit of a fully connected layer of a neural network according to claim 1 or 2, characterized in that, It includes at least one binary matrix vector multiplier. The binary matrix vector multiplier includes a data selector, a finite state machine, multiple up counters, a down counter and an accumulator. The data selector generates a random number sequence to the up counter under the control of the finite state machine. The up counter outputs a count value to the accumulator. The down counter is initialized according to the weight and starts counting. The down counter controls the up counter to stop counting.
4. The hardware circuit of a fully connected layer of a neural network according to claim 3, characterized in that, The number of bits of the data selector and the number of up counters are both equal to the length of the longest eigenvalue set in the corresponding re-grouped weight group after clustering.
5. A method for using a hardware circuit of a fully connected layer of a neural network according to claim 3 or 4, characterized in that, It includes the following steps: Step 1: Pad 0 to the eigenvalue set with a length less than the parallelism of the corresponding binary matrix vector multiplier to make the lengths of all eigenvalue sets equal to the parallelism of the corresponding binary matrix vector multiplier; Step 2: Send the weight and the corresponding eigenvalue to the corresponding binary matrix vector multiplier for calculation in sequence.
Citation Information
Patent Citations
FPGA-based sparsity neural network accelerating system
CN108932548A
Model compression method and system for deep neural network
CN111476366A