Convolution processor and method for spiking neural network, and electronic equipment

By designing a convolution processor for pulsed neural networks, using pre-stored pulse convolution results and multiple search circuits, the problems of large amount of adder usage and high resource utilization in the prior art are solved, and more efficient computing resource utilization is achieved.

CN120012851AActive Publication Date: 2025-05-16INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510487486.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-05-16
Estimated Expiration
2045-04-18

AI Technical Summary

Technical Problem

The general convolutional computing structure of existing pulse neural networks requires a large number of adders when performing convolutional calculations, resulting in a large amount of hardware computing resources occupied, a large fanout, and a low utilization rate of accelerator hardware computing resources.

Method used

A convolution processor for pulsed neural networks is designed to reduce dependence on the adder by pre-storing the weight-based pulse convolution results and using multiple search circuits to quickly find matching convolution results.

Benefits of technology

Without using adders and multipliers, this design can quickly find convolution results, reduce the usage of adders, reduce the use of hardware computing resources, increase resource utilization, and reduce the area of ​​hardware design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012851A_ABST
    Figure CN120012851A_ABST
Patent Text Reader

Abstract

The invention provides a convolution processor and method for a spiking neural network, and electronic equipment, and can be applied to the technical field of deep learning and artificial intelligence. The convolution processor for the spiking neural network comprises a first calculation unit, 2N1 pulse convolution results based on N1 first weights are pre-stored in the first calculation unit, and the first calculation unit is used for receiving N1 first features and searching pulse convolution results matched with the N1 first features from the 2N1 pulse convolution results as first convolution results; wherein the pulse convolution result comprises K parts, the first calculation unit comprises K search circuits, the kth search circuit stores the kth part of each pulse convolution result, and the kth search circuit is used for searching the kth part matched with the N1 first features from the kth part of each stored pulse convolution result to serve as the kth part of the first convolution result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of deep learning and artificial intelligence technology, and more specifically to a convolution processor, method and electronic device for a pulse neural network. Background Art

[0002] The pulse neural network simulates the transmission of brain neurons. The data it processes has sparse characteristics and event-triggered properties. Not all data at every moment needs to be calculated, which reduces the overall amount of calculation. It is an efficient brain-like computing structure. Therefore, pulse neural networks are widely used in image recognition, robot control and other fields.

[0003] In the related art, when the general convolution computing structure of the pulse neural network performs convolution calculations, it is necessary to use a large number of multi-level adders to accumulate the various weights involved in the convolution calculations. The accelerator hardware computing resources are more occupied, the fan-out is larger, and the utilization rate of the accelerator hardware computing resources is low. Summary of the invention

[0004] In view of the above problems, the present invention provides a convolution processor, method and electronic device for a pulse neural network.

[0005] According to a first aspect of the present invention, a convolution processor for a pulse neural network is provided, the convolution processor comprising: a first computing unit, wherein the first computing unit pre-stores 2 based on N1 first weights N1 The first calculation unit is used to receive N1 first features from the above 2 N1 The kth search circuit is used to search for a pulse convolution result that matches the N1 first features from the kth pulse convolution results as the first convolution result; wherein the pulse convolution result includes K parts, the first calculation unit includes K search circuits, the kth part of each pulse convolution result is stored in the kth search circuit, the kth search circuit is any search circuit among the K search circuits, the kth part is the part of the K parts corresponding to the kth search circuit, and the kth search circuit is used to search for the kth part that matches the N1 first features from the kth parts of the stored pulse convolution results as the kth part of the first convolution result, wherein 1≤k≤K and k is an integer, and N1 and K are integers greater than 1.

[0006] According to a second aspect of the present invention, a convolution processing method for a spiking neural network is provided, which is performed by the above-mentioned convolution processor, and the above-mentioned method comprises: a first computing unit receives N1 first features; the first computing unit extracts N1 first features from a pre-stored 2 based on N1 first weights; N1pulse convolution results, searching for a pulse convolution result matching the N1 first features as the first convolution result; wherein the pulse convolution result includes K parts, the first calculation unit includes K search circuits, the kth search circuit stores the kth part of each pulse convolution result, the kth search circuit is any search circuit among the K search circuits, the kth part is a part of the K parts corresponding to the kth search circuit, the 2 pre-stored weights based on the N1 first weights are selected. N1 The method of searching for a pulse convolution result that matches the N1 first features as the first convolution result includes: a k-th search circuit searches for a k-th part that matches the N1 first features from the k-th parts of the stored pulse convolution results as the k-th part of the first convolution result, wherein 1≤k≤K and k is an integer, and N1 and K are integers greater than 1.

[0007] According to a third aspect of the present invention, there is provided an electronic device comprising the above-mentioned convolution processor.

[0008] According to the convolution processor for a pulse neural network provided by an embodiment of the present invention, since the pulse convolution result includes K parts, the first computing unit includes K search circuits, and the kth search circuit stores the kth part of each pulse convolution result. Therefore, the kth search circuit can be used to search for the kth part matching the N1 first features from the kth parts of each stored pulse convolution result as the kth part of the first convolution result. Furthermore, without using an adder and a multiplier, the K search circuits included in the first computing unit can be used to quickly search for each part of the first convolution result, and the first computing unit can quickly obtain the kth part of the first convolution result from 2 N1 The first convolution result matching the N1 first features is found in the pulse convolution results, reducing the usage of adders in the related technology. Since the minimum logic unit of the accelerator hardware computing resources includes more search circuits and fewer adders, while reducing the usage of adders, the usage of the minimum logic unit can be reduced, the accelerator hardware computing resource usage is reduced, the fan-out is small, the design area of ​​the related hardware is reduced, and the utilization rate of the accelerator hardware computing resources is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The above contents and other objects, features and advantages of the present invention will become more apparent through the following description of the embodiments of the present invention with reference to the accompanying drawings.

[0010] Figure 1A A schematic diagram showing the principle of convolution calculation performed by a spiking neural network.

[0011] Figure 1B A schematic diagram showing the related technology of using multiple adders to accelerate the convolution calculation of pulse neural network.

[0012] Figure 2 A schematic diagram of the structure of a convolution processor for a pulse neural network according to a first embodiment of the present invention is shown.

[0013] Figure 3 A schematic diagram showing a pulse convolution result according to an embodiment of the present invention is shown.

[0014] Figure 4 A schematic diagram of the structure of a convolution processor for a pulse neural network according to a second embodiment of the present invention is shown.

[0015] Figure 5 A schematic diagram of the structure of a convolution processor for a pulse neural network according to a third embodiment of the present invention is shown.

[0016] Figure 6 A schematic diagram of a lookup table logic circuit according to an embodiment of the present invention is shown.

[0017] Figure 7 A schematic diagram of the structure of a convolution processor for a pulse neural network according to a fourth embodiment of the present invention is shown.

[0018] Figure 8 A schematic diagram showing a pulse convolution result according to another embodiment of the present invention is shown.

[0019] Fig. 9 A schematic diagram of the structure of a convolution processor for a pulse neural network according to a fifth embodiment of the present invention is shown.

[0020] Fig.10 A schematic diagram of the structure of a convolution processor for a pulse neural network according to a sixth embodiment of the present invention is shown.

[0021] Fig.11 A schematic diagram of the structure of a convolution processor for a pulse neural network according to the seventh embodiment of the present invention is shown.

[0022] Fig.12 A flow chart of a convolution processing method for a pulse neural network according to an embodiment of the present invention is shown.

[0023] Fig.13 A structural block diagram of an electronic device according to an embodiment of the present invention is shown. DETAILED DESCRIPTION

[0024] Below, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present invention. In the following detailed description, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of embodiments of the present invention. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessary confusion of concepts of the present invention.

[0025] The terms used herein are only for describing specific embodiments and are not intended to limit the present invention. The terms "comprise", "include", etc. used herein indicate the existence of the features, steps, operations and / or components, but do not exclude the existence or addition of one or more other features, steps, operations or components.

[0026] All terms (including technical and scientific terms) used herein have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0027] When using expressions such as "at least one of A, B, and C, etc.", they should generally be interpreted according to the meaning of the expression commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).

[0028] Figure 1A A schematic diagram showing the principle of convolution calculation performed by a spiking neural network.

[0029] like Figure 1A As shown, the feature matrix 110 input to the pulse neural network includes multiple feature data. Since the calculation characteristic of the pulse neural network is to use pulse single-bit data instead of multi-bit data to participate in the convolution calculation, the feature data included in the feature matrix 110 is 0 or 1. When the pulse neural network performs convolution calculation, the feature matrix 110 is multiplied by the weight matrix 120 to obtain the convolution matrix 130. Among them, the weight matrix 120 includes multiple weights corresponding to multiple feature data. The various values ​​in the convolution matrix 130 are accumulated and calculated to obtain the convolution result "w2+w3+w4+w5+w7+w9".

[0030] Depend on Figure 1A It is known that in Figure 1AWhen the feature matrix 110 and the weight matrix 120 in the 3×3 convolution calculation are performed, the feature data is used as a control signal for data selection to control whether the weight data or the 0 data is involved in the addition calculation. Therefore, when the pulse neural network performs convolution calculation, the multiplier can be omitted, and only the multi-stage addition unit is used to accumulate the weights involved in the convolution calculation to obtain the convolution result.

[0031] Figure 1B A schematic diagram showing the related technology of using multiple adders to accelerate the convolution calculation of pulse neural network.

[0032] like Figure 1B As shown, the pulse neural network convolution processor in the related art, when performing 3×3 convolution calculation on multiple feature data 111 included in the feature matrix and multiple weights 121 included in the weight matrix, omits the multiplier and only uses the multi-stage addition unit 140 to perform cumulative calculation on each weight involved in the convolution calculation to obtain the convolution result. Among them, the preset value can be 0.

[0033] according to Figure 1B It can be seen that although the pulse neural network convolution processor in the related art omits the multiplier and only uses the addition unit when performing 3×3 convolution calculations, it uses 4 levels of 9 adders, and uses more levels of addition trees. Among them, the addition unit can be a multi-level addition calculation structure implemented by logic gates and triggers in the accelerator hardware resources. Accelerator hardware resources usually include resources such as lookup tables, logic gates, and triggers. Therefore, the pulse neural network convolution processor in the related art uses a large number of logic gate and trigger resources, wasting a lot of resources such as lookup tables (LUT) in the accelerator hardware resources.

[0034] For example, the logic unit includes 4 6-input 2-output lookup tables and resources such as logic gates and triggers. When the logic gates included in the logic unit constitute an adder, the pulse neural network convolution processor in the related art uses 9 adders when performing 3×3 convolution calculations, which is equivalent to using 9 logic units. The fan-out will become 9, and the hardware computing resources are occupied more and the fan-out is larger. In the case of large-scale, multi-layer convolution calculations, the fan-out amount will become unacceptable, increasing the design area of ​​the related hardware. In addition, the related technology wastes multiple lookup table resources in the 9 lookup tables, and the utilization rate of the accelerator hardware computing resources is low. The larger the fan-out, the lower the peak frequency of the calculation. When the hardware computing resources of the related technology are occupied more and the fan-out is larger, the possible computing main frequency is reduced. Among them, the 6-input 2-output lookup table represents a lookup table including 6 input ports and 2 output ports.

[0035] In view of this, embodiments of the present invention provide a convolution processor, method and electronic device for a pulse neural network, which can be applied to the fields of deep learning and artificial intelligence technology.

[0036] Figure 2 A schematic diagram of the structure of a convolution processor for a pulse neural network according to a first embodiment of the present invention is shown.

[0037] like Figure 2 As shown, the convolution processor for a spiking neural network may include a first computing unit 210.

[0038] The first calculation unit 210 may pre-store 2 based on N1 first weights. N1 The first calculation unit 210 can be used to receive N1 first features F1, F2, ..., FN1, from 2 N1 The pulse convolution results are searched for the pulse convolution results matching the N1 first features F1, F2, ..., FN1 as the first convolution result, where N1 is an integer greater than 1.

[0039] According to the embodiment of the present invention, N1 can be selected according to actual conditions and is not limited here.

[0040] For example, when N1 is 4, 16 pulse convolution results based on 4 first weights may be pre-stored in the first calculation unit 210. The first calculation unit 210 may receive 4 first features, and use the 4 first features as addresses to search for pulse convolution results matching the 4 first features from the 16 pulse convolution results as the first convolution result.

[0041] For example, when N1 is 5, 32 pulse convolution results based on 5 first weights may be pre-stored in the first calculation unit 210. The first calculation unit 210 may receive 5 first features, and use the 5 first features as addresses to search for pulse convolution results matching the 5 first features from the 32 pulse convolution results as the first convolution result.

[0042] For example, when N1 is 9, 512 pulse convolution results based on 9 first weights may be pre-stored in the first calculation unit 210. The first calculation unit 210 may receive 9 first features, and use the 9 first features as addresses to search for a pulse convolution result matching the 9 first features from the 512 pulse convolution results as the first convolution result.

[0043] The pulse convolution result may include K parts. The first calculation unit 210 may include K search circuits, for example, the first search circuit LUT1, the second search circuit LUT2, ..., the Kth search circuit LUTK. The kth part of each pulse convolution result is stored in the kth search circuit. The kth search circuit is any search circuit among the K search circuits, and the kth part is the part of the K parts corresponding to the kth search circuit. For example, assuming that the pulse convolution result is 10-bit data, the 10-bit pulse convolution result can be divided into 5 parts, each part is 2 bits, and the 5 parts of the pulse convolution result can be stored in 5 search circuits respectively.

[0044] The kth search circuit is used to search for the kth part matching the N1 first features F1, F2, ..., FN1 from the kth part of each stored pulse convolution result as the kth part of the first convolution result, where 1≤k≤K and k is an integer, and K is an integer greater than 1. For example, the kth search circuit may store the kth part of each pulse convolution result by address, and when searching, the N1 first features F1, F2, ..., FN1 may be used as the target query address, and the kth part matching the target query address may be searched from the kth part of each stored pulse convolution result.

[0045] like Figure 2 As shown, the first search circuit LUT1 can store the first part of each pulse convolution result. The first search circuit LUT1 can search for the first part matching the N1 first features F1, F2, ..., FN1 from the stored first parts of each pulse convolution result as the first part S1 of the first convolution result.

[0046] The second search circuit LUT2 can store the second part of each pulse convolution result. The second search circuit LUT2 can search for the second part matching the N1 first features F1, F2, ..., FN1 from the stored second parts of each pulse convolution result as the second part S2 of the first convolution result.

[0047] By analogy, the kth search circuit can store the kth part of each pulse convolution result. The Kth search circuit LUTK can store the kth part S of each pulse convolution result from the stored K Find the Kth part that matches the N1 first features F1, F2, ..., FN1, as the Kth part S of the first convolution result K .

[0048] Thus, the first computing unit 210 generates the convolution results of the N1 first features F1, F2, ..., FN1 and the N1 first weights, i.e., the first convolution result, wherein the K search circuits in the first computing unit 210 respectively output the K parts S1, S2, ..., SK .

[0049] According to the convolution processor for a pulse neural network provided by an embodiment of the present invention, since the pulse convolution result includes K parts, the first computing unit includes K search circuits, and the kth search circuit stores the kth part of each pulse convolution result. Therefore, the kth search circuit can be used to search for the kth part matching the N1 first features from the kth parts of each stored pulse convolution result as the kth part of the first convolution result. Furthermore, without using an adder and a multiplier, the K search circuits included in the first computing unit can be used to quickly search for each part of the first convolution result, and the first computing unit can quickly obtain the kth part of the first convolution result from 2 N1 The first convolution result matching the N1 first features is found in the pulse convolution results, reducing the usage of adders in the related technology. Since the minimum logic unit of the accelerator hardware computing resources includes more search circuits and fewer adders, while reducing the usage of adders, the usage of the minimum logic unit can be reduced, the accelerator hardware computing resource usage is reduced, the fan-out is small, the design area of ​​the related hardware is reduced, and the utilization rate of the accelerator hardware computing resources is improved.

[0050] According to an embodiment of the present invention, since the convolution processor for a pulse neural network provided by the embodiment of the present invention has a smaller fan-out than that of the related art, it can balance the reduction in computing main frequency caused by low resource utilization and increased fan-out in the related art.

[0051] Combine the following Figure 3 and Figure 4 , to describe another embodiment of the convolution processor of an embodiment of the present invention. Figure 3 A schematic diagram of pulse convolution according to an embodiment of the present invention is shown. Figure 4 A schematic diagram of the structure of a convolution processor for a pulse neural network according to a second embodiment of the present invention is shown.

[0052] like Figure 3 As shown, both the feature matrix 110 and the weight matrix 120 are 3×3 matrices. The structure and working principle of the convolution processor when N1=4 are described below by taking the convolution of the four first features at positions 00, 01, 02 and 10 and the four first weights w1, w2, w3 and w4 as an example. 00 represents the first row and the first column, 01 represents the first row and the second column, 02 represents the first row and the third column, and 10 represents the second row and the first column. The four first weights w1, w2, w3 and w4 can generate a total of 16 impulse convolution results. The number of bits M of the impulse convolution result is 10. The 10-bit data of the 16 impulse convolution results can be divided into 5 parts, each with 2 bits. In this case, N1=4, M=10, and K=5.

[0053] like Figure 4 As shown, the convolution processor for the pulse neural network may include a first type first computing unit 210A. The first type first computing unit 210A may include 5 lookup circuits LUT1, LUT2, ..., LUT5, and the 5 parts of each pulse convolution result are respectively stored in the 5 lookup circuits LUT1, LUT2, ..., LUT5. For example, the pulse convolution result is 10-bit data, the first part of the pulse convolution result may be the 1st bit and the 2nd bit of the 10-bit data, the second part may be the 3rd bit and the 4th bit of the 10-bit data, and so on. The 1st bit and the 2nd bit of each of the 16 pulse convolution results may be stored in the first lookup circuit LUT1, the 3rd bit and the 4th bit may be stored in the second lookup circuit LUT2, and so on.

[0054] The search circuit may have N1 input terminals and m output terminals, where M=mK, M is the number of bits of the first convolution result, and m is the number of bits of each part of the first convolution result. The kth search circuit may be used to receive N1 first features from the N1 input terminals, respectively, and search the kth part that matches the N1 first features from the kth parts of each stored pulse convolution result as the kth part of the first convolution result, and output the m bits of the kth part of the first convolution result through the m output terminals respectively. Figure 4 In the example of , N=4, M=10, K=5, and therefore m=2. That is, each lookup circuit has 4 input terminals and 2 output terminals, the 4 input terminals are used to receive the 4 first features at positions 00, 01, 02, and 10 in the feature matrix 110, respectively, and the 2 output terminals are used to output the corresponding 2-bit portion of the first convolution result. For example, the first lookup circuit LUT1 outputs the 1st and 2nd bits of the first convolution result, the second lookup circuit LUT2 outputs the 3rd and 4th bits of the first convolution result, the third lookup circuit LUT3 outputs the 5th and 6th bits of the first convolution result, and so on.

[0055] According to an embodiment of the present invention, the kth search circuit is used to receive N1 first features from N1 input terminals respectively, and the kth part matching the N1 first features is searched from the kth part of each stored pulse convolution result as the kth part of the first convolution result, and the m bits of the kth part of the first convolution result are output through m output terminals respectively. Therefore, without using an adder and a multiplier, each search circuit can be used to quickly search for the m bits of the first convolution result, and K search circuits can be used to quickly search and obtain the M-bit first convolution result, thereby reducing the use of adders in related technologies, reducing the occupation of accelerator hardware computing resources, and improving the utilization rate of accelerator hardware computing resources.

[0056] Can be used Figure 4The convolution processor for the pulse neural network shown implements convolution calculations on the first features at positions 00, 01, 02 and 10 in the feature matrix 110 and the four first weights w1, w2, w3 and w4 at the corresponding positions in the weight matrix 120.

[0057] The kth search circuit can be used to receive four first features respectively through the four input terminals of the kth search circuit, search for the kth part matching the four first features from the kth parts of each stored pulse convolution result as the kth part of the first convolution result, and output the 2 bits of the kth part of the first convolution result through the two output terminals of the kth search circuit respectively.

[0058] According to an embodiment of the present invention, when the search circuit has 4 input terminals and 2 output terminals, N1=4, M=10, K=5, the 4 input terminals of the kth search circuit are used to receive the 4 first features respectively, and the kth part matching the 4 first features is searched from the kth parts of the stored pulse convolution results as the kth part of the first convolution result, and the 2 bits of the kth part of the first convolution result are output through the 2 output terminals of the kth search circuit respectively. This technical means can use each search circuit to quickly search for 2 bits of data in the first convolution result without using an adder and a multiplier, and use 5 search circuits to quickly search to obtain a 10-bit first convolution result, thereby reducing the use of adders in related technologies, reducing the occupation of accelerator hardware computing resources, and improving the utilization rate of accelerator hardware computing resources.

[0059] The following is also Figure 3 For example, refer to the 3×3 convolution Figure 5 To describe the structure and working principle of a convolution processor for a pulse neural network according to another embodiment of the present invention.

[0060] Figure 5 A schematic diagram of the structure of a convolution processor for a pulse neural network according to a third embodiment of the present invention is shown.

[0061] In this embodiment, the convolution processor for the pulse neural network may include multiple first computing units and a first adder, and the first adder may sum the first convolution results output by the multiple first computing units to obtain a first output. Figure 5 As shown, the convolution processor includes two first-type first calculation units 210A and a first adder 220. The first adder 220 can sum the first convolution results output by the two first-type first calculation units 210A to obtain a first output.

[0062] In some embodiments, the convolution processor may further include a second computing unit 230 and a second adder 240. The second computing unit 230 may be used to receive N2 second features, and perform convolution calculation on the received N2 second features and N2 second weights to obtain a second convolution result, where N2 is an integer greater than or equal to 1. The second adder 240 may be used to sum the first output with the second convolution result provided by the second computing unit 230 to obtain a second output.

[0063] Also refer to the following Figure 3 , taking 3×3 convolution as an example Figure 5 The convolution processor is described below.

[0064] N1=4, M=10, K=5. Each of the two first-type first calculation units 210A includes five lookup circuits LUT1 to LUT5. The lookup circuit may have four input terminals and two output terminals. Figure 5 The first type first computing unit 210A and Figure 4 The first type first computing unit 210A has similar structure and function, and for the sake of simplicity, it will not be repeated here.

[0065] Combination Figure 3 and Figure 5 , one of the two first-type first computing units 210A ( Figure 5 The first type first calculation unit 210A on the left side of the figure can perform convolution calculation on the four first features at positions 00, 01, 02, and 10 in the feature matrix 110 and the four first weights w1, w2, w3, and w4 at the corresponding positions in the weight matrix 120 to obtain a 10-bit first convolution result. Another first type first calculation unit 210A ( Figure 5 The first type first calculation unit 210A) on the right side of the figure can perform convolution calculation on the four first features at positions 11, 12, 20, and 21 in the feature matrix 110 and the first weights w5, w6, w7, and w8 at corresponding positions in the weight matrix 120 to obtain another 10-bit first convolution result.

[0066] According to an embodiment of the present invention, the first convolution results output by the plurality of first computing units are summed by using a first adder to obtain a first output. Compared with the related art, the number of adders used is greatly reduced, resource utilization is improved, the number of adder stages is reduced, and the inference calculation efficiency is improved.

[0067] In some embodiments, the second calculation unit 230 may include a selector. The selector may be used to select between the N2 second weights and the preset reference value according to the N2 second features to obtain the second convolution result. According to an embodiment of the present invention, the preset reference value may be selected according to actual conditions and is not limited here. For example, the preset reference value may be 0. For example, in combination with Figure 3 and Figure 5 , when N2=1, the second computing unit 230 can use a single selector to perform convolution calculation on the feature at position 22 in the feature matrix 110 (i.e., the second feature) and the weight w9 at position 22 in the weight matrix 120 (i.e., the second weight). Specifically, one input end of the selector receives the weight w9 at position 22 in the weight matrix, the other input end receives a predetermined reference value 0, the control end of the selector receives the feature at position 22 in the feature matrix (i.e., the second feature), and the output end of the selector outputs the second convolution result. When the feature received by the control end of the selector is 0, the selector outputs the predetermined reference value 0, and when the feature received by the control end of the selector is 1, the selector outputs the weight w9. Although Figure 5 2 is described by taking a single selector and a single second adder 240 as an example, however, the embodiments of the present invention are not limited thereto, and the number of selectors and second adders can be set as required. For example, in the case of N2=2, the second computing unit 230 may include two selectors, each used for convolution calculation of two features and two weights.

[0068] According to an embodiment of the present invention, a first output is obtained by summing the first convolution results output by multiple first computing units using a first adder, receiving N2 second features using a second computing unit, performing convolution calculation on the received N2 second features and N2 second weights to obtain a second convolution result, and summing the first output with the second convolution result provided by the second computing unit using a second adder to obtain a technical means for obtaining a second output. It is possible to use 2 adders and multiple search circuits to quickly perform convolution calculation on N1 first features and N2 second features to quickly obtain a second output. Compared with related technologies, fewer adders are used, which improves resource utilization, reduces the number of adder levels, and improves reasoning calculation efficiency.

[0069] According to an embodiment of the present invention, when a convolution processor for a pulse neural network includes two first computing units, a first adder, a second computing unit, and a second adder, N1=4, N2=1, M=10, K=5, and a search circuit has four input terminals and two output terminals, a convolution processor of a pulse neural network is used to implement 3×3 convolution calculations, and 10 search circuits, two adders, and a selector are used to quickly perform 3×3 convolution calculations. Compared with the related art that requires the use of 9 adders, that is, the use of 9 logic units, the adders used in the present invention are greatly reduced, and only 3 logic units are required, and the logic units used are greatly reduced. A large number of search circuit resources, logic gates, and trigger resources in the accelerator hardware computing resources are fully utilized. The influence of resource utilization and fan-out is balanced, and compared with the related art, the resource utilization rate is improved, the number of adder levels is reduced, and the reasoning calculation efficiency is improved.

[0070] The search circuit in the above embodiment may be a search table logic circuit. Figure 6 To describe the lookup table logic circuit.

[0071] Figure 6 A schematic diagram of a lookup table logic circuit according to an embodiment of the present invention is shown.

[0072] like Figure 6 As shown, the lookup table logic circuit may have 6 input terminals A1 to A6 and 2 output terminals O5 and O6. The lookup table logic circuit is the basic implementation unit of the field programmable gate array (FPGA). It is evenly distributed in all positions of the FPGA in large quantities and is the least scarce hardware resource on FPGA-type accelerators. The lookup circuit in the convolution processor according to an embodiment of the present invention can be implemented by a lookup table logic circuit. For example, if a 4-input 2-output lookup circuit is required in the convolution processor, then 4 of the 6 input terminals of the lookup table logic circuit (for example, A1 to A4) can be used as the input terminals of the lookup circuit, and the 2 output terminals O5 and O6 of the lookup table logic circuit can be used as the output terminals of the lookup circuit.

[0073] According to an embodiment of the present invention, when a convolution processor for a pulse neural network includes two first computing units, a first adder, a second computing unit, and a second adder, N1=4, N2=1, M=10, K=5, and the lookup circuit is a lookup table logic circuit, and the lookup circuit has four input terminals and two output terminals, a convolution processor of a pulse neural network is used to implement 3×3 convolution calculation, and 10 lookup table logic circuits, two adders, and a selector are used to quickly perform 3×3 convolution calculation. Compared with the related art that requires the use of 9 adders, the present invention reduces the number of adders to 2, 1, or even no adder is required. A large number of lookup table logic circuit resources, logic gates, and trigger resources in FPGA are fully utilized. The influence of resource utilization and fan-out is balanced, and compared with the related art, resource utilization is improved, the number of adder levels is reduced, and the efficiency of reasoning calculation is improved.

[0074] Figure 7 A schematic diagram of the structure of a convolution processor for a pulse neural network according to a fourth embodiment of the present invention is shown.

[0075] like Figure 7 As shown, the convolution processor for the pulse neural network may include at least one acceleration module, each of which may have Figure 5 The structure shown in the figure includes two first type first calculation units 210A, a first adder 220, a second calculation unit 230, and a second adder 240. Figure 5 The description of the first type first calculation unit 210A, the first adder 220, the second calculation unit 230, and the second adder 240 is also applicable to this embodiment and will not be repeated here.

[0076] You can use multiple Figure 7 The multiple acceleration modules shown perform 3×3 convolution calculations on different feature matrices 110 and corresponding weight matrices 120 in parallel, further improving the efficiency of convolution calculations.

[0077] like Figure 7 As shown, the convolution processor for the spiking neural network may further include a comparator 250. The comparator 250 may be used to compare the second output provided by the second adder 240 with a preset threshold, and output a pulse signal when the second output is greater than or equal to the preset threshold. For example, the comparator 250 may also compare the result of adding the plurality of second outputs with a preset threshold, and output a pulse signal when the sum is greater than or equal to the preset threshold.

[0078] In some embodiments, Figure 7The convolution processor for a spiking neural network may further include a cache unit 260. The cache unit 260 may be used to accumulate a plurality of consecutive second outputs and cache the accumulated result. The plurality of consecutive second outputs may be second outputs outputted by the same acceleration module for multiple consecutive times, or may be second outputs outputted sequentially by different acceleration modules.

[0079] The comparator 250 can add the second output of the current output and the cache accumulation result, and compare the summed result with a preset threshold, and when the summed result is greater than or equal to the preset threshold, output a pulse signal and set the cache unit 260 to 0.

[0080] According to the embodiment of the present invention, the preset threshold value can be selected according to actual conditions and is not limited here.

[0081] According to an embodiment of the present invention, a comparator is used to compare the second output with a preset threshold, and a pulse signal is output when the second output is greater than or equal to the preset threshold, thereby generating a pulse signal only for a second output that meets the requirements.

[0082] Reference below Figure 8 and Fig. 9 The structure and principle of a convolution processor according to another embodiment of the present invention are described in detail.

[0083] Figure 8 FIG. 4 is a schematic diagram of pulse convolution according to another embodiment of the present invention. Fig. 9 A schematic diagram of the structure of a convolution processor for a pulse neural network according to a fifth embodiment of the present invention is shown.

[0084] like Figure 8 As shown, the number of pulse convolution results corresponding to the five first features at positions 11, 12, 20, 21 and 22 in the feature matrix 110 and the five first weights w5, w6, w7, w8 and w9 at the corresponding positions in the weight matrix 120 can be 32. The number of bits of the maximum value among the 32 pulse convolution results can be 11 bits. The 11-bit data of the 32 pulse convolution results can be divided into 6 parts, respectively, and the first to fifth parts can all be 2 bits, and the sixth part can be 1 bit. The input signal of the search circuit can be a pulse signal.

[0085] Can be used Fig. 9 The convolution processor for the pulse neural network shown implements convolution processing on the five first features at positions 11, 12, 20, 21 and 22 in the feature matrix 110 and the five first weights w5, w6, w7, w8 and w9 at the corresponding positions in the weight matrix 120.

[0086] like Fig. 9 As shown, the convolution processor for the pulse neural network may include a second type first computing unit 210B. Figure 4 The first type of computing unit 210A is different in that, Fig. 9 The number of search circuits in the second type first calculation unit 210B and the number of input terminals and output terminals of the search circuits are different.

[0087] like Fig. 9 As shown, the second type first calculation unit 210B may include 6 lookup circuits LUT1 to LUT6. Each of the lookup circuits LUT1 to LUT6 may have 5 input terminals and 2 output terminals. N1=5, M=11, K=6, and M is the number of bits of the first convolution result.

[0088] The second type first calculation unit 210B may pre-store 32 pulse convolution results based on 5 first weights. The second type first calculation unit 210B may be used to receive 5 first features and search for a pulse convolution result matching the 5 first features from the 32 pulse convolution results as the first convolution result. The pulse convolution result may include 6 parts, which are respectively stored in 6 search circuits LUT1 to LUT6.

[0089] The first search circuit LUT1 can receive the five first features located at positions 11, 12, 20, 21 and 22 in the feature matrix, search for the first part matching the five first features from the first part of each stored pulse convolution result, and use it as the first part of the first convolution result, and output the 2 bits of the first part of the first convolution result through the two output terminals of the first search circuit respectively.

[0090] The second search circuit LUT2 can receive the five first features located at positions 11, 12, 20, 21 and 22 in the feature matrix, search for the second part matching the five first features from the second part of each stored pulse convolution result, and use it as the second part of the first convolution result, and output the 2 bits of the second part of the first convolution result through the two output terminals of the second search circuit respectively, and so on.

[0091] The sixth lookup circuit LUT6 can receive the five first features at positions 11, 12, 20, 21 and 22 in the feature matrix, and search for the sixth part matching the five first features from the sixth part of each stored pulse convolution result as the sixth part of the first convolution result, and output 1 bit of the sixth part of the first convolution result through an output terminal of the sixth lookup circuit. Another output terminal of the sixth lookup circuit can be set to be invalid, that is, no output is provided. Fig. 9 The lookup circuits LUT1 to LUT6 in the figure can also be lookup table logic circuits, which will not be described in detail here.

[0092] According to an embodiment of the present invention, in a convolution processor for a pulse neural network, the search circuit has 5 input terminals and 2 output terminals, N1=5, M=11, K=6, and M is the number of bits of the first convolution result. Without using an adder, the first 5 search circuits can be used to quickly search for 2 bits of the first convolution result, and the last search circuit can be used to quickly search for 1 bit of the first convolution result, thereby realizing a fast search using 6 search circuits to obtain an 11-bit first convolution result, reducing the use of adders in related technologies, reducing the amount of computing resources occupied by accelerator hardware, and improving the utilization rate of accelerator hardware computing resources.

[0093] Can be used Fig.10 The convolution processor for the pulse neural network shown implements 3×3 convolution processing on the feature matrix 110 and the weight matrix 120.

[0094] Fig.10 A schematic diagram of the structure of a convolution processor for a pulse neural network according to a sixth embodiment of the present invention is shown.

[0095] like Fig.10 As shown, the convolution processor for the pulse neural network may include a first type first computing unit 210A, a second type first computing unit 210B and a first adder 220. The first type first computing unit 210A may generate a 10-bit first convolution result based on 4 first features. The second type first computing unit 210B may generate an 11-bit first convolution result based on 5 first features. The first type first computing unit 210A may have the same structure as the first type first computing unit of any of the above-mentioned embodiments, and the second type first computing unit 210B may have the same structure as the second type first computing unit of any of the above-mentioned embodiments, which will not be repeated here. Figure 8 Taking the 3×3 convolution shown in FIG. 1 as an example, the first type first calculation unit 210A can perform table lookup convolution on the four features at positions 00, 01, 02 and 10 in the feature matrix 110 and the four weights w1, w2, w3 and w4 at the corresponding positions in the weight matrix 120 to obtain a 10-bit first convolution result. The second type first calculation unit 210B can perform table lookup convolution on the five features at positions 11, 12, 20, 21 and 22 in the feature matrix 110 to obtain an 11-bit first convolution result.

[0096] The first adder 220 may sum the first convolution results output by the first-type first computing unit 210A and the second-type first computing unit 210B to obtain a first output.

[0097] According to an embodiment of the present invention, a comparator may be used to compare the first output with a preset threshold, and output a pulse signal when the first output is greater than or equal to the preset threshold, thereby generating a pulse signal only for a second output that meets the requirements.

[0098] The first type first calculation unit 210A includes five lookup circuits having four input terminals and two output terminals. The second type first calculation unit 210B includes six lookup circuits having five input terminals and two output terminals.

[0099] According to an embodiment of the present invention, in a convolution processor for a pulse neural network, including two first computing units and a first adder, one of the two first computing units generates a 10-bit first convolution result according to four first features, the other of the two first computing units generates an 11-bit first convolution result according to five first features, one first computing unit includes five search circuits with four input terminals and two output terminals, and the other first computing unit includes six search circuits with five input terminals and two output terminals, a convolution processor for a pulse neural network is used to implement 3×3 convolution calculation, and 11 search circuits and one adder are used to quickly perform 3×3 convolution calculation. Compared with the related art that requires the use of nine adders, that is, the use of nine logic units, the amount of adder data used in the present invention is greatly reduced, and only three logic units are required, and fewer logic units are used. A large number of search circuit resources, logic gates, and trigger resources in the accelerator hardware computing resources are fully utilized. It balances the impact of resource utilization and fan-out, improves resource utilization, reduces the number of adder levels, and improves inference computing efficiency compared to related technologies.

[0100] According to an embodiment of the present invention, the lookup circuit included in one first computing unit may be a lookup table logic circuit, and the lookup circuit included in another first computing unit may be a lookup table logic circuit.

[0101] According to an embodiment of the present invention, in a convolution processor for a pulse neural network, including two first computing units and a first adder, one of the two first computing units generates a 10-bit first convolution result according to four first features, and the other of the two first computing units generates an 11-bit first convolution result according to five first features, and one first computing unit includes five lookup table logic circuits with four input terminals and two output terminals, and the other first computing unit includes six lookup table logic circuits with five input terminals and two output terminals. In this case, a convolution processor for a pulse neural network is used to implement 3×3 convolution calculation, and 11 lookup table logic circuits and one adder are used to quickly perform 3×3 convolution calculation. Compared with the related art that requires the use of nine adders, that is, the use of nine logic units, the amount of adder data used in the present invention is greatly reduced, and only three logic units are required, and fewer logic units are used. A large number of lookup circuit resources, logic gates, and trigger resources in the accelerator hardware computing resources are fully utilized. It balances the impact of resource utilization and fan-out, improves resource utilization, reduces the number of adder levels, and improves inference computing efficiency compared to related technologies.

[0102] The convolution processor for a pulse neural network provided in an embodiment of the present invention can be used for 3×3 convolution. When performing 3×3 convolution calculation, the number of the first calculation unit can be one.

[0103] When 3×3 convolution is implemented using 1 first computing unit, the number of pulse convolution results corresponding to 9 first features and 9 first weights can be 512. The number of bits of the maximum value among the 512 pulse convolution results is 12 bits.

[0104] When performing 3×3 convolution calculations, if the adder is not used and the lookup circuit is used entirely, some lookup circuits, such as the lookup table logic circuit in the FPGA, have a maximum of 6 input terminals, which does not meet the input width of 9 pulse signals, so a carry problem will occur. Therefore, it is necessary to Fig.11 The design shown is used to implement convolution calculations.

[0105] Fig.11 A schematic diagram of the structure of a convolution processor for a pulse neural network according to the seventh embodiment of the present invention is shown.

[0106] like Fig.11As shown, the convolution processor for the spiking neural network may include a third type first computing unit 210C. The third type first computing unit 210C includes K search circuits. The search circuit in the third type first computing unit 210C has a structure different from the search circuit in the aforementioned embodiment, which is referred to as a second type search circuit, represented by 211, to distinguish it from the search circuit in the aforementioned embodiment (also referred to as the first type search circuit). For the sake of simplicity, Fig.11 Only the first search circuit 211 is marked, and the other search circuits have similar structures, which will not be described here.

[0107] The third type first calculation unit 210C may pre-store 2 based on N1 first weights. N1 The first calculation unit 210 can be used to receive N1 first features from 2 N1 The pulse convolution results are searched for the pulse convolution results matching the N1 first features as the first convolution result, where N1 is an integer greater than 1.

[0108] N1 first features are divided into a first group and a second group. For example, the first group of 9 first features can be Figure 8 The first feature matrix 110 in FIG. 1 is located at 00, 01, 02 and 10. The second group of 9 first features can be Figure 8 The five first features at positions 11, 12, 20, 21 and 22 in the first feature matrix 110. The kth part of the first convolution result includes a first sub-part and a second sub-part.

[0109] The second type search circuit 211 may include a first sub-circuit 2111 and a second sub-circuit 2112 .

[0110] In the kth second type search circuit, the first subcircuit 2111 may store the first sub-portion in the kth portion of each pulse convolution result and the carry flag corresponding to the first sub-portion in the kth portion of each pulse convolution result. Fig.11 In the figure, "carry" represents a carry flag. The second subcircuit 2112 can store the second sub-portion in the kth part of each pulse convolution result and the carry flag corresponding to the second sub-portion in the kth part of each pulse convolution result. The first subcircuit 2111 can determine the first sub-portion in the kth part of the first convolution result and the carry flag according to the first group of first features. The second subcircuit 2112 can be used to determine the second sub-portion in the kth part of the first convolution result according to the carry flag output by the first subcircuit 2111 and the second group of first features.

[0111] According to an embodiment of the present invention, the first subcircuit determines the first sub-portion of the first convolution result and the carry flag according to the first group of first features, which may include: determining the first group of first features as the first target query address; according to the first target query address, searching the first sub-portion in the kth part that matches the first group of first features from the first sub-portion in the kth part of each stored pulse convolution result as the first sub-portion in the kth part of the first convolution result, and searching the carry flag corresponding to the first sub-portion in the kth part of each pulse convolution result for the carry flag that matches the first sub-portion in the kth part of the first convolution result.

[0112] According to an embodiment of the present invention, the second subcircuit determining the second subpart of the first convolution result based on the carry flag output by the first subcircuit and the second group of first features may include: determining the second group of first features and the carry flag output by the first subcircuit as the second target query address; and searching, according to the second target query address, the second subpart in the kth part that matches the second group of first features from the second subparts in the kth part of each stored pulse convolution result as the second subpart in the kth part of the first convolution result.

[0113] According to an embodiment of the present invention, a comparator can also be used to compare the first convolution result with a preset threshold, and output a pulse signal when the first convolution result is greater than or equal to the preset threshold, so as to generate a pulse signal only for the second output that meets the requirements.

[0114] According to an embodiment of the present invention, since N1 first features are divided into a first group and a second group, the kth part of the first convolution result includes a first sub-part and a second sub-part, and the kth search circuit includes a first sub-circuit and a second sub-circuit, the first sub-circuit can be used to determine the first sub-part and the carry flag in the kth part of the first convolution result according to the first group of first features, and the second sub-circuit can be used to determine the second sub-part in the kth part of the first convolution result according to the carry flag output by the first sub-circuit and the second group of first features. Furthermore, when N1 is greater than the number of input terminals of the minimum lookup table logic circuit in the accelerator hardware computing resources, without using adders and multipliers, two minimum lookup table logic circuits can be used to quickly search for the kth part of the first convolution result, and 2K minimum lookup table logic circuits can be used to quickly search for the first convolution result, thereby reducing the use of adders in the related art. Since the minimum logic unit of the accelerator hardware computing resources includes more lookup table logic circuits and fewer adders, while reducing the usage of adders, the usage of the minimum logic unit can be reduced, the accelerator hardware computing resource usage can be reduced, the fan-out is small, the design area of ​​​​the related hardware is reduced, and the utilization rate of the accelerator hardware computing resources is improved.

[0115] For example, Fig.10When the convolution processor of the pulse neural network is used for 3×3 convolution, N1=9, M=12, K=4, the first subcircuit 2111 has 4 input terminals and 2 output terminals, and the second subcircuit 2112 has 6 input terminals and 2 output terminals, where M is the number of bits of the first convolution result.

[0116] According to an embodiment of the present invention, the first subcircuit may have 4 input terminals and 2 output terminals. One output terminal of the first subcircuit outputs 1 bit of the first convolution result, and another output terminal of the first subcircuit outputs a carry flag. The second subcircuit may have 6 input terminals and 2 output terminals. The 2 output terminals of the second subcircuit output 2 bits of data in the first convolution result, so the first subcircuit and the second subcircuit can output 3 bits of data in the first convolution result, that is, the k-th search circuit can output 3 bits of data in the first convolution result.

[0117] According to an embodiment of the present invention, the first subcircuit may also have 5 input terminals and 2 output terminals. Alternatively, the first subcircuit may also have 6 input terminals and 2 output terminals. Wherein, when the first subcircuit has 5 input terminals or 6 input terminals, only 4 input terminals in the first subcircuit may be used.

[0118] According to an embodiment of the present invention, when the convolution processor of the pulse neural network includes a first computing unit, N1=9, M=12, K=4, the first subcircuit has 4 input terminals and 2 output terminals, the second subcircuit has 6 input terminals and 2 output terminals, and M is the number of bits of the first convolution result, the convolution processor of the pulse neural network is used to implement 3×3 convolution calculation, and 4 search circuits are used to quickly perform 3×3 convolution calculation. Compared with the related art that requires the use of 9 adders, that is, the use of 9 logic units, the present invention does not need to use adders, and only needs 2 logic units, and the number of logic units used is relatively small. A large number of search circuit resources in the accelerator hardware computing resources are fully utilized. Compared with the related art, the resource utilization rate is improved, the number of adder levels is reduced, and the reasoning calculation efficiency is improved.

[0119] For example, the first sub-circuit 2111 and the second sub-circuit 2112 may both be lookup table logic circuits. Figure 6 The lookup table logic circuit shown.

[0120] According to an embodiment of the present invention, the convolution processor of the pulse neural network includes a first computing unit, N1=9, M=12, K=4, the first subcircuit has 4 input terminals and 2 output terminals, the second subcircuit has 6 input terminals and 2 output terminals, M is the number of bits of the first convolution result, and the first subcircuit and the second subcircuit are both lookup table logic circuits. The convolution processor of the pulse neural network is used to implement 3×3 convolution calculations, and 4 lookup table logic circuits are used to quickly perform 3×3 convolution calculations. Compared with the related art that requires the use of 9 adders, that is, the use of 9 logic units, the present invention does not need to use adders, and only needs 2 logic units, and the use of fewer logic units. Fully utilize a large number of lookup circuit resources in the accelerator hardware computing resources. Compared with the related art, the resource utilization rate is improved, the number of adder levels is reduced, and the reasoning calculation efficiency is improved.

[0121] Based on the above-mentioned convolution processor for a spiking neural network, an embodiment of the present invention also provides a convolution processing method for a spiking neural network.

[0122] Fig.12 A flow chart of a convolution processing method for a pulse neural network according to an embodiment of the present invention is shown. Fig.12 The shown convolution processing method for a spiking neural network can be performed by the above-mentioned convolution processor for a spiking neural network.

[0123] like Fig.12 As shown, the convolution processing method for a pulse neural network may include operations S1210 to S1220.

[0124] In operation S1210, a first calculation unit receives N1 first features.

[0125] In operation S1220, the first calculating unit obtains a pre-stored 2 based on N1 first weights. N1 The pulse convolution results are searched for the pulse convolution results matching the N1 first features as the first convolution result. The pulse convolution result includes K parts, the first calculation unit includes K search circuits, the kth search circuit stores the kth part of each pulse convolution result, the kth search circuit is any search circuit among the K search circuits, and the kth part is a part of the K parts corresponding to the kth search circuit.

[0126] For operation S1220, the pre-stored 2 based on N1 first weights N1Searching for a pulse convolution result that matches N1 first features from among the kth parts of the stored pulse convolution results as the first convolution result may include: a kth search circuit searches for the kth part that matches N1 first features from the kth parts of the stored pulse convolution results as the kth part of the first convolution result, where 1≤k≤K and k is an integer, and N1 and K are integers greater than 1.

[0127] According to an embodiment of the present invention, the search circuit has N1 input terminals and m output terminals, wherein M=mK, M is the number of bits of the first convolution result, and m is the number of bits of each part of the first convolution result.

[0128] 2 from the pre-stored N1 based first weights N1 The method of searching for a pulse convolution result that matches N1 first features from N1 pulse convolution results as a first convolution result includes: a k-th search circuit receives N1 first features from N1 input terminals respectively, searches for a k-th part that matches N1 first features from the k-th parts of each stored pulse convolution result as the k-th part of the first convolution result, and outputs m bits of the k-th part of the first convolution result through m output terminals respectively.

[0129] According to an embodiment of the present invention, the lookup circuit is a lookup table logic circuit.

[0130] According to an embodiment of the present invention, there are multiple first computing units, and the convolution processor also includes a first adder. The convolution processing method for the pulse neural network also includes: the first adder sums the first convolution results output by the multiple first computing units to obtain a first output.

[0131] According to an embodiment of the present invention, the convolution processor also includes a second computing unit and a second adder. The convolution processing method for the pulse neural network also includes: the second computing unit receives N2 second features, performs convolution calculation on the received N2 second features and N2 second weights to obtain a second convolution result, wherein N2 is an integer greater than or equal to 1; the second adder sums the first output with the second convolution result provided by the second computing unit to obtain a second output.

[0132] According to an embodiment of the present invention, the convolution processor is used for 3×3 convolution, the number of the first calculation units is 2, N1=4, N2=1, M=10, K=5, and M is the number of bits of the first convolution result. The search circuit has 4 input terminals and 2 output terminals.

[0133] The kth search circuit searches for the kth part that matches N1 first features from the kth parts of each stored pulse convolution result, and the kth part of the first convolution result includes: the kth search circuit receives the four first features through the four input ends of the kth search circuit respectively, searches for the kth part that matches the four first features from the kth parts of each stored pulse convolution result, and uses it as the kth part of the first convolution result, and outputs the 2 bits of the kth part of the first convolution result through the two output ends of the kth search circuit respectively.

[0134] According to an embodiment of the present invention, N2=1, and the second computing unit includes a selector. The second computing unit receives N2 second features, and performs convolution calculation on the received N2 second features and N2 second weights to obtain a second convolution result, including: the selector selects between the N2 second weights and a preset reference value according to the N2 second features to obtain the second convolution result.

[0135] According to an embodiment of the present invention, N1 first features are divided into a first group and a second group, the kth part of the first convolution result includes a first sub-part and a second sub-part, and the kth search circuit includes a first sub-circuit and a second sub-circuit. The kth search circuit searches for the kth part that matches the N1 first features from the kth parts of each stored pulse convolution result, and the kth part of the first convolution result includes: the first sub-circuit determines the first sub-part and the carry flag in the kth part of the first convolution result according to the first group of first features, and the second sub-circuit determines the second sub-part in the kth part of the first convolution result according to the carry flag output by the first sub-circuit and the second group of first features.

[0136] According to an embodiment of the present invention, both the first sub-circuit and the second sub-circuit are look-up table logic circuits.

[0137] According to an embodiment of the present invention, the convolution processor is used for 3×3 convolution, N1=9, M=12, K=4, the first subcircuit has 4 input terminals and 2 output terminals, and the second subcircuit has 6 input terminals and 2 output terminals, wherein M is the number of bits of the first convolution result.

[0138] According to an embodiment of the present invention, a convolution processor is used for 3×3 convolution, the number of first computing units is 2, one of the two first computing units is used to generate a 10-bit first convolution result based on 4 first features, and the other of the two first computing units is used to generate an 11-bit first convolution result based on 5 first features. One first computing unit includes 5 search circuits with 4 input terminals and 2 output terminals, and the other first computing unit includes 6 search circuits with 5 input terminals and 2 output terminals.

[0139] According to an embodiment of the present invention, the convolution processor further comprises: a comparator. The convolution processing method for a pulse neural network further comprises: the comparator compares the second output with a preset threshold, and outputs a pulse signal when the second output is greater than or equal to the preset threshold.

[0140] Fig.13 A structural block diagram of an electronic device according to an embodiment of the present invention is shown.

[0141] like Fig.13 As shown, the electronic device 1300 may include a convolution processor 1310 for a pulse neural network. The electronic device 1300 may be a server for calculating convolution or a server cluster for calculating convolution.

[0142] The electronic device 1300 may include multiple convolution processors 1310 for pulse neural networks, which are used to perform parallel calculations on multiple first features and multiple first weights to further improve calculation efficiency.

[0143] It will be appreciated by those skilled in the art that the features described in the various embodiments of the present invention may be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features described in the various embodiments of the present invention may be combined and / or combined in various ways. All of these combinations and / or combinations fall within the scope of the present invention.

[0144] The embodiments of the present invention are described above. However, these embodiments are only for the purpose of illustration, and are not intended to limit the scope of the present invention. Although each embodiment is described above, it does not mean that the measures in each embodiment cannot be used in combination advantageously. Without departing from the scope of the present invention, those skilled in the art may make various substitutions and modifications, which should all fall within the scope of the present invention.

Claims

1. A convolution processor for a spiking neural network, characterized in that The convolution processor comprises: A first calculation unit is provided, wherein the first calculation unit pre-stores 2 based on N1 first weights. N1 The first computing unit is used to receive N1 first features from the 2 N1 Searching for a pulse convolution result matching the N1 first features from the pulse convolution results as the first convolution result; Wherein, the pulse convolution result includes K parts, the first calculation unit includes K search circuits, the kth search circuit stores the kth part of each pulse convolution result, the kth search circuit is any search circuit among the K search circuits, the kth part is the part of the K parts corresponding to the kth search circuit, and the kth search circuit is used to search for the kth part matching the N1 first features from the kth parts of each stored pulse convolution result as the kth part of the first convolution result, wherein 1≤k≤K and k is an integer, and N1 and K are integers greater than 1.

2. The convolution processor according to claim 1, characterized in that The search circuit has N1 input terminals and m output terminals, wherein M=mK, M is the number of bits of the first convolution result, and m is the number of bits of each part of the first convolution result; The kth search circuit is used to receive the N1 first features from N1 input terminals respectively, search for the kth part matching the N1 first features from the kth parts of each stored pulse convolution result as the kth part of the first convolution result, and output the m bits of the kth part of the first convolution result through m output terminals respectively.

3. The convolution processor according to claim 2, characterized in that The lookup circuit is a lookup table logic circuit.

4. The convolution processor according to claim 1, characterized in that There are multiple first computing units, and the convolution processor also includes: a first adder, used for summing the first convolution results output by the multiple first computing units to obtain a first output.

5. The convolution processor according to claim 4, characterized in that The convolution processor also includes: A second calculation unit is used to receive N2 second features, and perform convolution calculation on the received N2 second features and N2 second weights to obtain a second convolution result, wherein N2 is an integer greater than or equal to 1; The second adder is used to sum the first output with the second convolution result provided by the second computing unit to obtain a second output.

6. The convolution processor according to claim 5, characterized in that The convolution processor is used for 3×3 convolution, the number of the first computing units is 2, N1=4, N2=1, M=10, K=5, and M is the number of bits of the first convolution result; The search circuit has 4 input terminals and 2 output terminals, wherein the kth search circuit is used to receive 4 first features respectively through the 4 input terminals of the kth search circuit, search for the kth part matching the 4 first features from the kth parts of each stored pulse convolution result as the kth part of the first convolution result, and output the 2 bits of the kth part of the first convolution result respectively through the 2 output terminals of the kth search circuit.

7. The convolution processor according to claim 5, characterized in that N2=1, the second calculation unit includes: a selector, which is used to select between N2 second weights and a preset reference value according to N2 second features to obtain a second convolution result.

8. The convolution processor according to claim 1, characterized in that The N1 first features are divided into a first group and a second group, and the kth part of the first convolution result includes a first sub-part and a second sub-part; The kth search circuit includes a first sub-circuit and a second sub-circuit, the first sub-circuit is used to determine the first sub-part and the carry flag in the kth part of the first convolution result according to the first group of first features, and the second sub-circuit is used to determine the second sub-part in the kth part of the first convolution result according to the carry flag output by the first sub-circuit and the second group of first features.

9. The convolution processor according to claim 8, characterized in that The first sub-circuit and the second sub-circuit are both look-up table logic circuits.

10. The convolution processor according to claim 8, characterized in that The convolution processor is used for 3×3 convolution, N1=9, M=12, K=4, the first subcircuit has 4 input terminals and 2 output terminals, the second subcircuit has 6 input terminals and 2 output terminals, wherein M is the number of bits of the first convolution result.

11. The convolution processor according to claim 4, characterized in that: The convolution processor is used for 3×3 convolution, the number of the first computing units is 2, one of the two first computing units is used to generate a 10-bit first convolution result according to 4 first features, and the other of the two first computing units is used to generate an 11-bit first convolution result according to 5 first features. The one first computing unit includes 5 search circuits with 4 input terminals and 2 output terminals, and the other first computing unit includes 6 search circuits with 5 input terminals and 2 output terminals.

12. The convolution processor according to claim 5, characterized in that The convolution processor further includes a comparator for comparing the second output with a preset threshold and outputting a pulse signal when the second output is greater than or equal to the preset threshold.

13. A convolution processing method for a spiking neural network, executed by the convolution processor of claim 1, characterized in that: The method comprises: The first computing unit receives N1 first features; The first calculation unit obtains the pre-stored 2 based on N1 first weights N1 Among the pulse convolution results, searching for the pulse convolution result matching the N1 first features as the first convolution result; The pulse convolution result includes K parts, the first calculation unit includes K search circuits, the kth search circuit stores the kth part of each pulse convolution result, the kth search circuit is any search circuit among the K search circuits, the kth part is the part of the K parts corresponding to the kth search circuit, the 2 pre-stored N1 first weights are used to calculate the kth part of each pulse convolution result. N1 Searching for a pulse convolution result that matches the N1 first features from the N1 pulse convolution results as the first convolution result includes: a k-th search circuit searches for a k-th part that matches the N1 first features from the k-th parts of the stored pulse convolution results as the k-th part of the first convolution result, wherein 1≤k≤K and k is an integer, and N1 and K are integers greater than 1.

14. The method according to claim 13, characterized in that The search circuit has N1 input terminals and m output terminals, wherein M=mK, M is the number of bits of the first convolution result, m is the number of bits of each part of the first convolution result, and the pre-stored 2 based on N1 first weights is N1 The step of searching for a pulse convolution result matching the N1 first features as the first convolution result includes: The kth search circuit receives the N1 first features from N1 input terminals respectively, searches for the kth part matching the N1 first features from the kth parts of each stored pulse convolution result as the kth part of the first convolution result, and outputs the m bits of the kth part of the first convolution result through m output terminals respectively.

15. The method according to claim 13, characterized in that There are multiple first computing units, and the convolution processor also includes a first adder. The method also includes: the first adder sums the first convolution results output by the multiple first computing units to obtain a first output.

16. The method according to claim 15, characterized in that The convolution processor further includes a second computing unit and a second adder, and the method further includes: The second calculation unit receives N2 second features, and performs convolution calculation on the received N2 second features and N2 second weights to obtain a second convolution result, wherein N2 is an integer greater than or equal to 1; The second adder sums the first output and the second convolution result provided by the second computing unit to obtain a second output.

17. The method according to claim 16, characterized in that The convolution processor is used for 3×3 convolution, the number of the first computing units is 2, N1=4, N2=1, M=10, K=5, and M is the number of bits of the first convolution result; The search circuit has 4 input terminals and 2 output terminals, wherein the kth search circuit searches for the kth part matching the N1 first features from the kth parts of the stored pulse convolution results, as the kth part of the first convolution result, including: the kth search circuit receives the four first features respectively through the four input terminals of the kth search circuit, searches for the kth part matching the four first features from the kth parts of the stored pulse convolution results, as the kth part of the first convolution result, and outputs the 2 bits of the kth part of the first convolution result respectively through the two output terminals of the kth search circuit.

18. The method according to claim 17, characterized in that N2=1, the second computing unit includes a selector, the second computing unit receives N2 second features, performs convolution calculation on the received N2 second features and N2 second weights, and obtains a second convolution result including: The selector selects between the N2 second weights and a preset reference value according to the N2 second features to obtain a second convolution result.

19. The convolution processing method according to claim 13, characterized in that: The N1 first features are divided into a first group and a second group, the kth part of the first convolution result includes a first sub-part and a second sub-part, and the kth search circuit includes a first sub-circuit and a second sub-circuit; The kth search circuit searches for the kth part matching the N1 first features from the kth parts of each stored pulse convolution result, and the kth part of the first convolution result includes: a first subcircuit determines the first subpart and the carry flag in the kth part of the first convolution result according to the first group of first features, and a second subcircuit determines the second subpart in the kth part of the first convolution result according to the carry flag output by the first subcircuit and the second group of first features.

20. The convolution processing method according to claim 16, characterized in that: The convolution processor further includes: a comparator, and the method further includes: The comparator compares the second output with a preset threshold, and outputs a pulse signal when the second output is greater than or equal to the preset threshold.

21. An electronic device comprising the convolution processor according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Neural network calculation method and device, electronic equipment and storage medium

    CN112215338A

  • Programmable convolutional neural network processor, method, equipment, medium and terminal

    CN113435570A

  • Pulse neural network based on probability calculation and implementation method thereof

    CN115545190A

  • Control method and device of spiking neural network, equipment and storage medium

    CN116205274A

  • Pulse neural network training method and system, electronic equipment and storage medium

    CN117350355A