Convolution acceleration method, device, equipment and medium for pulse neural network

By adopting column-based pulse random access memory and row completion signal mechanism in the pulse neural network, parallel reading and single-cycle convolution of the pulse window and historical cache value are realized, which solves the problem of low efficiency of sparse data processing in the existing technology and improves computing efficiency and resource utilization.

CN120562489BActive Publication Date: 2025-09-19INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511064350.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-09-19
Estimated Expiration
2045-07-31

AI Technical Summary

Technical Problem

The existing spiking neural network computing architecture is inefficient when processing sparse data, resulting in increased system latency and an inability to effectively utilize the computing efficiency of the spiking neural network.

Method used

The column pulse random access memory and row completion signal mechanism are used to directly receive the coordinate events sent by the pulse camera. The pulse window and historical cache values ​​are read in parallel through the column pulse RAM and row completion signal mechanism. The pre-cured weights and storage table are used to perform single-cycle convolution calculations to filter out all-0 data.

Benefits of technology

It significantly reduces the amount of calculation and delay, improves the computational efficiency of the convolution accelerator, saves resources, reduces invalid operations, and improves processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120562489B_ABST
    Figure CN120562489B_ABST
Patent Text Reader

Abstract

This application discloses a convolution acceleration method, device, equipment, and medium for a pulse neural network. Because the accelerator card only receives events with changed pixel coordinates sent by the pulse camera, it can automatically filter out all-zero data. Compared with traditional frame-based processing that scans pixel by pixel, the amount of computation is significantly reduced, and decoding can be achieved directly using the accelerator / card without the need for CPU decoding. Through the column pulse RAM and row completion signal mechanism, the pulse window and historical cache value can be read in parallel in the same clock cycle, achieving single-cycle convolution. The weight and storage table solidify the convolution results of multiple encoded data inputs in advance. During operation, the convolution sum can be read by reading only the weight and storage table, saving resources and improving efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of hardware acceleration, and in particular to a convolution acceleration method, apparatus, device, and medium for pulse neural networks. Background Art

[0002] A spiking neural network (SNN) is an artificial neural network that uses pulses as basic signal units. Its neurons only emit discrete pulses when the membrane potential exceeds the threshold, transmitting information through time coding. It is an efficient brain-like computing structure.

[0003] In related technologies, a pulse neural network receives and processes data sent by a pulse camera. The processing mechanism of the pulse camera is to process only the images of areas with drastic changes into pulse data and send them out, while no data is sent for areas without changes. This ensures that the amount of data transmitted is minimized. However, the currently commonly used pulse neural network computing architecture usually uses a CPU (Central Processing Unit) to first restore the pulse sequence into pulse matrix data (that is, to restore sparse data to a complete matrix), and then passes it to the acceleration unit for matrix calculation. However, this method increases the data decoding time and the time to put the data into the CPU's DDR (Double Data Rate Synchronous Dynamic Random-Access Memory) and then retrieve it, resulting in lower computing efficiency, increased system latency, and affected processing efficiency.

[0004] Therefore, how to improve the computational efficiency of convolution accelerators is an urgent problem that needs to be solved. Summary of the Invention

[0005] The present application provides a convolution acceleration method, apparatus, device and medium for a pulse neural network to at least solve the problem that related technologies cannot dynamically adjust the grouping of path devices.

[0006] This application provides a convolution acceleration method for a spiking neural network, the method comprising:

[0007] Receive the coordinate event sent by the pulse camera, and write the pulse value of the coordinate event into the corresponding column pulse random access memory;

[0008] When the column pulse random access memory generates a row completion signal, storing the column coordinate value and the center row coordinate value corresponding to the row completion signal in the pulse buffer space, wherein the row completion signal is a signal generated when the column pulse random access memory includes three consecutive row pulse values;

[0009] According to the column coordinate value and the row coordinate value output by the pulse buffer space, the pulse value in the pulse window of the preset size and the historical buffer value in the pulse window of the preset size are read in parallel;

[0010] Encoding the pulse value of the pulse window of the preset size to obtain encoded data and a column address; the column address is the column address corresponding to the pulse window of the preset size;

[0011] Querying a weight and storage table according to the coded data to obtain a convolution result corresponding to the coded data; the weight and storage table includes a plurality of pre-calculated weight convolution sums of the coded data;

[0012] A historical cache value of a corresponding position is obtained according to the column address, and the convolution result is added to the historical cache value of the corresponding position to obtain a weighted cumulative sum.

[0013] The present application also provides a convolution acceleration device for a spiking neural network, comprising:

[0014] A receiving module, configured to receive a coordinate event sent by a pulse camera and write the pulse value of the coordinate event into a corresponding column pulse random access memory;

[0015] a storage module configured to store the column coordinate value and the center row coordinate value corresponding to the row completion signal and the center row coordinate value in a pulse buffer space when the column pulse random access memory generates a row completion signal, wherein the row completion signal is a signal generated when the column pulse random access memory includes three consecutive row pulse values;

[0016] a reading module, configured to read, in parallel, a pulse value within a pulse window of a preset size and a historical cache value within the pulse window of the preset size according to the column coordinate value and the row coordinate value output by the pulse buffer space;

[0017] an encoding module, configured to encode the pulse value of the pulse window of the preset size to obtain encoded data and a column address after encoding; the column address is the column address corresponding to the pulse window of the preset size; a table lookup module, configured to query a weight sum storage table according to the encoded data to obtain a convolution result corresponding to the encoded data; the weight sum storage table includes a plurality of pre-calculated weight convolution sums of the encoded data;

[0018] A calculation module is used to obtain a historical cache value of a corresponding position according to the column address, and add the convolution result to the historical cache value of the corresponding position to obtain a weighted cumulative sum.

[0019] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned convolution acceleration methods for a pulse neural network when executing the computer program.

[0020] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned convolution acceleration methods for pulse neural networks are implemented.

[0021] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned convolution acceleration methods for pulse neural networks when executed by a processor.

[0022] This application uses an accelerator card that only receives events with changed pixel coordinates from the pulse camera and automatically filters out all-zero data. This significantly reduces computational complexity compared to traditional frame-based processing, which scans pixel by pixel. Decoding can be performed directly on the accelerator card, eliminating the need for CPU decoding. Using a column pulse RAM and row completion signal mechanism, the pulse window and historical cache values ​​can be read in parallel during the same clock cycle, enabling single-cycle convolution. The weight and storage tables pre-store the convolution results for multiple 9-bit (encoded data) inputs. At runtime, the convolution sum can be read simply by reading the weight and storage tables, saving resources and improving efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0024] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0025] Figure 1 A schematic diagram of an end-to-end architecture from a pulse camera to an accelerator card provided in an embodiment of the present application;

[0026] Figure 2 A schematic diagram of the convolution calculation process of a pulse neural network in related technology;

[0027] Figure 3 A flow chart of a convolution acceleration method for a spiking neural network provided in an embodiment of the present application;

[0028] Figure 4A A schematic structural diagram of a column pulse random access memory provided in an embodiment of the present application;

[0029] Figure 4B A schematic diagram of a parallel column pulse RAM address decoding and caching architecture provided in an embodiment of the present application;

[0030] Figure 4C A coordinate event buffer decoding control block diagram of a pulse signal provided in an embodiment of the present application;

[0031] Figure 4D A schematic diagram of a pulse window encoding process provided in an embodiment of the present application;

[0032] Figure 4E A schematic diagram of a periodic convolution with a coding weight lookup table provided in an embodiment of the present application;

[0033] Figure 4F A schematic diagram of obtaining and judging a weighted cumulative sum provided in an embodiment of the present application;

[0034] Figure 4G A schematic diagram of a convolutional computation process framework for a spiking neural network according to an embodiment of the present disclosure;

[0035] Figure 5 A schematic diagram of the structure of a convolution acceleration device for a pulse neural network provided in an embodiment of the present application. DETAILED DESCRIPTION

[0036] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0037] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.

[0038] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0039] The rapid development of artificial intelligence has led to a surge in the amount of real-time data. However, truly useful information is often concentrated in a few areas, leaving much of the raw data unreliable. Therefore, efficiently processing this sparse data has become a hot topic of research. Spiking neural networks mimic the workings of neurons in the brain. Not only is data sparse, but they also employ event-triggered processing: computations are initiated only when an event occurs. This mechanism automatically filters out redundant operations, significantly reducing overall computational effort and establishing a highly efficient and energy-efficient paradigm for brain-inspired computing.

[0040] In related technologies, the pulse neural network receives and processes the data sent by the pulse camera. The processing mechanism of the pulse camera is to process only the image of the area with drastic changes as pulse data and send it out, and not send data for the area without change. This can ensure that the amount of data during transmission is minimized. However, the current commonly used pulse neural network computing architecture, refer to Figure 1 As shown, Figure 1 This is a diagram of the end-to-end architecture from a pulse camera to an accelerator card. Typically, the CPU (Central Processing Unit) is used to restore the pulse sequence to pulse matrix data (i.e., sparse data is restored to a complete matrix), which is then passed to the accelerator unit for matrix calculations. Although the convolution of a spiking neural network only requires addition and no multiplier, after the convolution is completed, the cached values ​​at the corresponding positions must be accumulated before determining whether a pulse can be sent. Figure 2 As shown, Figure 2 This is a diagram of the convolution calculation process of a spiking neural network in related technology. First, three rows of data are read out in parallel from the cache. The data in the read-out 3X3 data matrix are all 1 or 0 numbers. The positions with data 1 are replaced with the corresponding values ​​in the convolution weight matrix. The values ​​are then added and finally added to the previous cache. The result is then compared with the threshold. If it is greater than the threshold, the previous cache is cleared and a pulse is sent. If it is less than the threshold, the latest accumulated result is updated to the previous cache register and no pulse is sent. However, this method increases the data decoding time and the time to put the data into the CPU's DDR (Double Data Rate Synchronous Dynamic Random-Access Memory) and then retrieve it, resulting in lower computational efficiency, increased system latency, and affected processing efficiency.

[0041] Therefore, how to improve the computational efficiency of convolution accelerators is an urgent problem that needs to be solved.

[0042] Based on the above problems, an embodiment of the present application provides a convolution acceleration method for a pulse neural network, and describes the method in detail in combination with the execution process of the convolution acceleration method for a pulse neural network.

[0043] Reference Figure 3 As shown, the convolution acceleration method of the spiking neural network provided by the embodiment of the present invention includes the following steps:

[0044] S31 , receiving a coordinate event sent by a pulse camera, and writing the pulse value of the coordinate event into a corresponding column pulse random access memory.

[0045] The coordinate event includes the row coordinate value and the column coordinate value of the pulse signal.

[0046] Specifically, the accelerator card / device receives the coordinate event sent by the pulse camera and writes the pulse value of the coordinate event into the corresponding column pulse random access memory. It is understood that the accelerator card can simultaneously receive multiple coordinate events sent by the pulse camera and write each coordinate event into the corresponding column pulse random access memory.

[0047] For example, if 3×3 convolution is used, the depth of a single RAM only needs to be greater than or equal to 3. Since there is no RAM with a depth of only 3, the minimum RAM can be used here. Figure 4A As shown, Figure 4A This is a schematic diagram of the column pulse random access memory (RAM). Each RAM requires only three cells (height 3) to store three rows of pixels in a column. Since each RAM has four cells built in, it can be used directly without cascading to form a larger RAM. The initial value in the RAM is 0. When the accelerator card receives the pulse coordinates from the pulse camera, it simply updates the initial value in the RAM to the pulse value based on the pulse coordinates.

[0048] Optionally, the above step S31 (receiving the coordinate event sent by the pulse camera and writing the pulse value of the coordinate event into the corresponding column pulse random access memory) can also be implemented in the following way:

[0049] Receive the first row coordinate value and the first column coordinate value sent by the pulse camera;

[0050] determining a column pulse random access memory to be written into which the pulse signal is to be written according to the first column coordinate value;

[0051] The pulse value corresponding to the first row coordinate value and the first column coordinate value is written into the to-be-written column pulse random access memory.

[0052] Specifically, the first row coordinate value and the first column coordinate value sent by the pulse camera are received, and according to the first column coordinate value, the column pulse random access memory to be written of the pulse signal is determined, and the pulse value corresponding to the first row coordinate value and the first column coordinate value is written into the column pulse random access memory to be written.

[0053] For example, assuming that the first row coordinate value is 4 and the first column coordinate value is 1, based on the first column coordinate value "1", it is determined that the column pulse random access memory to be written by the pulse signal is RAM1, and the pulse value 1 corresponding to the first row coordinate value and the first column coordinate value is written into the column pulse random access memory RAM1 to be written.

[0054] In the disclosed embodiment, the column coordinates directly correspond to the "numbered column RAM", and there is no need for a row-column composite decoder. Each column has an independent write port, and multiple columns of RAM can be written in parallel within the same clock cycle, facilitating parallel computing.

[0055] In some embodiments, the above step S31 (receiving the coordinate event sent by the pulse camera and writing the pulse value of the coordinate event into the corresponding column pulse random access memory) can be implemented as follows:

[0056] Receive multiple coordinate events sent by the pulse camera.

[0057] Wherein, the coordinate event includes the row coordinate value and the column coordinate value of the pulse signal;

[0058] The pulse values ​​corresponding to the plurality of column coordinate values ​​are respectively written in parallel into the corresponding column pulse random access memory.

[0059] Specifically, the accelerator card / device receives multiple coordinate events sent by the pulse camera, and writes pulse values ​​corresponding to multiple column coordinate values ​​into corresponding column pulse random access memories in parallel.

[0060] For example, the pulse camera outputs coordinate events (x, y), which are written into the corresponding column pulse RAM by the pulse address decoding controller according to the column number. Figure 4B As shown, Figure 4B This is a schematic diagram of the parallel column pulse RAM address decoding and caching architecture. When the accelerator card receives the pulse coordinate value from the pulse camera, it only needs to update the initial value in the corresponding column RAM to the pulse value according to each column coordinate value.

[0061] In the disclosed embodiment, the pulse camera outputs coordinate events in the form of event packets, which can be written directly and in parallel to multiple column pulse RAMs by column number, eliminating the need to first decode sparse events into a full-frame 0 / 1 matrix, thus saving external bus bandwidth. Each column RAM has an independent port, allowing simultaneous reception of events from different columns. The write latency is constant and does not increase with increasing image width. The event-to-RAM write path has only a single-beat delay, allowing subsequent row detectors to immediately use the latest data, reducing the overall convolution latency and improving convolution computation efficiency.

[0062] S32 . When the column pulse random access memory generates a row completion signal, the column coordinate value and the center row coordinate value corresponding to the row completion signal are stored in a pulse buffer space.

[0063] The row completion signal is a signal generated when the column pulse random access memory includes three consecutive rows of pulse values.

[0064] Specifically, when the column spike random access memory generates a row completion signal, the column coordinate value corresponding to the row completion signal and the center row coordinate value are stored in the spike buffer space. The spike buffer space can be a FIFO (First-In-First-Out Buffer). In spiking neural networks (SNNs), FIFO buffers are often used to store spike events to avoid data loss or timing conflicts.

[0065] For example, refer to Figure 4C As shown, Figure 4C This is a pulse signal coordinate event cache decoding control block diagram. When a row completion signal is generated in the column RAM, the column coordinate value of the 3×3 data of the valid pulse position corresponding to the row completion signal and the center row coordinate value are stored in the pulse cache space.

[0066] In some embodiments, before executing the above step S32 (when the column pulse random access memory generates a row completion signal, storing the column coordinate value and the center row coordinate value corresponding to the row completion signal in the pulse buffer space), the following steps may also be executed:

[0067] detecting whether the column pulse random access memory includes three consecutive rows of pulse values;

[0068] If the same column of pulse random access memory includes three consecutive rows of pulse values, a row completion signal is generated.

[0069] Specifically, the row detector detects whether three consecutive rows of pulse values ​​are collected in the same column of the pulse random access memory. If three consecutive rows of pulse values ​​are collected in the same column of the pulse random access memory, a row completion signal is generated.

[0070] For example, taking the pulse camera output event (x, y) as an example, the row detector detects whether the pulse random access memory in the same column has collected three row events y-1, y, and y+1. If three consecutive rows of pulse values ​​are collected in the pulse random access memory in the same column, a row completion signal is generated.

[0071] S33 . Reading, in parallel, the pulse value within a pulse window of a preset size and the historical cache value within the pulse window of the preset size according to the column coordinate value and the row coordinate value output by the pulse buffer space.

[0072] The preset size may be a 3×3 window size or a window size of other reasonable values, and is not specifically limited here.

[0073] Specifically, the matrix address controller concurrently reads the pulse values ​​within a preset pulse window and the historical cache values ​​within that window based on the column and row coordinate values ​​output by the pulse buffer space. It should be noted that the data output by a pulse camera generally consists of pulse values ​​plus position information, and pulse cameras only transmit image regions where changes have occurred. Due to the sparse nature of pulse cameras, the matrix address controller only retrieves the 3×3 data points corresponding to the locations where changes have occurred, and ignores invalid zero-valued data.

[0074] In some embodiments, the above step S33 (parallel reading of the pulse value within a pulse window of a preset size and the historical cache value within the pulse window of the preset size according to the column coordinate value and the row coordinate value output by the pulse buffer space) can be implemented as follows:

[0075] According to the column coordinate values ​​and row coordinate values ​​output by the pulse cache space, three row read ports of three column pulse random access memories and three row read ports of the cumulative random access memory are opened in parallel to read the pulse values ​​within a pulse window of a preset size and the historical cache values ​​within the pulse window of a preset size.

[0076] The cumulative random access memory is used to store historical cache values.

[0077] Specifically, according to the column coordinate values ​​and row coordinate values ​​output by the pulse cache space, three row read ports of three column pulse random access memories and three row read ports of the cumulative random access memory are opened in parallel to read the pulse values ​​within a pulse window of a preset size and the historical cache values ​​within a pulse window of a preset size.

[0078] Exemplarily, the matrix address controller opens the three column RAMs col-1, col, and col+1 and the three row read ports of the accumulation RAM simultaneously according to the FIFO output to read the values ​​of the 3×3 pulse window and the 3×3 historical cache values.

[0079] In the disclosed embodiment, through the column-parallel and three-row read port structure, the read ports of three column pulse RAMs and three accumulation RAMs are simultaneously opened in the same clock cycle, so that a 3×3 pulse window and a 3×3 historical cache value can be read out at one time, compressing the traditional serial three-stage addition cycle to one clock cycle, significantly reducing latency.

[0080] S34 . Encode the pulse value of the pulse window of the preset size to obtain encoded data and column address.

[0081] The column address is a column address corresponding to the pulse window of the preset size.

[0082] Specifically, the encoder encodes the pulse values ​​within a preset pulse window to obtain the encoded data and column address. Data encoding combines nine 0 / 1 data points into a single 9-bit data point. Position encoding only requires storing the column coordinates; since all 3×3 data points in this example belong to the same row, there's no need to store the row data.

[0083] Optionally, the above step S34 (encoding the pulse value of the pulse window of the preset size and obtaining the encoded coded data and column address) can be implemented as follows:

[0084] Encoding the pulse window of the preset size into address data according to a fixed scanning order;

[0085] A column address corresponding to the pulse window of the preset size is obtained.

[0086] Specifically, a pulse window of a preset size is encoded into address data according to a fixed scanning sequence, and a column address corresponding to the pulse window of the preset size is obtained.

[0087] For example, refer to Figure 4D As shown, Figure 4D The following is a schematic diagram of the pulse window encoding process. For the 3×3 data read out, data encoding and position encoding are performed. Data encoding combines 9 0 / 1 data into a 9-bit data. Position encoding only needs to save the column coordinates. Since all the 3×3 data here are in the same row, there is no need to save the row data. Figure 4D shown.

[0088] In the disclosed embodiment, a 3×3 pulse window is directly assembled into a 9-bit binary address in a fixed scan sequence, ensuring a one-to-one correspondence between the address data of the 512 windows and the 512 weighted convolution sums, avoiding address conflicts or repeated table lookups. The column address is generated simultaneously with the 9-bit window address. The subsequent matrix address controller does not require secondary decoding and can directly open three column RAMs in parallel, saving one level of address decoding cycles. Furthermore, the fixed scan sequence can be expanded to larger windows, such as 4×4 and 5×5, simply by increasing the address bit width while maintaining the same logical structure, enabling smooth upgrades.

[0089] In some embodiments, after executing the above step S34 (encoding the pulse values ​​of the pulse window of the preset size and obtaining the encoded coded data and column address), the following steps may also be executed:

[0090] The address data and the column address are stored in a register with preset bits.

[0091] Among them, the preset bit number can be 9 bits.

[0092] Specifically, the address data and the column address are stored in a register with preset bits.

[0093] S35. Query the weight and storage table according to the encoded data to obtain the convolution result corresponding to the encoded data.

[0094] The weight and storage table includes a plurality of pre-calculated weight convolution sums of the encoded data.

[0095] Specifically, since the weight sum storage table includes a plurality of pre-calculated weight convolution sums of the coded data, the convolution result corresponding to the coded data can be directly queried and obtained by querying the weight sum storage table according to the coded data.

[0096] For example, refer to Figure 4E As shown, Figure 4E This is a diagram of a periodic convolution with a weighted lookup table. Since the weights of a single layer are essentially fixed, the convolution of a spiking neural network only involves addition. Here, a small RAM is used to store all weight results. By encoding the location and selecting the corresponding calculation result, the corresponding convolution addition calculation can be completed in just one clock cycle.

[0097] In the disclosed embodiment, the weight and storage table pre-stores the convolution results corresponding to 512 different 9-bit inputs. At runtime, a single ROM table lookup is required to complete an operation that would otherwise require nine multiplications and additions. This reduces the number of multipliers from nine to zero, and compresses the adder tree from three to one, significantly reducing the overall computational effort and improving computational efficiency. Because inputs are limited to 0 / 1, 512 table entries cover all valid cases. For sparse event streams, invalid inputs are automatically mapped to entries with a value of 0, eliminating the need for additional decision logic and further reducing invalid flips.

[0098] S36. Obtain a historical cache value of a corresponding position according to the column address, and add the convolution result to the historical cache value of the corresponding position to obtain a weighted cumulative sum.

[0099] Specifically, the accumulation RAM in the same column can be located based on the column address. Three rows of read ports output the historical cache value in parallel, and the adder completes the accumulation of the convolution result and the historical value within the same clock cycle, reducing overall latency. Furthermore, only the RAM column containing the valid event is accessed, while the remaining RAM columns remain static, reducing dynamic power consumption.

[0100] In some embodiments, after executing step S36 (obtaining a historical cache value at a corresponding position according to the column address, and adding the convolution result to the historical cache value at the corresponding position to obtain a weighted cumulative sum), the following steps may also be executed:

[0101] Determine whether the weighted cumulative sum is greater than or equal to a cumulative sum threshold;

[0102] If the weighted cumulative sum is greater than or equal to the cumulative sum threshold, a pulse signal is output, and the historical cache value at the corresponding position in the cumulative random access memory is cleared;

[0103] If the weighted cumulative sum is less than the cumulative sum threshold, the historical cache value of the corresponding position in the cumulative random access memory is updated according to the weighted cumulative sum.

[0104] The cumulative sum threshold can be set according to actual conditions and is not specifically limited here.

[0105] Specifically, it is determined whether the weighted cumulative sum is greater than or equal to the cumulative sum threshold. If the weighted cumulative sum is greater than or equal to the cumulative sum threshold, a pulse signal is output and the historical cache value of the corresponding position in the cumulative random access memory is cleared. If the weighted cumulative sum is less than the cumulative sum threshold, the historical cache value of the corresponding position in the cumulative random access memory is updated according to the weighted cumulative sum.

[0106] For example, refer to Figure 4F As shown, Figure 4F Schematic diagram for obtaining and judging the weighted cumulative sum.

[0107] In the embodiment of the present disclosure, by judging whether the weight cumulative sum is greater than or equal to the cumulative sum threshold, a pulse is output and the cache is cleared only when the weight cumulative sum is greater than or equal to the cumulative sum threshold. In other cases, only one addition is performed and the historical cache value is updated. This can avoid the meaningless operation of traditional frame convolution on the all-zero area, reduce power consumption, and shorten the overall convolution delay period.

[0108] Reference Figure 4G As shown, Figure 4G A schematic diagram of the convolution calculation process framework of a pulse neural network provided in an embodiment of the present disclosure, wherein the calculation process of the pulse calculation array can refer to Figure 4F shown. Figure 4G A convolution acceleration method for a pulse neural network is provided, which receives a coordinate event sent by a pulse camera and writes the pulse value of the coordinate event into a corresponding column pulse random access memory. When the column pulse random access memory generates a row completion signal, the column coordinate value corresponding to the row completion signal and the center row coordinate value are stored in a pulse buffer space, wherein the row completion signal is a signal generated when the column pulse random access memory includes three consecutive rows of pulse values. According to the column coordinate value and row coordinate value output by the pulse buffer space, the pulse value within a pulse window of a preset size and the historical cache value within the pulse window of the preset size are read in parallel, the pulse value of the pulse window of the preset size is encoded, and the encoded coded data and column address are obtained, wherein the column address is the column address corresponding to the pulse window of the preset size; a weight and storage table is queried according to the coded data to obtain a convolution result corresponding to the coded data, wherein the weight and storage table includes a plurality of pre-calculated weight convolution sums of the coded data; the historical cache value of the corresponding position is obtained according to the column address; the convolution result is added to the historical cache value of the corresponding position to obtain a weighted cumulative sum. Because the accelerator card only receives pixel coordinate events indicating changes from the pulse camera, it can automatically filter out all-zero data. This significantly reduces computational complexity compared to traditional frame-based processing that scans pixel by pixel, and decoding can be performed directly on the accelerator / card without requiring CPU decoding. Using column pulse RAM and a row completion signal mechanism, the pulse window and historical cache values ​​can be read in parallel during the same clock cycle, enabling single-cycle convolution. The weight and storage tables pre-store the convolution results for multiple 9-bit (encoded data) inputs. At runtime, the convolution sum can be read simply by reading the weight and storage tables, saving resources and improving efficiency. Furthermore, the position code (column address) allows direct targeting of valid pixels, skipping invalid zero values.

[0109] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0110] Figure 5 This is a schematic diagram of the structure of a convolution acceleration device 500 for a pulse neural network provided by the present disclosure, as shown in FIG. Figure 5 As shown, the device of this embodiment includes:

[0111] The receiving module 510 is configured to receive a coordinate event sent by a pulse camera and write the pulse value of the coordinate event into a corresponding column pulse random access memory;

[0112] a storage module 520 configured to store the column coordinate value and the center row coordinate value corresponding to a row completion signal generated by the column pulse random access memory in a pulse buffer space when the column pulse random access memory generates a row completion signal, wherein the row completion signal is a signal generated when the column pulse random access memory includes three consecutive row pulse values;

[0113] a reading module 530 configured to read, in parallel, the pulse value within a pulse window of a preset size and the historical cache value within the pulse window of the preset size according to the column coordinate value and the row coordinate value outputted by the pulse buffer space;

[0114] The encoding module 540 is used to encode the pulse value of the pulse window of the preset size to obtain encoded data and a column address after encoding; the column address is the column address corresponding to the pulse window of the preset size;

[0115] A table lookup module 550 is configured to query a weight sum storage table according to the coded data to obtain a convolution result corresponding to the coded data; the weight sum storage table includes a plurality of pre-calculated weight convolution sums of the coded data;

[0116] The calculation module 560 is configured to obtain a historical cache value of a corresponding position according to the column address, and add the convolution result to the historical cache value of the corresponding position to obtain a weighted cumulative sum.

[0117] As an optional implementation of the embodiment of the present disclosure, the receiving module 510 is specifically configured to:

[0118] Receive multiple coordinate events sent by the pulse camera; the coordinate events include row coordinate values ​​and column coordinate values ​​of the pulse signal;

[0119] The pulse values ​​corresponding to the plurality of column coordinate values ​​are respectively written in parallel into the corresponding column pulse random access memory.

[0120] As an optional implementation of the embodiment of the present disclosure, the reading module 530 is specifically configured to:

[0121] According to the column coordinate values ​​and row coordinate values ​​output by the pulse cache space, three row read ports of three column pulse random access memories and three row read ports of the cumulative random access memory are opened in parallel to read the pulse values ​​within a pulse window of a preset size and the historical cache values ​​within the pulse window of a preset size; the cumulative random access memory is used to store the historical cache values.

[0122] As an optional implementation of the embodiment of the present disclosure, the device further includes a detection module, which is specifically configured to:

[0123] detecting whether the column pulse random access memory includes three consecutive rows of pulse values;

[0124] If the same column of pulse random access memory includes three consecutive rows of pulse values, a row completion signal is generated.

[0125] As an optional implementation of the embodiment of the present disclosure, the device further includes a judgment module, which is specifically configured to:

[0126] Determine whether the weighted cumulative sum is greater than or equal to a cumulative sum threshold;

[0127] If the weighted cumulative sum is greater than or equal to the cumulative sum threshold, a pulse signal is output, and the historical cache value at the corresponding position in the cumulative random access memory is cleared;

[0128] If the weighted cumulative sum is less than the cumulative sum threshold, the historical cache value of the corresponding position in the cumulative random access memory is updated according to the weighted cumulative sum.

[0129] As an optional implementation of the embodiment of the present disclosure, the encoding module 540 is specifically configured to:

[0130] Encoding the pulse window of the preset size into address data according to a fixed scanning order;

[0131] A column address corresponding to the pulse window of the preset size is obtained.

[0132] As an optional implementation of the embodiment of the present disclosure, the device further includes a storage module, and the storage module is specifically configured to:

[0133] The address data and the column address are stored in a register with preset bits.

[0134] For the description of the features in the embodiment corresponding to the convolution acceleration device 500 of the pulse neural network, please refer to the relevant description of the embodiment corresponding to the convolution acceleration method of the pulse neural network, and will not be repeated here.

[0135] The convolution acceleration device of the pulse neural network provided by the embodiment of the present disclosure receives a coordinate event sent by a pulse camera and writes the pulse value of the coordinate event into the corresponding column pulse random access memory. When the column pulse random access memory generates a row completion signal, the column coordinate value corresponding to the row completion signal and the center row coordinate value are stored in the pulse buffer space, wherein the row completion signal is a signal generated when the column pulse random access memory includes three consecutive rows of pulse values; according to the column coordinate value and row coordinate value output by the pulse buffer space, the pulse value within a pulse window of a preset size and the historical cache value within the pulse window of the preset size are read in parallel, the pulse value of the pulse window of the preset size is encoded, and the encoded coded data and column address are obtained, wherein the column address is the column address corresponding to the pulse window of the preset size; the weight and storage table is queried according to the coded data to obtain the convolution result corresponding to the coded data, the weight and storage table includes a plurality of pre-calculated weight convolution sums of the coded data; the historical cache value of the corresponding position is obtained according to the column address; the convolution result is added to the historical cache value of the corresponding position to obtain the weighted cumulative sum. Because it only receives pixel coordinate events indicating changes from the pulse camera, it can automatically filter out all-zero data. This significantly reduces computational complexity compared to traditional frame-based processing that scans pixel by pixel. Decoding can be performed directly on the accelerator / card, eliminating the need for CPU decoding. Column pulse RAM and row completion signaling enable parallel reading of pulse windows and historical cache values ​​within the same clock cycle, enabling single-cycle convolution. The weight and storage tables pre-store the convolution results for various 9-bit (encoded data) inputs. At runtime, the convolution sum can be read simply by reading the weights and storage tables, saving resources and improving efficiency. Furthermore, position encoding (column address) allows direct targeting of valid pixels, skipping invalid zero values.

[0136] An embodiment of the present application also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above-mentioned convolution acceleration method embodiments of the pulse neural network.

[0137] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program, wherein the computer program is configured to execute the steps of any of the above-mentioned convolution acceleration method embodiments of the pulse neural network when running.

[0138] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0139] An embodiment of the present application also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the steps in any of the above-mentioned convolution acceleration method embodiments of the pulse neural network.

[0140] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps in any of the above-mentioned convolution acceleration method embodiments of the pulse neural network.

[0141] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0142] The above is a detailed introduction to the convolution acceleration method, device, equipment and medium of a pulse neural network provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.

Claims

1. A convolution acceleration method for a spiking neural network, characterized in that: The method comprises: Receive the coordinate event sent by the pulse camera, and write the pulse value of the coordinate event into the corresponding column pulse random access memory; When the column pulse random access memory generates a row completion signal, storing the column coordinate value and the center row coordinate value corresponding to the row completion signal in the pulse buffer space, wherein the row completion signal is a signal generated when the column pulse random access memory includes three consecutive row pulse values; According to the column coordinate value and the row coordinate value output by the pulse buffer space, the pulse value in the pulse window of the preset size and the historical buffer value in the pulse window of the preset size are read in parallel; Encoding the pulse value of the pulse window of the preset size to obtain encoded data and a column address; the column address is the column address corresponding to the pulse window of the preset size; Querying a weight and storage table according to the coded data to obtain a convolution result corresponding to the coded data; the weight and storage table includes a plurality of pre-calculated weight convolution sums of the coded data; A historical cache value of a corresponding position is obtained according to the column address, and the convolution result is added to the historical cache value of the corresponding position to obtain a weighted cumulative sum.

2. The convolution acceleration method for a spiking neural network according to claim 1, wherein: The receiving of the coordinate event sent by the pulse camera and writing the pulse value of the coordinate event into the corresponding column pulse random access memory includes: Receive multiple coordinate events sent by the pulse camera; the coordinate events include row coordinate values ​​and column coordinate values ​​of the pulse signal; The pulse values ​​corresponding to the plurality of column coordinate values ​​are respectively written in parallel into the corresponding column pulse random access memory.

3. The convolution acceleration method for a spiking neural network according to claim 1, wherein: The step of reading the pulse value in a pulse window of a preset size and the historical cache value in the pulse window of the preset size in parallel according to the column coordinate value and the row coordinate value outputted by the pulse buffer space comprises: According to the column coordinate values ​​and row coordinate values ​​output by the pulse cache space, three row read ports of three column pulse random access memories and three row read ports of the cumulative random access memory are opened in parallel to read the pulse values ​​within a pulse window of a preset size and the historical cache values ​​within the pulse window of a preset size; the cumulative random access memory is used to store the historical cache values.

4. The convolution acceleration method for a spiking neural network according to claim 1, wherein: When the column pulse random access memory generates a row completion signal, before storing the column coordinate value and the center row coordinate value corresponding to the row completion signal in the pulse buffer space, the method further includes: detecting whether the column pulse random access memory includes three consecutive rows of pulse values; If the same column of pulse random access memory includes three consecutive rows of pulse values, a row completion signal is generated.

5. The convolution acceleration method for a spiking neural network according to claim 3, wherein: After obtaining a historical cache value of a corresponding position according to the column address, and adding the convolution result to the historical cache value of the corresponding position to obtain a weighted cumulative sum, the method further includes: Determine whether the weighted cumulative sum is greater than or equal to a cumulative sum threshold; If the weighted cumulative sum is greater than or equal to the cumulative sum threshold, a pulse signal is output, and the historical cache value at the corresponding position in the cumulative random access memory is cleared; If the weighted cumulative sum is less than the cumulative sum threshold, the historical cache value of the corresponding position in the cumulative random access memory is updated according to the weighted cumulative sum.

6. The convolution acceleration method for a spiking neural network according to claim 1, wherein: The step of encoding the pulse value of the pulse window of the preset size to obtain encoded data and a column address includes: Encoding the pulse window of the preset size into address data according to a fixed scanning order; A column address corresponding to the pulse window of the preset size is obtained.

7. The convolution acceleration method for a spiking neural network according to claim 6, wherein: After encoding the pulse value of the pulse window of the preset size and obtaining the encoded coded data and column address, the method further includes: The address data and the column address are stored in a register with preset bits.

8. A convolution acceleration device for a spiking neural network, characterized in that: The device comprises: A receiving module, configured to receive a coordinate event sent by a pulse camera and write the pulse value of the coordinate event into a corresponding column pulse random access memory; a storage module configured to store the column coordinate value and the center row coordinate value corresponding to the row completion signal and the center row coordinate value in a pulse buffer space when the column pulse random access memory generates a row completion signal, wherein the row completion signal is a signal generated when the column pulse random access memory includes three consecutive row pulse values; a reading module, configured to read, in parallel, a pulse value within a pulse window of a preset size and a historical cache value within the pulse window of the preset size according to the column coordinate value and the row coordinate value output by the pulse buffer space; An encoding module, configured to encode the pulse value of the pulse window of the preset size and obtain encoded data and a column address after encoding; the column address is a column address corresponding to the pulse window of the preset size; A table lookup module, configured to query a weight and storage table according to the coded data to obtain a convolution result corresponding to the coded data; the weight and storage table includes a plurality of pre-calculated weight convolution sums of the coded data; A calculation module is used to obtain a historical cache value of a corresponding position according to the column address, and add the convolution result to the historical cache value of the corresponding position to obtain a weighted cumulative sum.

9. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the convolution acceleration method for a pulse neural network as claimed in any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the convolution acceleration method of the pulse neural network as claimed in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Pulse neural network accelerator based on time-space domain pulse convolutional coding

    CN120046660A

  • Sparse spiking neural network accelerator based on ping-pong architecture

    WO2024216857A1