Accelerator and acceleration method for spiking neural network
By designing an accelerator for pulsed neural networks, combining data acquisition, pulse acceleration and linear layer control, the problem of energy efficiency ratio and computational efficiency cannot be taken into account, efficient feature extraction and classification decisions are achieved, and the computing efficiency and energy efficiency ratio of pulsed neural networks are improved.
Patent Information
- Application Number
- CN202510517747.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-07-29
AI Technical Summary
The existing pulse neural network accelerators cannot take into account both the energy efficiency ratio and the computing efficiency. Traditional CPU architectures face performance bottlenecks. Mainstream architecture designs have problems such as high hardware resource occupancy, poor network adaptability or low computing efficiency.
An accelerator including a data acquisition module, a pulse acceleration layer and a pulse linear layer is designed. The input events are converted into event data through the data acquisition module, and the pulse acceleration layer is used to perform pulse convolution operations. Combined with the control unit to update the weight in the learning mode, generate classified decisions in the inference mode, and adopt a serial sliding and parallel thread collaboration architecture to improve computing efficiency and improve energy efficiency ratio.
While ensuring the accuracy of feature extraction, it significantly improves the computing efficiency and energy efficiency ratio, enhances the applicability and generalization capabilities of the model, and realizes efficient classification decision generation.
Smart Images

Figure CN120387488A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of hardware accelerators, and in particular to an accelerator and an acceleration method for a spiking neural network. Background Art
[0002] As specialized hardware for artificial intelligence computing, neural network accelerators aim to improve the computational performance of neural networks in both training and inference. With the rapid development of deep learning methods, the exponential growth in model computational complexity and data size has created performance bottlenecks for traditional CPU architectures, making specialized hardware optimized for deep learning an inevitable choice.
[0003] Unlike the static weighted computational model of traditional artificial neural networks (ANNs), spiking neural networks (SNNs) employ a biologically inspired information processing mechanism, encoding and transmitting information through dynamic pulse sequences between neurons. This event-driven nature gives them unique advantages in processing time-varying data, such as neural electrical signals and dynamic visual sensing. However, while the biologically inspired nature of SNNs offers advantages in time-series modeling, it also presents challenges in hardware implementation complexity and computational efficiency.
[0004] To meet the hardware requirements for large-scale SNN implementations, dedicated accelerator designs must strike a balance between computational accuracy, energy efficiency, and computational density. Currently, mainstream architectures fall into two categories: fully pipelined architectures and layer-by-layer acceleration architectures. The former achieves high throughput and low latency through deep inter-layer optimization, but suffers from limitations such as high hardware resource utilization and poor network adaptability. The latter utilizes reusable computing cores, offering improved platform portability but at the expense of computational efficiency. Summary of the Invention
[0005] In view of this, the present invention provides an accelerator and an acceleration method for a pulsed neural network to solve the technical problem that the energy efficiency and computing efficiency of the accelerator cannot be taken into account at the same time.
[0006] In a first aspect, the present invention provides an accelerator for a pulse neural network, the accelerator comprising: a data acquisition module, a pulse acceleration layer, a control unit, and a pulse linear layer, wherein the data acquisition module is used to convert the collected input events into event data; the pulse acceleration layer is used to perform pulse convolution operations based on the event data using a processing array and a neuron array, complete parallel-to-serial conversion, and determine a one-dimensional feature matrix; the control unit is used to send a learning enable signal; the pulse linear layer is used to perform online learning based on the one-dimensional feature matrix and update weights in response to the learning enable signal; wherein the control unit is further used to send an inference enable signal after the pulse linear layer completes the weight update to instruct the pulse linear layer to generate a classification decision.
[0007] In combination with the first aspect, in a possible implementation of the first aspect, the pulse acceleration layer includes: a spatio-temporal domain pulse coding unit for encoding and converting event data to determine spatio-temporal coding; and a convolutional pulse network layer for performing pulse convolution operations on the spatio-temporal coding in multi-dimensional channels by using a processing array and a neuron array to determine a one-dimensional feature matrix.
[0008] In combination with the first aspect, in a possible implementation of the first aspect, the convolutional pulse network layer includes: a processing sub-unit for establishing multi-dimensional channels, where each dimensional channel integrates a processing array and an input buffer, extracting peak data from the input buffer through a sliding window mechanism, and performing pulse convolution operations using the processing array to determine a calculation result; and a neuron sub-unit for traversing each channel and performing integration and firing processing on the calculation result to determine a one-dimensional feature matrix.
[0009] In combination with the first aspect, in a possible implementation of the first aspect, the data acquisition module includes: an event data acquisition sub-module for converting the acquired input events into pre-event data; and a pre-processing sub-module for cleaning the pre-event data to determine event data.
[0010] In combination with the first aspect, in a possible implementation of the first aspect, the pulse linear layer includes: an input layer for obtaining a one-dimensional feature matrix; a processing layer for dynamically updating synaptic weights based on the one-dimensional feature matrix by using the STDP mechanism, excitatory neurons, and inhibitory neurons; and a classification decision layer for connecting to the excitatory neurons and determining a classification decision based on the updated weights and a classifier.
[0011] In combination with the first aspect, in a possible implementation of the first aspect, the accelerator further includes: a post-processing engine, which includes: a confidence calibration sub-unit for performing semantic parsing and confidence calibration on the classification decision to determine a confidence calibration result; a probability calibration sub-unit for performing probability calibration on the classification decision based on a calibration model to determine a probability calibration result; and a comprehensive decision determination sub-unit for determining a comprehensive decision based on the confidence calibration result and the probability calibration result.
[0012] In combination with the first aspect, in a possible implementation of the first aspect, the post-processing engine further includes: a data sending sub-unit for feeding back the corresponding results of the confidence calibration result and the probability calibration result that are lower than the corresponding thresholds, so that the pulse linear layer updates the weights.
[0013] In a second aspect, the present invention provides an acceleration method for a pulsed neural network. The method includes: obtaining event data; extracting features from the event data to determine a one-dimensional feature matrix; and dynamically updating synaptic weights based on the one-dimensional feature matrix by using the STDP mechanism, excitatory neurons, and inhibitory neurons.
[0014] In combination with the first aspect, in a possible implementation manner of the first aspect, the method further includes: determining a classification decision based on the updated weights and the classifier.
[0015] In combination with the first aspect, in a possible implementation manner of the first aspect, the method further includes: performing semantic parsing and confidence calibration on the classification decision to determine a confidence calibration result; performing probability calibration on the classification decision based on a calibration model to determine a probability calibration result; determining a comprehensive decision based on the confidence calibration result and the probability calibration result.
[0016] The technical solution of the present invention has the following advantages: An accelerator and an acceleration method for a spiking neural network provided by the present invention. The accelerator, through a data acquisition module, a pulse acceleration layer, and a pulse linear layer, under the control of a control unit, updates weights when the accelerator is in a learning mode, and uses the pulse linear layer, under the control of the control unit, to generate a classification decision using the updated weights when the accelerator is in an inference mode. In this process, the data acquisition module is used to convert input events into event data, and through the pulse convolution operation of the pulse acceleration layer, using an architecture that combines serial sliding and parallel threads, serial-to-parallel conversion is completed to determine one-dimensional feature data, enabling the data to improve the operation efficiency while ensuring the accuracy of feature extraction, so as to complete the update of weights based on the pulse linear layer in the learning mode. Furthermore, under the control of the enable signal of the control unit, in the inference mode, based on the pulse linear layer, the generation of the classification decision is completed, enabling the accelerator to improve the energy efficiency ratio while enhancing the operation efficiency. Description of the Drawings
[0017] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0018] Figure 1 It is a schematic structural diagram of an application system of an accelerator for a spiking neural network provided according to an embodiment of the present invention; Figure 2 It is a schematic structural diagram of a spiking neural network of an accelerator for a spiking neural network provided according to an embodiment of the present invention; Figure 3 It is a schematic diagram of the neuron action mechanism of an accelerator for a spiking neural network provided according to an embodiment of the present invention; Figure 4Schematic diagram of a pulsed convolutional network structure for an accelerator for a spiking neural network according to an embodiment of the present invention; Figure 5 Schematic diagram of the application of the STDP mechanism in a pulsed linear layer of an accelerator for a spiking neural network according to an embodiment of the present invention; Figure 6 It is a schematic flowchart of an acceleration method for a spiking neural network according to an embodiment of the present invention. Detailed implementation manners
[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0020] It should also be noted that in the description of the present disclosure, unless otherwise clearly defined and limited, the terms "installed", "connected", and "coupled" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a direct connection or an indirect connection through an intermediate medium. For those of ordinary skill in the art, the specific meanings of the above terms in the present disclosure may be understood according to specific circumstances. When it is described that a specific device is between a first device and a second device, there may or may not be an intermediate device between the specific device and the first device or the second device.
[0021] All terms used in the present disclosure have the same meanings as understood by those of ordinary skill in the art to which the present disclosure pertains, unless otherwise specifically defined. It should also be understood that terms defined in a general dictionary, such as those, should be interpreted as having meanings consistent with their meanings in the context of the relevant art, and should not be interpreted in an idealized or overly formal sense, unless specifically defined as such here.
[0022] According to an embodiment of the present invention, an embodiment of an accelerator for a spiking neural network is provided, as Figure 1As shown, it is implemented based on the PL (Programmable Logic) / PS (Processing System) co - architecture of FPGA (Field Programmable Gate Array). It relies on the ARM processor on the PS side to communicate with the cloud, and at the same time connects to the linear layer of the accelerator for the spiking neural network to achieve dynamic parameter injection. And it transmits the events collected by the dynamic vision sensor through the AXI bus. After the event data is converted and cleaned through the time - data flow acquisition module and the pre - processing unit, the determined event data is input into the accelerator core to generate a classification decision. And under the processing of the post - processing engine, based on the classification decision, a comprehensive decision is determined.
[0023] According to an embodiment of the present invention, an embodiment of an accelerator for a spiking neural network is provided. This embodiment provides an accelerator for a spiking neural network, and the accelerator includes: a data acquisition module, a spiking acceleration layer, a control unit, and a spiking linear layer. The spiking linear layer includes a learning mode and an inference mode, where, The data acquisition module is used to convert the collected input events into event data; the spiking acceleration layer is used to perform spiking convolution operations based on the event data, using a processing array and a neuron array, complete the serial - to - parallel conversion, and determine a one - dimensional feature matrix; the control unit is used to send a learning enable signal; the spiking linear layer is used to perform online learning based on the one - dimensional feature matrix in response to the learning enable signal and update the weights; where, the control unit is further used to issue an inference enable signal after the spiking linear layer completes the weight update to instruct the spiking linear layer to generate a classification decision.
[0024] Specifically, the control unit generates a learning enable signal, an inference enable signal, and the address signal of the memory, and embeds a collaborative learning interface. It communicates with the cloud through the ARM processor on the PS side, enabling the spiking linear layer to achieve dynamic parameter injection when performing online learning using the STDP mechanism. That is, when in the learning mode, it interacts with the cloud large model through the federated learning framework, uploads the local learning results of the edge - end SNN, such as the change trend of synaptic weights, etc. to the cloud. After aggregating the data of multiple devices, global optimization parameters, such as the STDP learning rate, weight decay coefficient, etc. are generated and then sent down to the edge - end to improve the generalization ability of the model. And when in the inference mode, it enables the classifier connected to the excitatory neurons in the spiking linear layer to determine a classification decision based on the learned weights.
[0025] An accelerator and an acceleration method for a spiking neural network provided by the present invention. The accelerator, through a data acquisition module, a spiking acceleration layer, and a spiking linear layer, under the control of a control unit, updates weights when the accelerator is in the learning mode, and uses the spiking linear layer, under the control of the control unit, to generate classification decisions using the updated weights when the accelerator is in the inference mode. In this process, the data acquisition module is used to convert input events into event data, and through the spiking convolution operation of the spiking acceleration layer, using an architecture that coordinates serial sliding and parallel threads, serial-to-parallel conversion is completed to determine one-dimensional feature data, enabling the data to improve the operation efficiency while ensuring the accuracy of feature extraction, so as to complete the update of weights based on the spiking linear layer in the learning mode. Furthermore, under the control of the enable signal of the control unit, based on the spiking linear layer in the inference mode, the generation of classification decisions is completed, enabling the accelerator to improve the energy efficiency ratio while enhancing the operation efficiency.
[0026] In an alternative embodiment, the spiking acceleration layer includes: A spatio-temporal domain spiking coding unit for encoding and converting event data to determine spatio-temporal coding; a convolutional spiking network layer for performing spiking convolution operations on the spatio-temporal coding in multi-dimensional channels using a processing array and a neuron array to determine a one-dimensional feature matrix.
[0027] Specifically, the spatio-temporal domain spiking coding unit optimizes the complexity of the circuit architecture by introducing a reset control function, encodes the event data as parallel data, determines the spatio-temporal coding after completion of encoding, and completes the serial-to-parallel conversion of parallel data. Among them, the spatio-temporal coding mechanism reduces the frame rate output through the spatio-temporal domain image accumulation strategy while retaining the integrity of the original features, breaking through some limitations of traditional visible light image rate / delay coding. While improving the linearity and response speed of spiking coding, it significantly reduces the system power consumption, providing a data basis for subsequent improvement of the energy efficiency ratio.
[0028] Specifically, the structure of the convolutional spiking network layer is as Figure 2 shown. The processing array contains multiple processing elements (PEs). Among them, the processing array is as Figure 4 shown, and the mechanism of action of neurons is as Figure 3 shown. Using the processing array and the neuron array, spiking convolution operations are performed on the spatio-temporal coding in multi-dimensional channels, and through two convolution pooling operations, feature extraction is completed to determine a one-dimensional feature matrix. Among them, Figure 3 the shown SDTP online learning module corresponds to the process of interacting with the cloud large model through the federated learning framework. In this process, the convolutional spiking network layer will learn data such as STDP learning rate, weight decay coefficient, and synaptic weight change trend, and perform model parameter update, thereby improving the generalization ability of the model and enhancing the applicability of the model.
[0029] In an alternative embodiment, the convolutional spiking network layer includes: A processing subunit for establishing multi-dimensional channels, where each dimensional channel integrates a processing array and an input buffer, extracts peak data from the input buffer through a sliding window mechanism, performs spiking convolution operations using the processing array, and determines the calculation result; a neuron subunit for traversing each channel, performing integration and firing processing on the calculation result, and determining a one-dimensional feature matrix.
[0030] Specifically, the convolutional spiking network layer adopts an output channel parallelization design, establishing n independent calculation threads, that is, channels of n dimensions. Each thread integrates a spiking convolution PE, a LIF neuron array, and an output buffer. Peak data is extracted from the input buffer area through a sliding window mechanism, and the PE units of each thread perform spiking convolution operations according to weight address matching. After the sliding window traverses all the input channel data, the PE_OUT calculation result is transmitted to the neuron unit for integration and firing processing. This architecture of serial sliding and parallel thread cooperation significantly improves the operation efficiency while ensuring the feature extraction accuracy.
[0031] In an alternative embodiment, the data acquisition module includes: An event data acquisition sub-module for converting the acquired input events into pre-event data; a preprocessing sub-module for cleaning the pre-event data to determine the event data.
[0032] Specifically, as shown in Figure 1 , after the event stream input by the dynamic vision sensor is converted into streaming data, it is stored in the DDR3 memory in the AXI4 Memory Map format through the VDMA write channel; the VDMA reads data from the memory through the AXI Interconnect IP core and the AXI_HP port, and reconstructs it into a video format for output through the AXI4-Stream to Video Out IP core, and finally determines the pre-event data.
[0033] Specifically, as shown in Figure 1 , to further improve the data quality, a preprocessing unit is integrated on the PL side. This module is directly connected to the spatio-temporal domain spiking coding unit through the AXI-Stream interface, performs real-time data cleaning based on noise pulse filtering of the spatio-temporal statistical model, and uses pulse density normalization of the dynamic time window for feature enhancement to eliminate the pulse frequency distortion caused by light changes; at the same time, it fuses the semantic features extracted by the pre-trained vision large model, generates auxiliary coding vectors and performs feature fusion with the spatio-temporal domain spiking coding results, thereby enhancing the semantic parsing ability of the system for complex scenes. The entire processing chain realizes end-to-end optimization from raw event stream acquisition, noise suppression, feature enhancement to multi-modal coding.
[0034] In an alternative embodiment, the pulsed linear layer includes: An input layer for obtaining a one-dimensional feature matrix in the learning mode; a processing layer for dynamically updating synaptic weights based on the one-dimensional feature matrix by using the STDP mechanism, excitatory neurons and inhibitory neurons; and The pulsed linear layer further includes: a classification decision layer for connecting to the excitatory neurons and determining a classification decision based on the updated weights and a classifier in the inference mode.
[0035] Specifically, as Figure 5 shown, the input layer is used to obtain the one-dimensional feature matrix determined by the convolutional pulsed network layer, that is, the feature data. The input layer maps the high-dimensional feature matrix to the LIF neuron matrix. The processing layer configures an equal number of excitatory neuron and inhibitory neuron pairs. When in the learning mode, the weight dynamic adjustment is realized through the STDP mechanism and the synaptic connections of each neuron. Specifically, the excitatory neurons in the excitatory layer and the inhibitory neurons in the inhibitory layer form a fully interconnected inhibitory topology, that is, each excitatory neuron is directly connected only to the corresponding inhibitory neuron, while the inhibitory neurons send inhibitory signals to all the other excitatory neurons; when the excitatory neurons reach the firing threshold, the corresponding inhibitory neurons are triggered to release global inhibitory pulses, thereby realizing the dynamic adjustment of synaptic weights.
[0036] Specifically, as Figure 5 shown, when in the inference mode, the feature matrix conversion is performed through the fully connected layer, and a classifier connected to the excitatory neuron group is used to determine the classification decision. Among them, when in the inference mode, by analyzing historical task data, such as the classification error rate and the pulse firing frequency, the quantization bit width of the SNN, the STDP synaptic weight update rate or the leakage threshold of the LIF neurons is dynamically adjusted, so as to optimize the adaptability of the network in a specific scenario. By designing the pulsed linear layer to have a learning mode and an inference mode, through the synaptic plasticity mechanism and the neural dynamic balance strategy, efficient mode switching is realized while maintaining the network stability.
[0037] In an alternative embodiment, the accelerator further includes: a post-processing engine, and the post-processing engine includes: A confidence correction subunit for performing semantic parsing and confidence calibration on the classification decision to determine a confidence calibration result; a probability correction subunit for performing probability calibration on the classification decision based on a calibration model to determine a probability calibration result; a comprehensive decision determination subunit for determining a comprehensive decision based on the confidence calibration result and the probability calibration result.
[0038] Specifically, in complex tasks during the inference phase, the post-processing engine is independently deployed after the pulsed linear layer, receives the inference results through the AXI-Lite bus, and outputs them to an external display or an actuator. The confidence calibration subunit converts the sparse pulse sequence output by the SNN into an interpretable semantic result by parsing the semantics of the pulse sequence, calibrates the confidence, and determines the confidence calibration result. And the probability calibration subunit trains a lightweight calibration model based on historical task data, performs probability correction on the pulse frequency output by the SNN, and determines the probability calibration result. The comprehensive decision-making determination subunit integrates other sensor data, the probability calibration result, and the confidence calibration result to generate a comprehensive decision. Among them, the pulse sequence inference result of the SNN is fused with the semantic analysis result of the cloud large model, that is, the target trajectory prediction, through a gating mechanism to dynamically adjust the output weight and improve the decision-making reliability. The gating weight calculation can dynamically allocate the decision-making weights of the SNN and the cloud model according to the task complexity.
[0039] In an alternative embodiment, the post-processing engine further includes: The data sending subunit is used to feedback the corresponding results of the confidence calibration result and the probability calibration result that are lower than the corresponding thresholds, so that the pulsed linear layer updates the weights.
[0040] Specifically, feedbacking the corresponding results of the confidence calibration result and the probability calibration result that are lower than the corresponding thresholds means feedbacking misjudged samples or low-confidence results to the preprocessing sub-module or the collaborative learning interface embedded in the control unit, thereby triggering a data augmentation strategy or model parameter update to form a continuously optimized closed-loop system, that is, enabling the pulsed linear layer to continuously update the weights. Among them, misjudged samples are the results in the probability calibration result that are lower than the corresponding thresholds, and low-confidence results are the results in the probability calibration result that are lower than the corresponding thresholds. Specifically, feedbacking misjudged samples or low-confidence results to the preprocessing sub-module to clean the data during data cleaning, so as to clean up the wrong samples during data cleaning; feedbacking misjudged samples or low-confidence results to the collaborative learning interface means filtering out wrong samples during dynamic parameter injection.
[0041] According to an embodiment of the present invention, an embodiment of an acceleration method for a spiking neural network is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described herein can be executed in a different order.
[0042] This embodiment provides an acceleration method for a spiking neural network, as Figure 6 shown, the method includes the following steps: S101. Obtain event data. The specific process can refer to the relevant description in the above embodiment and will not be elaborated here.
[0043] S102. Extract features from the event data to determine a one-dimensional feature matrix. The specific process can be referred to the relevant description in the above embodiments and will not be elaborated here.
[0044] S103. Based on the one-dimensional feature matrix, use the STDP mechanism, excitatory neurons, and inhibitory neurons to dynamically update the synaptic weights. The specific process can be referred to the relevant description in the above embodiments and will not be elaborated here.
[0045] In an alternative embodiment, the method further includes: Determine a classification decision based on the updated weights and the classifier. The specific process can be referred to the relevant description in the above embodiments and will not be elaborated here.
[0046] In an alternative embodiment, the method further includes: Perform semantic parsing and confidence calibration on the classification decision to determine the confidence calibration result; perform probability calibration on the classification decision based on the calibration model to determine the probability calibration result; determine a comprehensive decision based on the confidence calibration result and the probability calibration result. The specific process can be referred to the relevant description in the above embodiments and will not be elaborated here.
[0047] Although the embodiments of the present invention are described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations fall within the scope defined by the appended claims.
[0048] Finally, it should be noted that those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium of the program can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc. The embodiments of the above computer program can achieve the same or similar effects as the corresponding foregoing method embodiments.
[0049] Those skilled in the art will also understand that the various exemplary logical blocks, modules, circuits, and algorithm steps described in connection with the disclosure herein can be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this interchangeability of hardware and software, functions have been generally described for various illustrative components, blocks, modules, circuits, and steps. Whether such functions are implemented as software or hardware depends on the particular application and the design constraints imposed on the overall system. The functions that can be implemented in various ways for each specific application by those skilled in the art, but such implementation decisions should not be construed as causing a departure from the scope of the disclosure of the embodiments of the present invention.
[0050] The above are exemplary embodiments of the disclosure of the present invention. However, it should be noted that various changes and modifications can be made without departing from the scope of the disclosure of the embodiments of the present invention as defined by the claims. The functions, steps, and / or actions of the method claims according to the disclosed embodiments herein need not be performed in any particular order. The above serial numbers of the disclosed embodiments of the present invention are only for description and do not represent the superiority or inferiority of the embodiments. In addition, although the elements of the disclosed embodiments of the present invention can be described or claimed in an individual form, they can also be understood as plural unless explicitly limited to the singular.
[0051] It should be understood that, as used herein, unless the context clearly supports exceptions, the singular form "a" is also intended to include the plural form. It should also be understood that the "and / or" used herein refers to any and all possible combinations of one or more of the associated listed items.
[0052] Those of ordinary skill in the art should understand that: the discussion of any of the above embodiments is only exemplary and is not intended to imply that the scope of the disclosure of the embodiments of the present invention (including the claims) is limited to these examples; under the concept of the embodiments of the present invention, the technical features between the above embodiments or different embodiments can also be combined, and there are many other variations in different aspects of the above embodiments of the present invention, which are not provided in detail for the sake of brevity. Therefore, any omission, modification, equivalent replacement, improvement, etc. made within the spirit and principle of the embodiments of the present invention shall be included within the protection scope of the embodiments of the present invention.
Claims
1. An accelerator for a spiking neural network, characterized in that, The accelerator includes: a data acquisition module, a pulse acceleration layer, a control unit, and a pulse linear layer. Among them, the data acquisition module is used to convert the acquired input events into event data; the pulse acceleration layer is used to perform pulse convolution operations based on the event data, using a processing array and a neuron array, complete serial-to-parallel conversion, and determine a one-dimensional feature matrix; the control unit is used to send a learning enable signal; the pulse linear layer is used to perform online learning based on the one-dimensional feature matrix in response to the learning enable signal and update the weights; wherein, the control unit is further used to issue an inference enable signal after the pulse linear layer completes weight update to instruct the pulse linear layer to generate a classification decision.
2. The accelerator according to claim 1, characterized in that, The pulse acceleration layer includes: a spatio-temporal domain pulse coding unit, which is used to perform coding conversion on event data to determine spatio-temporal coding; a convolutional pulse network layer, which is used to perform pulse convolution operations on the spatio-temporal coding in multiple-dimensional channels using a processing array and a neuron array to determine the one-dimensional feature matrix.
3. The accelerator according to claim 2, wherein The convolutional pulse network layer includes: a processing sub-unit, which is used to establish multiple-dimensional channels. Each dimensional channel integrates a processing array and an input buffer, extracts peak data from the input buffer through a sliding window mechanism, and uses the processing array to perform pulse convolution operations to determine the calculation result; a neuron sub-unit, which is used to traverse each channel, perform integration and firing processing on the calculation result, and determine a one-dimensional feature matrix.
4. The accelerator according to claim 1, characterized in that, The data acquisition module includes: an event data acquisition sub-module, which is used to convert the acquired input events into pre-event data; a preprocessing sub-module, which is used to clean the pre-event data to determine event data.
5. The accelerator according to claim 1, characterized in that The pulse linear layer includes: an input layer, which is used to obtain a one-dimensional feature matrix; a processing layer, which is used to dynamically update the synaptic weights based on the one-dimensional feature matrix, using the STDP mechanism, excitatory neurons, and inhibitory neurons; and a classification decision layer, which is used to connect to the excitatory neurons and determine a classification decision based on the updated weights and a classifier.
6. The accelerator according to claim 1, characterized in that, The accelerator further includes a post-processing engine, and the post-processing engine includes: a confidence calibration sub-unit, which is used to perform semantic parsing and confidence calibration on the classification decision to determine a confidence calibration result; a probability calibration sub-unit, which is used to perform probability calibration on the classification decision based on a calibration model to determine a probability calibration result; a comprehensive decision determination sub-unit, which is used to determine a comprehensive decision based on the confidence calibration result and the probability calibration result.
7. The accelerator according to claim 6, characterized in that The post-processing engine further includes: a data sending sub-unit, which is used to feedback the corresponding results of the confidence calibration result and the probability calibration result that are lower than the corresponding thresholds, so that the pulse linear layer performs weight update.
8. An acceleration method for a spiking neural network, characterized in that, Applied to the accelerator for a pulse neural network according to any one of claims 1 to 7, the acceleration method includes: obtaining event data; performing feature extraction on the event data to determine a one-dimensional feature matrix; dynamically updating the synaptic weights based on the one-dimensional feature matrix, using the STDP mechanism, excitatory neurons, and inhibitory neurons.
9. The method according to claim 8, wherein The method further includes: Based on the updated weights and classifier, determine the classification decision.
10. The method according to claim 8, characterized in that, The method further includes: Perform semantic parsing and confidence calibration on the classification decision to determine the confidence calibration result; Perform probability calibration on the classification decision based on the calibration model to determine the probability calibration result; Based on the confidence calibration result and the probability calibration result, determine the comprehensive decision.