Self-timed feedforward synthesizable and technology-scalable mixed-signal neurons

The spike capture and local trigger generation module in spiking neural networks addresses scalability and power efficiency issues by using pulse latching and local trigger generation, enhancing the performance of spiking neural networks through asynchronous operation.

JP2026504101APending Publication Date: 2026-02-03INNATERA NANOSYSTEMS BV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025541657
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-01-17
Filing Date
2024-01-17
Publication Date
2026-02-03

AI Technical Summary

Technical Problem

Existing spiking neural networks face challenges in scalability and power efficiency due to reliance on global clock signals, which introduce latency and high power consumption, while analog implementations suffer from signal-to-noise ratio issues.

Method used

A spike capture and local trigger generation module for spiking neural networks that uses pulse latching and local trigger generation to synchronize operations without a global clock, combined with weight encoding and accumulation modules for asynchronous operation.

Benefits of technology

This approach reduces power consumption and latency, enabling scalable and efficient operation of spiking neural networks by eliminating the need for global clock signals and optimizing timing through local event-based triggering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026504101000001_ABST
    Figure 2026504101000001_ABST
Patent Text Reader

Abstract

The present invention relates to a spike capture and local trigger generation module for neurons of a spiking neural network. The module comprises spike input ports each configured to receive a spike input signal. The module further comprises a pulse latch module configured to capture spikes in the spike input signal by pulse latching. The pulse latch module outputs, for each spike input port, a latched signal indicating whether a spike has been captured. The module further comprises a local trigger generator module configured to generate a trigger signal used to synchronize and / or control the timing of operations in the neuron, the local trigger generator module configured to use the latched signal to generate the trigger signal such that the trigger signal comprises pulses numerically equal to the number of spikes captured. Finally, the module may comprise a trigger signal output port for outputting the generated trigger signal.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to automatic signal recognition techniques, and more particularly to a system and method for a hardware-resilient deep learning inference accelerator using spiking neural networks. [Background technology]

[0002] Spiking neural networks (SNNs) are signal processing systems whose design is inspired by biological neural networks: information is encoded in patterns of spike signals distributed across a complex network of neurons and synapses.

[0003]

[0003] Synapses store information and use it to perform basic operations on incoming signals. Neurons generate new spikes based on the signals coming from synapses. The connections between synapses and neurons determine the shape of the SNN, which, together with the information stored in the synapses, determines the signal processing function of the SNN.

[0004]

[0004] There are many digital implementations in the literature of typical functions performed by neurons in spiking neural networks. See, for example, P Kampf, P Koch, K Roy, M Sullivan, Z Delalic, and S DasGupta, "Optimization of a Digital Neuron Design," in 1990 Eastern Multiconference. Record of Proceedings. The 23rd Annual Simulation Symposium, 1990, pp. 73-80. Another example is D. Lee, G. Lee, D. Kwon, S. Lee, Y. Kim, and J. Kim, "Flexon: A Flexible Digital Neuron for Efficient Spiking Neural Network Simulations," in 2018 ACM / IEEE 45th Annual International Symposium on Computer Architecture (ISCA), 2018, pp. 275-288. These functions can include operations such as addition / subtraction, multiplication, division, and / or complex state machines.

[0005]

[0005] On the other side of the spectrum, there are many analog implementations of neurons in SNNs. See, for example, A Joubert, B Belhadj, O Temam, and R Heliot, "Hardware spiking neurons design: Analog or digital?", 2012 International Joint Conference on Neural Networks (IJCNN), 2012, pp. 1-5. Another example is given in JM Zurada, "Analog implementation of neural networks," IEEE Circuits and Devices Magazine, vol. 8, no. 5, pp. 36-41, 1992, doi:10.1109 / 101.158511. Also noteworthy is A. Rubino, M. Payvand, and G. Indiveri, "Ultra-Low Power Silicon Neuron Circuit for Extreme-Edge Neuromorphic Intelligence," in the 2019 26th IEEE International Conference on Electronics, Circuits, and Systems (ICECS), 2019, pp. 458-461. These have the advantage of encoding more value states in a more power-efficient manner, sometimes at the expense of signal-to-noise ratio (SNR). Phase-encoded neurons also exist, which use phase as a continuous-time variable [6] for computation. See, for example, A. Madhavan, T. Sherwood, and D. Strukov, "Race logic," ACM SIGARCH Computer Architecture News, vol. 42, pp. 517-528, December 2014, doi:10.1145 / 2678373.2665747. They also suffer from the same trade-off as analog neurons, i.e., phase-encoded neurons can suffer from a worse SNR.

[0006]

[0006] The benefit of digital neurons comes from the fact that whatever function they implement has a defined state that is stored every clock cycle. The state is a well-defined value that does not get corrupted over time. The problem with digital neurons comes from the fact that every time an operation is performed, a well-defined concept of time is created using clocking elements such as phase-locked loop-based voltage-controlled oscillators. The distribution of this clock network ends up consuming a major portion of the power. See, for example, S. Ali Butt, S. Schmermbeck, J. Rosenthal, A. Pratsch, and E. Schmidt, System Level Clock Tree Synthesis for Power Optimization. [IEEE Computer Society], 2007, p. 1682.

[0007]

[0007] An extension of digital neurons that alleviates the clock distribution problem is based on arbitrating incoming signals using multiple phases of oscillators. See, for example, J. Stuijt, M. Sifalakis, A. Yousefzadeh, and F. Corradi, "μBrain: An Event-Driven and Fully Synthesizable Architecture for Spiking Neural Networks," Frontiers in Neuroscience, vol. 15, December 2021, doi:10.3389 / fnins.2021.664208. This extension solves the problems of avoiding phase locking and global clock distribution. The system is not scalable because scaling the system would require faster multiphase oscillators with more clock distribution. It has better power characteristics but suffers from the same tradeoffs as conventional digital neuron systems.

[0008] Another style of digital neuron implementation addresses this tradeoff using a handshaking mechanism, approaching the problem from an asynchronous perspective. See, for example, A. Yousefzadeh et al., "Asynchronous Spiking Neurons, the Natural Key to Exploit Temporal Sparsity," IEEE Journal on Emerging and Selected Topics in Circuits and Systems, vol. PP, December 2019, doi:10.1109 / JETCAS.2019.2951121. The problem with this approach is that scalability introduces latency, and power from the handshaking dominates the system. Summary of the Invention

[0009]

[0009] In order to solve the above-mentioned problems, the subject matter of the present claims is proposed.

[0010]

[0010] In a first aspect of the present invention, a spike capture and local trigger generation module for a neuron of a spiking neural network is disclosed. The module may include a plurality of spike input ports. Each of the plurality of spike input ports may be configured to receive a spike input signal from the spiking neural network. Further, a pulse latch module may be included within the module. The pulse latch module may be configured to capture spikes in the spike input signal received by the spike input port through pulse latching. Pulse latching may be described as digital circuitry that captures and holds a state in response to a trigger pulse, storing the state until reset by a subsequent pulse. The state may be transferred to other portions of the module. The pulse latch module may output a latched signal for each of the plurality of spike input ports indicating whether a spike has been captured. Further, a local trigger generator module may be included within the module. The local trigger generator module may be configured to generate a trigger signal used by the neuron to synchronize and / or control the timing of operations in the neuron. The local trigger generator module may be configured to use the latched signals to generate a trigger signal, such that the trigger signal comprises one or more pulses numerically equal to the number of spikes captured by the pulse latch module. The trigger signal thus obtained may be used to trigger operations performed by other portions of the neuron, thus providing a local event-based triggering mechanism. Thus, the present invention does not require a global clock signal to operate those portions of the neuron. To use the trigger signal, the module may comprise a trigger signal output port configured to output the generated trigger signal.

[0011] In one embodiment of the first aspect, the trigger signal may be generated such that operations in neurons may be triggered by the trigger signal rather than by a global clock signal of the spiking neural network.

[0012] In one embodiment of the first aspect, the trigger signal may be a delay-based asynchronous counter.

[0013] In one embodiment of the first aspect, the one or more pulses provided in the trigger signal may have an interval between them that may be determined by a minimum time taken to perform an operation in the neuron. Preferably, the operation may be an accumulation of one or more weights.

[0014] In one embodiment of the first aspect, the module may further include a plurality of spike output ports each configured to output a spike output signal. Each of the plurality of spike output ports may correspond to one of the plurality of spike input ports. A particular spike output port may be configured to output a pulse when the corresponding spike input port receives a spike.

[0015] In one embodiment of the first aspect, the pulses output by the spike output port may be spaced in time substantially similar to the pulses provided in the clock signal.

[0016] In one embodiment of the first aspect, the pulse latch module may include a plurality of latch elements, each of which may be configured to capture one or more spikes in a spike input signal received by a particular spike input port by pulse latching.

[0017] In one embodiment of the first aspect, the latch element may be a CMOS D flip-flop. The CMOS D flip-flop may be triggered by a spike input signal. The latched signal output by the CMOS D flip-flop may be a bit value indicating whether the spike is latched or not.

[0018] In one embodiment of the first aspect, the local trigger generator module may include a pulse generator that generates a pulse. The local trigger generator module may include delay logic for each spike input port. The generated pulse may then be sent through each of the delay logics.

[0019]

[0019] In one embodiment of the first aspect, each of the delay logics may include a delay element that delays the pulse signal. The generated pulse may be sent through the delay element of a particular delay logic if the corresponding latched signal indicates that a spike was captured by the corresponding spike input port. The generated pulse may not be sent through the delay element of a particular delay logic if the corresponding latched signal indicates that a spike was not captured by the corresponding spike input port.

[0020]

[0020] In one embodiment of the first aspect, the generated pulse may be output by a specific delay logic to generate a trigger signal when the corresponding latched signal indicates that a spike has been captured by the corresponding spike input port. The generated pulse may not be output by a specific delay logic to generate a trigger signal when the corresponding latched signal indicates that a spike has not been captured by the corresponding spike input port.

[0021] In one embodiment of the first aspect, the local trigger generator module may be configured to generate a plurality of trigger signals used by the neuron to synchronize and / or control the timing of operations in respective portions of the neuron. The local trigger generator module may be configured to use the plurality of latched signals to generate the plurality of trigger signals, each trigger signal comprising one or more pulses numerically equal to the number of spikes captured by the pulse latch module for a particular set of spike input ports. The module may further comprise a plurality of trigger signal output ports configured to output the generated trigger signals.

[0022]

[0022] In one embodiment of the first aspect, the module may further include a multi-level trigger generator module having a multi-level trigger output port configured to generate a multi-level trigger signal. The multi-level trigger signal may indicate the total time it will take for the neuron to perform operations triggered by the trigger signal. Preferably, when there are multiple trigger signals generated and output, the multi-level trigger signal may correspond to the maximum time it will take for the neuron to perform all sets of operations triggered by different trigger signals.

[0023]

[0023] In a second aspect of the present invention, a weight encoding and accumulation module for accumulating spike input signals in neurons of a spiking neural network is disclosed. The module may include a plurality of spike input ports. Each of the plurality of spike input ports may be configured to receive a spike input signal. Further, the module may include a trigger signal input port configured to receive a trigger signal. Pulses included in the trigger signal may indicate the number of spikes received by the plurality of spike input ports. Further, the module may include a spike-to-weight encoder module. The spike-to-weight encoder module may be configured to receive the spike input signals and output corresponding weights associated with particular spike input signals. This may be done by encoding the received spike input signals as weights. Further, an accumulator module may be included in the module. The accumulator module may be configured to receive weights from a memory module and a trigger signal and accumulate weights for received spikes. The timing of one or more of the operations involved in the accumulation may be governed by the trigger signal.

[0024] In one embodiment of the second aspect, each of the spike input signals may be an analog signal, and the weights may be digital values.

[0025]

[0025] In one embodiment of the second aspect, the accumulator module may include an accumulation element and a loop circuit that takes the output of the accumulation element as an input and outputs a value back to the accumulation element each time the loop circuit receives a pulse from the trigger signal. When the spike-to-weight encoder module outputs a first weight to the accumulation element, the first weight may be added by the accumulation element to a second weight output by the loop circuit.

[0026] In one embodiment of the second aspect, the loop circuit comprises a CMOS D flip-flop that can be triggered by the trigger signal.

[0027]

[0027] In one embodiment of the second aspect, the spike-to-weight encoder module may comprise a local memory, preferably an SRAM or a non-volatile memory. The spike-to-weight encoder may encode a pulse received at a particular spike input port as a particular memory address used to request a weight stored at that memory address. A weight may be associated with a particular spike input port.

[0028]

[0028] In one embodiment of the second aspect, the spike-to-weight encoder module may include a distributed memory including distributed weight storage elements. Weights associated with spike input ports may be stored in one of the distributed weight storage elements, respectively. When a spike is present in the spike input signal, the corresponding weight may be output to the accumulator module by the distributed weight storage element in which the weight is stored.

[0029] In one embodiment of the second aspect, the spike-to-weight encoder module may be configured to change the order in which the weights are output based on a priority rule.

[0030]

[0030] In a third aspect of the present invention, a spiking neuron of a spiking neural network is disclosed. The neuron may comprise a spike capture and local trigger generation module according to the first aspect. The neuron may additionally or alternatively comprise a weight encoding and accumulation module according to the second aspect. A trigger signal output port of the spike capture and local trigger generation module may output a trigger signal to a trigger signal input port of the weight encoding and accumulation module. This may be done in such a way that the timing of one or more of the operations involved in the accumulation is governed by the trigger signal.

[0031] In one embodiment of the third aspect, the neuron may further include a comparator module that compares the total accumulated weight with a predetermined threshold. The total accumulated weight may be the sum of all weights of the spike input ports from which the spike of the corresponding spike input signal is captured. The spiking neuron may be configured to output a spike to the spiking neural network when the total accumulated weight exceeds the predetermined threshold.

[0032]

[0032] In one embodiment of the third aspect, a spiking neuron may include a plurality of weight encoding and accumulation modules according to the second aspect. The spike capture and local trigger generation module may be according to the first aspect. Each of the plurality of trigger signal output ports may output a respective trigger signal to a trigger signal input port of a corresponding weight encoding and accumulation module, and thus, timing of one or more of the operations involved in the accumulation performed in the corresponding weight encoding and accumulation module may be governed by the respective trigger signal.

[0033] In one embodiment of the third aspect, the neuron may further comprise a digital weight accumulator module that accumulates outputs from at least some of the plurality of weight encoding and accumulation modules and outputs an accumulated value.

[0034] In one embodiment of the third aspect, the spiking neuron may include a plurality of digital weight accumulator modules, at least one of which may accumulate output cumulative values ​​of other digital weight accumulator modules.

[0035]

[0035] In one embodiment of the third aspect, delay-based pipelining and / or handshake-based pipelining may be used to synchronize accumulation operations among multiple weight encoding and accumulation modules and / or one or more digital weight accumulator modules.

[0036] In one embodiment of the third aspect, the spike capture and local trigger generation module may be according to the first aspect, and the multi-level trigger signal may be used to synchronize accumulation operations among multiple weight encoding and accumulation modules and / or one or more digital weight accumulator modules.

[0037] In a fourth aspect, a method for spike capture and local trigger generation in neurons of a spiking neural network is disclosed. The method may comprise receiving a plurality of spike input signals from the spiking neural network, capturing spikes in the spike input signals by pulse latching, generating a latch signal for each of the spike input signals, wherein the latched signals may indicate whether a spike was captured, generating a trigger signal using the plurality of latched signals such that the trigger signal may comprise one or more pulses numerically equal to the number of captured spikes, and outputting the generated trigger signal, wherein the trigger signal may be used by the neuron to synchronize and / or control the timing of operations in the neuron.

[0038]

[0038] In an embodiment of the fourth aspect, the method may further comprise outputting a spike output signal. The trigger signal may be generated such that operations in the neuron may be triggered by the trigger signal rather than by a global clock signal of the spiking neural network.

[0039] In a fifth aspect, a method for weight encoding and accumulation in neurons of a spiking neural network is disclosed. The method may comprise receiving a plurality of spike input signals, receiving a trigger signal comprising one or more pulses, wherein the one or more pulses indicate a number of spikes received in the plurality of spike input signals, obtaining or generating a weight associated with one of the plurality of spike input signals, accumulating weights of the spike input signals comprising the spikes, and outputting the accumulated weight, wherein timing of one or more of the operations involved in the accumulation may be governed by the trigger signal.

[0040]

[0040] In an embodiment of the fifth aspect, the weight may be obtained or generated by encoding a pulse received in one of a plurality of spike input signals as a specific memory address that can be used to request a weight stored in local memory at that memory address.

[0041]

[0041] In an embodiment of the fifth aspect, when a spike is present in the spike input signal, a corresponding weight may be retrieved from a distributed memory comprising a plurality of weight storage elements. The weights associated with the spike input signal may be stored in one of the distributed weight storage elements, respectively. When a spike is present in the spike input signal, the corresponding weight may be output by a distributed weight storage element where the weight is stored and used during weight accumulation.

[0042] In a sixth aspect, a method for spike capture and accumulation in a neuron of a spiking neural network is disclosed. The method may comprise receiving a plurality of spiking input signals from the spiking neural network, capturing spikes in the spiking input signals by pulse latching, outputting a latched signal for each of the spiking input signals, wherein the latched signal may indicate whether a spike has been captured, generating a trigger signal using the plurality of latched signals such that the trigger signal may comprise one or more pulses numerically equal to the number of captured spikes, wherein the trigger signal may be used by the neuron to synchronize and / or control the timing of operations in the neuron, obtaining or generating weights associated with one or more of the plurality of spiking input signals, accumulating the weights of the spiking input signals with the spikes, and outputting the accumulation weights, wherein the timing of one or more of the operations involved in the accumulation may be governed by the trigger signal.

[0043]

[0043] In one embodiment of the sixth aspect, the method may further comprise comparing the total accumulated weight with a predetermined threshold. The total accumulated weight may be a sum of weights of all of the spike input signals from which the spike was captured. Furthermore, the method may further comprise outputting a spike to the spiking neural network when the total accumulated weight exceeds the predetermined threshold.

[0044] In one embodiment of the sixth aspect, the steps of obtaining or generating weights and accumulating weights may be parallelized by dividing the plurality of spike input signals into one or more subsets of the plurality of spike input signals, and further, the steps of obtaining or generating weights and accumulating weights for these subsets may be performed simultaneously to obtain accumulated subset weights for each subset.

[0045] In one embodiment of the sixth aspect, the accumulated subset weights may be accumulated digitally, preferably in a parallelized manner.

[0046]

[0046] In one embodiment of the sixth aspect, delay-based pipelining and / or handshake-based pipelining are used to synchronize accumulation operations between different subsets and / or one or more digital accumulations.

[0047]

[0047] Embodiments will now be described, by way of example only, with reference to the accompanying drawings in which corresponding reference symbols indicate corresponding parts, and in which: [Brief explanation of the drawings]

[0048] [Figure 1]

[0048] Schematic of reproducible design blocks connected in an MxN structure. [Figure 2]

[0049] 1 is a schematic diagram of a design method according to the present invention; [Figure 3]

[0050] Schematic of the high-level architecture of a neuron with m spiking inputs. [Figure 4]

[0051] Schematic diagram of the capture of m spike inputs and associated timing generation. [Figure 5]

[0052] Schematic of the summation of spike weights from local memory in a unit cell. [Figure 6]

[0053] Schematic of the summation of spike weights from distributed memory in a unit cell. [Figure 7]

[0054] Schematic of the design of a multilevel self-timed neuron with regenerative unit cells. [Figure 8]

[0055] Schematic of delay-based pipelining. [Figure 9]

[0056] Schematic diagram of handshake control logic based pipelining. DETAILED DESCRIPTION OF THE INVENTION

[0049]

[0057] In the following, some embodiments will be described in more detail. However, it should be appreciated that these embodiments should not be construed as limiting the scope of protection of the present disclosure.

[0050]

[0058] Self-timed circuits have been explored in the literature for various systems and high-performance designs. See, for example, S. Fairbanks, "High precision timing using self-timed circuits," 2009. Implementations can be made using various circuit design techniques, such as self-resetting logic (see Bloker et al., "US Pat. No. 5,864,251"), pulsed self-timed logic (see M. Miller, C. Segal, D. McCarthy, A. Dalakoti, P. Mukim, and F. Brewer, "Impolite High Speed ​​Interfaces with Asynchronous Pulse Logic," December 2018, pp. 99-104), etc. This style of logic design has no clock associated with it. The idea is to combine multiple repeatable patterns of logic in a systematic way so that each logic block can complete within a given delay.

[0051]

[0059] To achieve these designs, a design methodology must be established to achieve larger designs.

[0052]

[0060] Source-coupled logic has been shown to be an efficient way to reduce power in digital designs and operate designs at extremely low supply voltages near transistor device thresholds. See, for example, S. Roy and K. Nipun, "Understanding sub-threshold source coupled logic for ultra-low power applications." The integration of source-coupled logic in neuron logic and delay design can further lead to area and power optimization in the design. Self-timing logic has also been shown to optimize timing performance in CMOS designs. See, for example, A. Dalakoti, M. Miller, and F. Brewer, "Pulse Ring Oscillator Tuning via Pulse Dynamics," 2017 IEEE International Conference on Computer Design (ICCD), 2017, pp. 469-472. The inventors of the present invention have recognized that this can help optimize designs that create neurons at different time scales. Delay-based race calculations can also be made repeatable, which can help create power-efficient logic. See, for example, A. Madhavan, T. Sherwood, and D. Strukov, "Race logic," ACM SIGARCH Computer Architecture News, vol. 42, pp. 517-528, December 2014, doi:10.1145 / 2678373.2665747.

[0053]

[0061] The inventors of the present invention have recognized that techniques such as source-coupled design, self-timed logic, and delay-based race logic can all be incorporated into the design methodology to create complex neuron structures. The benefits of the design methodology stem from the fact that these structures can be synthesized based on use case requirements under the hypothesis of repeatable blocks that are pre-verified and integrated into the design methodology.

[0054]

[0062] All traditional digital techniques such as time division multiplexing, memory optimization, etc. can be applied to the design, and further, the design can be optimized for area, power, or timing tradeoffs. The design can also be augmented with complex connectivity to create complex SNN structures based on the needs of each use case.

[0055]

[0063] Figure 1 shows a schematic diagram of reproducible design blocks 101 connected in an MxN structure 100, and Figure 2 shows a schematic diagram of a design method 200 in accordance with the present invention. In particular, Figure 1 shows design blocks 101 that can be interpreted through the design method 200 shown in Figure 2 to automatically create the final connectivity of neurons.

[0056]

[0064] The reproducible design blocks 101 connected in the M×N structure 100 shown in FIG. 1 refer to a modular architecture, where M represents the number of rows of interconnected design blocks 101 and N represents the number of columns, with the M×N structure 100 collectively forming (at least a portion of) an SNN. The design blocks 101 can comprise a single neuron, multiple connected neurons, and / or more complex circuits. This M×N structure 100 can facilitate parallel processing, and the use of design blocks can follow modular design principles. The connections between the design blocks 101 can represent synaptic connections in the network. The design blocks can be connected using, for example, wire connections, crossbar buses, and logic such as NAND trees, tri-state buses, and / or routers. Note that the design blocks 101 used in the M×N structure 100 can be similar or different from one another and can vary by row and / or column.

[0057]

[0065] The M×N structure 100 promotes modular design and enables easy replication and scalability. Each block 101 can be a self-contained module with a specific function, and the entire network can be constructed by replicating and connecting these modules. For example, different design blocks 101 can be specialized for several specific tasks or functions. The overall SNN can then be adapted to various applications by adjusting the connectivity and properties of the individual design blocks. The same design block 101 can be replicated across different parts of the network, promoting consistency and simplifying the design process. This is beneficial for both hardware and software implementations of SNNs. Because design blocks (e.g., neurons) 101 in the same row or column can operate simultaneously, the grid structure encourages parallelism within the network and can enhance the overall computational efficiency of the network.

[0058]

[0066] Depending on the specific requirements of a task or application, the connectivity pattern between design blocks 101 in the M×N structure 100 can vary. Common connectivity patterns include nearest neighbor connections, exhaustive connections, or more complex connectivity schemes. The modular and reproducible nature of the design allows for easy scalability. As SNN requirements grow, additional rows or columns of design blocks can be added to increase the capacity of the SNN without fundamentally changing the architecture.

[0059]

[0067] The design block may comprise unit cells 102 with delay circuits 103 and / or logic circuits 104. These circuits 103, 104 are further described below and may each have input ports, output ports, and / or feedback loops. The unit cells 102 may form, for example, neurons.

[0060]

[0068] We now describe the design method 200 shown in Fig. 2. Generally, the design method 200 according to the invention uses inputs 201 and performs design steps 202 on these inputs 201, thus obtaining an output 203 of the design method 200. The output 203 may disclose, for example, the configured network architecture, in particular, for example, the MxN structure shown in Fig. 1.

[0061]

[0069] Inputs 201 to the design methodology may be, for example, primitives 201A, logic blocks RTL 201B, delay blocks 201C, connectivity requirements 201D, and / or timing constraints 201E.

[0062]

[0070] Primitive 201A is a design block that can be used to design complex logic circuits for register transfer level (RTL) blocks. RTL is an abstraction level of digital circuit design that represents a circuit in terms of the flow of data between registers and the logical operations performed between those registers. An RTL block may represent a portion of the hardware design of a neural network that is described at the register transfer level. The RTL block may capture the data flow and the operations performed on the data in hardware. Various types of primitive designs are possible depending on the style of logic used based on power-performance-area requirements.

[0063]

[0071] Some exemplary primitive designs 201A that can be used in the design method 200 as inputs 201 include (a) SR CMOS (see, e.g., Bloker et al., U.S. Pat. No. 5,864,251), (b) pulse logic (see, e.g., D. Mc Carthy, M. Miller, and F. Brewer, “Automated Timing Constraint Generation for PulseGate Circuits,” IEEE International Conference on Program Comprehension, 2022, vol. 2022-March, pp. 36-47), (c) standard CMOS (see, e.g., I. Sourikopoulos et al., “A 4-fJ / spike artificial neuron in 65 nm CMOS technology,” Frontiers in Neuroscience, vol. 11, no. MAR, Mar. 2017, doi:10.3389 / fnins.2017.00123), and / or (d) source-coupled CMOS logic (see, e.g., S. Roy and (See K. Nipun, "Understanding sub-threshold source coupled logic for ultra-low power applications").

[0064]

[0072] Another input 201 may be a logic block RTL 201B. The neuron block may consist of various mathematical operations. These operations may be captured as standard RTL descriptions and passed to the synthesis flow. In an RTL design, registers store data between different stages of a computation. In the case of an SNN, these registers may represent the state of a neuron, synaptic weights, or other related information at different time steps. The RTL description may detail how data is transferred to and from the registers and how various computational operations (e.g., spike generation, synaptic updates) are performed. The RTL description may also include timing information, specifying when various operations occur. In an SNN, this may be important for capturing the temporal dynamics of spikes and synapses. The RTL block may specify operations specific to spiking neurons in an SNN, such as spike generation, spike propagation, synaptic weight updates, and any other computations related to the neuron's spiking behavior. Depending on the design, the RTL description may also capture parallel processing within the SNN, especially if the hardware implementation involves parallel processing of multiple neurons or synapses.

[0065]

[0073] Another input 201 may be a delay block 201C. The delay block 201C may be implemented in various forms. The delay may be derived from standard CMOS gates or through specialized analog designs to create a variable delay. The delay block is used to model the transmission delays associated with synaptic connections between neurons. These delays can help capture the temporal dynamics of spike events. The delay block 201C may be implemented in various ways, and the choice of implementation depends on the specific requirements of the neural network and hardware platform. Examples of digital delay blocks include shift registers and digital counters. Each stage of the shift register represents a unit of time, and the signal is shifted through the register to introduce the delay. A digital counter may be employed to count the number of time steps before allowing the signal to pass. This allows for a programmable delay based on the count value. Examples of analog delay blocks include RC circuits and transmission lines. An analog delay block may be used to introduce a variable delay. An RC circuit comprising a resistor (R) and a capacitor (C) can be used to create a variable delay by adjusting the values ​​of the resistor and capacitor. Charging and discharging the capacitor introduces a time constant that determines the delay. Analog delay lines, such as transmission lines, can also be used to introduce delay. The signal travels along the transmission line, and the length of the line determines the delay.

[0066]

[0074] Another input may be connectivity requirements 201D. In real use cases, the connectivity between different logic blocks may vary, which needs to be captured in a separate file, e.g., a regular expression-based connectivity description file, for correct implementation of the design.

[0067]

[0075] Another input may be timing constraints 201E. To make a design synthesizable, all designs must be constrained. For example, all logic block paths must be constrained. Timing constraints 201E may be implemented by limiting the time it takes for logic to perform a certain operation. Examples of timing constraints 201E may be, for example, (a) minimum and maximum delays in a path, (b) minimum and maximum capacitances in a path (e.g., charging and discharging a capacitor introduces a time constant that determines the delay), and / or (c) floorplanning of a structure that satisfies the constraints. Regarding the latter point (c), different design choices such as delay, complexity, and logic size, in particular, result in blocks of different sizes. It may not be feasible to place all blocks next to each other to obtain maximum area utilization in a design. Intelligent floorplanning is needed to find combinations of blocks that can be placed next to each other to obtain the best area utilization, which can be achieved by the described method.

[0068]

[0076] The different method steps 202 of the design method 200 are then performed based on the input 201 .

[0069]

[0077] As a first step, the primitives used in complex neuron logic need to be characterized for delay. This is preferably done automatically during primitive characterization 202A to create the necessary timing files needed by synthesis and place-and-route tools. A scaled voltage library needs to be characterized for near-threshold calculations. A scaled voltage library is generally a set of libraries in (digital) integrated circuit designs that contain components with different threshold voltages (Vt) or supply voltages (Vdd). These libraries are designed to support the creation of circuits that operate at multiple voltage scales. Near-threshold calculations are the operation of digital circuits at or near their threshold voltage levels. The threshold voltage is the voltage at which a transistor switches from an off state to an on state. Operating near the threshold voltage is often associated with low-power design strategies in digital integrated circuits. That is, lowering the supply voltage and operating near the threshold voltage can result in significant power savings compared to traditional higher-voltage operation. Therefore, near-threshold calculations are particularly relevant in mixed-signal and analog / mixed-signal designs, where power efficiency is important.

[0070]

[0078] Next, as a second step, the logic for the digital neuron needs to be characterized during logic characterization 202B. The goal is to create an optimized neuron logic block and then use it in a time- and space-multiplexed manner. To use the logic block multiple times, a macro-IP of the logic block is fabricated and implemented according to the connectivity of the use case. IP blocks, often called IP cores or simply IP, are pre-designed, pre-verified, and reusable functional units that can be integrated into larger systems or designs. These IP blocks encapsulate specific functions or logic, and they are designed to be easily integrated into various projects to save time and effort. The design process is performed for block-level and top-level automation.

[0071]

[0079] As a third step, because there are various types of delay design methods, each needs to be characterized during delay characterization 202C. The correct delay is chosen based on the logic style and logic delay. Different delays have different tradeoffs associated with them. Some examples are (a) CMOS delay, which is a standard delay element in processing development kits, or (b) current-starved delay, which is an area-optimized delay controlled using a widely adjustable digital-to-analog converter. CMOS delay refers to the time it takes a signal to propagate through a complementary metal-oxide semiconductor (CMOS) logic gate from the input of the CMOS gate to its output. The size and characteristics of n-type transistors (NMOS) and p-type transistors (PMOS) affect the delay. For example, larger transistors generally have higher drive strength, but may also have higher capacitance. The capacitance at the output of a gate, also called load capacitance, which includes the parasitic capacitance of connected wires and the input capacitance of subsequent gates, can also affect the delay. Current-starved delay is a design technique in which delay is primarily controlled by adjusting the current flowing through the circuit and intentionally limiting this current to achieve the desired delay. For example, it can be applied to ring oscillators and delay-locked loops present in the network.

[0072]

[0080] As a fourth step, time and space multiplexing 202D can be performed. That is, the logic and connectivity need to be converted to a hardware implementation using time and space multiplexing. Time multiplexing increases latency and power usage. Spatial distribution increases the area in which the network is implemented. The correct design point needs to be chosen (specific to the design technology), and an automatic compiler can generate the blocks placed in the network, which will be routed. Quantification can be based on the power, performance, and area (PPA) number we are trying to reach. Generally, we try to find the optimal solution using optimization algorithms such as convex optimization or simulated annealing.

[0073]

[0081] As a fifth step, power, performance, and area (PPA) are optimized during PPA optimization 202E. These three parameters can be optimized using, for example, (a) multiplexing dependent area, (b) technology node usage, (c) mixed-signal design controls such as current-starved delay elements, (d) design methods such as source-coupled logic or self-resetting logic, (e) pipelining, and / or (f) precisely chosen spatial and temporal signaling between multiplexing designs.

[0074]

[0082] After performing these method steps, we obtain output 203 of the design method 200, which may include: (a) a final layout 203A, which is the final GDS with the complete digital neuron structure connected in the required network structure; (b) a gate-level netlist 203B, which is a gate-level netlist with all the primitives characterized in the flow; and / or (c) a PPA report 203C, which is a report of the expected performance, power, and area for the created use case-driven design.

[0075]

[0083] The final layout 203A can be a physical representation of the SNN design on a semiconductor substrate. It includes the placement of transistors, interconnects, metal layers, and other physical components that form synapses and neurons. It provides a map of where each component of the SNN will be located on the chip. "GDS" stands for "graphical data system" or "graphical design system" and refers to a standard file format used to describe the geometric layout of the various components and layers of an IC. GDS files can be used by semiconductor foundries to manufacture the designed IC.

[0076]

[0084] The gate-level netlist 203B is a textual representation of the SNN design at the gate level. It describes the logical connectivity and functionality of the design in terms of basic logic gates (AND, OR, NOT, etc.). Each line in the netlist generally represents a gate or flip-flop, and the connections between these elements are explicitly specified. The netlist serves as an intermediate representation that can be used for simulation, verification, and synthesis.

[0077]

[0085] The PPA report 203C provides detailed information about how an SNN design performs in terms of power consumption, speed (performance), and the physical area it occupies on the chip. Power metrics may include dynamic power consumption (related to switching activity), static power (leakage), and total power consumption. Performance metrics may involve delay characteristics and clock frequency. Area metrics provide insight into the size of the design.

[0078]

[0086] The following describes some example designs that may be used with or obtained through the design method 200 described above.

[0079]

[0087] One of the basic logical operations performed by neurons is accumulation. Accumulation refers to the process by which a neuron accumulates incoming input spikes over time and ultimately determines whether the neuron will generate an output spike. Accumulation captures the temporal dynamics of information processing in the network. Accumulation is performed based on weights associated with particular input spikes received from the network interconnect or from encoders or sensors via the interconnect.

[0080]

[0088] One common neuron design is the leaky integrate-and-fire (LIF) neuron. LIF neurons have a membrane potential that evolves over time based on incoming synaptic input and a leakage term. Accumulation occurs in the neuron's membrane potential. The membrane potential represents the electrical potential across the neuron's membrane and is influenced by the synaptic inputs it receives. For example, each incoming spike from a connected neuron contributes to the accumulation of the membrane potential. Synaptic weights determine the strength of these contributions. Excitatory synapses increase the membrane potential, while inhibitory synapses decrease it. Leakage can cause the membrane potential to decay slowly over time. This causes the neuron's membrane potential to return to its resting state in the absence of input. When the accumulated membrane potential reaches a certain threshold, the neuron fires or generates an output spike. This threshold is an important parameter that affects the responsiveness and sensitivity of the neuron. After firing, the membrane potential may reset to a resting or reset value, and the neuron may go through a refractory period during which it cannot immediately fire again.

[0081]

[0089] The accumulations that occur within neurons can take different forms depending on the interpretation of the weights within the neuron logic. Each bit of the weight can have a different logical function encoded within it. Some examples of logical operations within neuron accumulations are (a) addition, (b) subtraction, and / or (c) left and right shifts.

[0082]

[0090] Addition and subtraction logic operations model the accumulation of excitatory and inhibitory spikes, respectively. Left-shifting involves updating an accumulator or shift register by discarding the least significant bit or oldest value in the register. Left-shifting models the natural decay or leakage of accumulated information over time. In the case of leaky integrate-and-fire (LIF) neurons, it simulates the gradual decrease in membrane potential when no new input spikes are received. Right-shifting involves updating an accumulator or shift register by discarding the most significant bit or most recent value in the register. Right-shifting operations can be used in certain scenarios to introduce a time delay during the accumulation process. It effectively delays the contribution of the most recent input value to the accumulation, allowing the neuron to respond to past input patterns.

[0083]

[0091] Figure 3 shows a schematic diagram of the high-level architecture of a neuron 300 with m spiking inputs, specifically using the self-timing method described in the previous section. The goal is to capture m input spikes 306 coming from the network and generate timing information locally in a feedforward manner to enable the accumulation of weights associated with the spiking inputs. The weight summation is sent to an accumulator 302, which is connected to comparator logic 304 with a threshold 305. When the accumulator 302 reaches a programmed threshold 305 and an accumulator reset 308, an output spike 307 is generated.

[0084]

[0092] As shown, m input spikes 306 enter neuron 300, after which spike capture and timing generation take place in spike capture and timing generation module 301. Each incoming spike carries information and is generally associated with a synaptic weight. The neuron's spike capture mechanism involves detecting the arrival of the incoming spike. This detection may be based on a comparison with a voltage threshold. Timing generation involves creating a local train of timing pulses. The capture mechanism and timing generation are further described below. Weights 303 are then applied to the different captured spikes, and the spikes are accumulated in accumulator 302. The accumulator sends an accumulated signal to comparator logic 304, where the accumulated signal is compared with threshold 305. Threshold 305 may be fixed or variable. When the accumulated signal reaches threshold 305, an output spike 307 is generated, which is sent to (another part of) the network or to the output layer. When an output spike 307 is generated and the neuron "fires," the accumulator 307 may be reset using, for example, a reset signal sent from the comparator 304. This brings the accumulated signal in the accumulator 302 back to zero or some other predetermined level.

[0085]

[0093] FIG. 4 shows a schematic diagram of the capture of m spike inputs 401 and associated timing generation in a spike capture and timing generation module 400, in particular an example of an implementation of capturing spikes received from a network, and more particularly an example of how the spike capture and timing generation module 301 may be implemented.

[0086]

[0094] Because the system need not be based on or utilize a clock, the captured spikes can be used to generate local and / or multi-level timing information in the neuron itself. Spikes can be captured using pulse latching in the pulse capture module 402, for example, using m latches. One example of such a pulse latch is a standard CMOS D flip-flop 403. However, spike capture can be implemented using any other logic method; this is merely one example. In general, pulse latching spike capture can generate a logic state from the pulse, which thus indicates whether a pulse occurred or not. Multiple different spike capture methods can be used in the same pulse capture module 402, within the same neuron, or within a network.

[0087]

[0095] The CMOS D flip-flop (DFF) 403 is a digital circuit element that stores a single bit of information. The "D" in CMOS D flip-flop stands for "data." The flip-flop may have a data input (D) 4032, a clock input (CLK or CP) 4033, a set input (S), a reset input (R) 4031, and / or complementary outputs (Q and Q-bar) 4034a, 4034b. The DFF outputs the data input signal of the D input at one or both complementary outputs Q and Q-bar when triggered by a clock signal received at the CLK input.

[0088]

[0096] The CMOS D flip-flop 403 can be used to capture spike signals for a neuron by serving as part of a circuit that processes and integrates incoming spike events. The D flip-flop serves as a spike capture element, where the flip-flop's D (data) input 4032 is set to logic 1 and the CLK input 4033 is connected to a particular spike input signal 401 (e.g., coming into the neuron from a particular input synapse).

[0089]

[0097] The clock inputs of the flip-flops are driven by a clock signal, where each spike input signal corresponding to each DFF is used as the clock signal for that DFF. That is, when a particular spike signal is sent to a neuron, the corresponding DFF captures a logic 1 signal at a particular edge of the spike input signal and passes it through to its outputs Q and Q-bar.

[0090]

[0098] Q and Q-bar 4034a, 4034b are logical complements of each other. When Q is in a high state (logic 1), Q-bar is in a low state (logic 0), and vice versa. The two outputs always have opposite logical values. The complementary nature of Q and Q-bar 4034a, 4034b allows for convenient generation of inverted signals without the need for additional logic gates. The S and R inputs are asynchronous inputs that can force the Q output 4034a to a high or low state, respectively, regardless of the clock and data inputs.

[0091]

[0099] The output of each of the latches, e.g., the Q output 4034a of each CMOS D flip-flop, may be connected to a control input of a demultiplexer (DMUX) 405. A demultiplexer is a digital circuit that takes a single input and directs it to one of several possible outputs based on one or more control signals. Each DMUX may be controlled by the output of a DFF, i.e., the value of the output of the DFF determines to which output the DMUX directs its input signal. In this case, the input signal 404 of the first demultiplexer, controlled by a first latch connected to a first spike input port, may be a signal comprising a single pulse.

[0092]

[0100] If a spike is captured via the ith input (i∈{1,...,m}), the output of the ith DMUX may be sent through the delay element 406 and then forward to the (i+1)th demultiplexer controlled by the latch connected to the (i+1)th input. If a spike is not captured via the ith input (i∈{1,...,m}), the output of the ith DMUX may be sent forward to the (i+1)th demultiplexer controlled by the latch connected to the (i+1)th input without passing through the delay element 406.

[0093]

[0101] For i∈{1,...,m}, if a spike is captured, the incoming signal 410i from the corresponding ith delay element is also forwarded to a shared OR gate element 407, which combines the incoming signals from all delay elements whose spikes were captured into a train of timing pulses, the local timing signal 408. Preferably, this shared OR gate element is a pulseor tree element, which combines signals comprising pulses carried via different input lines into a single output line that pulses whenever one of the different input lines pulses. The local timing signal 408 is used by the neuron to synchronize and / or control the timing of operations within the neuron and may act as a local trigger signal. The local timing signal 408 comprises a number of pulses numerically equal to the number of spikes captured by the pulse latch module 402. The interval between pulses created by the delay element 406 may be based on the minimum time it takes to accumulate for a particular weight associated with a particular input spike line, as explained further below.

[0094]

[0102] A delay element 406 may be present to delay the local timing signal 408 so that it arrives at a fixed time at the next portion of the neuron.

[0095]

[0103] In other words, this embodiment discloses a spike capture and local trigger generation module 400 for a neuron of a spiking neural network. The module may include a plurality of spike input ports 401. Each of the plurality of spike input ports may be configured to receive a spike input signal from the spiking neural network. Furthermore, a pulse latch module 402 may be included within the module. The pulse latch module may be configured to capture spikes in the spike input signal received by the spike input port through pulse latching. Pulse latching may be described as digital circuitry that captures and holds a state in response to a trigger pulse, storing the state until reset by a subsequent pulse. The state may be transferred to other parts of the module. The pulse latch module may output a latched signal indicating whether a spike has been captured for each of the plurality of spike input ports. Furthermore, a local trigger generator module may be included within the module. The local trigger generator module may be configured to generate a trigger signal used by the neuron to synchronize and / or control the timing of operations in the neuron. The local trigger generator module may be configured to use the latched signals to generate a trigger signal, such that the trigger signal comprises one or more pulses numerically equal to the number of spikes captured by the pulse latch module. The trigger signal thus obtained may be used to trigger operations performed by other portions of the neuron, thus providing a local event-based triggering mechanism. Thus, the present invention does not require a global clock signal to operate those portions of the neuron. To use the trigger signal, the module may comprise a trigger signal output port 408 configured to output the generated trigger signal.

[0096]

[0104] Each of the signals 410i coming from the delay element may also be output by the spike capture and timing generation module 400 to a different portion of the neuron. For example, the output signal 410i representing that the corresponding spike input 401 has spiked may be i may be output by the spike capture and timing generation module 400. In this embodiment, these output signals 410i are delayed by the delay element 406, so that i will each have a different delay and are therefore output one after the other by the spike capture and timing generation module 400.

[0097]

[0105] Another option is for the spike capture and timing generation module 400 to generate the output signal 410 i Each latch element 4033 can also control a pulse generator to output a pulse in a particular output 410i if the latch element 4033 determines that a spike was present in the input signal. i can then have their own separate pulse generators, or the pulse generator can have multiple outputs 410 i can be shared between

[0098]

[0106] The local timing signal may also be determined in other ways. For example, different latch elements 4033 may detect whether a spike occurs in a particular input channel, and the number of latch elements that detect a spike may be determined, for example, via a counter element. The pulse generator may then be controlled to output the local timing signal 408 to the spike capture and timing generation module 400 with the number of pulses equal to the number of detected input spikes.

[0099]

[0107] The signal inputs 401 may also be grouped into different sets of varying numbers of signal inputs 401. A local timing signal 408 may then be generated for each set of signal inputs 401 in, for example, one of the exemplary manners described above.

[0100]

[0108] The spike capture and timing generation module 400 may also not output a spike output signal, but instead may only generate a local timing signal 408 .

[0101]

[0109] If i=m, the output of the mth DMUX (or corresponding delay element) may be sent through an OR gate element 407 and output as a multi-level timing signal 409, which will be described in more detail below. The multi-level timing signal 409 is generated to fractionate the end of the timing generation for a particular level of summation. The delay of the multi-level timing must be greater than the delay of all summations performed at that particular level. Levels are described further below. Two or more multi-level timing signals may need to be generated depending on how the accumulation process is implemented in the neuron.

[0102]

[0110] FIG. 5 shows a schematic diagram of the summation of spike weights from local memory in a unit cell 500. In particular, the diagram shows k spike inputs 5101-5102 of a total of m spike inputs of a neuron. k (k can be less than or equal to m). The unit cell 500 in this case is therefore the part of the neuron responsible for accumulation, which in this example has k of the m spike inputs. The spike inputs 5101-510 in this case k is the respective spike output 410 shown in FIG. 4 coming from the spike capture and timing generation module 301. iThus, by using these unit cells, the m inputs to the accumulator are split into one or more sets of k. The reason for splitting the m inputs into k sets is a trade-off between:

[0103]

[0111] That is, the more gates a combinational logic operation, such as accumulation, has, the greater the uncertainty regarding the maximum time it will take to complete it. Design methods may be based on maximizing time, and therefore the greater the uncertainty, the slower the overall design will be, as delays in the design must be maintained for the maximum possible time based on timing uncertainty. Furthermore, the larger the delay, the more uncertain the delay. Thus, this uncertainty adds on top of the uncertainty in the combinational logic delay. The objective is to divide the design into logic blocks that can then be used as reproducible design blocks (shown in FIG. 1). These reproducible blocks need to be timed within the same layer and between layers, as further shown below.

[0104]

[0112] In the case of summing incoming spikes as shown in FIG. 5, it is expected that the input that had the spike will be high (whether or not the spike occurred can be determined by the latching module using the logic of FIG. 4) and the correct number of pulses will be generated to time the accumulation operation of all weights associated with the input that had the spike.

[0105]

[0113] Input 5101~510 k The inputs 5101-5100 go to an encoder 501, which encodes the spikes from the inputs into memory addresses. The encoder can also reorder the addition sequence if necessary, for example based on some rules of priority. The encoder 501 can also operate to delay the addition for one delay cycle so that the addition can be performed on the next set of incoming inputs. The encoder 501 encodes the received pulse inputs 5101-5100 into a local memory 502, for example a non-volatile memory such as an SRAM or ReRAM.k A memory address can be passed for each.

[0106]

[0114] Received Inputs 5101-510 k The weight stored in memory corresponding to the address for each pulse in the accumulator is passed onto accumulation element 503. The accumulator performs the function of adding / subtracting the incoming input to the last stored value. The combination of 503 and 504 accomplishes this. In this case, 503 can be an adder, a subtractor, or both in standard CMOS logic. The signal from accumulation element 503 is then passed to, for example, the D input 5042 of CMOS D flip-flop 504. A local timing signal 508 can also be passed as an input to CMOS D flip-flop 504, i.e., at the CLK input 5043. DFF 504 captures the incoming accumulated weight signal at a particular edge of the clock signal corresponding to a pulse in local timing signal 508.

[0107]

[0115] The DFF has a feedback loop from its output, for example from Q port 5044a, back to accumulation element 503. In this way, an accumulated weight is output by the DFF and used as a second input to accumulation element 503. The accumulation element, meanwhile, gets the next weight from local memory 502 and adds the new weight with the already accumulated weight to produce a new accumulated weight. This new accumulated weight is then sent again to DFF 504. Since local timing signal 508 comprises a number of pulses equal to the number of times a signal is received from accumulation element 503, all signal inputs 5101-510 at which a spike is received will be synchronized. k The accumulation of weights is performed in this way.

[0108]

[0116] In other words, the local timing signal 508 comprises a number of pulses corresponding to the incoming spike signal. Each pulse is used to determine when a particular (accumulated) weight coming from the local memory 502 has been captured. The captured weights are integrated over time through successive clock cycles. The DFF 504 accumulates the presence of spikes over multiple clock cycles (e.g., after each trigger from the local timing signal 508), providing a form of temporal integration. In this way, operations relating to the logic encoded in the accumulator corresponding to the weight can be performed in this manner. As mentioned above, this process is performed for all captured input spikes 5101-510. k Multiple unit cells may perform this function within a neuron, thus repeating for all m spike signals 4101-410. m 4. The local timing pulse 408, 508 is expected to be delayed enough from the start of the pulse for the spike to be processed in the accumulator, e.g., in DFF 504. The start of the pulse was used to convey that the local timing pulse 408 has an additional delay in it. This additional delay can be used to control the relative timing of when logic reaches, for example, local memory 502 or accumulation element 503, relative to the timing of the local timing pulse 508.

[0109]

[0117] Figure 6 shows a schematic diagram of the summation of spike weights from a distributed memory in a unit cell. In particular, Figure 6 shows an alternative implementation of weight accumulation to that shown in Figure 5.

[0110]

[0118] Input 6101~610 k goes to a priority encoder 601, which in this case has a size of j bits and is k The pulse input 6101-610 received by one of the distributed memory elements 607a stores a weight corresponding to k607a and 607b. If an input spike is not received at an input, this corresponds to a logic 0 signal through the corresponding channel 607b. If an input spike is received at an input, this corresponds to a logic 1 signal through the corresponding channel 607b. The weight from 607a and the signal from channel 607b are then passed through an AND gate element 607c. If channel 607b carries a logic 1, the input weight is forwarded to multiplexer 602. All inputs 6101-610 k All outputs from AND gate elements 607c corresponding to are then passed through multiplexer 602. Multiplexer 602 can pass through the channels one by one and forward the weights to the accumulator logic.

[0111]

[0119] Therefore, the received pulse inputs 6101 to 610 k The weights stored in the distributed memory corresponding to the addresses of each input are passed to an accumulator. In one example, the stored weights are passed to an accumulation element 603. The signal from accumulation element 603 is then passed to, for example, a D input 6042 of a CMOS D flip-flop 604. A local timing signal 608 may also be passed as an input to the accumulator, for example, to CMOS D flip-flop 604, i.e., at a CLK input 6043. DFF 604 captures the incoming weight signal at a particular edge of the clock signal corresponding to a pulse in the local timing signal 608.

[0112]

[0120] The DFF has a feedback loop from its output, e.g., from Q port 6044a, back to accumulation element 603. In this way, an accumulated weight is output by the DFF and used as a second input to accumulation element 603. The accumulation element, meanwhile, gets the next weight from the distributed memory, e.g., from multiplexer 602, and adds the new weight with the already accumulated weight to produce a new accumulated weight. This new accumulated weight is then sent again to DFF 604. Since local timing signal 608 comprises a number of pulses equal to the number of times a signal is received from accumulation element 603, all signal inputs 6101-610 at which a spike is received will be synchronized. k The accumulation of weights is performed in this way.

[0113]

[0121] In other words, the local timing signal 608 comprises a number of pulses corresponding to the incoming spike signal. Each pulse is used to determine when a particular (accumulated) weight coming from the distributed memory 602 is captured by the DFF. The captured weight is integrated over time through successive clock cycles triggered by the local timing signal 608. The DFF 604 accumulates the presence of spikes over multiple clock cycles (e.g., after each trigger from the local timing signal 608), providing a form of temporal integration. In this way, an operation relating to the logic encoded in the accumulator corresponding to the weight can be performed accordingly. As mentioned above, this process is performed for all captured input spikes 6101-610. k Multiple unit cells may perform this function within a neuron, thus repeating for all m spike signals 4101-410. m The local timing pulse 408, 608 is expected to be delayed in the accumulator sufficiently from the start of the pulse for the spike to be processed, for example, in DFF 604.

[0114]

[0122] In this way, input spikes 6101 to 610 kThe weights associated with can be instantiated using distributed weight storage elements, comprising, for example, latches and / or flip-flops. A benefit of a distributed weight storage mechanism is that conventional place-and-route tools (e.g., Innovus, ICCII, Openroad, all readily available software solutions) can optimize the combinational logic around the weights, and the generation of modules using the design methodology shown in FIG. 2 can be automated with minimal human intervention for all the different logic blocks required in this manner. Furthermore, logic storage elements such as latches or flip-flops offer the designer the option of undervoltageing their supplies, which potentially significantly improves dynamic power consumption, although this leads to leakage values ​​comparable to memories like SRAM and ReRAM. Undervoltage mitigation involves intentionally reducing the operating voltage (VDD) below the standard or nominal level specified for the latch or flip-flop. This can be done for a variety of reasons, including power optimization and energy efficiency.

[0115]

[0123] Thus, a weight encoding and accumulation module for accumulating spike input signals in neurons of a spiking neural network is disclosed in FIGS. 5 and 6. The module may include multiple spike input ports 510, 610. Each of the multiple spike input ports may be configured to receive a spike input signal. Furthermore, the module may include trigger signal input ports 508, 608 configured to receive a trigger signal. Pulses included in the trigger signal may indicate the number of spikes received by the multiple spike input ports. Furthermore, the module may include a spike-to-weight encoder module. The spike-to-weight encoder module may be configured to receive the spike input signals and output corresponding weights associated with particular spike input signals. This may be done by encoding the received spike input signals as weights. Furthermore, an accumulator module may be included in the module. The accumulator module may be configured to receive weights from the memory module and a trigger signal and accumulate weights for the received spikes. The timing of one or more of the operations involved in the accumulation may be governed by the trigger signal. In this case, the loop circuit formed by the CMOS D flip-flop is governed by the trigger signal. Other ways of using the trigger signal to trigger the accumulation can also be envisioned.

[0116]

[0124] With the delay-based serialization method shown in Figures 4-6, only captured spikes are serialized and added into the adder. Using a combination of delay-based timing (for the worst-case setup time of the adder) and the feed-forward nature of the design, only certain constraints (e.g., the delay of each logic block in the design and the delay of the delay cells) need to be met for the design to function correctly. Note that if the design is feed-forward and the logic paths are well-balanced, then from a timing perspective, we can only target solving the setup time problem, not the stall time problem. If the specific timing signal leading to the capture of data into the latch is a pulse rather than an edge (hence pulse-based latching rather than edge-based latching), then stall time is also mitigated. Therefore, it should be understood that if we are only solving setup time problems, we can always fix the setup time problem by slowing down the design. Variable distributed delays can be used to solve setup time problems in any part of the design.

[0117]

[0125] Furthermore, the interval between pulses must be greater than the time required to access a particular memory plus the accumulation time. As noted above, the pulses reaching the accumulator can be further delayed compared to those going to the memory (as is done at local timing signal output 408). Because the SNN is a spare network with periods of high activity, the above design approach is able to handle high activity periods while at the same time not having unnecessary dynamic power consumption during periods of low activity.

[0118]

[0126] Note that the unit cells 500, 600 shown in Figures 5 and 6, for example, see input analog signals and output digital values ​​that are the cumulative sum of the weights of the input signals with spikes. Also, a second type of unit cell can be designed that simply adds the digital values ​​of the weights together. These can be used in further layers of additional neuron architectures, where, through parallelization, multiple unit cells see input analog signals and convert them to digital signals, and the second type of unit cell can be used to add a digital signal (representing the cumulative weights) to the cumulative weights.

[0119]

[0127] FIG. 7 shows a schematic diagram of the design of a multilevel self-timed neuron with regenerative unit cells. In particular, a high-level architecture of how multiple inputs are divided across different levels at a given summation level is shown. A tradeoff between logic complexity and timing uncertainty divides the m inputs in the neuron into smaller sets of inputs, e.g., all k inputs, with the same or different number of inputs across the set. For each set, each of the k inputs is processed in a unit cell. The unit cell may have the architecture described in FIG. 5 or FIG. 6.

[0120]

[0128] Each unit cell may generate a timing signal for when it completes its computation, as shown in FIG. 4 with a multi-level spike timing signal. As an example, the done signal may be generated by an additional delay signal operating in parallel with the logic. As long as the delay signal is longer than the worst-case delay of the logic, the multi-level timing signal will always indicate timing (e.g., through a pulse) after the worst-case delay of the logic. The done signal of multiple unit cells in a level may be the OR of the individual done signals. This may be applied in a multi-stage pipeline design implemented in a feedforward manner through a delay-based timing approach. For an example, see FIG. 8 below.

[0121]

[0129] Accumulation can be performed at different levels. When accumulation is divided into different levels, it means that the accumulation process occurs independently or hierarchically at each level of the system. In neural networks or signal processing systems, information is often processed at different levels of abstraction. Each level may represent features or patterns with varying degrees of complexity.

[0122]

[0130] There are multiple reasons for splitting the addition into different levels. First, pipelining improves throughput at the expense of latency. Furthermore, with delay-based timing, when one attempts to improve latency by increasing the logic in a single level, there is an additional tradeoff: increased timing uncertainty. This increase in timing uncertainty degrades latency.

[0123]

[0131] For example, a multi-level spike timing signal 409 generated using spike capture and timing generation as shown in FIG. 4 may be used to perform timing between different levels using feed-forward timing as shown in FIG. 8 and / or using handshaking as shown in FIG. 9.

[0124]

[0132] Input 701 m In m, m different inputs to neuron 700 enter neuron 700. These may correspond, for example, to input synapses to neuron 700. The different input spikes are latched in latch 7020. This may be done, for example, using the spike capture and timing generation modules 301, 400 shown in FIG.

[0125]

[0133] Next, the input signal is divided into k inputs 701 k and the input signal 701 k5 and 6. At the end of level 0, the accumulated signal coming from unit cell 703 is latched into level 1 register set 7021. Region 711 indicates that unit cell 703 in this region 711 is a spike-driven unit cell that takes analog signals with spikes as input, e.g., the unit cells shown in FIGS. 5 and 6. Henceforth, in neuron logic, accumulation unit cell 703 can be both spike-driven or data-driven. As an example, region 712 indicates that unit cell in this region 712 is a data-driven unit cell, e.g., a unit cell that takes digital signals as input and accumulates the values ​​of these digital signals, as described above.

[0126]

[0134] Since the implementation generates a multi-level timing signal indicating when level 0 completes its computation (as shown with respect to FIG. 4 ), this can be used as an input change signal; the remaining unit cells do not need to be internally timed; they can be made from conventional CMOS and timed using static timing analysis (STA). Note that standard CMOS logic can be made to operate without a clock. The design can be done using, for example, maximum and minimum delay constraints. Standard software EDA synthesis tools such as Design Compiler and Genus can determine these constraints. They ensure that the logic created always has a delay longer than the minimum delay and shorter than the maximum delay. Because timing between levels is still delay-driven and timing uncertainty at longer delays is still a reality, further division of levels may be required based on the timing tradeoffs previously described.

[0127]

[0135] A better tradeoff can be achieved at a particular level versus latency for timing uncertainty by using time multiplexing within a given level. If it is known that a previous level will take longer than the current level (due to logical tradeoffs within the previous level), the current level can be optimized to process more information by using time multiplexing. Timing in that case would still be done via delay-based signals, and standard time multiplexing techniques could be used.

[0128]

[0136] At the first level, the latched accumulated signals latched in the level 1 register set 7021 are the sum of the k input signals 701 k1 , which are then transferred to respective accumulation unit cells 703 which may operate by accumulating digital accumulation signal values.

[0129]

[0137] At the last level, the outputs of the unit cells from the second layer to the last layer are latched by the level N register set 702N. The spike input 701kN is then sent through element 705 to an accumulator 706, e.g., DFF 706, and in particular, the spike input 701kN is sent to the D port of DFF 706. Thus, at the end of the last level, accumulation is performed on the previous addition. The timing is still delay-based feedforward or delay-based handshaking.

[0130]

[0138] When the final level accumulator 706 reaches a preset threshold 709, an output spike 710a is generated, for example, by a comparator 708 that checks whether the accumulated signal has reached the preset threshold 709, and the accumulated value in the accumulator 706 is reset to a predetermined value. The accumulator 706 can send an accumulated signal to the comparator 708, where the accumulated signal is compared with the threshold 709. Additionally, the accumulator 706 can output an accumulated sum 710b as a separate output, which can be used as a real-time signal used in training and inference in neural networks.

[0131]

[0139] Pipelining of multiple small accumulators, such as the k-input unit cell accumulator 703, must be done to be able to create an m-input high fan-in accumulator. A pipeline such as that shown in Figure 7 can be implemented through a variety of techniques. The logic resides in the control logic 704.

[0132]

[0140] 8 and 9 show two potential methods for the pipeline implementation of FIG.

[0133]

[0141] Figure 8 implements a multi-stage pipeline design in a feedforward manner using delay-based timing techniques.

[0134]

[0142] The input signal enters the pulse latch 804, and the pulser tree element 801 combines the signals comprising the pulses carried through the different input lines into a single output line that pulses whenever one of the different input lines pulses. The output from the pulse latch 804, the weights stored in the weight memory 803, and / or the output of the pulser tree element 801 can be used as inputs to the L0 logic 8050. The output of the L0 logic 8050 is then sent to the register 806. The output from the pulser tree element 801 is sent through a delay, which generates a delayed signal equal to or greater than the processing time of the L0 logic 8050. This delayed signal is used as an input to the register 806 to determine when the register should transfer its latched signal to the next level of logic, in this case, the L1 logic 8051. The delayed signal is also used as an input to the level logic, the L1 logic 8051. The delayed signal is also sent to a different delay element 802 corresponding to the level logic L1 logic 8051, so that the delay it produces is equal to or greater than the processing time of the level logic L1 logic 8051. This method is repeated for the remaining layers.

[0135]

[0143] Figure 9 implements the pipeline structure using a handshaking mechanism between stages. For handshaking, correct loop breaking needs to be implemented for the proposed constraints and methods for working.

[0136]

[0144] The input signal enters pulse latch 904 and pulser tree element 901. The output from pulse latch 904, the weight stored in weight memory 903, can be used as an input to L0 logic 9050. The output of L0 logic 9050 is then sent to register 906. The output from pulser tree element 901 is also sent to control logic 902, which initiates a handshake procedure with L0 logic 9050 by sending a synchronize (SYN) signal to which L0 logic 9050 will respond when an acknowledge (ACK) signal is ready.

[0137]

[0145] When the L0 logic 9050 is ready, it outputs an output signal to a register 906, which latches the output signal. It then signals the next control logic 902, which performs a handshake procedure with the next logical layer L1 logic 9051. This control logic 902 then signals the register 906 to output the latched data as an input signal to the logical layer L1 logic 9051, based on the signal received from the first L0 logic 9050. This method is repeated for the remaining layers.

[0138]

[0146] Depending on the use case traffic, each stage of the pipeline can be multiplexed, so proper optimization can help further optimize the throughput of the pipeline. This gives added flexibility as the network traffic can be high throughput, low latency, or a combination of the two. The correct design topology can be determined by a mix of flexible pipeline depth and delay-based time division multiplexing.

[0139]

[0147] The designs described in this work are more at the high architectural and logical function levels. Depending on the power, area, and performance requirements, the actual implementation can be done in high-speed design methods such as self-resetting CMOS, pulse-based design, or in cases where the requirements are less stringent than in standard CMOS by using static timing analysis modified for delay-based timing closure as described in the Methods section.

[0140]

[0148] Neurons designed using the above examples can be further optimized for specific use cases.

[0141]

[0149] It has been found that if a use case does not utilize all possible input lines, the delay at a particular stage can be significantly reduced. This leads to a trade-off improvement in both latency and throughput, but controls the delay at each stage. This differs from traditional clock-based design techniques, where the granularity of control is more global due to the more complex design of the clock network at the local level. In general, latency and throughput are inverses of each other and are fixed by design. Using this method, they can both be improved and made use-case dependent.

[0142]

[0150] Since the delays can be implemented using mixed signal design techniques, similar calibration techniques can be used to mitigate any process, temperature and voltage issues.

[0143]

[0151] The AND gates herein may be implemented using semiconductor devices such as transistors to perform logical ANDs, while OR gates use similar components for logical ORs. Pulse latches may be constructed using flip-flops and pulse generators. Accumulators may employ operational amplifiers and / or digital counters. Subtractor circuits may consist of arithmetic units capable of calculating differences. Input and output ports are designed with interface modules to enable seamless communication between neurons.

[0144]

[0152] The present invention particularly discloses the implementation of spiking neurons and spiking neural networks in hardware, e.g., on a chip. Silicon is the predominant material used in semiconductor fabrication. Metal-oxide-semiconductor field-effect transistors (MOSFETs) are the basic building blocks of digital integrated circuits. They can be used to implement logic gates, memory cells, and other essential components. SRAM is commonly used for fast, volatile storage of synaptic weights, neuron states, and other dynamic information within SNNs. For more persistent storage needs, non-volatile memory technologies such as Flash or resistive random-access memory (RRAM®) can be employed. Multiple metal layers can be used to interconnect different components on a chip. These layers facilitate signal routing between neurons, synapses, and other functional units. Through-silicon vias (TSVs) can be used to enable vertical stacking of multiple layers, improving overall connectivity and reducing the physical footprint.

[0145]

[0153] Energy-efficient design techniques such as voltage scaling (dynamically adjusting operating voltage to balance power consumption and performance), clock gating (temporarily suspending the clock signal to idle portions of the circuit during inactive periods), and / or power gating (completely cutting off power to certain sections of the chip when not in use) may be used. According to the present application, local timing is generated based on some trigger event (in this case, receiving a spike), which can help to obtain an energy-efficient design.

[0146]

[0154] It should be noted that the features of any of the embodiments disclosed herein may be combined in any suitable manner.

Claims

1. 1. A spike capture and local trigger generation module 400 for neurons of a spiking neural network, comprising: a plurality of spike input ports 401, wherein each spike input port 401 is configured to receive a spike input signal from said spiking neural network; a pulse latch module 402 configured to capture spikes in one or more of the spike input signals, wherein the pulse latch module 402 outputs a signal indicating whether a spike has been captured at one of the spike input ports 401; a local trigger generator module configured to generate a trigger signal for use by the neuron to synchronize and / or control the timing of operations in the neuron, wherein the local trigger generator module is configured to generate the trigger signal based on the signal from the pulse latch module 402, wherein the trigger signal indicates the number of spikes captured by the pulse latch module 402; a trigger signal output port 408 configured to output the generated trigger signal; A spike capture and local trigger generation module 400 comprising:

2. The spike capture and local trigger generation module of claim 1 , wherein the trigger signal comprises a number of pulses equal to the number of spikes captured by the pulse latch module.

3. 3. The spike capture and local trigger generation module of claim 1, wherein the trigger signal indicates the number of spikes captured by the pulse latch module during a predetermined time period.

4. 4. The spike capture and local trigger generation module 400 of claim 1, wherein the trigger signal is generated such that the operations in the neurons can be triggered by the trigger signal rather than by a global clock signal of the spiking neural network.

5. The spike capture and local trigger generation module 400 of claim 1 , wherein the trigger signal is a delay-based asynchronous counter.

6. 6. The spike capture and local trigger generation module 400 of claim 1, wherein the one or more pulses provided in the trigger signal have an interval between them that is determined by the minimum time it takes to perform the operation in the neuron, preferably wherein the operation is the accumulation of one or more weights.

7. 7. The spike capture and local trigger generation module 400 of claim 1, further comprising: a plurality of spike output ports each configured to output a spike output signal, wherein each of the plurality of spike output ports corresponds to one of the plurality of spike input ports, and wherein a particular spike output port is configured to output a pulse when a corresponding spike input port receives a spike.

8. 8. The spike capture and local trigger generation module 400 of claim 7, wherein the pulses output by the spike output port are spaced in time substantially similarly to the pulses provided in the clock signal.

9. 9. The spike capture and local trigger generation module 400 of claim 1, wherein the pulse latch module 402 comprises a plurality of latch elements 403, each latch element configured to capture one or more spikes in the spike input signal received by a particular spike input port by pulse latching.

10. 10. The spike capture and local trigger generation module 400 of claim 9, wherein the latch element 403 is a CMOS D flip-flop 403, wherein the CMOS D flip-flop is triggered by a spike input signal, and wherein the latched signal output by the CMOS flip-flop is a bit value indicating whether a spike has been latched.

11. The local trigger generator module includes a pulse generator 404 for generating pulses. wherein the local trigger generator module comprises delay logic for each spike input port, wherein the generated pulse is then sent through each of the delay logics. A spike capture and local trigger generation module 400 according to any one of claims 1 to 10.

12. each of the delay logics comprises a delay element that delays the pulse signal; wherein the generated pulse is sent through the delay element of the particular delay logic if the corresponding latched signal indicates that a spike has been captured by the corresponding spike input port, and wherein the generated pulse is not sent through the delay element of the particular delay logic if the corresponding latched signal indicates that a spike has not been captured by the corresponding spike input port. The spike capture and local trigger generation module 400 of claim 11.

13. the generated pulse is output by a particular delay logic to generate the trigger signal when the corresponding latched signal indicates that a spike has been captured by the corresponding spike input port, and wherein the generated pulse is not output by the particular delay logic to generate the trigger signal when the corresponding latched signal indicates that a spike has not been captured by the corresponding spike input port. A spike capture and local trigger generation module 400 according to claim 11 or 12.

14. the local trigger generator module is configured to generate a plurality of trigger signals used by the neuron to synchronize and / or control the timing of operations in respective portions of the neuron, wherein the local trigger generator module is configured to use the plurality of latched signals to generate the plurality of trigger signals, each trigger signal comprising one or more pulses numerically equal to the number of spikes captured by the pulse latch module 402 for a particular set of spike input ports 401; a plurality of trigger signal output ports 408 configured to output the generated trigger signals; A spike capture and local trigger generation module 400 according to any one of claims 1 to 13.

15. 15. The spike capture and local trigger generation module 400 of any one of claims 1 to 14, further comprising: a multi-level trigger generator module having a multi-level trigger output port 409 configured to generate a multi-level trigger signal, wherein the multi-level trigger signal indicates the total time it will take the neuron to perform the operations triggered by a trigger signal, preferably wherein, if there are multiple trigger signals generated and output, the multi-level trigger signal corresponds to the maximum time it will take the neuron to perform all the sets of operations triggered by the different trigger signals.

16. 1. A weight encoding and accumulation module for accumulating spike input signals in neurons of a spiking neural network, comprising: a plurality of spike input ports, wherein each spike input port is configured to receive a spike input signal; a trigger signal input port configured to receive a trigger signal, wherein the trigger signal indicates a number of spikes received by the spike input port; a spike-to-weight encoder module configured to receive the spike input signal and output a corresponding weight associated with the spike input signal when a spike is received in the spike input signal; an accumulator module configured to receive the trigger signal and the weights from the spike-to-weight encoder module and to accumulate the weights, wherein the timing of one or more of the operations involved in the accumulation is based on the trigger signal; a weight encoding and accumulation module comprising:

17. 17. The weight encoding and accumulation module of claim 16, wherein the trigger signal comprises a number of pulses equal to the number of spikes received by the spike input signal port.

18. 18. The weight encoding and accumulating module of claim 16 or 17, wherein each of the spike input signals is an analog signal, and wherein the weights are digital values.

19. 19. The weight encoding and accumulation module of claim 16, wherein the accumulator module comprises an accumulation element 503 and a loop circuit that takes the output of the accumulation element as an input and outputs a value back to the accumulation element each time the loop circuit receives a pulse from the trigger signal, wherein when the spike-to-weight encoder module outputs a first weight to the accumulation element 503, the first weight is added by the accumulation element with a second weight output by the loop circuit.

20. 20. The weight encoding and accumulating module of claim 19, wherein the loop circuit comprises a CMOS D flip-flop triggered by the trigger signal.

21. 21. A weight encoding and accumulation module according to any one of claims 16 to 20, wherein the spike-to-weight encoder module comprises a local memory, preferably an SRAM or a non-volatile memory, wherein the spike-to-weight encoder encodes a pulse received at a particular spike input port as that particular memory address being used to request a weight stored in that memory address, the weight being associated with the particular spike input port.

22. 22. The weight encoding and accumulation module of claim 16, wherein the spike-to-weight encoder module comprises a distributed memory comprising distributed weight storage elements, wherein the weights associated with the spike input ports are respectively stored in one of the distributed weight storage elements, and wherein when a spike is present in the spike input signal, the corresponding weight is output to the accumulator module by the distributed weight storage element in which the weight is stored.

23. 23. The weight encoding and accumulation module of any one of claims 16 to 22, wherein the spike-to-weight encoder module is configured to change the order in which the weights are output based on a priority rule.

24. 24. A spiking neuron of a spiking neural network, the neuron comprising a spike capture and local trigger generation module 400 according to any one of claims 1 to 15 and a weight encoding and accumulation module according to any one of claims 16 to 23, wherein a trigger signal output port of the spike capture and local trigger generation module outputs the trigger signal to the trigger signal input port of the weight encoding and accumulation module, such that the timing of one or more of the operations involved in the accumulation is governed by the trigger signal.

25. 25. The spiking neuron of claim 24, further comprising: a comparator module that compares a total accumulated weight with a predetermined threshold, wherein the total accumulated weight is the sum of weights of all spiking input ports from which spikes of a corresponding spiking input signal are captured, and wherein the spiking neuron is configured to output a spike to the spiking neural network when the total accumulated weight exceeds the predetermined threshold.

26. 26. The spiking neuron of claim 24 or 25, wherein the spiking neuron comprises a plurality of weight encoding and accumulation modules according to any one of claims 16 to 23, wherein the spike capture and local trigger generation module 400 is according to claim 14, wherein each of the plurality of trigger signal output ports outputs a respective trigger signal to the trigger signal input port of a corresponding weight encoding and accumulation module, and wherein the timing of one or more of the operations involved in the accumulation performed in the corresponding weight encoding and accumulation module is therefore governed by the respective trigger signal.

27. 27. The spiking neuron of claim 26, further comprising a digital weight accumulator module that accumulates the outputs of at least some of the plurality of weight encoding and accumulation modules and outputs the accumulated value.

28. 28. The spiking neuron of claim 27, wherein the spiking neuron comprises a plurality of digital weight accumulator modules, wherein at least one digital weight accumulator module accumulates the output cumulative values ​​of other digital weight accumulator modules.

29. 29. The spiking neuron of claims 26 to 28, wherein delay-based pipelining and / or handshake-based pipelining is used to synchronize the accumulation operations among the weight encoding and accumulation modules and / or the one or more digital weight accumulator modules.

30. 30. The spiking neuron of claim 29, wherein the spike capture and local trigger generation module is according to claim 15, and wherein a multi-level trigger signal is used to synchronize the accumulation operations among the plurality of weight encoding and accumulation modules and / or the one or more digital weight accumulator modules.

31. 1. A method for spike capture and local trigger generation in neurons of a spiking neural network, comprising: receiving a plurality of spiking input signals from the spiking neural network; capturing spikes in said spike input signal by pulse latching; generating a latch signal for each of said spike input signals, wherein said latch signal indicates whether a spike has been captured; generating a trigger signal based on the latch signal, wherein the trigger signal indicates the number of spikes captured, and wherein the trigger signal is used by the neuron to synchronize and / or control the timing of operations in the neuron. outputting the generated trigger signal; A method comprising:

32. 32. The method of claim 31 , further comprising outputting a spiking output signal, wherein the trigger signal is generated such that the operations in the neuron can be triggered by the trigger signal rather than by a global clock signal of the spiking neural network.

33. 1. A method for weight encoding and accumulation in neurons of a spiking neural network, comprising: receiving a plurality of spike input signals; receiving a trigger signal indicative of a number of spikes received in the spike input signal; obtaining or generating a weight associated with one of the spike input signals when a spike is received in the spike input signal; accumulating the weights associated with the spike input signal, wherein the timing of one or more of the operations involved in the accumulation is based on the trigger signal. outputting the cumulative weight; A method comprising:

34. 34. The method of claim 33, wherein the weight is obtained or generated by encoding a pulse received in the one of the plurality of spiking input signals as a memory address used to request a weight stored in a local memory.

35. 34. The method of claim 33, wherein when a spike is present in the spike input signal, the corresponding weight is retrieved from a distributed memory comprising a plurality of weight storage elements, wherein the weights associated with the spike input signal are stored in one of the distributed weight storage elements respectively, and wherein when a spike is present in the spike input signal, the corresponding weight is output by the distributed weight storage element in which the weight was stored and used during the accumulation of the weights.

36. 1. A method for spike capture and accumulation in neurons of a spiking neural network, comprising: receiving a plurality of spiking input signals from the spiking neural network; capturing spikes in said spike input signal by pulse latching; outputting a latch signal for each of said spike input signals, wherein said latch signal indicates whether a spike has been captured. generating a trigger signal based on the latch signal, wherein the trigger signal indicates the number of spikes captured, and wherein the trigger signal is used by the neuron to synchronize and / or control the timing of operations in the neuron. obtaining or generating a weight associated with one of the plurality of spike input signals when a spike is received in the spike input signal; accumulating the weights associated with the spike input signal, wherein the timing of one or more of the operations involved in the accumulation is based on the trigger signal. outputting the cumulative weight; A method comprising:

37. 37. The method of claim 36, further comprising: comparing a total accumulated weight with a predetermined threshold, wherein the total accumulated weight is the sum of weights of all of the spiking input signals from which a spike was captured; and outputting a spike to the spiking neural network when the total accumulated weight exceeds the predetermined threshold.

38. 38. The method of claim 36 or 37, wherein the steps of obtaining or generating weights and accumulating the weights are parallelized by dividing the plurality of spiking input signals into one or more subsets of the plurality of spiking input signals, and simultaneously performing the steps of obtaining or generating weights for these subsets to obtain accumulated subset weights for each subset, and accumulating the weights.

39. 39. The method of claim 38, wherein the accumulated subset weights are accumulated digitally, preferably in a parallelized manner.

40. 40. The method of claim 38 or 39, wherein delay-based pipelining and / or handshake-based pipelining is used to synchronize the accumulation operations between different subsets and / or the one or more digital accumulations.