Self-timed feed-forward synthesizable and technologically scalable mixed signal neurons

Through spike capture and local trigger generation modules, combined with weight encoding and accumulation modules, the problems of high power consumption, poor scalability and delay in spike neural networks are solved, and low-power consumption and efficient neuron design is achieved.

CN120569730APending Publication Date: 2025-08-29INNATERA NANOSYSTEMS BV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202480008198.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-01-17
Filing Date
2024-01-17
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

The digital neuron design of existing spike neural networks has problems with high power consumption, poor scalability and delay, while simulated neurons face the problem of insufficient signal-to-noise ratio.

Method used

The spike capture and local trigger generation module is used to capture the spike input signal through the pulse latch module, and the local trigger generation generator is used to generate a trigger signal to synchronize the operation of neurons, avoid global clock signals, and combine the weight encoding and accumulation module for information processing.

Benefits of technology

It realizes a low-power, scalable neuron design, reduces latency and improves signal-to-noise ratio, and enhances the computing efficiency of the neural network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120569730A_ABST
    Figure CN120569730A_ABST
Patent Text Reader

Abstract

The invention relates to a spike capture and local trigger generation module for neurons of a spiking neural network. The module includes spiking input ports, each port configured to receive a spiking input signal. Further, the module includes a pulse latch module configured to capture a spike in the spike input signal by pulse latching. The pulse latch module outputs a latch signal for each spike input port, the signal indicating whether a spike is captured. The module further comprises a local trigger generator module configured to generate a trigger signal to synchronize and / or control timing of operations in the neurons and configured to generate the trigger signal using the latched signal such that the trigger signal contains a number of pulses equal to the number of spikes captured. Finally, the module can further comprise a trigger signal output port which outputs the generated trigger signal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to automatic signal recognition technology, and more particularly, to a system and method for a hardware-resilient deep learning inference accelerator using spiking neural networks. Background Art

[0002] A spiking neural network (SNN) is a signal processing system whose design is inspired by biological neural networks. Information is encoded in patterns of spikes distributed across a complex network of neurons and synapses.

[0003] Synapses store information and use it to perform basic operations on incoming signals. Neurons generate new spikes based on signals from synapses. The connections between synapses and neurons determine the shape of the SNN, and the shape of the SNN, together with the information stored in the synapses, determines the signal processing function of the SNN.

[0004] There is a large amount of digital implementation of typical functions performed in neurons of spiking neural networks in the literature. See, for example, P Kampf, P Koch, K Roy, M Sullivan, Z Delalic, and S Das Gupta, "Optimization of A Digital Neuron Design," published in the 1990 Eastern Multiconference. Record of Proceedings. The 23rd Annual Simulation Symposium, 1990, pp. 73–80. Another example is D. Lee, G. Lee, D. Kwon, S. Lee, Y. Kim, and J. Kim, "Flexon: A Flexible Digital Neuron for Efficient Spiking Neural Network Simulations," published in the 2018 ACM / IEEE 45th Annual International Symposium on Computer Architecture (ISCA), 2018, pp. 275–288. These functions can include operations like addition / subtraction, multiplication, division, and / or complex state machines.

[0005] On the other hand, many neurons in SNNs are analog. See, for example, A. Joubert, B. Belhadj, O. Temmam, and R. Héliot, “Hardware spiking neurons design: Analog or digital?” in The 2012 International Joint Conference on Neural Networks (IJCNN), 2012, pp. 1–5. Another example is J. M. Zurada, “Analog implementation of neural networks,” in IEEE Circuits and Devices Magazine, vol. 8, no. 5, pp. 36–41, 1992, doi:10.1109 / 101.158511. Also see A. Rubino, M. Payvand, and G. Indiveri, “Ultra-Low Power Silicon Neuron Circuit for Extreme-Edge Neuromorphic Intelligence,” in 2019 26th IEEE International Conference on Electronics, Circuits, and Systems (ICECS), 2019, pp. 458–461. Their advantage is that they can encode more value states in a smaller, more power-efficient way, but sometimes at the expense of signal-to-noise ratio (SNR). There are also phase-encoding neurons that take advantage of using phase as a continuous-time variable for computation [6]. See, for example, A. Madhavan, T. Sherwood, and D. Strukov, “Race logic,” ACM SIGARCH Computer Architecture News, vol. 42, pp. 517–528, Dec. 2014, doi: 10.1145 / 2678373.2665747. They also face the same trade-off as analog neurons, namely that phase-encoding neurons can suffer from worse SNR.

[0006] The benefits of digital neurons stem from the fact that, regardless of the function they implement, they store a defined state at every clock cycle. These states are well-defined values ​​that do not deteriorate over time. The problem with digital neurons stems from the fact that every time an operation is performed, a timing element, such as a voltage-controlled oscillator based on a phase-locked loop, is used to create a well-defined concept of time. The distribution of this clock network ultimately consumes the majority of the power. See, for example, S. Ali Butt, S. Schmermbeck, J. Rosenthal, A. Pratsch, and E. Schmidt, "System-Level Clock Tree Synthesis for Power Optimization," [IEEE Computer Society], 2007, p. 1682.

[0007] One extension of digital neurons that alleviates the clock distribution problem is based on using multiple phases of an oscillator to arbitrate incoming signals. See, for example, J. Stuijt, M. Sifalakis, A. Yousefzadeh, and F. Corradi, “μBrain: An Event-Driven and Fully Synthesizable Architecture for Spiking Neural Networks,” Frontiers in Neuroscience, vol. 15, Dec. 2021, doi: 10.3389 / fnins.2021.664208. This extension addresses the issues of avoiding phase lock and global clock distribution. This system is not scalable because expanding the system requires faster multiphase oscillators and more clock distribution. It has better power characteristics but also faces the same trade-offs as traditional digital neuron systems.

[0008] Another digital neuron implementation uses a handshake mechanism to address this trade-off and solve the problem from an asynchronous perspective. See, for example, A. Yousefzadeh et al., "Asynchronous Spiking Neurons, the Natural Key to Exploit Temporal Sparsity," IEEE Journal on Emerging and Selected Topics in Circuits and Systems, vol. PP, Dec. 2019, doi:10.1109 / JETCAS.2019.2951121. The problem with this approach is that scalability leads to latency, and the handshake consumes a significant amount of power. Summary of the Invention

[0009] In order to solve the above-mentioned problems, the subject matter of the present claims is proposed.

[0010] In a first aspect of the present invention, a spike capture and local trigger generation module for neurons in a spiking neural network is disclosed. The module may include multiple spike input ports. Each of the multiple spike input ports may be configured to receive a spike input signal from a spiking neural network. Furthermore, the module may include a pulse latch module. The pulse latch module may be configured to capture spikes in the spike input signal received by the spike input port through pulse latching. A pulse latch can be described as a digital circuit mechanism that captures and holds a state in response to a trigger pulse, thereby storing the state until reset by a subsequent pulse. This state can be forwarded to other parts of the module. The pulse latch module outputs a latch signal for each of the multiple spike input ports, indicating whether a spike was captured. Furthermore, the module may include a local trigger generator module. The local trigger generator module may be configured to generate a trigger signal used by the neurons to synchronize and / or control the timing of operations in the neurons. The local trigger generator module may be configured to generate a trigger signal using the multiple latch signals, such that the trigger signal contains one or more pulses equal to the number of spikes captured by the pulse latch module. The trigger signal thus obtained can be used to trigger operations performed by other parts of the neuron, thereby providing a trigger mechanism based on local events. Therefore, the present invention does not require a global clock signal to operate these parts of the neuron. In order to use the trigger signal, the module may include: a trigger signal output port configured to output the generated trigger signal.

[0011] In an embodiment of the first aspect, a trigger signal may be generated such that operations in the neurons can be triggered by the trigger signal rather than by a global clock signal of the spiking neural network.

[0012] In an embodiment of the first aspect, the trigger signal may be a delay-based asynchronous counter.

[0013] In an embodiment of the first aspect, one or more pulses included in the trigger signal may have a certain interval between pulses, and the interval may be determined by the minimum time required for the neuron to perform an operation. Preferably, the operation may be the accumulation of one or more weights.

[0014] In an embodiment of the first aspect, the module may further include a plurality of spike output ports, each port being configured to output a spike output signal. Each of the plurality of spike output ports may correspond to one of the plurality of spike input ports. If a particular spike input port receives a spike, the corresponding spike output port may be configured to output a pulse.

[0015] In an embodiment of the first aspect, a time interval of the pulses outputted by the spike output port may be substantially the same as a time interval of the pulses included in the clock signal.

[0016] In an embodiment of the first aspect, the pulse latch module may include a plurality of latch elements, each of which may be configured to capture one or more spikes in a spike input signal received by a specific spike input port through pulse latching.

[0017] In an embodiment of the first aspect, the latch element may be a CMOSD flip-flop. The CMOSD flip-flop may be triggered by a spike input signal. The latch signal output by the CMOSD flip-flop may be a bit value indicating whether the spike is latched.

[0018] In an embodiment of the first aspect, the local trigger generator module may include a pulse generator that generates a pulse. The local trigger generator module may include delay logic for each spike input port. The generated pulse may then be sent through each delay logic.

[0019] In an embodiment of the first aspect, each delay logic may include a delay element that delays the pulse signal. If the corresponding latch signal indicates that the corresponding spike input port has captured a spike, the generated pulse may be sent through the delay element of the specific delay logic. If the corresponding latch signal indicates that the corresponding spike input port has not captured a spike, the generated pulse may not be sent through the delay element of the specific delay logic.

[0020] In an embodiment of the first aspect, if the corresponding latch signal indicates that the corresponding spike input port has captured a spike, the specific delay logic may output the generated pulse to generate the trigger signal. If the corresponding latch signal indicates that the corresponding spike input port has not captured a spike, the specific delay logic may not output the generated pulse to generate the trigger signal.

[0021] In an embodiment of the first aspect, the local trigger generator module can be configured to generate a plurality of trigger signals used by neurons to synchronize and / or control the timing of operations of corresponding portions of the neurons. The local trigger generator module can be configured to generate the plurality of trigger signals using a plurality of latch signals such that, for a particular set of spike input ports, each trigger signal includes one or more pulses equal to the number of spikes captured by the pulse latch module. The module can also include a plurality of trigger signal output ports configured to output the generated trigger signals.

[0022] In an embodiment of the first aspect, the module may further include a multi-level trigger generator module having a multi-level trigger output port, configured to generate a multi-level trigger signal. The multi-level trigger signal may indicate the total time required for the neuron to perform an operation triggered by the trigger signal. Preferably, if multiple trigger signals are generated and output, the multi-level trigger signal may correspond to the maximum time required for the neuron to perform all sets of operations triggered by different trigger signals.

[0023] In a second aspect of the present invention, a weight encoding and accumulation module for accumulating spike input signals in neurons of a spiking neural network is disclosed. The module may include multiple spike input ports. Each of the multiple spike input ports may be configured to receive a spike input signal. In addition, the module may also include a trigger signal input port, which may be configured to receive a trigger signal. The pulse contained in the trigger signal may indicate the number of spikes received by the multiple spike input ports. In addition, the module may also include a spike-weight encoder module. The spike-weight encoder module may be configured to receive a spike input signal and output a corresponding weight associated with a particular spike input signal. This may be accomplished by encoding the received spike input signal into a weight. In addition, the module may include an accumulator module. The accumulator module may be configured to receive weights and a trigger signal from a storage module and accumulate the weights of the received spikes. The timing of one or more operations involved in the accumulation may be controlled by the trigger signal.

[0024] In an embodiment of the second aspect, each peak input signal may be an analog signal, and the weight may be a digital value.

[0025] In an embodiment of the second aspect, the accumulator module may include an accumulator element and a loop circuit, wherein the loop circuit receives an output of the accumulator element as an input and outputs a value back to the accumulator element each time the loop circuit receives a pulse from the trigger signal. When the spike-weight encoder module outputs the first weight to the accumulator element, the accumulator element may add the first weight to a second weight output by the loop circuit.

[0026] In an embodiment of the second aspect, the loop circuit includes a CMOS D flip-flop, which can be triggered by a trigger signal.

[0027] In one embodiment of the second aspect, the spike-weight encoder module may include a local memory, preferably an SRAM or a non-volatile memory. The spike-weight encoder may encode a pulse received at a specific spike input port into a memory address, the address being used to request a weight stored at the specific memory address. The weight may be associated with the specific spike input port.

[0028] In an embodiment of the second aspect, the spike-weight encoder module may include a distributed memory including distributed weight storage elements. Weights associated with the spike input ports may each be stored in one of the distributed weight storage elements. If a spike is present in the spike input signal, the corresponding weight may be output to the accumulator module by the distributed weight storage element storing the weight.

[0029] In an embodiment of the second aspect, the spike-weight encoder module may be configured to change the order of weight outputs based on a priority rule.

[0030] In a third aspect of the present invention, a spiking neuron of a spiking neural network is disclosed. The neuron may include the spike capture and local trigger generation module according to the first aspect. The neuron may additionally or alternatively include the weight encoding and accumulation module according to the second aspect. The trigger signal output port of the spike capture and local trigger generation module may output a trigger signal to the trigger signal input port of the weight encoding and accumulation module. This may be achieved by controlling the timing of one or more operations involved in the accumulation by the trigger signal.

[0031] In an embodiment of the third aspect, the neuron may further include a comparator module that compares a total accumulated weight to a predetermined threshold. The total accumulated weight may be the sum of all weights of the spike input ports that capture a spike corresponding to the spike input signal. The spiking neuron may be configured to output a spike to the spiking neural network when the total accumulated weight exceeds the predetermined threshold.

[0032] In an embodiment of the third aspect, the spiking neuron may include multiple weight encoding and accumulation modules according to the second aspect. The spike capture and local trigger generation module may be based on the first aspect. Each of the multiple trigger signal output ports may output a corresponding trigger signal to a trigger signal input port of a corresponding weight encoding and accumulation module, so that the timing of one or more operations involved in the accumulation performed in the corresponding weight encoding and accumulation module can be controlled by the corresponding trigger signal.

[0033] In an embodiment of the third aspect, the neuron may further include a digital weight accumulator module that accumulates outputs of at least a portion of the plurality of weight encoding and accumulation modules and outputs an accumulated value.

[0034] In an embodiment of the third aspect, the spiking neuron may include a plurality of digital weight accumulator modules, wherein at least one digital weight accumulator module may accumulate accumulated values ​​output by other digital weight accumulator modules.

[0035] In an embodiment of the third aspect, a delay-based pipeline and / or a handshake-based pipeline can be used to synchronize accumulation operations between multiple weight encoding and accumulation modules and / or one or more digital weight accumulation modules.

[0036] In an embodiment of the third aspect, the spike capture and local trigger generation module may be based on the first aspect and may use a multi-level trigger signal to synchronize accumulation operations between multiple weight encoding and accumulation modules and / or one or more digital weight accumulation modules.

[0037] In a fourth aspect, a method for spike capture and local trigger generation in a neuron of a spiking neural network is disclosed. The method may include: receiving a plurality of spike input signals from the spiking neural network; capturing spikes in the spike input signals by pulse latching; generating a latch signal for each spike input signal, wherein the latch signal may indicate whether a spike was captured; generating a trigger signal using the plurality of latch signals, such that the trigger signal may include one or more pulses equal in number to the number of captured spikes, wherein the trigger signal may be used by the neuron to synchronize and / or control the timing of operations in the neuron; and outputting the generated trigger signal.

[0038] In an embodiment of the fourth aspect, the method may further include: outputting a spike output signal. A trigger signal may be generated such that operations in the neuron can be triggered by the trigger signal rather than by a global clock signal of the spike neural network.

[0039] In a fifth aspect, a method for weight encoding and accumulation in neurons of a spiking neural network is disclosed. The method may include: receiving a plurality of spiking input signals; receiving a trigger signal comprising one or more pulses, wherein the one or more pulses indicate a number of pulses received in the plurality of spiking input signals; obtaining or generating a weight associated with a spiking input signal in the plurality of spiking input signals; accumulating the weights of the spiking input signal comprising the spikes, wherein the timing of one or more operations involved in the accumulation may be governed by the trigger signal; and outputting the accumulated weights.

[0040] In an embodiment of the fifth aspect, the weight may be obtained or generated by encoding a pulse received in one of the plurality of spike input signals as a memory address, which memory address may be used to request a weight stored at the particular memory address in a local memory.

[0041] In an embodiment of the fifth aspect, if a spike is present in the spike input signal, a corresponding weight may be obtained from a distributed memory comprising a plurality of weight storage elements. The weights associated with the spike input signal may each be stored in one of the distributed weight storage elements. If a spike is present in the spike input signal, the corresponding weight may be output by the distributed weight storage element storing the weight and used for weight accumulation.

[0042] In a sixth aspect, a method for spike capture and accumulation in a neuron of a spiking neural network is disclosed. The method may include: receiving a plurality of spiking input signals from a spiking neural network; capturing spikes in the spiking input signals by pulse latching; outputting a latch signal for each spiking input signal, wherein the latch signal may indicate whether a spike is captured; using the plurality of latch signals to generate a trigger signal, such that the trigger signal may include one or more pulses equal in number to the number of captured spikes, wherein the neuron may use the trigger signal to synchronize and / or control the timing of operations in the neuron; obtaining or generating weights associated with one or more of the plurality of spiking input signals; accumulating weights of the spiking input signals containing the spikes, wherein the timing of one or more operations involved in the accumulation may be governed by the trigger signal; and outputting the accumulated weights.

[0043] In an embodiment of the sixth aspect, the method may further include comparing a total accumulated weight to a predetermined threshold. The total accumulated weight may be the sum of all weights of the spike input signal that captured the spike. Furthermore, the method may further include outputting the spike to the spiking neural network when the total accumulated weight exceeds the predetermined threshold.

[0044] In an embodiment of the sixth aspect, the steps of obtaining or generating weights and accumulating weights can be parallelized by dividing the plurality of spike input signals into one or more subsets of the plurality of spike input signals. Furthermore, the steps of obtaining or generating weights and accumulating weights for the subsets can be performed simultaneously, so that an accumulated subset weight is obtained for each subset.

[0045] In an embodiment of the sixth aspect, the accumulated subset weights may be accumulated digitally, preferably in parallel.

[0046] In an embodiment of the sixth aspect, a delay-based pipeline and / or a handshake-based pipeline is used to synchronize accumulation operations between different subsets and / or one or more digital accumulations. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Embodiments will now be described, by way of example only, with reference to the accompanying drawings, in which corresponding reference numerals indicate corresponding parts, and in which:

[0048] Figure 1 A schematic diagram showing repeatable design blocks connected in an M×N structure;

[0049] Figure 2 A schematic diagram of a design method according to the present invention is shown;

[0050] Figure 3 Schematic diagram showing the high-level architecture of a neuron with m spiking inputs;

[0051] Figure 4 A schematic diagram of capturing m spike inputs and their associated timing generation is shown;

[0052] Figure 5 Schematic diagram showing summing spike weights from local memory in a unit cell;

[0053] Figure 6 Schematic diagram showing summing spike weights from distributed memory in a unit cell;

[0054] Figure 7 Schematic diagram showing a multi-level self-timed neuron design with repeatable unit cells;

[0055] Figure 8 A schematic diagram showing a delay-based pipeline; and

[0056] Figure 9 A schematic diagram of a pipeline based on handshake control logic is shown. DETAILED DESCRIPTION

[0057] Some embodiments will be described in more detail below. However, it should be understood that these embodiments should not be construed as limiting the scope of protection of the present disclosure.

[0058] Self-timed circuits have been explored in the literature for various systems and high-performance designs. See, for example, S. Fairbanks, “High Precision Timing Using Self-Timed Circuits,” 2009. These can be implemented using various circuit design techniques, such as self-resetting logic (see Bloker et al., “US Pat. No. 5,864,251”) and pulsed self-timed logic (see M. Miller, C. Segal, D. McCarthy, A. Dalakoti, P. Mukim, and F. Brewer, “Impolite High-Speed ​​Interfaces with Asynchronous Pulse Logic,” December 2018, pp. 99–104). This type of logic design has no associated clock. The idea is to combine multiple repeatable logic patterns in a systematic way so that each logic block can complete within a given delay.

[0059] In order to realize these designs, a design methodology must be established to carry out larger designs.

[0060] In order to reduce power consumption in digital designs and operate designs near the threshold of transistor devices and at extremely low supply voltages, source-coupled logic has been proven to be an effective method. See, for example, S. Roy and K. Nipun, “Understanding sub-threshold source coupled logic for ultra-low power application.” Integrating source-coupled logic into neuron logic and performing delay design can further optimize the area and power consumption of the design. Self-timed logic has also been shown to optimize the timing performance of CMOS designs. See, for example, A. Dalakoti, M. Miller, and F. Brewer, “Pulse Ring Oscillator Tuning via Pulse Dynamics,” published in the 2017 IEEE International Conference on Computer Design (ICCD), 2017, pp. 469–472. The inventors of the present invention realized that this can help optimize the design to create neurons of different time scales. Delay-based competitive calculations can also be repeatable and can help create power-efficient logic. See, eg, A. Madhavan, T. Sherwood, and D. Strukov, “Race logic,” ACM SIGARCH Computer Architecture News, vol. 42, pp. 517–528, Dec. 2014, doi: 10.1145 / 2678373.2665747.

[0061] The inventors of the present invention have recognized that techniques such as source-coupled design, self-timed logic, and delay-based competition logic can be incorporated into a design methodology to create complex neuron structures. The benefit of this design methodology comes from the fact that these structures can be synthesized based on use case requirements, using repeatable blocks that have been pre-verified and integrated into the design methodology.

[0062] All traditional digital techniques, such as time multiplexing, memory optimization, etc., can be applied to the design and can be further optimized with respect to area, power or timing trade-offs. The design can also be enhanced with complex connections to create complex SNN structures based on the needs of each use case.

[0063] Figure 1 Schematic diagram showing a repeatable design block 101 connected in an M×N structure 100, Figure 2 Schematic diagram of a design method 200 according to the present invention is shown. Specifically, Figure 1 A design block 101 is shown, which can be Figure 2The design method 200 shown is obtained to create the final connections of the neurons in an automated manner.

[0064] Figure 1 The repeatable design blocks 101 shown connected in an M×N structure 100 refer to a modular architecture, where M represents the number of rows and N represents the number of columns of interconnected design blocks 101, and the entire M×N structure 100 constitutes an SNN (at least a portion thereof). The design blocks 101 may include a single neuron, multiple connected neurons and / or more complex circuits. The M×N structure 100 can facilitate parallel processing and can follow modular design principles through the use of design blocks. The connections between the design blocks 101 can represent synaptic connections in the network. The design blocks can be connected using, for example, wire connections, crossbar buses with logic (e.g., NAND trees), tri-state buses and / or routers. Note that the design blocks 101 used in the M×N structure 100 can be similar or different from each other, and can also vary per row and / or per column.

[0065] The M×N structure 100 promotes modular design and is easy to copy and expand. Each block 101 can be a self-contained module with a specific function, and the entire network can be constructed by copying and connecting these modules. For example, different design blocks 101 can be specialized for certain specific tasks or functions. Then, by adjusting the connectivity and properties of the individual design blocks, the entire SNN can be adapted to various applications. The same design block 101 can be replicated on different parts of the network, thereby promoting consistency and simplifying the design process. This is beneficial to both the hardware and software implementations of the SNN. The grid structure can promote parallel processing within the network because design blocks (e.g., neurons) 101 in the same row or column can operate simultaneously, thereby improving the overall computational efficiency of the network.

[0066] Depending on the specific requirements of the task or application, the connection pattern between design blocks 101 within the M×N structure 100 can vary. Common connection patterns include nearest neighbor connections, all-to-all connections, or more complex connection schemes. The modularity and repeatability of the design make it easy to expand. As the requirements of the SNN grow, additional rows or columns of design blocks can be added to increase the capacity of the SNN without fundamentally changing the architecture.

[0067] A design block may include a unit cell 102 with a delay circuit 103 and / or a logic circuit 104. These circuits 103, 104 will be discussed further below, and each circuit may have an input port, an output port, and / or a feedback loop. The unit cell 102 may, for example, form a neuron.

[0068] Next, we will discuss how Figure 2The design method 200 shown. In general, the design method 200 according to the present invention uses inputs 201, performs design steps 202 on these inputs 201, and obtains outputs 203 of the design method 200. The output 203 may, for example, disclose the configured network architecture, in particular, for example Figure 1 The M×N structure shown.

[0069] The input 201 to the design methodology may be, for example, primitives 201A, logic blocks RTL 201B, delay blocks 201C, connectivity requirements 201D, and / or timing constraints 201E.

[0070] Primitive 201A is a design block that can be used to design complex logic circuits for register transfer level (RTL) blocks. RTL is an abstraction level in digital circuit design that represents the circuit in terms of data flow between registers and the logical operations performed between these registers. An RTL block can represent a portion of a neural network hardware design described at the register transfer level. An RTL block can capture the flow of data within the hardware and the operations performed on that data. Different types of primitive designs are possible, depending on the type of logic used based on power, performance, and area requirements.

[0071] Some example primitive designs 201A that can be used as input 201 in the design method 200 are: (a) SR CMOS (see, for example, Bloker et al., “US5864251”); (b) pulse logic (see, for example, D. McCarthy, M. Miller, and F. Brewer, “Automated Timing Constraint Generation for PulseGate Circuits,” in IEEE International Conference on Program Comprehension, 2022, vol. 2022-March, pp. 36–47); (c) standard CMOS (see, for example, I. Sourikopoulos et al., “A 4-fJ / spikeartificial neuron in 65nm CMOS technology,” in Frontiers in Neuroscience, vol. 11, no. MAR, Mar. 2017, doi: 10.3389 / fnins.2017.00123); and / or (d) source-coupled CMOS logic (see, for example, S. Roy and K. Nipun, “Understanding sub-threshold source coupled logic for ultra-low power application”).

[0072] Another input 201 can be a logic block RTL 201B. Neuron blocks can be composed of various mathematical operations. These operations can be captured as standard RTL descriptions and can be passed to the synthesis process. In RTL design, registers store data between different computational stages. In the case of SNNs, these registers can represent neuron states, synaptic weights, or other relevant information at different time steps. The RTL description can detail how data is transferred between registers and how various computational operations (e.g., spike generation, synaptic updates) are performed. The RTL description can also include timing information, indicating when various operations occur. In SNNs, this is very important for capturing the temporal dynamics of spikes and synapses. RTL blocks can specify operations specific to spiking neurons within the SNN, such as spike generation, spike propagation, synaptic weight updates, and any other computations related to the spiking behavior of neurons. Depending on the design, the RTL description can also capture parallelism within the SNN, especially when the hardware implementation involves parallel processing of multiple neurons or synapses.

[0073] Another input 201 may be a delay block 201C. Delay block 201C can be implemented in a variety of ways. Delays can be derived from standard CMOS gates or through specialized analog designs to create variable delays. Delay blocks are used to model the transmission delays associated with synaptic connections between neurons. These delays can be useful in capturing the temporal dynamics of spike events. Delay block 201C can be implemented in a variety of ways, with the choice of implementation depending on the specific requirements of the neural network and hardware platform. Examples of digital delay blocks include shift registers and digital counters. Each stage of a shift register represents a unit of time, and the signal shifts through the register to introduce a delay. A digital counter can be used to count the number of time steps before allowing the signal to pass. This allows for programmable delays based on the count value. Examples of analog delay blocks include RC circuits and transmission lines. Analog delay blocks can be used to introduce variable delays. An RC circuit consisting of a resistor (R) and a capacitor (C) can be used to create a variable delay by adjusting the values ​​of the resistor and capacitor. The charging and discharging of the capacitor introduces a time constant, which determines the delay. Analog delay lines (e.g., transmission lines) can also be used to introduce delays. The signal travels along a transmission line, and the length of the transmission line determines the delay.

[0074] Another input may be connection requirements 201D. In real use cases, the connections between different logic blocks may vary. In order to correctly implement the design, the connection requirements need to be captured in a separate file, such as a connection description file based on regular expressions.

[0075] Another input may be timing constraints 201E. In order for a design to be synthesizable, all designs need to be constrained. For example, all logic block paths need to be constrained. Timing constraints 201E can be implemented by bounding the time required for the logic to perform a particular operation. Examples of timing constraints 201E may be: (a) minimum and maximum delays in a path; (b) minimum and maximum capacitance in a path, e.g., the charging and discharging of a capacitor introduces a time constant that determines the delay; and / or (c) a structural floorplan that satisfies the constraints. Regarding the latter point (c), different design choices (e.g., delay size, complexity, logic, etc.) will result in blocks of different sizes. It is possible that not all blocks are compatible to be placed adjacent to each other to achieve maximum utilization of area in the design. In order to find a combination of blocks that can be placed adjacent to each other to achieve the best area utilization, intelligent floorplanning is required, which can be achieved by the described method.

[0076] Then, based on the input 201 , the different method steps 202 of the design method 200 are performed.

[0077] As a first step, the primitives used in complex neuron logic need to be characterized with respect to delay. This is done during primitive characterization 202A, preferably in an automated manner, to create the necessary sequence files required for synthesis and placement and routing tools. It is necessary to characterize the scaled voltage library with respect to near-threshold calculations. A scaled voltage library is typically a set of libraries in (digital) integrated circuit design that includes components with different threshold voltages (Vt) or supply voltages (Vdd). These libraries are designed to support the creation of circuits that operate at multiple voltage levels. Near-threshold calculations refer to the operation of a digital circuit at or near its threshold voltage level. The threshold voltage is the voltage when a transistor switches from an off state to an on state. Operating near the threshold voltage is often associated with low-power design strategies in digital integrated circuits. That is, compared to traditional high-voltage operation, reducing the supply voltage and operating near the threshold voltage can significantly save power. Therefore, near-threshold calculations are particularly relevant in mixed-signal and analog / mixed-signal designs where power efficiency is crucial.

[0078] Next, as a second step, the logic of the digital neuron needs to be characterized during logic characterization 202B. The goal is to create an optimized neuron logic block and then use it in a time- and space-reuse manner. In order to use the logic block multiple times, a macro IP of the logic block is built based on the use case connectivity and then implemented. IP blocks (often called IP cores or simply IP) are pre-designed, pre-verified, and reusable functional units that can be integrated into larger systems or designs. These IP blocks encapsulate specific functions or logic and are designed to be easily integrated into various projects to save time and effort. The design process implements block-level and top-level automation.

[0079] As a third step, since there are different types of delay design methods, each method needs to be characterized during Delay Characterization 202C. The correct delay is selected based on the logic type and logic delay. Different delays have different trade-offs associated with them. Some examples are: (a) CMOS delays, which are standard delay elements in process development kits; or (b) current-starved delays, which are area-optimized delays controlled by widely tunable digital-to-analog converters. CMOS delay refers to the time it takes for a signal to propagate through a complementary metal-oxide-semiconductor (CMOS) logic gate, from its input to its output. The size and characteristics of n-type (NMOS) and p-type (PMOS) transistors affect delay. For example, larger transistors generally have higher drive strength, but may also have higher capacitance. The capacitance at the gate's output (also known as load capacitance, which includes the parasitic capacitance of the connecting leads and the input capacitance of the subsequent gate) can also affect delay. Current-starved delay is a design technique in which delay is primarily controlled by regulating the current flowing through the circuit and intentionally limiting this current to achieve the desired delay. This can be applied, for example, to ring oscillators and delay-locked loops within networks.

[0080] As a fourth step, time and space multiplexing 202D can be performed. That is, the logic and connections need to be converted into a hardware implementation with time and space multiplexing. Time multiplexing leads to increased latency and power usage. Spatial distribution leads to an increase in the area of ​​the network implementation. The right design point needs to be selected (design technology specific) and an automatic compiler can generate blocks to be placed within the routed network. Quantization can be based on the power, performance and area (PPA) values ​​we are trying to achieve. Typically, people try to use optimization algorithms (e.g., convex optimization or simulated annealing) to find the optimal solution.

[0081] As a fifth step, power, performance, and area (PPA) are optimized during PPA optimization 202E. These three parameters can be optimized, for example, by: (a) multiplexing-dependent area; (b) technology node; (c) mixed-signal design control (e.g., current-hungry delay elements); (d) design methodology (e.g., source-coupled logic or self-reset logic); (e) pipelining; and / or (f) properly selected signal transmission between spatially and temporally multiplexed designs.

[0082] After performing these method steps, outputs 203 of the design method 200 are obtained, which may include: (a) a final layout 203A, which is a final GDS with a complete digital neuron structure connected in a desired network structure; (b) a gate-level netlist 203B, which is a gate-level netlist with all primitives characterized in the flow; and / or (c) a PPA report 203C, which is a report on the expected performance, power, and area for the created use case-driven design.

[0083] The final layout 203A can be a physical representation of the SNN design on the semiconductor substrate. It includes the placement of transistors, interconnects, metal layers, and other physical components to form synapses and neurons. It provides a map of where the various components of the SNN will be located on the chip. "GDS" stands for "Graphics Data System" or "Graphics Design System," which refers to a standard file format for describing the geometric layout of various components and layers of an IC. Semiconductor foundries can use GDS files to manufacture the designed IC.

[0084] The gate-level netlist 203B is a textual representation of the SNN design at the gate level. It describes the logical connectivity and functionality of the design using basic logic gates (AND, OR, NOT, etc.). Each row in the netlist typically represents a gate or flip-flop, and the connections between these elements are explicitly specified. The netlist serves as an intermediate representation that can be used for simulation, verification, and synthesis.

[0085] The PPA report 203C provides detailed information on how the SNN design performs in terms of power consumption, speed (performance), and the physical area it occupies on the chip. Power consumption metrics can include dynamic power (associated with switching activity), static power (leakage), and total power consumption. Performance metrics can involve delay characteristics and clock frequency. Area metrics provide insight into the size of the design.

[0086] Several design examples that may be used with or obtained through the design method 200 described above are discussed below.

[0087] One of the fundamental logical operations performed by neurons is accumulation. Accumulation is the process by which a neuron integrates incoming input spikes over time, ultimately determining whether the neuron generates an output spike. Accumulation captures the temporal dynamics of information processing within the network. Accumulation is performed based on the weights associated with a particular input spike received from the network interconnect or via the interconnect from an encoder or sensor.

[0088] A common neuron design is the leaky integrate-and-fire (LIF) neuron. LIF neurons have a membrane potential that evolves over time based on incoming synaptic input and a leak term. Accumulation occurs in the neuron's membrane potential. The membrane potential represents the electrical potential across the neuron's membrane and is influenced by the synaptic input it receives. Each incoming spike, for example from a connected neuron, contributes to the accumulation of the membrane potential. The synaptic weights determine the strength of these contributions. Excitatory synapses increase the membrane potential, while inhibitory synapses decrease it. Leakage can cause the membrane potential to decay slowly over time. This causes the neuron's membrane potential to return to its resting state in the absence of input. When the accumulated membrane potential reaches a certain threshold, the neuron fires, or generates an output spike. This threshold is a key parameter that influences the neuron's responsiveness and sensitivity. After firing, the membrane potential can be reset to its resting or reset value, and the neuron can undergo a refractory period, during which it cannot immediately fire again.

[0089] The accumulation performed within a neuron can take different forms depending on the interpretation of the weights within the neuron's logic. Each bit of the weight can have a different logical function encoded within it. Some examples of logical operations in neuron accumulation include (a) addition, (b) subtraction, and / or (c) left and right shifts.

[0090] The addition and subtraction logic operations mimic the accumulation of excitatory and inhibitory spikes, respectively. A left shift involves updating an accumulator or shift register by discarding the least significant bit or the oldest value in the register. A left shift mimics the natural decay or leakage of accumulated information over time. For leaky integrate-and-fire (LIF) neurons, it simulates the gradual decline in membrane potential when no new input spikes are received. A right shift involves updating an accumulator or shift register by discarding the most significant bit or the most recent value in the register. The right shift operation can be used in specific scenarios to introduce a time delay in the accumulation process. It effectively delays the contribution of the latest input value to the accumulation, allowing the neuron to respond to past input patterns.

[0091] Figure 3 A schematic diagram of the high-level architecture of a neuron 300 with m spiking inputs is shown, specifically using the self-timing method described in the previous section. The goal is to capture the m input spikes 306 from the network and generate timing information locally in a feedforward manner so that the weights associated with the spikes' inputs can be accumulated. The sum of the weights is sent to an accumulator 302, which is connected to comparator logic 304 with a threshold 305. When the accumulator 302 reaches the programmed threshold 305, an output spike 307 is generated and the accumulator is reset 308.

[0092] As shown, m input spikes 306 enter neuron 300, followed by spike capture and timing generation in spike capture and timing generation module 301. Each incoming spike carries information and is typically associated with a synaptic weight. The spike capture mechanism of a neuron involves detecting the arrival of an incoming spike. This detection can be based on a comparison with a certain voltage threshold. Timing generation involves creating a local sequence of timing pulses. The capture mechanism and timing generation will be explained further below. Next, weights 303 are applied to the different captured spikes, and these spikes are accumulated in accumulator 302. The accumulator sends the accumulated signal to comparator logic 304, where the accumulated signal is compared with a threshold 305. Threshold 305 can be fixed or variable. When the accumulated signal reaches threshold 305, an output spike 307 is generated, which is sent to (another part of) the network or the output layer. Once the output pulse 307 is generated and the neuron "fires", the accumulator 307 may be reset using, for example, a reset signal from the comparator 304. This will cause the accumulated signal in the accumulator 302 to return to zero or other predetermined level.

[0093] Figure 4 A schematic diagram of the capture of m spike inputs 401 and the associated timing generation in the spike capture and timing generation module 400 is shown, specifically, an example of implementing the capture of spikes received from the network, and more specifically, an example of how the spike capture and timing generation module 301 can be implemented.

[0094] The captured spikes can be used to generate local and / or multi-level timing information at the neuron itself, since the system does not need to be based on or use a clock. Spikes can be captured using a pulse latch in the pulse capture module 402, such as an m-latch. An example of such a pulse latch is a standard CMOS D flip-flop 403. However, spike capture can be performed using any other logic method, and this is merely an example. In general, the spike capture of the pulse latch can generate a logic state from the pulse, such that the generated logic state indicates whether the pulse occurred. A variety of different spike capture methods can be used in the same pulse capture module 402, within the same neuron, or within a network.

[0095] A CMOS D flip-flop (DFF) 403 is a digital circuit element that stores a single bit of information. The "D" in the CMOS D flip-flop stands for "data." The flip-flop may include a data input (D) 4032, a clock input (CLK or CP) 4033, a set input (S), a reset input (R) 4031, and / or complementary outputs (Q and Q-bar) 4034a and 4034b. When triggered by a clock signal received at the CLK input, the DFF outputs the data input signal of the D input at one or both of the complementary outputs Q and Q-bar.

[0096] A CMOS D flip-flop 403 can be used to capture spike signals for a neuron by serving as part of a circuit that processes and integrates incoming spike events. The D flip-flop acts as a spike capture element, where the D (data) input 4032 of the flip-flop is set to logic 1 and the CLK input 4033 is connected to a specific spike input signal 401 (e.g., from a specific input synapse of a neuron).

[0097] The clock input of the flip-flop is driven by a clock signal, in this case the corresponding spike input signal corresponding to the corresponding DFF is used as the clock signal of that DFF. This means that if a specific spike signal is sent to the neuron, the corresponding DFF captures the logic 1 signal at the specific edge of the spike input signal and passes it to the output Q and Q-bar.

[0098] Q and Q-bar 4034a, 4034b are the logical complements of each other. If Q is high (logic 1), Q-bar is low (logic 0), and vice versa. The two outputs always have opposite logic values. The complementary nature of Q and Q-bar 4034a, 4034b allows for the convenient generation of inverted signals without the need for additional logic gates. The S and R inputs are asynchronous inputs that can force Q output 4034a to a high or low state, respectively, regardless of the clock and data inputs.

[0099] The output of each latch, such as the Q output 4034a of each CMOS D flip-flop, can be connected to the control input of a demultiplexer (DMUX) 405. A demultiplexer is a digital circuit that receives a single input and directs it to one of several possible outputs based on one or more control signals. Each DMUX can be controlled by the output of a DFF, i.e., the output value of the DFF determines to which output the DMUX directs its input signal. In this case, the input signal 404 of the first demultiplexer (controlled by the first latch connected to the first spike input port) can be a signal containing a single pulse.

[0100] If a spike is captured via the i-th input (where i∈{1,…,m}), the output of the i-th DMUX may be sent through the delay element 406 and then continue to the (i+1)th demultiplexer controlled by the latch connected to the (i+1)th input. If a spike is not captured via the i-th input (where i∈{1,…,m}), the output of the i-th DMUX may be sent to the (i+1)th demultiplexer controlled by the latch connected to the (i+1)th input without passing through the delay element 406.

[0101] For i∈{1,…,m}, if a spike is captured, the signal 410 from the corresponding i-th delay element iIt is also forwarded to a shared OR gate element 407, which combines the signals from all delay elements that captured the spike into a series of timing pulses, namely the local timing signal 408. Preferably, the shared OR gate element is a pulse tree element that combines signals containing pulses carried on different input lines into a single output line that pulses whenever one of the different input lines pulses. The local timing signal 408 can be used by neurons to synchronize and / or control the timing of operations in the neuron, thereby acting as a local trigger signal. The local timing signal 408 contains a number of pulses equal to the number of spikes captured by the pulse latch module 402. The spacing between the pulses generated by the delay element 406 can be based on the minimum time required to accumulate for a specific weight associated with a specific input spike line, as will be explained further below.

[0102] A delay element 406 may be present to delay the local timing signal 408 so that the local timing signal 408 reaches the next portion of the neuron at a determined time.

[0103] In other words, in this embodiment, a spike capture and local trigger generation module 400 for neurons in a spiking neural network is disclosed. The module may include multiple spike input ports 401. Each of the multiple spike input ports may be configured to receive a spike input signal from a spiking neural network. In addition, the module may include a pulse latch module 402. The pulse latch module may be configured to capture spikes in the spike input signal received by the spike input port through pulse latching. A pulse latch can be described as a digital circuit mechanism that captures and holds a state in response to a trigger pulse, thereby storing the state until reset by a subsequent pulse. The state can be forwarded to other parts of the module. The pulse latch module may output a latch signal for each of the multiple spike input ports, indicating whether a spike was captured. In addition, the module may include a local trigger generator module. The local trigger generator module may be configured to generate a trigger signal used by the neurons to synchronize and / or control the timing of operations in the neurons. The local trigger generator module can be configured to generate a trigger signal using multiple latch signals, such that the trigger signal includes one or more pulses equal to the number of spikes captured by the pulse latch module. The trigger signal thus obtained can be used to trigger operations performed by other parts of the neuron, thereby providing a trigger mechanism based on local events. Therefore, the present invention does not require a global clock signal to operate these parts of the neuron. To use the trigger signal, the module can include a trigger signal output port 408 configured to output the generated trigger signal.

[0104] Each signal 410 from the delay element iThe spike capture and timing generation module 400 may also output to different parts of the neuron. For example, the spike capture and timing generation module 400 may output an output signal 410 indicating that the corresponding spike input 401 contains a spike. i In this embodiment, since these output signals 410 i are delayed by the delay element 406, so these output signals 410 i Each will have a different delay and therefore be output one after another by the spike capture and timing generation module 400 .

[0105] Another option is that the spike capture and timing generation module 400 directly forwards the input signal 401 to the output signal 410 i It is also possible that each latch element 4033 controls a pulse generator, and if the latch element 4033 determines that there is a spike in the input signal, the pulse generator switches to a specific output 410. i Therefore, each output 410 i May have their own separate pulse generator, or multiple outputs 410 i A pulse generator can be shared between them.

[0106] The local timing signal can also be determined in other ways. For example, different latch elements 4033 can detect whether a spike occurs in a particular input channel, and the number of latch elements that detected a spike can be determined, for example, via a counter element. The pulse generator can then be controlled to output a local timing signal 408 to the spike capture and timing generation module 400, with a number of pulses equal to the number of detected input spikes.

[0107] The signal inputs 401 may also be grouped into different groups of various numbers of signal inputs 401. Then, a local timing signal 408 may be generated for each group of signal inputs 401, for example, in one of the above-described exemplary ways.

[0108] The spike capture and timing generation module 400 may also not output the spike output signal, but instead only generate the local timing signal 408 .

[0109] If i=m, the output of the mth DMUX (or the corresponding delay element) can be sent via the OR gate element 407 to be output as a multi-level timing signal 409, which will be described in more detail below. The multi-level timing signal 409 is generated to mark the end of the timing generation of the summation at a particular level. The delay of the multi-level timing must be greater than the delay of all summations completed in a particular level. Each level will be further explained below. Depending on how the accumulation process is performed in the neuron, it may be necessary to generate more than one multi-level timing signal.

[0110] Figure 5Schematic diagram showing the summation of spike weights from local memory in a unit cell 500, specifically, the figure shows k spike inputs 5101-510 out of a total of m spike inputs to the neuron k (k can be less than or equal to m). Thus, in this case, unit cell 500 is the part of the neuron responsible for summation and has k of the m spiking inputs in this example. In this case, the spiking inputs 5101-510 k Connect to Figure 4 The corresponding spike output 410 from the spike capture and timing generation module 301 is shown in FIG. i Thus, by using these unit cells, the m inputs of the accumulator are divided into one or more groups of k inputs. The reason for dividing the m inputs into multiple groups of k inputs is the trade-off between the following.

[0111] That is, the larger the gate count of a combinational logic operation (such as addition), the greater the uncertainty in the maximum time required to complete the operation. Design methodologies are based on maximizing time, so the greater the uncertainty, the slower the overall design becomes because the delays in the design need to be kept for the maximum possible time based on the timing uncertainty. Furthermore, the larger the delay, the more uncertain the delay becomes. So, this uncertainty is superimposed on the uncertainty of the combinational logic delay. The goal is to break the design into logic blocks and then use them as repeatable design blocks (such as Figure 1 These repeatable blocks need to be timed within the same layer and between different layers as shown below.

[0112] For Figure 5 The incoming spikes are summed as shown, and the input expected to have a spike is high (whether a spike occurs can be determined using Figure 4 , determined by the logic in the latch module), and the correct number of pulses is generated to time the accumulation operation of all weights associated with the input with the spike.

[0113] Enter 5101-510 k Entering encoder 501, encoder 501 encodes the spikes from the input into memory addresses. The encoder can also reorder the summation sequence (if necessary), for example based on certain priority rules. Encoder 501 can also perform operations such as delaying the summation by a delay period so that the summation can be performed with the next set of inputs to arrive. Encoder 501 can input each pulse received into 5101-510 k The memory address is passed to the local memory 502, such as SRAM or non-volatile memory like ReRAM.

[0114] Stored in the received input 5101-510 kThe weight in the memory corresponding to the address of each pulse in is passed to the accumulation element 503. The accumulator performs the function of adding / subtracting the input from the last saved value. The combination of 503 and 504 achieves this. In this case, 503 can be an adder, a subtractor, or both in standard CMOS logic. Thereafter, the signal from the accumulation element 503 is passed to the D input 5042 of the CMOS D flip-flop 504, for example. The local timing signal 508 can also be passed as an input to the CMOS D flip-flop 504, namely, the CLK input 5043. The DFF 504 captures the incoming accumulated weight signal at a specific edge of the clock signal corresponding to the pulse in the local timing signal 508.

[0115] The DFF has a feedback loop from its output (e.g., from the Q port 5044a) back to the summing element 503. In this way, the accumulated weight is output by the DFF and used as the second input of the summing element 503. At the same time, the summing element obtains the next weight from the local memory 502 and adds the new weight to the accumulated weight to obtain a new accumulated weight. The new accumulated weight is then sent to the DFF 504 again. Since the local timing signal 508 contains a number of pulses equal to the number of times the signal is received from the summing element 503, all signal inputs 5101-5103 that receive spikes are equal. k The weights are accumulated in this way.

[0116] In other words, the local timing signal 508 contains multiple pulses corresponding to the incoming spike signal. Each pulse is used to determine when to capture a specific (accumulated) weight from the local memory 502. The captured weight is integrated over time through successive clock cycles. The DFF 504 accumulates the presence of the spike over multiple clock cycles (e.g., after each trigger from the local timing signal 508), thereby providing a form of time integration. In this way, operations corresponding to the weights can be performed on the logic encoded in the accumulator. As described above, all captured input spikes 5101-510 k This process is repeated. Multiple unit cells can perform this function within a neuron, ensuring that all m spike signals 4101-410 m Accumulation occurs. Local timing pulses 408 and 508 are expected to be delayed enough from the start of the pulse to allow the spike to be processed in the accumulator (e.g., DFF 504). The start of the pulse is used to convey that local timing pulse 408 has additional delay in it. This additional delay can be used to control the relative timing of the time at which logic arrives at, for example, local memory 502 or accumulation element 503 relative to the timing of local timing pulse 508.

[0117] Figure 6Schematic diagram showing the summation of spike weights from distributed memory in a unit cell. Specifically, Figure 6 Shown Figure 5 Another implementation of weight accumulation is shown.

[0118] Enter 6101-610 k into the priority encoder 601, but in this case each received pulse may be input 6101-610 k The address is passed to one of the distributed storage elements 607a, which is j bits in size and stores therein the address corresponding to the particular input 6101-610 k If no input spike is received at a particular input, this corresponds to a logic 0 signal through the corresponding channel 607b. If an input spike is received at a particular input, this corresponds to a logic 1 signal through the corresponding channel 607b. Next, the weight from 607a and the signal from channel 607b pass through AND gate element 607c. If channel 607b carries a logic 1, the weight of the input is forwarded to multiplexer 602. The signal from AND gate element 607c is then ANDed to all inputs 6101-610 k All corresponding outputs pass through the multiplexer 602. The multiplexer 602 can pass through the channels one by one and forward the weights to the accumulator logic.

[0119] Therefore, each pulse input 6101-610 stored in the distributed memory is received k The weight corresponding to the address is passed to the accumulator. In one example, the stored weight is passed to the accumulator element 603. Thereafter, the signal from the accumulator element 603 is passed, for example, to the D input 6042 of the CMOS D flip-flop 604. The local timing signal 608 can also be passed as an input to the accumulator, for example, to the CMOS D flip-flop 604, namely at the CLK input 6043. The DFF 604 captures the incoming weight signal at a specific edge of the clock signal corresponding to the pulse in the local timing signal 608.

[0120] The DFF has a feedback loop from its output (e.g., from the Q port 6044a) back to the summing element 603. In this way, the accumulated weight is output by the DFF and used as the second input of the summing element 603. At the same time, the summing element obtains the next weight from the distributed memory (e.g., from the multiplexer 602) and adds the new weight to the accumulated weight to obtain a new accumulated weight. This new accumulated weight is then sent to the DFF 604 again. Since the local timing signal 608 contains a number of pulses equal to the number of times the signal is received from the summing element 603, all signal inputs 6101-6103 that receive spikes are equal. kThe accumulation of weights is performed in this way.

[0121] In other words, the local timing signal 608 contains multiple pulses corresponding to the incoming spike signal. Each pulse is used to determine when the DFF captures a specific (accumulated) weight from the distributed memory 602. The captured weight is integrated over time through consecutive clock cycles, triggered by the local timing signal 608. The DFF 604 accumulates the presence of the spike over multiple clock cycles (e.g., after each trigger from the local timing signal 608), thereby providing a form of time integration. In this way, operations corresponding to the weight can be performed on the logic encoded in the accumulator. As described above, this process can be performed on all captured input spikes 6101-610 k Repeatedly. Multiple unit cells can perform this function within a neuron, ensuring that all m spike signals 4101-410 m The local timing pulse 408, 608 is expected to be delayed enough from the start of the pulse so that the spike is processed in the accumulator (e.g., DFF 604).

[0122] In this manner, distributed weight storage elements (eg, including latches and / or flip-flops) may be used to instantiate a signal that is consistent with input spikes 6101-610. k The benefit of the distributed weight storage mechanism is that traditional place-and-route tools (e.g., Innovus, ICCII, Openroad; all off-the-shelf software solutions) can optimize the combinational logic around the weights, and for Figure 2 All the different logic blocks required for the design methodology shown, using this approach to generate modules can be automated with minimal manual intervention. Furthermore, logic storage elements (e.g., latches or flip-flops) offer designers the option of lowering their supply voltage, bringing leakage values ​​comparable to memories like SRAM and ReRAM while potentially significantly improving dynamic power. Bucking involves intentionally lowering the operating voltage (VDD) below the standard or nominal voltage specified for the latch or flip-flop. This can be done for a variety of reasons, including power optimization and energy efficiency.

[0123] therefore, Figure 5 and Figure 6A weight encoding and accumulation module for accumulating spike input signals in neurons of a spiking neural network is disclosed. The module may include multiple spike input ports 510, 610. Each of the multiple spike input ports may be configured to receive a spike input signal. Furthermore, the module may include trigger signal input ports 508, 608, which may be configured to receive a trigger signal. A pulse included in the trigger signal may indicate the number of spikes received by the multiple spike input ports. Furthermore, the module may include a spike-to-weight encoder module. The spike-to-weight encoder module may be configured to receive spike input signals and output corresponding weights associated with a particular spike input signal. This may be achieved by encoding the received spike input signals into weights. Furthermore, the module may include an accumulator module. The accumulator module may be configured to receive weights and a trigger signal from a memory module and accumulate the weights of the received spikes. The timing of one or more operations involved in the accumulation may be controlled by a trigger signal. In this case, a loop circuit composed of CMOS flip-flops is controlled by the trigger signal. Other methods of triggering the accumulation using a trigger signal are also contemplated.

[0124] Through delay-based serialization methods (such as Figure 4-6 As shown, only the captured spikes are serialized and summed in the adder. Using a combination of delay-based timing (targeting the worst-case setup time of the adder) and the feed-forward nature of the design, correct functionality only requires meeting specific constraints (e.g., the delay of each logic block and the delay of the delay cell in the design). Note that if the design is feed-forward and the logic paths are well-balanced, then from a timing perspective, only the setup time issue needs to be addressed, not the hold time issue. Hold time can also be mitigated if specific timing signals cause data to be captured into the latch on a pulse rather than an edge (i.e., pulse-based latching rather than edge-based latching). So, if we only address the setup time issue, it is understandable that the setup time issue can always be fixed by slowing down the design. Using variable distributed delays, the setup time issue can be addressed for every part of the design.

[0125] Furthermore, the interval between pulses must be greater than the time required to access a particular memory plus the accumulation time. As previously mentioned, pulses arriving at the accumulator can be delayed further than pulses arriving at the memory (as in the local timing signal output 408). Since the SNN is a standby network with high activity periods, the above design approach can handle high activity periods while avoiding unnecessary dynamic power consumption during low activity periods.

[0126] Notice, Figure 5 and Figure 6The unit cells 500, 600 shown in the figure take, for example, an analog signal as input and output a digital value that is the cumulative sum of the weights of the input signals containing spikes. A second type of unit cell is also designed that simply adds together the digital values ​​of the weights. These unit cells can be used in additional layers of a neuron addition scheme, in which, through parallelization, multiple unit cells take analog signals as input and convert them into digital signals, and the second type of unit cell can be used to add the digital signals (representing the accumulated weights) to obtain the accumulated weights.

[0127] Figure 7 A schematic diagram of a multi-level self-timed neuron design with repeatable unit cells is shown. Specifically, a high-level architecture is shown of how to partition multiple inputs at a given summation level and across different levels. Due to the trade-off between logic complexity and timing uncertainty, the m inputs in a neuron are partitioned into smaller sets of inputs, such as k inputs, with each set having the same or different number of inputs. For each input set, each of the k inputs is processed in a unit cell. The unit cell may include Figure 5 or Figure 6 The architecture described in .

[0128] Each unit cell can generate a multi-level spike timing signal to complete its calculation (such as Figure 4 As an example, a completion signal can be generated by an additional delayed signal running in parallel with the logic. As long as the delayed signal is greater than the worst-case delay of the logic, the multi-level sequential signal will always indicate the timing after the worst-case delay of the logic (for example, by a pulse). The completion signals of multiple unit cells in a layer can be the OR of the individual completion signals. This can be applied to multi-stage pipeline designs implemented in a feedforward manner by a delay-based timing method. See the example below Figure 8 .

[0129] Accumulation can be performed at different levels. When accumulation is split across different levels, it means that the integration process occurs independently, or hierarchically, at each level of the system. In neural networks or signal processing systems, information is often processed at different levels of abstraction. Each level can represent features or patterns of varying complexity.

[0130] There are several reasons to split addition into different levels. First, pipelining improves throughput, but at the expense of latency. Furthermore, in latency-based timing, there's an additional trade-off: when attempting to improve latency by adding logic to a single level, timing uncertainty increases. This increase in timing uncertainty worsens latency.

[0131] For example, use Figure 4The spike capture and timing generation shown generates a multi-level spike timing signal 409 that can be used to generate spikes using a method such as Figure 8 The feed forward timing shown and / or Figure 9 The handshake shown is timed between the different layers.

[0132] Enter 701 m , m different inputs of neuron 700 enter neuron 700. These inputs may correspond, for example, to the input synapses of neuron 700. The different input spikes are latched in latch 7020. This can be done using, for example, Figure 4 The spike capture and timing generation modules 301 and 400 are shown to accomplish this.

[0133] Next, the input signal can be split into k smaller input sets 701 k , and used as input for corresponding unit cells, which accumulate the input signal 701 k These unit cells can be as above Figure 5 or Figure 6 At the end of level 0, the accumulated signal from the unit cell 703 is latched into the level 1 register bank 7021. The region 711 indicates that the unit cell 703 in the region 711 is a spike driven unit cell, e.g. Figure 5 and Figure 6 The unit cell shown takes as input an analog signal containing spikes. From this point forward in the neuron logic, the accumulation unit cell 703 can be either spike-driven or data-driven. As an example, region 712 indicates that the unit cells in region 712 are data-driven unit cells, e.g., unit cells that take as input digital signals and accumulate the values ​​of those digital signals as described above.

[0134] Since this implementation generates an indication that level 0 has completed its computation (e.g. Figure 4 The multi-level timing signal (as shown) can therefore be used as an input change signal and the rest of the unit cells do not need to be timed internally, they can be made of conventional CMOS and timed using static timing analysis (STA). Note that standard CMOS logic can be made to operate without a clock. For example, a design can be done using maximum and minimum delay constraints. Standard software EDA synthesis tools such as Design Compiler and Genus can determine these constraints. They ensure that the logic created always has a delay greater than the minimum delay and less than the maximum delay. Since the timing between the levels is still delay driven and the timing uncertainty for longer delays still holds, it may be necessary to further divide the levels based on the timing trade-offs described previously.

[0135] By time-division multiplexing within a given layer, a better trade-off between latency and timing uncertainty at a particular layer can be achieved. If we know that the previous layer took longer than the current layer (due to logic trade-offs within the previous layer), we can optimize the current layer to handle more information by time-division multiplexing. In this case, timing is still performed via delay-based signals, and standard time-division multiplexing methods can be used.

[0136] For the first level, the latched accumulated signal latched into the level 1 register bank 7021 is divided into k1 input signals 701 k1 The groups are then each separately forwarded to the accumulation unit 703, which can operate by accumulating the digital accumulation signal values.

[0137] At the last level, the output of the second-to-last unit cell is generated by the level N register bank 702. N Then, the spike input 701 kN Sent to accumulator 706 (e.g., DFF 706) via element 705, specifically, spike input 701 kN is sent to the D port of DFF 706. Therefore, at the end of the last level, the previous summation is accumulated. The timing is still based on delay feedforward or delay handshake.

[0138] When the last-level accumulator 706 reaches a preset threshold 709, an output spike 710a is generated, for example, by a comparator 708. The comparator 708 checks whether the accumulated signal has reached the preset threshold 709 and resets the accumulated value in the accumulator 706 to a predetermined value. The accumulator 706 can send the accumulated signal to the comparator 708, where it is compared with the threshold 709. In addition, the accumulator 706 can output an accumulated sum 710b as a separate output. This can be used as a real-time signal in the training and inference of a neural network.

[0139] In order to be able to create a high fan-in accumulator with m inputs, multiple small accumulators (e.g., k-input unit cell accumulator 703) need to be pipelined. Figure 7 As shown, this pipeline can be implemented in a variety of ways. Its logic exists in the control logic 704.

[0140] Figure 8 and Figure 9 Shows the implementation Figure 7 There are two potential approaches to pipeline.

[0141] Figure 8 A delay-based timing method is used to implement a multi-level pipeline design in a feed-forward manner.

[0142] The input signal enters the pulse latch 804 and the pulse tree element 801, which combines the signals including the pulses transmitted on the different input lines into a single output line that generates a pulse whenever one of the different input lines generates a pulse. The output from the pulse latch 804, the weights stored in the weight memory 803 and / or the output of the pulse tree element 801 can be used as inputs to the L0 logic 8050. Next, the output of the L0 logic 8050 is sent to the register 806. The output from the pulse tree element 801 is sent after a delay that generates a delayed signal that is equal to or greater than the processing time of the L0 logic 8050. This delayed signal is used as an input to the register 806 to determine when the register should forward its latch signal to the next level of logic (in this case, the L1 logic 8051). The delayed signal is also used as an input to the hierarchical logic L1 logic 8051. The delayed signal is also sent to different delay elements 802 corresponding to the hierarchical logic L1 logic 8051 so that the delay it generates is equal to or greater than the processing time of the hierarchical logic L1 logic 8051. The method is repeated for the remaining layers.

[0143] Figure 9 The pipeline structure is implemented using an inter-stage handshaking mechanism. For handshaking, correct loop interruption needs to be implemented so that the proposed constraints and methods work.

[0144] The input signal enters the pulse latch 904 and the pulse tree element 901. The output from the pulse latch 904 and the weight stored in the weight memory 903 can be used as the input of the L0 logic 9050. Next, the output of the L0 logic 9050 is sent to the register 906. The output from the pulse tree element 901 is also sent to the control logic 902, which initiates the handshake process with the L0 logic 9050 by sending a synchronization (SYN) signal. The L0 logic 9050 will respond with an acknowledgement (ACK) signal when ready.

[0145] When the L0 logic 9050 is ready, it outputs an output signal to the register 906, which latches the output signal. Furthermore, it sends a signal to the next control logic 902, which performs a handshake process with the next logic layer L1 logic 9051. Furthermore, based on the signal received from the first L0 logic 9050, the control logic 902 sends a signal to the register 806 to output the latched data as an input signal to the logic layer L1 logic 9051. This method is repeated for the remaining layers.

[0146] Proper optimization can help further optimize pipeline throughput, as each pipeline stage can be reused depending on the use case. This provides additional flexibility, as network traffic can be high-throughput, low-latency, or both. The right design topology can be determined by a combination of flexible pipeline depth and latency-based time multiplexing.

[0147] The design described in this article is more at the architectural and logic function level. Depending on the power, area, and performance requirements, the actual implementation can be done using fast design methods such as self-resetting CMOS, pulse-based design, or, if the requirements are less stringent than standard CMOS, a modified static timing analysis with delay-based timing closure can be used, as described in the Methodology section.

[0148] Neurons designed using the examples above can be further optimized for specific use cases.

[0149] If a use case doesn't utilize all possible input lines, we know that the latency of a specific stage can be significantly reduced. This results in an improved trade-off between latency and throughput, but with control over latency at each stage. This differs from traditional clock-based design techniques, where the granularity of control is more global as the clock network design becomes more complex at the local level. Typically, latency and throughput are inversely proportional and fixed by design. With this approach, they can both be improved and adjusted based on the use case.

[0150] Since mixed-signal design techniques can be used to implement the delay, similar calibration techniques can also be used to mitigate any process, temperature, and voltage issues.

[0151] The AND gates in this specification can be implemented using semiconductor devices (e.g., transistors) to perform logical AND, while the OR gates use similar components to implement logical OR. A pulse latch can be constructed using a flip-flop and a pulse generator. The accumulator can be an operational amplifier and / or a digital counter. The subtractor circuit can be composed of an arithmetic unit capable of calculating the difference. The input and output ports are designed using interface modules to allow seamless communication between neurons.

[0152] The present invention discloses, among other things, the implementation of spiking neurons and spiking neural networks in hardware (e.g., chips). Silicon is the primary material used in semiconductor manufacturing. Metal oxide semiconductor field effect transistors (MOSFETs) are the basic building blocks of digital integrated circuits. They can be used to implement logic gates, memory cells, and other basic components. SRAM is typically used to quickly and volatilely store synaptic weights, neuron states, and other dynamic information within SNNs. For more persistent storage requirements, non-volatile memory technologies such as flash memory or resistive random access memory (RRAM) can be used. Multiple metal layers can be used to interconnect different components on a chip. These layers facilitate signal routing between neurons, synapses, and other functional units. Through silicon vias (TSVs) can be used to achieve vertical stacking of multiple layers, thereby improving overall connectivity and reducing physical footprint.

[0153] Energy-saving design techniques can be used, such as voltage scaling (dynamically adjusting the operating voltage to balance power consumption and performance), clock gating (temporarily stopping the clock signal to idle parts of the circuit during inactive periods), and / or power gating (completely cutting power to specific parts of the chip when not in use). Through the present application, local timing is generated based on certain trigger events (in this case, receive spikes), which can help achieve energy-saving designs.

[0154] Note that the features of any of the embodiments disclosed herein may be combined in any suitable manner.

Claims

1. A spike capture and local trigger generation module 400 for neurons of a spiking neural network, comprising: a plurality of spike input ports 401 , wherein each spike input port 401 is configured to: receive a spike input signal from a spiking neural network; a pulse latch module 402 configured to: capture a spike in one or more spike input signals, wherein the pulse latch module 402 outputs a signal indicating whether a spike is captured at one of the spike input ports 401; a local trigger generator module configured to generate a trigger signal for use by the neuron to synchronize and / or control the timing of operations in the neuron, wherein the local trigger generator module is configured to generate the trigger signal based on a signal from the pulse latch module 402, wherein the trigger signal indicates the number of spikes captured by the pulse latch module 402; The trigger signal output port 408 is configured to output the generated trigger signal.

2. The spike capture and local trigger generation module 400 according to claim 1, wherein: The trigger signal includes a number of pulses equal to the number of spikes captured by the pulse latch module 402 .

3. The spike capture and local trigger generation module 400 according to claim 1 or 2, wherein: The trigger signal indicates the number of spikes captured by the pulse latch module 402 during a predetermined timing period.

4. The spike capture and local trigger generation module 400 according to any one of the preceding claims, wherein: The trigger signal is generated such that operations in the neurons can be triggered by the trigger signal rather than by a global clock signal of the spiking neural network.

5. The spike capture and local trigger generation module 400 according to any one of the preceding claims, wherein: The trigger signal is based on a delay-based asynchronous counter.

6. The spike capture and local trigger generation module 400 according to any one of the preceding claims, wherein: The one or more pulses contained in the trigger signal have a certain spacing therebetween, and the spacing is determined by the minimum time required for the execution of an operation in the neuron, preferably, wherein the operation is the accumulation of one or more weights.

7. The spike capture and local trigger generation module 400 according to any one of the preceding claims, further comprising a plurality of spike output ports, each port being configured to output a spike output signal, wherein Each of the plurality of spike output ports corresponds to one of the plurality of spike input ports, and The specific spike output port is configured to output a pulse if the corresponding spike input port receives a spike.

8. The spike capture and local trigger generation module 400 according to claim 7, wherein: The time intervals of the pulses outputted by the peak output port are substantially the same as those of the pulses contained in the clock signal.

9. The spike capture and local trigger generation module 400 according to any one of the preceding claims, wherein: The pulse latch module 402 includes a plurality of latch elements 403 , each of which is configured to capture one or more spikes in a spike input signal received by a specific spike input port through pulse latching.

10. The spike capture and local trigger generation module 400 according to claim 9, wherein: The latch element 403 is a CMOS D flip-flop 403, The CMOSD flip-flop is triggered by a peak input signal, and The latch signal output by the CMOSD flip-flop is a bit value indicating whether the peak is latched.

11. The spike capture and local trigger generation module 400 according to any one of the preceding claims, wherein: The local trigger generator module comprises a pulse generator 404 for generating pulses, wherein the local trigger generator module includes delay logic for each spike input port, and Therein, the generated pulse is then sent through each delay logic.

12. The spike capture and local trigger generation module 400 according to claim 11, in, Each delay logic includes a delay element for delaying a pulse signal; Wherein, if the latch signal indicates that a specific spike input port has captured a spike, the corresponding generated pulse is sent through the delay element of the corresponding delay logic, and If the latch signal indicates that a specific spike input port has not captured a spike, the corresponding generated pulse is not sent through the delay element of the corresponding delay logic.

13. The spike capture and local trigger generation module 400 according to claim 11 or 12, in, If the latch signal indicates that a specific spike input port captures a spike, the corresponding delay logic outputs the corresponding generated pulse to generate the trigger signal, and If the latch signal indicates that a specific spike input port has not captured a spike, the corresponding delay logic does not output the corresponding generated pulse to generate the trigger signal.

14. The spike capture and local trigger generation module 400 according to any one of the preceding claims, wherein: The local trigger generator module is configured to generate a plurality of trigger signals used by the neurons to synchronize and / or control the timing of operations in corresponding parts of the neurons, The local trigger generator module is configured to generate multiple trigger signals using multiple latch signals, such that for a specific set of spike input ports 401, each trigger signal includes one or more pulses equal in number to the number of spikes captured by the pulse latch module 402, and The system further comprises: a plurality of trigger signal output ports 408 configured to output the generated trigger signals.

15. The spike capture and local trigger generation module 400 according to any one of the preceding claims, further comprising: A multi-level trigger generator module having a multi-level trigger output port 409 configured to generate a multi-level trigger signal, The multi-level trigger signal indicates the total time required for the neuron to perform the operation triggered by the trigger signal. Preferably, if there are multiple trigger signals generated and output, the multi-level trigger signal corresponds to the maximum time required for the neuron to perform all sets of operations triggered by different trigger signals.

16. A weight encoding and accumulation module for accumulating spike input signals in neurons of a spiking neural network, comprising: a plurality of spike input ports, wherein each spike input port is configured to receive a spike input signal; a trigger signal input port configured to receive a trigger signal, wherein the trigger signal indicates a number of spikes received by the spike input port; a spike-weight encoder module configured to: receive the spike input signal and, when a spike is received in the spike input signal, output a corresponding weight associated with the spike input signal; An accumulator module is configured to receive a trigger signal and weights from the spike-weight encoder module and accumulate the weights, wherein a timing of one or more operations involved in the accumulation is based on the trigger signal.

17. The weight encoding and accumulation module according to claim 16, wherein: The trigger signal includes a number of pulses equal to the number of spikes received by the spike input port.

18. The weight encoding and accumulation module according to claim 16 or 17, wherein: Each spike input signal is an analog signal, and The weight is a numerical value.

19. The weight encoding and accumulation module according to any one of claims 16 to 18, wherein: The accumulator module includes an accumulating element 503 and a loop circuit. The loop circuit takes the output of the accumulating element as input and outputs a value back to the accumulating element whenever the loop circuit receives a pulse from the trigger signal. When the peak-weight encoder module outputs a first weight to the accumulation element 503 , the accumulation element adds the first weight to the second weight output by the cyclic circuit.

20. The weight encoding and accumulation module according to claim 19, wherein: The loop circuit includes a CMOS D trigger triggered by the trigger signal.

21. The weight encoding and accumulation module according to any one of claims 16 to 20, wherein: The spike-weight encoder module includes a local memory, preferably SRAM or non-volatile memory, and The spike-weight encoder encodes a pulse received at a specific spike input port into a memory address, where the memory address is used to request a weight stored in the specific memory address, and the weight is associated with the specific spike input port.

22. The weight encoding and accumulation module according to any one of claims 16 to 21, wherein: The spike-weight encoder module includes a distributed memory including distributed weight storage elements, and wherein the weights associated with the spike input ports are each stored in one of the distributed weight storage elements, and If there is a spike in the spike input signal, the corresponding weight is output to the accumulator module by the distributed weight storage element storing the weight.

23. The weight encoding and accumulation module according to any one of claims 16 to 22, wherein: The spike-weight encoder module is configured as follows: Based on the priority rules, the order of weight output is changed.

24. A spiking neuron of a spiking neural network, wherein: The neuron comprises a spike capture and local trigger generation module 400 according to any one of claims 1-15, and a weight encoding and accumulation module according to any one of claims 16-23, and Among them, the trigger signal output port of the spike capture and local trigger generation module outputs the trigger signal to the trigger signal input port of the weight encoding and accumulation module, so that the timing of one or more operations involved in the accumulation is controlled by the trigger signal.

25. The spiking neuron of claim 24, further comprising: A comparator module compares the total accumulated weight with a predetermined threshold, Wherein, the total accumulated weight is the sum of the weights of all spike input ports that capture the spike of the corresponding spike input signal, and The spiking neuron is configured to output a spike to the spiking neural network when the total accumulated weight exceeds a predetermined threshold.

26. The spiking neuron according to claim 24 or 25, wherein The spiking neuron comprises a plurality of weight encoding and accumulation modules according to any one of claims 16-23, and Wherein, the peak capture and local trigger generation module 400 is based on 14, and Among them, each of the multiple trigger signal output ports outputs a corresponding trigger signal to the trigger signal input port of the corresponding weight coding and accumulation module, so that the timing of one or more operations involved in the accumulation performed in the corresponding weight coding and accumulation module is controlled by the corresponding trigger signal.

27. The spiking neuron of claim 26, further comprising: A digital weight accumulator module accumulates outputs of at least a portion of the plurality of weight encoding and accumulation modules and outputs an accumulated value.

28. The spiking neuron of claim 27, wherein The spiking neuron includes a plurality of digital weight accumulator modules, and Among them, at least one digital weight accumulator module accumulates the accumulated values ​​output by other digital weight accumulator modules.

29. The spiking neuron according to claims 26-28, wherein Accumulation operations between the plurality of weight encoding and accumulation modules and / or the one or more digital weight accumulator modules are synchronized using a delay-based pipeline and / or a handshake-based pipeline.

30. The spiking neuron of claim 29, wherein The spike capture and local trigger generation module according to claim 15, and The multi-level trigger signal is used to synchronize the accumulation operations between the multiple weight encoding and accumulation modules and / or the one or more digital weight accumulator modules.

31. A method for spike capture and local trigger generation in neurons of a spiking neural network, comprising: receiving a plurality of spiking input signals from the spiking neural network; Capturing the peak in the peak input signal by pulse latching; generating a latch signal for each spike input signal, wherein the latch signal indicates whether a spike is captured; generating a trigger signal based on the latch signal, wherein the trigger signal indicates a number of captured spikes, wherein the trigger signal is used by the neuron to synchronize and / or control timing of operations in the neuron; and Outputs the generated trigger signal.

32. The method of claim 31 , further comprising: Output peak output signal, The trigger signal is generated so that the operation in the neuron can be triggered by the trigger signal rather than by the global clock signal of the spiking neural network.

33. A method for weight encoding and accumulation in neurons of a spiking neural network, comprising: receiving a plurality of spike input signals; receiving a trigger signal indicative of a number of spikes received in the spike input signal; When a spike is received in one of the spike input signals, obtaining or generating a weight associated with the spike input signal; Accumulating weights associated with the spike input signal, wherein a timing of one or more operations involved in the accumulation is based on the trigger signal; as well as Output the accumulated weight.

34. The method according to claim 33, wherein The weights are obtained or generated by encoding pulses received in the one of the plurality of spike input signals as a memory address for requesting a weight to be stored at that particular memory address in a local memory.

35. The method of claim 33, wherein: If there is a spike in the spike input signal, the corresponding weight is obtained from a distributed memory comprising a plurality of weight storage elements, wherein the weights associated with the spike input signals are each stored in one of the distributed weight storage elements, and If there is a spike in the spike input signal, the corresponding weight is output by the distributed weight storage element storing the weight and is used for weight accumulation.

36. A method for spike capture and accumulation in neurons of a spiking neural network, comprising: receiving a plurality of spiking input signals from the spiking neural network; Capturing the peak in the peak input signal by pulse latching; outputting a latch signal for each spike input signal, wherein the latch signal indicates whether a spike is captured; generating a trigger signal based on the latch signal, wherein the trigger signal indicates a number of captured spikes, wherein the trigger signal is used by the neuron to synchronize and / or control timing of operations in the neuron; When a spike is received in a spike input signal among the plurality of spike input signals, obtaining or generating a weight associated with the spike input signal; accumulating weights associated with the spike input signal, wherein a timing of one or more operations involved in the accumulation is based on the trigger signal; and Output the accumulated weight.

37. The method of claim 36, further comprising: comparing a total accumulated weight with a predetermined threshold, wherein the total accumulated weight is the sum of all weights of the spike input signal that captured the spike; and When the total accumulated weight exceeds a predetermined threshold, a spike is output to the spiking neural network.

38. The method according to claim 36 or 37, wherein The steps of obtaining or generating weights and accumulating weights are parallelized by the following steps: The plurality of spike input signals are divided into one or more subsets of the plurality of spike input signals, and the steps of obtaining or generating weights and accumulating weights are performed on the subsets simultaneously, so that an accumulated subset weight is obtained for each subset.

39. The method according to claim 38, wherein The accumulated subset weights are accumulated digitally, preferably in a parallelized manner.

40. The method according to claim 38 or 39, wherein Accumulation operations between different subsets and / or the one or more digital accumulations are synchronized using a delay-based pipeline and / or a handshake-based pipeline.

Citation Information

Patent Citations

  • Method and apparatus for self-resetting logic circuitry

    US5864251A