Neuromorphic Computing System and Computing Method Based on Spiking Neural Network
By designing a neuromorphic computing system based on pulsed neural networks, using the approximate biological neuron model and RISC-V instruction architecture, the problems of large hardware resource consumption, low computing efficiency and limited online learning capabilities in the existing technology are solved, and efficient, accurate and flexible neuromorphic computing effects are achieved.
Patent Information
- Application Number
- CN202510066018.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-01-16
AI Technical Summary
Existing neuromorphic processors face problems such as excessive hardware resource consumption, low computing efficiency, limited online learning capabilities, and insufficient routing scheduling, which leads to restrictions on the implementation and popularization of applications in actual scenarios.
A neuromorphic computing system based on pulsed neural networks was designed, and physical neurons were constructed using an approximate biological neuron model, combined with a central processor with RISC-V instruction architecture to achieve online learning capabilities and efficient computing.
It achieves a more efficient, more accurate, lighter and more flexible neuromorphic calculation effect, achieving a high balance of recognition accuracy, response speed and hardware resource consumption, and is suitable for changing practical scenarios.
Smart Images

Figure CN119494377B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial neural networks, and particularly to a neuromorphic computing system and a computing method based on a spiking neural network. Background Art
[0002] By simulating the activity mechanism of brain neurons, combining biological principles and hardware design, a neuromorphic processor based on a spiking neural network (SNN for short) aims to achieve efficient computing and task execution, especially with the characteristics of online learning and event-driven. Such processors are widely used in fields such as medical image processing, autonomous driving, intelligent monitoring, and intelligent transportation, and are an important research direction in the field of neuromorphic computing. Different from traditional artificial neural network circuits (such as convolutional neural networks, abbreviated as CNN), it usually relies on continuous signal processing and fixed training models. For the processing of dynamic changes and data sparsity, these methods usually have poor effects. In addition, the hardware design based on CNN fails to fully reproduce the biological characteristics and dynamic learning ability of the human brain, has poor approximate computing simulation ability, and its complex calculations and large amount of calculations have poor adaptability in environments with limited hardware resources and low energy efficiency, usually relying on offline training of real-time applications and environments. The lack of the function of online learning means that once the model is deployed into the hardware circuit, its computing ability can only perform inference based on the fixed pre-trained model, and cannot adapt to further learning and accurate inference of new things.
[0003] Currently, pulse neuromorphic processors face a series of key problems, such as excessive consumption of hardware resources, low computational efficiency of neurons, limited online learning methods, and lack of flexibility in data transmission between chips. These problems restrict its implementation and popularization in actual scenarios and also limit the deployment of neuromorphic processors in computer systems. Among these problems, the internal computing core, that is, the resource overhead of spiking neurons, and neuromorphic chips with low computational efficiency have become important factors restricting the performance improvement. For example, current-based models make neuron behavior adjustable, but their complex membrane potential update process brings additional computational costs. Typical representatives include the Leaky Integrate-and-Fire (LIF) neuron model and the Izh neuron model with up to dozens of neuron event states. The LIF neuron model has low computational efficiency because it is not optimized for hardware systems; while the Izh neuron model can simulate up to dozens of neuron behaviors, but it introduces complex control logic, resulting in a large amount of additional resource consumption. Considering the complexity of neuron computing, existing research has applied approximate computing technologies such as weight quantization and pruning to neuron model design to reduce the difficulty of hardware implementation and improve the approximate simulation performance of artificial neurons, but these methods still face the problem of affecting the recognition accuracy of neuron network models. Secondly, the lack of online learning ability or insufficient self-optimization is a common problem in existing SNN processors. Considering computational complexity and cost factors, many existing SNN processors, such as TrueNorth, Tsinghua Tianjic, SpiNNaker, etc., fail to provide online learning functions on-chip and can only deploy weights offline, limiting their adaptability in changing actual scenarios. Another type of neuromorphic chip with online learning ability, such as chips using the Spike-Timing-Dependent Plasticity (STDP) method, although different learning rules can be designed through microcode, the learning rules of this method are more complex and will occupy more hardware resources when implemented in hardware.
[0004] The design of neuromorphic computing systems in the prior art usually lacks sufficient flexibility. For computing systems deployed with SNN processors, there are various technical problems. In particular, during the design of time windows, due to the differences in processing times of each network layer, problems of time mismatch will occur. Such mismatch will cause, after a certain layer of computation is completed, it is necessary to wait for the previous network layer to finish computing, thus wasting computing resources. Therefore, how to design accurate artificial physical neurons closer to biological characteristics and combine multiple physical neurons to form a spiking neural network to achieve precise target recognition is the main challenge faced by the current technology. How to dynamically adapt artificial physical neurons to multi-scale image inputs and support spike trains with different time windows and the computation of large-scale spiking neural networks, so as to achieve an accuracy close to biological characteristics and save hardware computing resources as much as possible, is the core goal for current intelligent computer systems to achieve efficient and accurate image recognition. Summary of the Invention
[0005] There are technical problems such as high hardware resource overhead, low computing efficiency, limited online learning ability, and insufficient routing and scheduling in existing neuromorphic processors and computing systems. To solve the problems existing in the above prior art, inspired by biological models and aiming at applying neuromorphic computing technology to real-time intelligent computer systems, the present invention designs an approximate neural computing model that can efficiently simulate the complex computing functions of biological nervous systems. Specifically, the present invention proposes an online learnable neuromorphic computing system that can dynamically update and efficiently process complex tasks, achieving a more efficient, accurate, lightweight, and flexible neuromorphic computing effect, achieving a high balance among recognition accuracy, response speed, and hardware resource consumption, and finally realizing the processing of intelligent image target recognition tasks. The technical solution provided by the present invention is applicable to various computer systems or embedded system technical fields that require artificial neural network computing, such as autonomous driving, Internet of Things, robotics, drones, medical image processing, intelligent monitoring systems, and intelligent transportation, etc., to solve the problems existing in the above prior art.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] A neuromorphic computing system based on a spiking neural network, comprising:
[0008] A neuromorphic processor, which is used to construct a hardware-based spiking neural network with a number of physical neurons, and perform spike encoding and target recognition on real-time input images through the spiking neural network; the physical neurons are constructed using an approximate biological neuron model; the approximate biological neuron model pre-stores synaptic weights and activation thresholds;
[0009] A central processing unit implemented based on the RISC-V instruction architecture, which communicates with the neuromorphic processor and is used to execute control instructions required for constructing the spiking neural network;
[0010] The neuromorphic processor is further used to perform online learning on the spiking neural network under the control of the control instructions, and update the synaptic weights and activation thresholds associated with each physical neuron of the spiking neural network;
[0011] The memory is used to store program instructions required for the operation of the central processing unit, and the program instructions include control instructions required for constructing the spiking neural network; and is used to store data transmitted between the central processing unit and the neuromorphic processor.
[0012] Preferably, the central processing unit is a single-issue out-of-order execution RISC-V processor; the RISC-V processor includes a processor front end, an instruction buffer unit, and a processor back end;
[0013] The processor front end includes a program decoder, an instruction fetch unit, and a dynamic branch predictor;
[0014] The program decoder is used to convert a computer program for controlling the operation of the spiking neural network into a RISC-V instruction stream, and store the RISC-V instruction stream in the memory;
[0015] The instruction fetch unit is used to access the memory, read a single RISC-V instruction corresponding to the program count value of the current clock cycle, and send it to the instruction buffer unit; the program count value is used as the read address for accessing the memory;
[0016] The dynamic branch predictor is used to predict the program count value of the next clock cycle;
[0017] The instruction buffer unit is used to buffer the RISC-V instructions read by the instruction fetch unit, and isolate the front-end calculation thread and the back-end calculation thread of the RISC-V processor;
[0018] The processor back end is a five-stage pipeline architecture composed of a decoding unit, an issue unit, a physical register file, an execution unit, and a retirement unit, which is used to read the RISC-V instructions of the instruction buffer unit and perform decoding, issuing, reading and writing the physical register file, out-of-order calculation, and instruction retirement operations on them in sequence, so as to realize data interaction between the RISC-V processor and the neuromorphic processor.
[0019] Preferably, the approximate biological neuron model includes a charging model and a discharging model, where:
[0020] The charging model is as follows:
[0021] Reference decay charging model: or
[0022] Incremental decay charging model: ,
[0023] The discharging model is as follows:
[0024] Hard reset discharging model: or
[0025] Soft reset discharging model: ,
[0026] Wherein, V(t) is the membrane potential of the physical neuron at the current time step, V(t - 1) is the membrane potential of the physical neuron at the previous time step, X(t) is the membrane potential increment formed by the pulse signal input at the current time step, is the preset membrane potential decay factor; is the reset reference potential of the physical neuron; V(t+) is the membrane potential of the physical neuron after reset at the current time step, V(t -) is the membrane potential of the physical neuron before reset at the current time step; Vth is the activation threshold of the current physical neuron.
[0027] Based on the same inventive concept, further, the present invention also provides an image target recognition method based on a spiking neural network, including: writing a computer program; inputting the computer program into any one of the neuromorphic computing systems based on a spiking neural network for execution, and using the spiking neural network to perform target recognition on a real-time input image.
[0028] The beneficial effects of the present invention are as follows:
[0029] The present invention proposes a neuromorphic computing system and an image target recognition method based on a spiking neural network with optimized structure. Combining the computing characteristics of the spiking neural network, it realizes a hardware-friendly neuron structure design and high-precision approximate calculation, and can achieve efficient online network learning, obtaining a neuromorphic computing system with low computing resources and power consumption, high recognition accuracy, and high computing efficiency. And based on this neuromorphic computing system, it provides an accurate, efficient, and flexible real-time image target recognition method for various mobile scenarios and edge computing scenarios with high requirements for timeliness and accuracy but limited resources. Description of the Drawings
[0030] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0031] Figure 1 A schematic internal structure diagram of a neuromorphic computing system according to an embodiment of the present invention;
[0032] Figure 2 A schematic structural diagram of a central processing unit implemented based on the RISC-V instruction architecture according to an embodiment of the present invention;
[0033] Figure 3 A schematic structural diagram of an online learnable neuromorphic processor according to an embodiment of the present invention;
[0034] Figure 4 A schematic structural diagram of a spiking neural computing core in the hidden layer of a neuromorphic processor according to an embodiment of the present invention;
[0035] Figure 5 A schematic structural diagram of a spiking neural decision-making core in the output layer of a neuromorphic processor according to an embodiment of the present invention;
[0036] Figure 6 A preferred structural diagram of a spiking neural computing core in the hidden layer of a neuromorphic processor according to an embodiment of the present invention;
[0037] Figure 7 A schematic internal structure diagram of a physical neuron according to an embodiment of the present invention;
[0038] Figure 8 A schematic structural diagram of routers at all levels according to an embodiment of the present invention. Detailed implementation manners
[0039] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0040] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the drawings and specific implementation manners.
[0041] This embodiment provides a neuromorphic computing system based on a spiking neural network, as Figure 1As shown in the figure, the system includes: a neuromorphic processor 100, a central processor 200 implemented based on the RISC-V instruction architecture, and a memory 300. Among them:
[0042] The neuromorphic processor 100 is used to construct a hardware-implemented spiking neural network with a number of physical neurons, perform spiking encoding and target recognition on real-time input images through the spiking neural network; the physical neurons are constructed using an approximate biological neuron model; the approximate biological neuron model pre-stores synaptic weights and activation thresholds. In particular, the hardware implementation of the spiking neural network is based on the CMOS digital process technology.
[0043] The central processor 200 implemented based on the RISC-V instruction architecture communicates with the neuromorphic processor 100 and is used to execute the control instructions required to construct the spiking neural network. The neuromorphic processor 100 is also used to perform online learning on the spiking neural network under the control of the control instructions, and update the synaptic weights and activation thresholds associated with each physical neuron of the spiking neural network. The memory 300 is used to store the program instructions required for the operation of the central processor 200, and the program instructions include the control instructions required to construct the spiking neural network; and, it is used to store the data transmitted between the central processor 200 and the neuromorphic processor 100.
[0044] For the neuromorphic processor 100, most traditional neuron models are of two types: Integrate and Fire (IF for short) neurons and Leaky Integrate and Fire (LIF for short) neurons. However, the IF neuron model is an idealized integrator model, and its characteristic is that the membrane potential of the neuron remains constant without the influence of external inputs. For the LIF neuron model, in addition to the influence of the input current, the role of the leakage conductance is also an important factor that needs to be considered. In addition, for the neuromorphic processor implemented by digital circuit hardware, the input signal of the spiking neural network it depends on is a discrete pulse signal, which makes it difficult to accurately implement the continuous input current signal in the biological model on the digital circuit, and differential operations cannot be performed through traditional continuous signal calculation methods.
[0045] In order to make the spiking neural network model adaptable to the implementation of digital circuit hardware, the embodiment of the present invention improves the LIF neuron model, and based on the improved LIF neuron model, constructs a hardware-friendly spiking neural network to implement the neuromorphic processor 100. The basic charging equation (i.e., the neuron dynamic equation) of the original LIF neuron is:
[0046]
[0047] Among them, g is the leakage conductance and E is the resting potential of the neuron. This equation reflects the relationship between the input current, the membrane potential, and the decay process. The differential equation of the LIF neuron charging model (neuron dynamic equation) can be approximated as a recurrence equation of discrete-time signals.
[0048] In the embodiment of the present invention, since the pulse sequence is serially input according to the overall clock pace of the neuromorphic computing circuit system. At the same time, multiple pulse signals (0 or 1) can be input into the physical neuron in parallel. After being combined into a multi-bit digital signal, it is sent to each pulse core in the hidden layer of the pulse neural network for parallel computing. Within the accumulated discrete time step Δt, according to the computing model of the pulse neural network, the input current I(t) of the pulse sequence has a linear relationship with the increment X(t) of the neuron membrane potential. In specific implementation, the membrane potential increment X(t) is expressed as the product of the input current I(t) and the time step Δt during this computing process, that is: X(t) = I(t)∙Δt, where the statistically discrete time step Δt is also called the time window T. In specific implementation, the length of the discrete time window T is affected by the coding method used when the image is pulse-coded; within the discrete time window, multiple time steps are required to achieve the accumulation of the membrane potential. The membrane potential V(t) accumulated over time on the physical neuron is determined by a specifically constructed approximate biological neuron model.
[0049] This embodiment improves the LIF neuron model and provides the following optimized approximate biological neuron model. The approximate biological neuron model includes a charging model and a discharging model, where:
[0050] The charging model is:
[0051] Benchmark decay charging model: Or,
[0052] Incremental decay charging model: ,
[0053] The discharging model is:
[0054] Hard reset discharging model: Or,
[0055] Soft reset discharging model: ,
[0056] Among them, V(t) is the membrane potential of the physical neuron at the current time step t, V(t - 1) is the membrane potential of the physical neuron at the previous time step t - 1, and X(t) is the increment of the membrane potential formed by the pulse signal input at the current time step t. is a preset membrane potential decay factor; is the reset reference potential of the physical neuron; V(t+) is the membrane potential after reset of the physical neuron at the current time step t, and V(t-) is the membrane potential before reset of the physical neuron at the current time step t; Vth is the activation threshold of the current physical neuron. The spiking neural network constructed by the above approximate biological neuron model in this embodiment can support various types of synaptic connections, can simulate neurons of different models, and thus can more accurately reflect the diversity in the biological nervous system. This makes the spiking neural network more flexible in information processing and can achieve a higher level of neural computation.
[0057] In the implementation process of the present invention, the charging model is used as the dynamic equation of the physical neuron to accumulate the membrane potential of the physical neuron. The improved LIF neuron model provided in this embodiment takes into account the natural decay characteristic of the membrane potential. When there is no input, the membrane potential will gradually recover to the resting potential. This model is closer to the actual behavior of biological neurons, can better simulate the signal processing process of the biological nervous system, and has a certain anti-noise ability, and can effectively cope with unforeseen interference in the environment. When the reference decay charging model is used to charge the physical neuron, the input membrane potential increment is not decayed, and the increment of the membrane potential will directly adjust the membrane potential of the neuron. Then, according to the preset decay factor , the membrane potential of the previous time step is decayed relative to the reset reference potential , specifically expressed as generating the difference between the membrane potential of the previous time step and the reset reference potential . When using the LIF neuron charging model that decays the input membrane potential increment , the input membrane potential increment also participates in the decay according to the preset decay factor , so that the change of the neuron membrane potential is affected by both the increment of the input signal and the natural decay process. For images with more detailed changes and complex structures, using the improved LIF neuron model provided in this embodiment can simulate the dynamic characteristics of biological neurons, such as the natural decay and refractory period of the membrane potential, which makes it excellent in extracting the time-series related features in the input image, thereby improving the sensitivity to subtle differences and the recognition accuracy.
[0058] In the above optimized LIF neuron charging model, the time constant in the decay factor ( The selection of () can be adjusted according to the requirements of the neuron model, biological rationality, etc. Preferably, if the neuron needs to respond to rapidly changing input pulse signals, a smaller time constant (i.e., a larger decay factor) can be selected to accelerate the decay of the membrane potential; otherwise, if the neuron does not need to respond to input pulses too quickly, in order to delay the decay process of the membrane potential, a larger time constant (i.e., a smaller decay factor) should be selected. This adjustment method is modeled after the actual reaction mechanism of biological neurons. For example, fast-paced neurons in the cerebral cortex usually have a smaller time constant, so that they can respond to the rapid changes of input signals; while neurons related to learning and memory usually have a larger time constant in order to better maintain the response characteristics to input. The selection of the time constant is consistent with the characteristics of the behavior of actual neurons in biological neuroscience. For the reset reference potential For the setting of, preferably, it is set to a value lower than the resting membrane potential E to ensure that the neuron enters the refractory period after excitation. This setting can simulate the behavior of biological neurons, prevent neurons from being over-activated, and conforms to the inhibitory mechanism commonly existing in biological nervous systems.
[0059] Furthermore, in this embodiment, a central processing unit 200 based on the RISC-V instruction architecture is used to control the construction of a spiking neural network in the neuromorphic processor 100 to realize the transmission of the input image to the neuromorphic processor 100 for target recognition. RISC-V is an open-source instruction set architecture (ISA) based on the principle of reduced instruction set, which can run various programs by controlling the computer to execute instructions. Due to its characteristics such as simplicity, open source (different from traditional monopolistic instruction set architectures such as x86 and ARM), modularity, and strong instruction extensibility, RISC-V can be easily extended to different application scenarios from microcontrollers to high-performance computer processors. However, the RISC-V instruction set itself only serves as an instruction set architecture, providing the basic design principles of a specific RISC-V processor. Its actual hardware implementation needs to be designed according to the actual application scenarios and resource conditions, and instruction extension is carried out when necessary. Therefore, the circuit structures, resource consumption, computing performance, etc. of different RISC-V processors vary greatly.
[0060] Preferably, in this embodiment, the central processing unit 200 is a single-issue out-of-order execution RISC-V processor. As Figure 2 shown, the RISC-V processor includes a processor front end 210, an instruction buffer unit 220, and a processor back end 230.
[0061] The processor front end 210 includes a program decoder 211, an instruction fetch unit 212, and a dynamic branch predictor 213.
[0062] The program decoder 211 is configured to convert a computer program for controlling the operation of the spiking neural network into a RISC-V instruction stream, and send the RISC-V instruction stream to the memory 300 for storage;
[0063] The instruction fetch unit 212 is configured to access the memory 300, read a single RISC-V instruction corresponding to the program counter value (PC) of the current clock cycle, and send it to the instruction buffer unit 220; the program counter value is used as the read address for accessing the memory 300;
[0064] The dynamic branch predictor 213 is configured to predict the program counter value (PC) of the next clock cycle. In most cases, the value of the PC will automatically increment as instructions are executed, thus ensuring that the program is executed in sequence. However, there are unconditional jump instructions JAL (jump and link) and J (direct jump) instructions, as well as conditional jump instructions such as BEQ (branch if equal), BNE (branch if not equal), and BLT (branch if less than) in the RISC-V instruction set. Therefore, the RISC-V processor designed in this embodiment can implement the execution of such branch instructions (also known as jump instructions or branch jump instructions) through the dynamic branch predictor 213. Specifically, since it is impossible to know whether the instruction to be executed in the current cycle is a jump instruction before executing the instruction, during the instruction fetch stage, the dynamic branch predictor 213 needs to perform branch prediction on the PC value corresponding to each instruction to be executed, and it can be known whether the PC value is predicted successfully during the instruction decoding stage or the execution stage; if successful, the execution of the instruction in the next cycle can be quickly entered, and if not, it is necessary to flush the pipelines at all levels of the processor and restore the state of the RISC-V processor before the instruction jump. The RISC-V processor provided in the embodiment of the present invention has the functions of fast branch prediction and branch prediction failure recovery.
[0065] The instruction buffer unit 220 is used to buffer the RISC-V instructions read by the instruction fetch unit 212, and isolate the front-end computing thread and the back-end computing thread of the RISC-V processor 200. In specific implementation, the instruction buffer unit 220 is composed of a sequential first-in-first-out (FIFO) buffer with a depth of d (preferably d = 8 or d = 16). The RISC-V processor 200 is divided into a processor front-end and a processor back-end through this first-in-first-out queue. Among them, the processor front-end is responsible for generating program instructions and fetching instructions, while the processor back-end is responsible for processing the fetched instructions. The FIFO setting with a depth of 8 or 16 is a compromise solution, which can provide sufficient instruction capacity to cope with the rate differences between the front-end and the back-end and the requirements of out-of-order execution, while avoiding excessive increase in hardware resources and complexity. In traditional technical solutions, a general control logic is usually adopted to control the operation of the entire RISC-V processor. Due to the large number of states of the front-end and the back-end, the control logic is huge and becomes the critical path of the kernel. In this embodiment, the instruction buffer unit 220 is adopted. On the one hand, it can decouple the front-end / back-end of the processor. The control logic of the decoupled front-end / back-end can be implemented separately, which simplifies the control logic of the front-end / back-end and achieves a faster working frequency. On the other hand, it can solve the problem of inconsistent computing rates between the front-end and the back-end. The RISC-V processor may be blocked when accessing peripherals due to various factors (such as too long clock cycles required to access the off-chip memory 300 during front-end instruction fetching, filling the emission queue of the back-end due to multi-cycle instruction execution, etc.). The instruction buffer unit 220 can match the problem of inconsistent rates between the front-end and the back-end. When the back-end is blocked, the front-end can still continue to run, and the throughput of the pipeline is increased. When the execution of the processor back-end is blocked, if the instruction buffer is not full at this time, the instruction fetch front-end can still fetch instructions normally. The decoupling of the front-end and the back-end not only simplifies the design of the control logic but also can improve the pipeline throughput.
[0066] The processor back-end 230 is a five-stage pipeline architecture composed of a decoding unit 231, an emission unit 232, a physical register file 233, an execution unit 234, and a retirement unit 235. It is used to read the RISC-V instructions of the instruction buffer unit 220 and perform decoding, emission, reading and writing the physical register file, out-of-order calculation, and instruction retirement operations on them in sequence, so as to realize the data interaction between the RISC-V processor 200 and the neuromorphic processor 100.
[0067] In this embodiment, the neuromorphic processor 100 is an online learning neuromorphic processor based on a spiking neural network. The neuromorphic processor 100 includes: a spike encoder 130, a plurality of parallel computing spiking neural computing cores 110, at least one spiking neural decision-making core 120, and a first-level router 140. As Figure 3As shown in the figure, it is a schematic diagram of a preferred structure of the neuromorphic processor provided in this embodiment. Specifically, Figure 3 A preferred implementation of the neuromorphic processor 100 implemented based on four pulsed neural computing cores and one pulsed neural decision-making core is schematically drawn. Among them, the four pulsed neural computing cores 110 implement the function of the hidden layer of the pulsed neural network in parallel, and the pulsed neural decision-making core 120 is used to implement the function of the output layer of the pulsed neural network. The input image generates multiple frames of binary images based on a preset pulse coding method after passing through the pulse encoder 130. Each frame of binary image is converted into a multi-bit pulse sequence and then distributed to the four pulsed neural computing cores 110 for parallel computing; the parallel computing results of the four pulsed neural computing cores 110 are summarized to the pulsed neural decision-making core 120 through a first-level router 140 for inference and recognition.
[0068] As Figure 4 and Figure 5 shown, Figure 4 It is a schematic diagram of a specific structure of the pulsed neural computing core provided in the embodiment of the present invention; Figure 5 It is a schematic diagram of a specific structure of the pulsed neural decision-making core provided in the embodiment of the present invention. Among them, Figure 4 The input terminals (virtual neurons) of the input layer in [] are only for schematic signal display and are not included in the pulsed neural computing core; similarly, Figure 5 The physical neurons of the hidden layer in [] are not included in the pulsed neural decision-making core 120 either, and are only used to illustrate the signal source of the physical neurons of the output layer. Figure 6 It is a schematic diagram of a preferred structure of a pulsed neural computing core provided in the embodiment of the present invention. For the pulsed neural decision-making core 120, the structures of its components identical to those of the pulsed neural computing core can be the same, so no further drawing is shown. Specifically, based on the structure of the pulsed neural computing core shown in Figure 6 , the structure of the pulsed neural decision-making core 120 further includes an output classification decision maker, which is used to judge the recognition result of the output pulsed neural network according to the calculation results of the physical neurons.
[0069] Regarding the common part of the pulsed neural computing core and the pulsed neural decision-making core, as Figure 6 shown, each of the pulsed neural computing cores and the pulsed neural decision-making cores is internally provided with a synaptic computing module 111, a neuron computing module 112, and a routing interface 113. The neuron computing module 112 includes a transient data memory 1121, a neuron arbiter 1122, and a plurality of physical neurons NE0~NE(N-1) cascaded in a pipeline architecture. The synaptic computing module 111 pre-stores the initial synaptic weights obtained by offline training of each physical neuron; the transient data memory 1121 pre-stores the initial activation thresholds of each physical neuron.
[0070] The pulse encoder 130 is configured to perform pulse encoding on an input image based on a convolutional neural network, and generate a string of pulse signals at each time step within a preset time window to form a pulse sequence to be input into the pulse neural network.
[0071] The multiple pulse neural computing cores 110 that perform parallel computing are used to form the hidden layer of the pulse neural network, batch-connect the pulse sequence to each physical neuron in the hidden layer, and control the activation states of the physical neurons according to the approximate biological neuron model.
[0072] The pulse neural decision-making core 120 is used to form the output layer of the pulse neural network, and distribute the new pulse signals generated by the activated physical neurons in the hidden layer to the physical neurons in the output layer for target prediction.
[0073] The first-level router 140 is used to construct a communication route between the pulse encoder 130 and the pulse neural computing cores 110, distribute the pulse sequence to the multiple pulse neural computing cores 110 in the hidden layer for parallel computing; and is also used to construct a communication route between the multiple pulse neural computing cores 110 and the pulse neural decision-making core 120, and transfer the calculation results of the hidden layer to the output layer.
[0074] Preferably, the pulse encoder 130 includes a plurality of convolutional filters, a pooling layer, and a spatio-temporal encoder. Among them:
[0075] The plurality of convolutional filters are connected in sequence, and are used to perform multiple convolutional operations on the input image through trainable convolutional kernels, extract image features at different levels, and fuse them into an image feature map; the pooling layer is used to perform data dimensionality reduction on the image feature map and reduce the spatial information of the image feature map; the spatio-temporal encoder is used to divide the downsampled image feature map into multiple local spatial regions, and encode each local spatial region in sequence at a preset time interval, and quantize the intensity values of the pixels in the local spatial region into a pulse sequence that fires at a frequency.
[0076] Since spiking neural networks are driven by spike trains, this characteristic is different from that of traditional artificial neural networks which directly receive image pixel values. Therefore, it is necessary to perform spike encoding operations on images in order to input two-dimensional images into spiking neural networks for calculation. In this embodiment, the spike encoder 130 will perform rate encoding on the input image according to a preset time step (TimeStep) and time window (Time Window), and then output a multi-frame binary image spike train that matches the length of the time window. Specifically, the spike encoder 130 can batch-transport each frame of the binary image spike train of the entire row or column to the corresponding physical neurons in the hidden layer; or, transport each frame of the entire binary image spike train at once to the corresponding physical neurons in the hidden layer, and all the spike signals of the spike train will be used to accumulate the membrane potential of the corresponding physical neurons.
[0077] Specifically, when implementing, set the time window to T, and each pixel point encodes the input image for multiple time steps (Time Step) with a probability proportional to the magnitude of its pixel intensity value, respectively obtain the spike train generated by each pixel point at each time step, and finally merge them into a complete spike train corresponding to the input image.
[0078] The time window generator 114 is used to generate time steps t according to preset RISC-V instructions. The combination of all time steps forms a time window T. Within this time window T, the pulse sequence will be scheduled and transmitted to the input terminals (also known as virtual neurons) of the input layer at the rhythm of the time steps by the global controller 115. The pulse coding length of each pixel in the input image after rate coding can be adjusted by modifying the number of time steps within the time window, which will cause the expression accuracy of each pixel point to change, and ultimately affect the training and inference accuracy of the entire spiking neural network. Specifically, the entire encoded frame of the image can be processed into a multi-frame binary image pulse sequence with a corresponding number of time steps in the above process. Subsequently, under the control of instructions from the RISC-V processor, the global controller 115 distributes the pulse sequence to the hidden layers of each spiking neural computing core in a row-by-row, column-by-column, or frame-by-frame manner for calculation. The physical neurons in the hidden layer calculate the real-time change of the membrane potential of the physical neurons according to the preset neuron charging model, and determine whether to send a pulse to the physical neurons in the next layer (output layer) by comparing with the activation threshold. In each time step within the same time window, if no activation event occurs in the neuron, its membrane potential will be temporarily stored after being leaked. These temporarily stored leaked membrane potential data will be provided as input for the next neuron event together with the membrane potential of the next time step. According to the time change within the time window, the neuron membrane potential will be updated in real time with the number and frequency of input pulses. Specifically, when the membrane potential increment X(t) caused by the input pulse according to the neuron charging model makes the neuron membrane potential reach the neuron activation threshold, the neuron will emit a pulse (signal "1"), that is, an activation event occurs. Then, the approximate biological neuron model will simulate the behavior of biological neurons and reset the membrane potential state of the physical neurons, that is, the firing behavior described in the firing model occurs. After firing, the reset of the physical neurons can be started in two modes: in the hard reset mode, the reset of the neuron membrane potential is directly set to the preset reset reference potential; while in the soft reset mode, the membrane potential does not need to be completely reset, but is adjusted by the attenuation factor and the activation threshold of the current membrane potential, and gradually returns to the reset potential.
[0079] The classic network structure of a spiking neural network includes an input layer, at least one hidden layer, and an output layer. Among them, the input layer includes multiple input terminals, and each terminal is used to connect an input spiking signal. Since the input terminals themselves are not real neurons, they can also be called virtual neurons. The number of input terminals needs to match the scale of the input image in some way (such as in parallel or serial ways) to receive the spiking sequence of the binarized image after rate encoding of the input image. The spiking sequence of the multi-scale binarized image after spiking encoding received by the input layer is transmitted batch by batch and in parallel to each physical neuron involved in the hidden layer for calculation according to its charge / discharge model. Preferably, the hidden layer includes four spiking neural computation cores in parallel processing, and the output layer includes at least one spiking neural decision core. The spiking neural computation cores and the spiking neural decision cores contain multiple physical neurons, and these physical neurons are constructed through the aforementioned neuron model. The structures of the spiking neural computation cores in each hidden layer are the same, and specifically, it can be implemented using the structure as shown in Figure 6 ; The structural design principles of the spiking neural computation cores and the spiking neural decision cores are generally the same, but the structure of the spiking neural decision core needs to add an output classification decision maker (as shown in Figure 6 ) on the basis of the spiking neural computation core of Figure 5 . In specific implementation, the number and functions of the physical neurons actually participating in the neural network calculation will change due to the input signals and arrangement methods of the network layer.
[0080] Taking the spiking neural computation core shown in Figure 6 as an example to introduce the core working principles of each spiking neural computation core and the spiking neural decision core. The synaptic weight data memory in the synaptic computation module 111 pre-stores the synaptic weight W obtained through pre-training, and this synaptic weight W may be updated during future online learning. The neuron computation module pre-stores the neuron activation threshold Vth obtained through pre-training. Similarly, this neuron activation threshold Vth can be learned and updated. The neuron computation module has multiple physical neurons (NE) built in, and these physical neurons (NE) transmit the received binarized image spiking sequence in turn according to the pipeline structure and perform calculations according to the approximate neuron model. In specific implementation, the number of physical neurons (NE) in the spiking neural computation cores and the spiking neural decision cores should be set according to the needs of the actual recognition task.
[0081] Specifically, in the task of recognizing handwritten digits (from 0 to 9), the number of physical neurons NE in a pulsed neural computing core can be 20 or other numbers, corresponding to some or all of the neurons in the hidden layer of the pulsed neural network; while the number of physical neurons NE in the pulsed neural decision-making core is consistent with the number of physical neurons required for the output layer and also with the number of target categories to be recognized. For example, for the handwritten digits from 0 to 9 in the MNIST dataset, the number of physical neurons NE in the pulsed neural decision-making core should be set to 10.
[0082] During the inference process of the neuromorphic processor 100 using the pulsed neural network model, the pulsed neural computing core and the pulsed neural decision-making core perform statistics on the approximate calculation results of multiple time steps according to the pre-stored or updated synaptic weights and activation thresholds after online learning, so as to generate the recognition result of the multi-scale input image.
[0083] During the online learning process of the neuromorphic processor 100 using the pulsed neural network model, multiple images can be input into the pulsed neural computing core and the pulsed neural decision-making core, and the synaptic weights and neuron activation thresholds can be updated using a specific online learning method. Different online learning mechanisms or architectures will determine the performance of the neuromorphic processor 100 in different recognition tasks.
[0084] In the prior art, the learning method widely used to implement the training of spiking neural network models is the Spike Timing Dependent Plasticity (STDP) mechanism. However, the STDP learning mechanism still has two major defects: on the one hand, STDP relies on the precise spike time difference to adjust the synaptic weights, which is difficult to achieve precise control in circuit implementation; on the other hand, exponential and logarithmic operations need to be performed to obtain the various values required for updating the weights. These mathematical operations are difficult to accurately implement in digital circuits and often require truncating and quantifying the values. At the same time, their approximate implementation requires a large amount of hardware computing resources and energy consumption. For example, a large amount of storage space is required to store the spike information of each neuron at each time step, which is undoubtedly a significant challenge for edge devices with limited resources and costs. In addition, in a large number of experiments of the present invention, it is observed that for a trained spiking neural network model, when the spiking neural network model is trained for a new task, the knowledge of the previous old tasks cannot be guaranteed. That is, while the recognition accuracy of the new task is improved, the recognition accuracy of the old task may decrease. The reason for this phenomenon is that in the case of learning these two similar tasks, a large number of the same neurons are activated. That is, when performing these two tasks, the synaptic connection behavior between neurons in the network is very similar, and the technical problem of catastrophic forgetting occurs in the network model. Further, in a large number of studies of the present invention, it is found that the connection strength (weight) of neuron synapses is affected by the spike firing timings of the pre-neurons and post-neurons to which they are connected. Specifically, when the current neuron (pre) fires a spike before the post-neuron (post), the synaptic weight increases; while when the current neuron (pre) fires a spike after the post-neuron (post), the synaptic weight decreases.
[0085] Therefore, in the embodiments of the present invention, in order to reduce the computational complexity of the online training of the hardware-implemented spiking neural network model, rationally allocate the consumption of hardware resources, and overcome the problem of catastrophic forgetting during model training, the embodiments of the present invention propose a segmented online learning method, which updates the synaptic weights and activation thresholds of the current physical neurons by judging the strength of the correlation between the pre-neuron and the post-neuron.
[0086] Specifically, each of the spiking neural computing cores and the spiking neural decision cores further includes an online learning unit 116. The online learning unit 116 includes: a forgetting inactivator 1161, a freezer 1162, a random learner 1163, and a parameter updater 1164.
[0087] The forgetting inactivator 1161 is used to randomly inactivate the physical neurons involved in the hidden layer based on the random forgetting learning method.
[0088] The freezer 1162 is used to freeze the synaptic weights and activation thresholds of the physical neurons involved in the output layer.
[0089] The random learner 1163 is used to train the physical neurons that are not inactivated in the hidden layer, and is also used to train the physical neurons that are thawed in the output layer;
[0090] The parameter updater 1164 is used to update the synaptic weights and activation thresholds of the physical neurons after being trained by the random learner 1163. Specifically:
[0091] If the difference (|Wcs|) between the membrane potential of the physical neuron at the current time step and the activation threshold at the current time step is greater than a preset pseudo-random number (P), then the following equations are used to update the synaptic weight and activation threshold of the current physical neuron respectively:
[0092]
[0093]
[0094] Wherein, represents the synaptic weight at the next time step, represents the synaptic weight at the current time step, represents the neuron activation threshold at the next time step, represents the neuron activation threshold at the current time step, represents the time difference between the pre-neuron and the post-neuron firing pulses, a and b represent two different synaptic weight update amplitudes, and 0 < a < b; c represents the update amplitude of the neuron activation threshold, and c > 0.
[0095] Otherwise (i.e., |Wcs| ≤ P), the synaptic weight and neuron activation threshold of the physical neuron at the next time step remain unchanged. That is, when |Wcs| ≤ P, the physical neuron skips the online learning process, and the synaptic weight and activation threshold at the next time step remain unchanged, that is, W(t) = W(t - 1), Vth(t) = Vth(t - 1). Specifically in implementation, the online learning unit 116 further has a historical data buffer 1165 for temporarily storing the synaptic weights and activation thresholds of the previous time step involved in the physical neurons that need to be online trained.
[0096] In the segmented online learning process provided by the present invention, in the first stage, the forgetting inactivator 1161 is first used to randomly inactivate the physical neurons involved in the hidden layer, and the random learning device 1163 and the parameter updater 1164 are used to implement online real-time updates of the synaptic weights and neuron activation thresholds connected to each physical neuron in the hidden layer by using the random plasticity method based on historical weights; in the second stage, the freezer 1162 is used to freeze the synaptic weights and activation thresholds of the physical neurons involved in the output layer, and after selecting some physical neurons or neuron parameters in the output layer to be unfrozen, the random learning device 1163 and the parameter updater 1164 are used to implement the training of the output layer by using the random plasticity method based on historical weights, and online real-time updates of the synaptic weights and neuron activation thresholds connected to the corresponding physical neurons are performed.
[0097] Specifically, in the online learning process of the first stage, the forgetting inactivator 1161 can randomly set the output pulses of some neurons in the hidden layer to "0", so as to randomly forget the connections between some neurons in the hidden layer and the neurons in the output layer. Considering the premise of being hardware-friendly, the forgetting inactivator 1161 can be implemented by the random inactivation pulse neuron method or the random inactivation synaptic connection method. Among them, in the random inactivation pulse neuron method, after calculating the membrane potential of the current physical neuron, the membrane potential of the physical neuron is set to "0" (that is, the output pulse is set to "0") with a preset probability; in the random inactivation synaptic connection method, between two layers of physical neurons, the synapses of the neurons to be forgotten are randomly inactivated, and only the activated synaptic inputs are accumulated when calculating the membrane potential of the physical element. In the specific hardware circuit design, the random inactivation pulse neuron method can be implemented by using the method of randomly turning off the enable terminal of the physical neuron; the random inactivation synaptic connection method can be implemented by randomly turning off the read enable terminal of the synaptic weight, without too much additional circuit logic, which conforms to the characteristics of being hardware-friendly. Through the embodiments of the present invention, by dynamically adjusting the enable states of neurons or synaptic weights, while maintaining the memory of the original task, new synaptic connection paths are randomly built for new tasks through the synaptic connections between some neurons, endowing the network with higher adaptability, enabling it to efficiently learn new tasks under limited hardware resources, and at the same time avoiding excessive interference with the learned content, thereby realizing the ability of continuous learning, and improving the adaptability and generality of the neural network to dynamic changes in the environment.
[0098] During the online learning process of the second stage, the physical neurons in the output layer are trained using the method of frozen learning. Since there are no synaptic connections from the output layer to the next network layer, the above-mentioned forgetting learning is no longer applicable to the online learning of the physical neurons in the output layer. To overcome the catastrophic forgetting of the physical neurons in the output layer, the present invention adopts a learning method of parameter freezing. For a spiking neural network that has been trained for an old task, when a new task arrives, all the parameters of some physical neurons in the output layer are frozen, or, some parameters (such as synaptic weights) of all the physical neurons in the output layer are frozen and only other trainable parameters (such as activation thresholds) are fine-tuned. This frozen learning method limits the interference of the new task on the synaptic weights of the original task, thereby reducing catastrophic forgetting. The synaptic weights and activation thresholds between the frozen neurons do not undergo any update, and only the synaptic connections of the neurons related to this task are fine-tuned. This non-global update method can avoid the situation of catastrophic forgetting caused by large-scale weight adjustment in a well-trained neural network, and is very friendly to hardware implementation. When learning a new task, it significantly reduces the computational core storage overhead.
[0099] In the model training of the above two stages, the present embodiment adopts a stochastic plasticity learning method based on historical information to train the corresponding physical neurons. The update of the synaptic weights and activation thresholds of the physical neurons is mainly realized by the stochastic learner 1163 and the parameter updater 1164. The working principle of this stochastic plasticity learning method based on historical information is as follows: If the relative difference (|W cs |) between the membrane potential and the activation threshold of a physical neuron at the current time step (t - 1) is greater than a preset pseudo-random number (P), then the synaptic weight and activation threshold of the current physical neuron at the next time step (t) are updated according to the aforementioned mathematical model; otherwise, the synaptic weight and activation threshold of the current physical neuron at the next time step (t) remain the current values. This method can significantly reduce the computational complexity and improve the accuracy of the neuromorphic processor implemented based on hardware during the training and inference processes, making the use efficiency of hardware resources better. This stochastic learning method that requires information from previous time steps is called the stochastic plasticity learning method based on historical information in the embodiments of the present invention.
[0100] The neuromorphic processor 100 with online learning ability can achieve more application-scenario-targeted object recognition. By inputting images in real time to implement model training and parameter adjustment, the neuromorphic processor 100 can still maintain a high recognition accuracy when facing different recognition tasks.
[0101] The neuron arbiter 1122 in the neuron computing module 112 is used to deliver the relevant network parameters of the most active physical neurons to the online learning unit 116 and the synaptic computing module 111 for learning and updating. This decision-making process will determine the neuron targets that ultimately participate in online learning, taking into account not only factors such as the membrane potential and synaptic weights of the current neurons, but also evaluating their contributions to the overall network performance. Once the neuron arbiter 1122 determines the neurons that need to be updated, in order to trigger the learning operations of specific weights and activation thresholds, it will send update signals to the online learning unit 116 and the synaptic computing module 111. Among them, the most active physical neurons are obtained through a competitive learning mechanism. The more active a neuron is, the more likely it is to be activated, and thus the synaptic weights and activation thresholds are updated, which conforms to the basic characteristics of biological learning.
[0102] The random learner 1163, based on a linear feedback shift register (LFSR), generates a pseudo-random number P as a reference value for whether to perform random learning. The LFSR includes a set of cascaded registers. The values in each register are shifted according to each clock cycle, and a new input bit needs to be generated based on the feedback signals output from some registers. Based on this principle, the random learner 1163 uses the LFSR to generate a cyclic pseudo-random sequence through the combined action of the shift register and the feedback logic. The random learner 1163 compares the relative difference W cs of the neurons mentioned above cs and the pseudo-random number P. If W
[0103] is greater than P, then the next step will be to control the parameter updater 1164 to start the online learning logic; otherwise, the learning process for this time step will be skipped.
[0104] This embodiment takes the working principle of a single spiking neural computing core in the hidden layer as an example to introduce the online learning process of physical neurons. During the online learning process of the spiking neural network, the spike sequence of each frame of binary image after spike coding is input into the neuromorphic processor 100 for training in a predetermined order, and at the same time, the number of training rounds is determined. The training process is triggered once per time step. The global controller 115 is responsible for controlling the operation of the entire online learning process during the training process, coordinating the work of each sub-module, receiving the time step information from the time window generator 114 and the feedback from other modules, and ensuring the correct timing and function execution of the training process. When the online learning function is executed, the neuron computing module 112 will send the historical synaptic weights and historical input spikes that have participated in the accumulation within the same time window into the online learning unit 116 to participate in the learning update process, receive the updated neuron activation threshold, and control the output of the updated synaptic weights to the synaptic weight memory 1111 of the synaptic computing module 111 for storage. The transient data memory 1121 in the neuron computing module 112 is used to store parameters such as the membrane potential, activation threshold, and activation state of physical neurons.
[0105] Specifically, according to the modeling equation of the foregoing parameter updater 1164, the neuron online learning process provided by the embodiments of the present invention includes: in the first piecewise function, and are linearly related in each segment, and their change relationship is affected by two parameter values a and b, which are empirical values obtained through multiple sufficient experiments. Preferably, 0 < a < b, and the historical weight is selected as the synaptic weight value one time step away. . Compared with the in the conventional STDP online learning method, the time difference between the spikes emitted by the pre-neuron and the post-neuron adopted in the embodiments of the present invention only needs to be used to indicate the order of spike release of the pre-neuron and the post-neuron. Specifically, greater than 0 indicates that the post-neuron releases spikes before the pre-neuron; on the contrary, it indicates that the pre-neuron releases spikes before the post-neuron. Therefore, in this embodiment, there is no need to store a large amount of historical spike release information, and only the spike release order of the pre-neuron and the post-neuron at the current time step needs to be obtained. At the same time, the embodiments of the present invention can reflect the influence of different historical weights on the synaptic weight adjustment amplitude by reasonably setting the values of a and b. In the embodiments of the present invention, particularly, a can be set to 1 and b can be set to 2 to represent different synaptic weight update amplitudes.
[0106] Furthermore, in this embodiment, an amplitude update table of the change amplitude c of the activation threshold can be obtained through experiments, and the activation threshold can be dynamically adjusted according to the situation of historical physical neurons emitting spikes at each time step. Specifically, if representing an increase in synaptic weight, the activation threshold is increased; if within the next time window it persists all the time, the change amplitude c of the activation threshold will increase according to the value preset in the amplitude update table. Conversely, if representing a decrease in synaptic weight, the activation threshold is decreased; and if within the next time window it persists all the time, the change amplitude c of the activation threshold will decrease according to the value preset in the amplitude update table. In particular, a large number of experiments show that the value of the change amplitude c can be set in stages to 1, 2, and 3 in the amplitude update table. Note that the change amplitude of c will be initialized due to the change of
[0107] In the biological nervous system, the update of synaptic weight is a non-deterministic process, which shows significant randomness and probability (such as neuron activity, voltage change, neurotransmitter release, etc.), and is affected by various factors. By introducing a probability mechanism to control the update of synaptic weight, the characteristics of biological learning can be more realistically simulated. In addition, this method helps to prevent the problem of overfitting of the neural network model to a certain type of data due to frequent updates. Overfitting usually makes the network perform excellently on the training data, but lacks good generalization ability in practical applications. The update of each neuron and the update of synapses are independently completed in the neural network. By setting the update probability, more diverse synaptic connections can be formed, thereby enhancing the ability to predict unknown data and helping to more comprehensively adapt to the input space.
[0108] It should be noted that although the structures of each spiking neural computing core and spiking neural decision-making core are the same or similar, the synaptic weights and their own activation thresholds are stored in the circuit structures of each physical neuron. These model parameters will be dynamically updated according to the input pulses during the training process, so that computational differences occur in each spiking neural computing core and spiking neural decision-making core. That is, the membrane potential and activation state of each physical neuron have different statuses in the spiking neural network.
[0109] In a preferred embodiment, the embodiments of the present invention can implement the spiking neural computing core 110 and the spiking neural decision-making core 120 based on a digital integrated circuit (IC) chip to achieve high-speed and low-power neuron activation and reset.
[0110] Furthermore, as Figure 7 shown, each physical neuron (abbreviation: NE) includes: a PE approximate computing array, an approximate adder tree, and a spike generation decision unit.
[0111] Specifically, the PE approximate calculation array is composed of multiple PE (Processing Element) units adopting a parallel pipelined calculation architecture, and is used to access the pulse sequences in batches and distribute each pulse signal accessed in each batch to the corresponding PE unit in parallel for pipelined calculation; each PE unit includes an accumulation control logic and at least one first approximate adder; each PE unit is used to, under the control of the accumulation control logic, use the first approximate adder to accumulate the local pulse sequences accessed in each batch to obtain the local membrane potential increment within a preset time window; and according to the charging model, calculate the local membrane potential of the physical neuron in the current PE unit by using the local membrane potential increment; the approximate adder tree is hierarchically composed of multiple second approximate adders adopting a binary tree structure, and is used to superimpose the local membrane potentials output by each PE unit to obtain the complete membrane potential accumulated by the pulse sequence on the current physical neuron; the pulse firing decision unit is used to compare the complete membrane potential with the activation threshold of the current physical neuron to determine whether the current physical neuron is activated; if the current physical neuron is not activated, new pulse signals are accessed and the membrane potential of the current physical neuron is charged according to the charging model; if the current physical neuron is activated, new pulses are fired, and after the new pulses are fired, the membrane potential of the current physical neuron is reset according to the discharge model.
[0112] Within a neuron time step, the PE array calculates the local current of some of the input pulses, and the approximate adder tree accumulates the local currents to obtain the cumulative membrane potential, which determines the behavior of the current physical neuron at this time step: when the cumulative membrane potential of the physical neuron exceeds the activation threshold, a pulse signal is emitted and the membrane potential is reset; otherwise, the membrane potential decays according to the neuron model and is stored for the next time step. Thus, in this embodiment, the training process of the spiking neural network includes all the modules and calculation steps in the inference process, that is, the training subsystem calls each hardware module of the inference subsystem to update parameters such as the membrane potential, synaptic weight, and activation threshold of the neurons, thereby optimizing the performance of the spiking neural network. The PE approximate calculation array accesses the local pulse sequence of the binarized image in batches at each time step, and uses this local pulse sequence to screen the input synaptic weight sequence. If the pulse in the sequence is "1", the synaptic weight corresponding to the address is selected. After the corresponding synaptic weights are selected for each pulse respectively, the synaptic weights are sent to the approximate adder tree for summation and temporary storage. By time-division multiplexing the PE approximate calculation array, the synaptic weights corresponding to the addresses of multiple local pulse sequences (rows or columns) of the binarized image can be sent to the PE array in batches according to the input image scale for selection. Each PE obtains the local current at each batch time step according to the preset neuron model and uses the local pulse sequence input at each batch time step to calculate the local membrane potential increment of the physical neuron at the current time step. And at each time step, the local pulse sequence is sent to each cascaded physical neuron NE through each stage of the pipeline for separate calculation. The approximate adder tree is used to sum the local currents output by the PE approximate calculation array in a pipelined manner to calculate the complete membrane potential of the physical neuron corresponding to the accessed binarized image pulse sequence.
[0113] Such as Figure 7As shown, the pulse generation decision unit is a configurable hardware architecture, which incorporates a third approximate adder required for the neuron charging model and the firing model, and a logic circuit that includes the membrane potential and the pulse sequence of the entire physical neuron output. Among them, the third approximate adder includes each adder required for calculating the optimized approximate neuron charging model and firing model described above. Using the optimized LIF neuron model based on approximate calculation, the pulse generation decision unit calculates the activity value of the current physical neuron, which is used to simulate the membrane potential change of the biological neuron and participates in data preparation during the online learning stage. Specifically, the pulse generation decision unit is used to compare the membrane potential summarized by the approximate adder tree with the neuron activation threshold to determine whether the current physical neuron needs to be activated; if the current physical neuron is not activated, its membrane potential is decayed according to the preset neuron charging model; if the current physical neuron is activated, a new pulse is sent to the next layer of neurons, and after the pulse is sent, the membrane potential of the current neuron is reset according to the preset neuron firing model. In addition, the pulse generation decision unit is also used to select the corresponding neuron model according to the preset conditions and calculate the activity value of the physical neuron, which will participate in the functional logic of online learning. The synaptic computing module 111 stores the synaptic weights of the synapses connected to each physical neuron through the synaptic weight memory 1111, and transmits the synaptic weights corresponding to each physical neuron to the neuron computing module through the synaptic arbiter 1112 to participate in inference or training. Within each preset time window, the PE approximate computing array and the approximate adder tree work together to calculate the complete membrane potential of the current physical neuron, and transmit the membrane potential value to the configurable pulse generation decision unit at each time step. The configurable pulse generation decision unit compares with the preset neuron model (optimized LIF model) and the neuron activation threshold to determine whether the current physical neuron emits a pulse, and at the same time uses the synaptic arbiter 1112 to orderly store the reset or leaked neuron membrane potential.
[0114] In this embodiment, a pipeline parallel computing structure is formed among multiple physical neurons (NE) built in the neuron computing module, and the pulse sequence of the binary image accessed is transmitted to each NE along with the clock of the computing system. Time-division multiplexing technology is applied inside each physical neuron to distribute the multiple input pulse signals corresponding to each clock cycle to each PE for parallel computing during each clock cycle, so as to use the approximate adder tree to accumulate the local membrane potential of the current physical neuron according to the approximate neuron model. As Figure 7As shown, the one-dimensional PE array is composed of multiple PEs. By using time-division multiplexing, the array circuit correctly realizes the function of membrane potential accumulation, significantly reducing the circuit area and power consumption. To further improve the computing efficiency, the time-division multiplexing technology is adopted to batch the pulse sequence of the input image into the physical neuron for parallel superposition according to the circuit clock as the time step (T). Specifically, the pulse sequence of each row or each column (i.e., M or N pixels) of the input image will be passed to the PE array of the physical neuron through each input. The PE array is composed of M PE units in parallel. By multiplexing the partial synaptic weights W0,0~W32,T N times, the pulse signals of the M×N size of the input image can be all passed to the physical neuron for membrane potential calculation. Through reasonable instruction control, different sizes of input images can be processed with limited hardware resources.
[0115] In the pulsed neural computing core and the pulsed neural decision-making core, the adder is the core component for realizing the PE accumulation in the physical neuron, the approximate adder tree, and the configurable pulse generation decision unit. To obtain a hardware-friendly neuron circuit with low power consumption and high speed and better simulate the operation mechanism of biological neurons, in a preferred implementation manner, the adders adopted in this embodiment are all speculative carry adders with an error correction mechanism. Therefore, the above-mentioned first approximate adder and the second approximate adder are both speculative carry adders with an error correction mechanism. Specifically, for two input signals A and B to be added with a bit width of n, the signals A and B can be divided into m segments, and the bit width of each sub-signal segment is k bits, that is, n = m·k.
[0116] The speculative carry adder includes m k-bit sub-adders for respectively accessing and accumulating m segments of k-bit input sub-signals, where the m segments of k-bit input sub-signals are divided from the input signal with a total bit width of n = m·k;
[0117] Each of the sub-adders includes a speculative carry generator; the speculative carry generator is used to speculate the approximate carry signal of the corresponding sub-adder according to the following equation:
[0118]
[0119] where i represents the number of the sub-adder and 2 ≤ i ≤ m, represents the approximate carry signal of the i-th sub-adder; the approximate carry signal of the first sub-adder is = 0; k represents the bit width of each sub-adder, represents the carry identification signal of the (k - 1)-th bit generated by the i-th sub-adder, represents the carry identification signal of the (k - 1)-th bit generated by the (i - 1)-th sub-adder, represents the carry propagation signal of the j-th bit output by the i-th sub-adder;
[0120] A multiplexer is also provided between each of the sub - adders, which is used to select the carry flag signal of the previous sub - adder or the approximate carry signal of the previous sub - adder as the carry input signal of the current sub - adder according to the state value of the carry propagation signal of the current sub - adder.
[0121] The extension of the critical path of the adder will directly lead to an increase in the calculation time of the entire processor, thereby reducing the calculation efficiency. Therefore, an embodiment of the present invention proposes an approximate adder, which adopts a carry speculation technique to shorten the carry propagation path and reduce the length of the carry chain.
[0122] When the speculated carry signal of the sub - adder is "0" but the actual carry signal is "1", such an error will cause an accumulated error of the 2i level, which will have an irreversible and serious impact on the entire neural network inference or training process, resulting in the collapse of the entire computing system. In this embodiment, to prevent the above - mentioned error, an error - correction mechanism is introduced: a "two - to - one" multiplexer is inserted between each sub - adder block. Among them, the carry flag signal of the previous sub - adder and the approximate carry signal generated by the previous sub - adder serve as the input signals of the multiplexer, and the carry propagation signal of the current sub - adder serves as the selection signal of the multiplexer. In this embodiment, there is no need to design an additional circuit to generate the carry propagation signal because it has already been generated in the carry generator of each sub - adder. In this embodiment, the function of the error - correction circuit is: if the carry propagation signal is "1", the output of the multiplexer is selected as , otherwise is selected, and the output signal serves as the carry input signal of the current sub - adder. Through this improvement, although the area of the adder increases slightly, the accumulated error caused by carry errors can be significantly reduced, thereby enhancing the reliability of the circuit.
[0123] In addition, it cannot be ignored that in the addition operation of signed numbers, an error in the sign bit will cause a large calculation deviation. Although the approximate adder allows a certain range of precision loss, an error in the sign bit belongs to an unacceptable error type. To solve this error, a second - level error - correction mechanism is introduced in this embodiment: by calculating the logical "AND" result of the carry propagation signal and as the sign error - correction signal . When it is detected that any level of signal is "1", the sign - bit error - correction mechanism is triggered, and it is judged whether to select the inferred approximate carry signal As the final carry signal. This error correction mechanism can correct the sign bit to avoid irreversible excessive deviation or major calculation errors. The maximum relative error of the approximate adder in this embodiment is , when k is greater than or equal to 3, due to the strong robustness of the spiking neural network to approximate calculations, this error hardly affects the spike distribution of neurons.
[0124] In summary, the speculative carry adder with an error correction mechanism provided in this embodiment effectively reduces the delay of the critical path by optimizing the critical path delay, significantly improving the maximum operating frequency of the processor. Moreover, the dual error correction mechanism designed in this embodiment helps to ensure the addition accuracy during the membrane potential accumulation calculation in the spiking neural network, avoiding the failure of the spiking neural network due to approximate calculations and ensuring the reliability of the neuromorphic processor 100. Adopting the technical method of approximate calculation in the entire neuromorphic processor 100 can effectively reduce the operating power consumption of the entire processor.
[0125] In addition, this embodiment further reduces unnecessary computational overhead according to the characteristics of the pulse sequence after pulse coding of the input image. Further, the neuromorphic processor 100 further includes: a pulse density discriminator.
[0126] The pulse density discriminator is used to discriminate the pulse sequence as a dense pulse sequence or a sparse pulse sequence according to the pulse density of the pulse sequence.
[0127] The PE approximate calculation array is further used to, when the pulse sequence is a dense pulse sequence, calculate the complete membrane potential of the pulse sequence on the current physical neuron by using the approximate adder tree; when the pulse sequence is a sparse pulse sequence, calculate the complete membrane potential of the pulse sequence on the current physical neuron by using a logic OR gate circuit. Experimental data show that the resource consumption of the logic "OR" gate circuit is much less than that of the approximate adder, and the calculation efficiency of the logic gate is much higher than that of the approximate adder.
[0128] After a large number of experiments in the embodiments of the present invention, it is found that since the pulse encoder performs multiple encodings of the same input image at different times (time steps), there are often many repeated pulse sequences at the corresponding positions of the multiple frames of binary input images obtained. Specifically, when starting the spiking neural network for image recognition, each spiking neural computing core receives the input pulse sequences of different time steps of the same input image. Since they are from the same image, the synaptic weights are exactly the same (in the case of not performing online learning), and there are also partial repetitions of the input pulse sequences in different spiking neural computing cores. Therefore, during the inference process, partial neuron membrane potentials and all synaptic weights can be reused. Preferably, as Figure 7As shown, the complete pulse sequence of a frame of input image can be decomposed into a non-multiplexed partial sequence and a multiplexed partial sequence. The synaptic weights are multiplexed into all four spiking neural computing cores in the hidden layer. In particular, the complete input pulse sequence is input into one of the spiking neural computing cores, while the non-multiplexed partial sequence is correspondingly input into the other spiking neural computing cores. After the complete neuron membrane potential is calculated in the first spiking neural computing core, the neuron membrane potential corresponding to the multiplexed partial sequence is separated from it and added to the membrane potential of the non-multiplexed partial sequence in their respective computing cores, so as to obtain the complete neuron membrane potential of multiple frames of binary input images. Then, the complete neuron membrane potentials corresponding to the four cores are transmitted to the first-level router, and the neuron computing results of the hidden layer are sent to the output layer for inference decision-making. The advantage of such sparse activation of the physical neurons in the hidden layer is that it can further reduce the redundant computational overhead in the repeated part during the inference process, achieving the design goal of low power consumption. This embodiment utilizes the sparse characteristics of the pulse sequence, reduces unnecessary computational overhead, improves the learning efficiency and effect, and at the same time reduces the energy consumption, which is of great help to the improvement of the stability and accuracy of the model.
[0129] In addition, the communication system of the neuromorphic processor 100 may further include a second-level router 150, which is used to construct a large-scale spiking neural network through multiplexing multiple spiking neural computing cores and spiking neural decision-making cores to achieve target recognition of big data. The first-level router 140 and the second-level router 150 are hierarchical routing architectures, which can realize the routing control of data communication in the inference process and training process of a large-scale neuromorphic computing system, and the routing control of multi-level data transmission between each spiking neural computing core and spiking neural decision-making core.
[0130] The internal cores of the first-level router 140 and the second-level router 150 have basically the same composition and working principle, but through applying data communication at different levels, the data they process and the direction of data transfer are different. Figure 8 Shows the internal core components and working principle of the first-level router 140 and the second-level router 150. Specifically, the router internally includes at least four stages: input buffering, path calculation and channel switching, virtual channel allocation, and transmission. The first stage is input buffering, which is used to distribute the data packets accessed from various external directions to the buffers of multiple channels inside it, and then extract specific information segments (filts) according to the encapsulation information carried by the data packets and send them to the corresponding channels or paths. Since the information segments (filts) may need to be distributed to multiple spiking neural computing cores or multiple overall approximate neuromorphic processors for calculation, the router internally needs to have a corresponding path calculation and channel switching mechanism (the second stage). The third and fourth stages are mainly used to realize the data interaction and external transmission of the information segments.
[0131] Each of the spiking neural computing cores and spiking neural decision cores on the neuromorphic processor transmits spiking data packets through the internal routing interface 113 and the first-level router 140, thus forming an on-chip routing system with a star topology. The first-level router 140 communicates with the second-level router 150 to implement data interaction between the current neuromorphic processor and an external neuromorphic processor. In the routing mechanism of the first-level router 140, when a spiking data packet is received, it is decoded into corresponding input signals and transmitted to the spiking computing core and the spiking decision core for processing; after the processing is completed, the output signals will be re-encoded into spiking data packets and used for data communication between different spiking cores through the first-level router 140. The routing mechanism of the second-level router 150 is similar and will not be elaborated here.
[0132] In summary, based on ensuring efficient network transmission, the embodiments of the present invention have successfully implemented a neuromorphic computing system with significantly improved computing speed and recognition accuracy.
[0133] Furthermore, based on the same inventive concept as the above neuromorphic computing system based on a spiking neural network, the embodiments of the present invention also provide a neuromorphic computing method based on a spiking neural network. The method includes:
[0134] Writing a computer program; inputting the computer program into the neuromorphic computing system based on a spiking neural network for execution to achieve real-time object recognition of an input image. Since the main technical innovation points of this image object recognition method are the same as those of the neuromorphic computing system, the step process of the image object recognition method will not be elaborated here.
[0135] Specifically, in this embodiment, the neuromorphic computing system based on the spiking neural network is further implemented based on the 40-nanometer CMOS semiconductor process. Among them, on the MNIST handwritten digit dataset (including 70,000 handwritten digit images of ten categories from 0 to 9, including 60,000 training images and 10,000 test images), the accuracy of the neuromorphic computing system using the neuromorphic processor for handwritten digit recognition is verified. This neuromorphic computing system constructs 330 physical neurons on the neuromorphic processor 100 based on digital integrated circuits. Each physical neuron is connected to a total of 128,640 synapses, and the model parameters of the spiking neural network (mainly synaptic weights) are quantized to 4 bits for online learning. Finally, under the 40-nanometer CMOS digital process, the area of the neuromorphic processor is only 0.95 mm2, the operating voltage is 0.8 V, the maximum operating frequency is 500 MHz, and the dynamic power consumption is between 75.6 and 93.52 μW / MHz, demonstrating the performance advantages of the neuromorphic processor in terms of recognition accuracy, computing efficiency, etc. At the same time, the classification accuracy of this embodiment on the MNIST dataset reaches 95%, verifying the image target recognition ability of this neuromorphic computing system. Compared with existing recognition processors, the neuromorphic processor provided in this embodiment requires a larger number of physical neurons, and relying on the designed multi-layer routing can support the needs of various scale neural computing tasks. Compared with other processors, this embodiment better balances the trade-offs among area, power consumption, and accuracy of the neuromorphic computing system. Through careful design and optimization, the neuromorphic computing system in this embodiment shows significant advantages in terms of the number of neurons, computing accuracy, power consumption efficiency, and area utilization. This enables it to outperform existing similar processors in terms of performance, achieving an excellent balance between power consumption control and performance integration, providing strong support for future applications of neuromorphic computing.
[0136] This embodiment combines a bio-inspired computing model with a hardware architecture, and proposes a design scheme of a neuromorphic processor that can flexibly adapt to and efficiently complete complex tasks. This processor is based on a spiking neural network, has low power consumption, high precision, and online learning ability, and can efficiently simulate the dynamic computing process of the biological nervous system. The neuromorphic computing system provided in this embodiment not only improves the computing speed and energy efficiency, but also effectively reduces the consumption of hardware resources. Its application scope is wide, suitable for fields such as medical image analysis, autonomous driving, intelligent monitoring, and smart city, aiming to solve the deficiencies of existing technologies in terms of performance, power consumption, and flexibility.
[0137] The embodiments described above are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.
Claims
1. A neuromorphic computing system based on a spiking neural network, characterized in that: The system includes: A neuromorphic processor is used to construct a hardware spiking neural network using a plurality of physical neurons, and to perform pulse encoding and target recognition on a real-time input image through the spiking neural network; the physical neurons are constructed using an approximate biological neuron model; the approximate biological neuron model has pre-stored synaptic weights and activation thresholds; A central processing unit implemented based on the RISC-V instruction architecture, communicating with the neuromorphic processor, for executing control instructions required to construct the spiking neural network; The neuromorphic processor is further used to perform online learning on the spiking neural network under the control of the control instruction, and update the synaptic weight and activation threshold associated with each physical neuron of the spiking neural network; A memory for storing program instructions required for the operation of the central processing unit, the program instructions including control instructions required for constructing the spiking neural network; and for storing data transmitted between the central processing unit and the neuromorphic processor; The approximate biological neuron model includes a charging model and a discharging model, wherein: The charging model is: Baseline decay charging model: ,or, Incremental decay charging model: , The discharge model is: Hard reset discharge model: ,or, Soft reset discharge model: , Among them, V(t) is the membrane potential of the physical neuron at the current time step, V(t-1) is the membrane potential of the physical neuron at the previous time step, and X(t) is the membrane potential increment formed by the pulse signal input at the current time step. is the preset membrane potential attenuation factor; is the reset reference potential of the physical neuron; V(t + ) is the reset membrane potential of the physical neuron at the current time step, V(t - ) is the membrane potential of the physical neuron before reset at the current time step; V th is the activation threshold of the current physical neuron; The pulse neural network constructed using the above approximate biological neuron model can support multiple types of synaptic connections and can simulate neurons of different models.
2. The neuromorphic computing system based on a spiking neural network according to claim 1, characterized in that: The central processing unit is a single-issue out-of-order RISC-V processor; the RISC-V processor includes a processor front end, an instruction buffer unit and a processor back end; The processor front end includes a program decoder, an instruction fetch unit and a dynamic branch predictor; The program decoder is used to convert the computer program that controls the operation of the pulse neural network into a RISC-V instruction stream, and transmit the RISC-V instruction stream to the memory for storage; The instruction fetch unit is used to access the memory, read a single RISC-V instruction corresponding to a program counter value of a current clock cycle and transmit it to the instruction buffer unit; the program counter value is used as a read address for accessing the memory; The dynamic branch predictor is used to predict the program counter value of the next clock cycle; The instruction buffer unit is used to buffer the RISC-V instructions read by the instruction fetch unit, and isolate the front-end computing thread from the back-end computing thread of the RISC-V processor; The processor backend is a five-stage pipeline architecture consisting of a decoding unit, an emission unit, a physical register stack, an execution unit and a retirement unit, which is used to read the RISC-V instructions of the instruction buffer unit and perform decoding, emission, reading and writing of the physical register stack, out-of-order calculation, and instruction retirement operations on them in sequence, so as to realize data interaction between the RISC-V processor and the neuromorphic processor.
3. The neuromorphic computing system based on a spiking neural network according to claim 2, characterized in that: The neuromorphic processor includes: a pulse encoder, a plurality of pulse neural computing cores for parallel computing, at least one pulse neural decision core, and a first-level router; Each of the pulse neural computing core and the pulse neural decision core has a built-in synaptic computing module, a neuron computing module and a routing interface; the neuron computing module includes a transient data storage, a neuron arbitrator, and a plurality of physical neurons cascaded using a pipeline architecture; the synaptic computing module pre-stores the initial synaptic weights obtained by offline training of each physical neuron; the transient data storage pre-stores the initial activation threshold of each physical neuron; The pulse encoder is used to pulse encode the input image based on the convolutional neural network, generate a series of pulse signals every other time step within a preset time window, and form a pulse sequence to be input into the pulse neural network; The plurality of parallel computing pulse neural computing cores are used to form a hidden layer of the pulse neural network, connect the pulse sequence to each physical neuron in the hidden layer in batches, and control the activation state of each physical neuron according to the approximate biological neuron model; The pulse neural decision core is used to form the output layer of the pulse neural network, and distribute the new pulse signals generated by each physical neuron activated in the hidden layer to the physical neurons in the output layer for target prediction; The first-level router is used to construct a communication route between the pulse encoder and the pulse neural computing core, and distribute the pulse sequence to the multiple pulse neural computing cores in the hidden layer for parallel calculation; it is also used to construct a communication route between the multiple pulse neural computing cores and the pulse neural decision core, and transmit the calculation results of the hidden layer to the output layer.
4. The neuromorphic computing system based on a spiking neural network according to claim 3, characterized in that: Each of the spiking neural computing core and the spiking neural decision core also includes an online learning unit; the online learning unit includes: a forgetting deactivator, a freezer, a random learner and a parameter updater; The forgetting deactivator is used to randomly deactivate the physical neurons involved in the hidden layer based on a random forgetting learning method; The freezer is used to freeze the synaptic weights and activation thresholds of the physical neurons involved in the output layer; The random learner is used to train the physical neurons that are not inactivated in the hidden layer, and to train the physical neurons that are unfrozen in the output layer; A parameter updater is used to update the synaptic weights and activation thresholds of the physical neurons trained by the stochastic learner, specifically: If the difference between the membrane potential of the physical neuron at the current time step and the activation threshold at the current time step is greater than the preset pseudo-random number, the following equations are used to update the synaptic weight and activation threshold of the current physical neuron respectively: Among them, represents the synaptic weight at the next time step, represents the synaptic weight at the current time step, represents the neuron activation threshold at the next time step, represents the neuron activation threshold at the current time step, represents the time difference between the pre - neuron and the post - neuron firing pulses. a and b represent two different synaptic weight update amplitudes, and 0 < a < b; c represents the update amplitude of the neuron activation threshold, and c > 0; Otherwise, the synaptic weights and neuron activation thresholds of the physical neuron at the next time step remain unchanged.
5. The neuromorphic computing system based on a spiking neural network according to claim 3, characterized in that: Each of the physical neurons comprises: a PE approximate calculation array, an approximate addition tree and a pulse emission decision unit; The PE approximate calculation array is composed of multiple PE units using a parallel pipeline calculation architecture, which is used to access the pulse sequence in batches and issue each batch of accessed pulse signals in parallel to the corresponding PE units for pipeline calculation; Each PE unit includes accumulation control logic and at least one first approximate adder; Each of the PE units is used to accumulate the local pulse sequences connected in each batch by using the first approximate adder under the control of the accumulation control logic to obtain the local membrane potential increment within the preset time window; and according to the charging model, calculate the local membrane potential of the physical neuron in the current PE unit by using the local membrane potential increment; The approximate addition tree is composed of a plurality of second approximate adders hierarchically formed by a binary tree structure, and is used to superimpose the local membrane potential output by each PE unit to obtain the complete membrane potential accumulated by the pulse sequence on the current physical neuron; The pulse emission decision unit is used to compare the complete membrane potential with the activation threshold of the current physical neuron to determine whether the current physical neuron is activated; if the current physical neuron is not activated, a new pulse signal is connected and the membrane potential of the current physical neuron is charged according to the charging model; if the current physical neuron is activated, a new pulse is emitted, and after the new pulse is emitted, the membrane potential of the current physical neuron is reset according to the discharge model.
6. The neuromorphic computing system based on a spiking neural network according to claim 5, characterized in that: The neuromorphic processor further includes: a pulse density discriminator; The pulse density discriminator is used to discriminate the pulse sequence as a dense pulse sequence or a sparse pulse sequence according to the pulse density of the pulse sequence; The PE approximate calculation array is also used to calculate the complete membrane potential of the pulse sequence on the current physical neuron using the approximate addition tree when the pulse sequence is a dense pulse sequence; and to calculate the complete membrane potential of the pulse sequence on the current physical neuron using a logic or gate circuit when the pulse sequence is a sparse pulse sequence.
7. The neuromorphic computing system based on a spiking neural network according to claim 3, characterized in that: The pulse encoder includes a plurality of convolution filters, a pooler, and a spatiotemporal encoder; The multiple convolution filters are connected in sequence, and are used to perform multiple convolution operations on the input image through a trainable convolution kernel, extract image features at different levels and fuse them into an image feature map; The pooler is used to perform data dimensionality reduction on the image feature map to reduce the spatial information of the image feature map; The spatiotemporal encoder is used to divide the image feature map after dimensionality reduction into multiple local spatial regions, and encode each of the local spatial regions in sequence according to a preset time interval, and quantize the intensity value of the pixel in the local spatial region into a pulse sequence emitted at a frequency.
8. The neuromorphic computing system based on a spiking neural network according to claim 5, characterized in that: The first approximate adder and the second approximate adder are both speculative carry adders with error correction mechanisms; The speculative carry adder includes m k-bit sub-adders, which are used to accumulate the corresponding m-segment k-bit input sub-signals, wherein the m-segment k-bit input sub-signals are divided by the input signal with a total bit width of n=m·k; Each of the sub-adders includes a speculative carry generator; the speculative carry generator is used to infer the approximate carry signal of the corresponding sub-adder according to the following equation: Where i represents the number of the sub-adder, represents the approximate carry signal of the ith sub-adder; the approximate carry signal of the first sub-adder is =0; k represents the bit width of each sub-adder, represents the carry flag signal of the k-1th bit generated by the i-th sub-adder, represents the carry flag signal of the k-1th bit generated by the i-1th sub-adder, The carry propagation signal representing the j-th bit of the output of the i-th sub-adder; A multiplexer is also provided between each of the sub-adders, for selecting the carry identification signal of the previous sub-adder or the approximate carry signal of the previous sub-adder as the carry input signal of the current sub-adder according to the state value of the carry propagation signal of the current sub-adder.
9. A neuromorphic computing method based on a spiking neural network, characterized in that: include: Write computer programs; The computer program is input into a neuromorphic computing system based on a spiking neural network as described in any one of claims 1 to 8 for execution, so as to realize real-time target recognition of an input image.
Citation Information
Patent Citations
Chip architecture supporting RISC-V instruction set and spiking neural network dedicated extended instruction set
CN118504628A