Global and Local Time Step Determination Schemes for Neural Networks
Through event-driven time-hopping calculation and global time-step communication scheme, the high energy consumption problem of neural networks under sparse activity conditions is solved, and energy consumption and performance optimization is achieved.
Patent Information
- Application Number
- CN201811130578.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2017-09-29
- Filing Date
- 2018-09-27
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2038-09-27
AI Technical Summary
Existing neural networks consume high memory access energy under sparse active conditions, affecting computing efficiency and performance.
The event-driven time-hop calculation method is adopted to update the neural unit state only at the active time step, reduce memory access through the global time step communication scheme, and optimize the energy consumption of the neural network based on the local and global time step determination schemes.
Effectively reduces the number of memory accesses, reduces energy consumption, while maintaining computational accuracy and performance, suitable for neural networks with sparse workloads.
Smart Images

Figure CN109583578B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to the field of computer development, and more particularly, to global and local time step determination schemes for neural networks. Background Art
[0002] A neural network may include a group of neural units that are loosely modeled after the structure of a biological brain including a large cluster of neurons connected by synapses. In a neural network, neural units are connected via links to other neural units, and the links may have an excitatory or inhibitory effect on the activation state of the connected neural units. A neural unit may perform a function using the values of its inputs to update the membrane potential of the neural unit. When a threshold associated with the neural unit is exceeded, the neural unit may propagate a pulse signal to the connected neural units. A neural network may be trained or otherwise adapted to perform various data processing tasks, such as computer vision tasks, speech recognition tasks, or other suitable computing tasks. Brief Description of the Drawings
[0003] Figure 1 A block diagram of a processor including a network-on-chip (NoC) system that may implement a neural network is shown in accordance with certain embodiments.
[0004] Figure 2 An example portion of a neural network is shown in accordance with certain embodiments.
[0005] Figure 3A An example progression of the membrane potential of a neural unit is shown in accordance with certain embodiments.
[0006] Figure 3B An example progression of the membrane potential of a neural unit in an event-driven and time-hopping neural network is shown in accordance with certain embodiments.
[0007] Figure 4A An example progression of the membrane potential of an integrate-and-fire neural unit is shown in accordance with certain embodiments.
[0008] Figure 4B An example progression of the membrane potential of a leaky integrate-and-fire neural unit is shown in accordance with certain embodiments.
[0009] Figure 5 Communication of the local next pulse time across the NoC is shown in accordance with certain embodiments.
[0010] Figure 6 Communication of the global next pulse time across the NoC is shown in accordance with certain embodiments.
[0011] Figure 7 Logic for calculating the local next pulse time is shown in accordance with certain embodiments.
[0012] Figure 8 illustrates an example process for calculating the next pulse time and receiving the global pulse time according to certain embodiments.
[0013] FIG. 9 illustrates an allowable relative time step between two connected neuron cores for localizing a time step determination scheme according to certain embodiments.
[0014] Figures 10A - 10D illustrates a sequence of connection states between multiple cores according to certain embodiments.
[0015] Figure 11 illustrates an example neuron core controller 1100 for tracking the time step of a neuromorphic core according to certain embodiments.
[0016] Figure 12 illustrates a neuromorphic core 1200 according to certain embodiments.
[0017] Figure 13 illustrates a process for processing pulses of various time steps and incrementing the time step of a neuromorphic core according to certain embodiments.
[0018] Figure 14A is a block diagram illustrating a exemplary in-order pipeline and an exemplary register renaming, out-of-order issue / execution pipeline according to certain embodiments.
[0019] Figure 14B is a block diagram illustrating an exemplary embodiment of an out-of-order issue / execution architecture core, an exemplary register renaming, and an in-order architecture core to be included in a processor according to certain embodiments;
[0020] Figure 15A -B illustrates a block diagram of a more specific exemplary in-order core architecture according to certain embodiments, the core being one of several logic blocks in a chip (potentially including other cores of the same type and / or different types);
[0021] Figure 16 is a block diagram of a processor according to certain embodiments that may have more than one core, may have an integrated memory controller, and may have integrated graphics;
[0022] Figure 17 、 18 、19 and 20 are block diagrams of exemplary computer architectures according to certain embodiments; and
[0023] Figure 21 is a block diagram contrasting the use of a software instruction converter for converting binary instructions in a source instruction set to binary instructions in a target instruction set according to certain embodiments.
[0024] Like reference numerals and names in the various figures indicate like elements. Detailed Description
[0025] In the following description, numerous specific details are set forth, such as examples of specific types of processors and system configurations, specific hardware architectures, specific architectural and microarchitectural details, specific register configurations, specific instruction types, specific system components, specific measurements / height, specific processor pipeline stages and operations, etc., in order to provide a thorough understanding of the present disclosure. However, it will be apparent to those skilled in the art that the present disclosure may be practiced without these specific details. In other instances, well-known components or methods, such as specific and alternative processor architectures, specific logic circuits / codes for the described algorithms, specific firmware codes, specific interconnection operations, specific logic configurations, specific manufacturing techniques and materials, specific compiler implementations, specific expressions of algorithms in code, specific power-down and gating techniques / logics, and other specific operating details of computer systems are not described in detail so as not to unnecessarily obscure the present disclosure.
[0026] Although the following embodiments may be described with reference to specific integrated circuits such as computing platforms or microprocessors, other embodiments are applicable to other types of integrated circuits and logic devices. Similar techniques and teachings of the embodiments described herein may be applied to other types of circuits or semiconductor devices. For example, the disclosed embodiments may be used in various devices, such as server computer systems, desktop computer systems, handheld devices, tablet computers, other thin laptops, system-on-chip (SOC) devices, and embedded applications. Some examples of handheld devices include cellular phones, Internet protocol devices, digital cameras, personal digital assistants (PDAs), and handheld PCs. Embedded applications typically include microcontrollers, digital signal processors (DSPs), system-on-chip, network computers (NetPCs), set-top boxes, network hubs, wide area network (WAN) switches, or any other system capable of performing the functions and operations taught below. Additionally, the devices, methods, and systems described herein are not limited to physical computing devices, but may also relate to software optimizations for energy conservation and efficiency.
[0027] Figure 1 A block diagram of a processor 100 including a network-on-chip (NoC) system that can implement a neural network is shown in accordance with certain embodiments. Processor 100 may include any processor or processing device, such as a microprocessor, embedded processor, digital signal processor (DSP), network processor, handheld processor, application processor, coprocessor, SoC, or other device that executes code. In a specific embodiment, processor 100 is implemented on a single die.
[0028] In the depicted embodiment, the processor 100 includes a plurality of network elements 102 arranged in a mesh network and coupled to each other by bi-directional links. However, the NoC according to various embodiments of the present disclosure can be applied to any suitable network topology (e.g., hierarchical network or ring network), size, bus width, and process. In the depicted embodiment, each network element 102 includes a router 104 and a core 108 (which can be a neuromorphic core in some embodiments), however, in other embodiments, multiple cores from different network elements 102 can share a single router 104. The routers 104 can be communicatively linked to each other in a network, such as a packet-switched network and / or a circuit-switched network, thus enabling communication between components of the NoC connected to the routers, such as cores, storage elements, or other logic blocks. In the depicted embodiment, each router 104 is communicatively coupled to its own core 108. In various embodiments, each router 104 can be communicatively coupled to a plurality of cores 108 (or other processing elements or logic blocks). As used herein, reference to a core can also apply to other embodiments that use different logic blocks in place of a core. For example, various logic blocks can include hardware accelerators (e.g., graphics accelerators, multimedia accelerators, or video encoding / decoding accelerators), I / O blocks, memory controllers, or other suitable fixed-function logic. The processor 100 can include any number of processing elements or other logic blocks that can be symmetric or asymmetric. For example, the cores 108 of the processor 100 can include asymmetric cores or symmetric cores. The processor 100 can include logic for operating as either or both a packet-switched network and a circuit-switched network to provide on-die communication.
[0029] In a specific embodiment, the resources of a packet-switched network can be used to transfer packets between various routers 104. That is to say, the packet-switched network can provide communication between routers (and their associated cores). A packet can include a control part and a data part. The control part can include the destination address of the packet, and the data part can contain specific data to be transferred on the processor 100. For example, the control part can include a destination address corresponding to one of the network elements or cores of the die. In some embodiments, the packet-switched network includes buffering logic because a dedicated path from the source to the destination cannot be ensured, and so if two or more packets need to traverse the same link or interconnect, the packets may need to be temporarily stopped. As an example, when a packet travels from the source to the destination, the packet can be buffered (e.g., by a flip-flop) at each corresponding router. In other embodiments, the buffering logic can be omitted, and the packets can be discarded when a conflict occurs. Packets can be received, transmitted, and processed by the routers 104. The packet-switched network can use point-to-point communication between adjacent routers. The control part of the packet can be transmitted between routers based on a packet clock (such as a 4 GHz clock). The data part of the packet can be transmitted between routers based on a similar clock (such as a 4 GHz clock).
[0030] In an embodiment, the routers of the processor 100 can be provided differently in two networks or communicate in two networks, such as a packet-switched network and a circuit-switched network. Such a communication method can be called a hybrid packet / circuit-switched network. In such embodiments, the resources of the packet-switched network and the circuit-switched network can be used to transfer packets differently between various routers 104. To transmit a single data packet, the circuit-switched network can allocate the entire path, while the packet-switched network can only allocate a single segment (or interconnect). In some embodiments, the packet-switched network can be utilized to reserve the resources of the circuit-switched network for transmitting data between the routers 104.
[0031] Router 104 may include multiple sets of ports to couple to and communicate with adjacent network elements 102 differently. For example, circuit - switched and / or packet - switched signals may be passed through these sets of ports. The sets of ports of router 104 may be logically partitioned, for example, according to the orientation of adjacent network elements and / or the traffic - switching direction with such elements. For example, router 104 may include a north set of ports having input (“IN”) and output (“OUT”) ports configured to receive communications from and send communications to network elements 102 located in the “north” direction relative to router 104 (respectively). Additionally or alternatively, router 104 may include similar sets of ports to interface with network elements located in the south, west, east, or other directions. In the depicted embodiment, router 104 is configured for X - first, Y - second routing, where data first moves in the east / west direction and then in the north / south direction. In other embodiments, any suitable routing scheme may be used.
[0032] In various embodiments, router 104 also includes another set of ports that includes input ports and output ports configured to receive communications from and send communications to another agent of the network. In the depicted embodiment, this set of ports is shown at the center of router 104. In one embodiment, these ports are used for communication with logic that is adjacent to, communicates with, or is otherwise associated with router 104, such as the logic of “local” core 108. Herein, this set of ports will be referred to as the “core set of ports,” although in some implementations it may interface with logic other than cores. In various embodiments, the core set of ports may interface with multiple cores (e.g., when multiple cores share a single router), or router 104 may include multiple core sets of ports (each interfacing with a corresponding core). In another embodiment, this set of ports is used for communication with network elements at the next network tier higher than the network tier of router 104. In one embodiment, the east and west direction links are on one metal layer, the north and south direction links are on a second metal layer, and the core links are on a third metal layer. In an embodiment, router 104 includes cross - bar switching and arbitration logic to provide paths for inter - port communication such as Figure 1 shown. The logic (e.g., core 108) in each network element may have a unique clock and / or voltage or may share a clock and / or voltage with one or more other components of the NoC.
[0033] In a specific embodiment, the core 108 of the network element may include a neuromorphic core (including one or more neural units). The processor may include one or more neuromorphic cores. In various embodiments, each neuromorphic core may include one or more computational logic blocks that are time-multiplexed across neural units of the neuromorphic core. The computational logic blocks may be operable to perform various computations for the neural units, such as updating the membrane potential of the neural unit, determining whether the membrane potential exceeds a threshold, and / or other operations associated with the neural unit. Herein, references to neural units may refer to the logic for implementing neurons of a neural network. Such logic may include storage for one or more parameters associated with the neuron. In some embodiments, the logic for implementing a neuron may overlap with the logic for implementing one or more other neurons (in some embodiments, the neural units corresponding to neurons may share computational logic with other neural units corresponding to other neurons and control signals, and the control signals may determine which neural unit is currently using the logic for processing).
[0034] Figure 2 An example portion of a neural network 200 is shown in accordance with certain embodiments. The neural network 200 includes neural units X1-X9. The neural units X1-X4 are input neural units that respectively receive primary inputs I1-I4 (which may remain constant while the neural network 200 processes the output). Any suitable primary inputs may be used. As an example, when the neural network 200 performs image processing, the primary input values may be the values of pixels from an image (and the values of the primary inputs may remain constant while processing the image). As another example, when the neural network 200 performs speech processing, the primary input values applied to a particular input neural unit may change over time based on changes to the input speech.
[0035] Although a particular topology and connectivity scheme are shown in Figure 2 , the teachings of the present disclosure may be used in neural networks having any suitable topology and / or connectivity. For example, the neural network may be a feedforward neural network, a recurrent network, or other neural network having any suitable connectivity between neural units. In the depicted embodiment, each link between two neural units has a synaptic weight that indicates the strength of the relationship between the two neural units. The synaptic weights are depicted as WXY, where X indicates the presynaptic neural unit and Y indicates the postsynaptic neural unit. The effect of the link between neural units on the activation state of the connected neural units may be excitatory or inhibitory. For example, depending on the value of W15, the impulse propagated from X1 to X5 may increase or decrease the membrane potential of X5. In various embodiments, the connections may be directed or undirected.
[0036] Generally, during each time step of a neural network, a neuron may receive any suitable input, such as a bias value or one or more input pulses from one or more neurons (this set of neurons is referred to as the fan-in neurons of the neuron) connected to the neuron via corresponding synaptic connections. The bias value applied to a neuron may be a function of the main input applied to the input neuron and / or some other value applied to the neuron (e.g., a constant value that may be adjusted during other operations or training of the neural network). In various embodiments, each neuron may be associated with its own bias value, or the bias value may be applied to multiple neurons.
[0037] A neuron can perform functions using its input values and its current membrane potential. For example, the input can be added to the current membrane potential of the neuron to generate an updated membrane potential. As another example, a non-linear function (such as a sigmoid transfer function) can be applied to the input and the current membrane potential. Any other suitable function can be used. Then the neuron updates its membrane potential based on the output of the function. When the membrane potential of a neuron exceeds a threshold, the neuron can send a pulse to each of its fan-out neurons (i.e., the neurons connected to the output of the spiking neuron). For example, when X1 spikes, the pulse can propagate to X5, X6, and X7. As another example, when X5 spikes, the pulse can propagate to X8 and X9 (and in some embodiments to X1, X2, X3, and X4). In various embodiments, when a neuron spikes, the pulse can propagate to one or more connected neurons residing on the same neuromorphic core and / or be grouped and transmitted to a neuromorphic core through one or more routers 104, the neuromorphic core including one or more of the fan-out neurons of the spiking neuron. The neurons to which a pulse is sent when a particular neuron spikes are referred to as the fan-out neurons of the neuron.
[0038] In a particular embodiment, one or more memory arrays may include memory cells that store synaptic weights, membrane potentials, thresholds, outputs (e.g., the number of times a neuron has spiked), bias amounts, or other values used during the operation of the neural network 200. The number of bits for each of these values may vary depending on the implementation. In the examples shown below, a specific bit length may be described relative to a particular value, but in other embodiments, any suitable bit length may be used. Any suitable volatile and / or non-volatile memory can be used to implement the memory arrays.
[0039] In a specific embodiment, the neural network 200 is a spiking neural network (SNN) that includes a plurality of neurons, each neuron tracking its respective membrane potential over a plurality of time steps. The membrane potential is updated for each time step by adjusting the membrane potential of the previous time step using a bias term, a leakage term (e.g., if the neuron is a leaky integrate-and-fire neuron), and / or a contribution for incoming spikes. A transfer function applied to the result can generate a binary output.
[0040] Although the degree of sparsity in various SNNs for typical pattern recognition workloads is very high (e.g., for a specific input pattern, 5% of the entire population of neurons may spike), the amount of energy consumed in memory accesses for updating the neural state (even without input spikes) is quite large. For example, memory accesses for fetching synaptic weights and updating neuron states can be the main component of the total energy consumption of a neuromorphic core. In neural networks with sparse activity (e.g., SNNs), many neuron state updates perform very little useful computation.
[0041] In various embodiments of the present disclosure, a global time step communication scheme for event-driven neural networks that utilizes time-hopping computation is provided. The various embodiments described herein provide systems and methods for reducing the number of memory accesses (without compromising the accuracy or performance of the computational workload of a neuromorphic computing platform). In a specific embodiment, the neural network computes neuron state changes only at time steps when spike events are being processed (i.e., active time steps). When the membrane potential of a neuron is updated, the contribution to the membrane potential due to time steps in which the state of the neuron is not updated (i.e., idle time steps) is determined and aggregated with the contribution to the membrane potential due to active time steps. The neuron can then remain idle (i.e., skip the membrane potential update) until the next active time step, thus improving performance while reducing memory accesses to minimize energy consumption (due to skipping memory accesses for idle time steps). The next active time step of the neural network (or a sub-part thereof) can be determined at a central location and passed to the various neuromorphic cores of the neural network.
[0042] The event-driven time-hopping neural network can be used to perform any suitable workload, such as sparse coding of an input image or other suitable workloads (e.g., workloads where the frequency of spikes is relatively low). Although the various embodiments herein are discussed in the context of SNNs, the concepts of the present disclosure can be applied to any suitable neural network, such as a convolutional neural network or other suitable neural networks.
[0043] Figure 3AShows an example progression of the membrane potential 302A of a neuron according to certain embodiments. The depicted progression is based on time-step based neural computations, where the membrane potential of the neuron is updated at each time step 308. Figure 3A Depicts an example membrane potential progression of an integrate-and-fire neuron (without leakage) through the integration of an arbitrary input pulse pattern. 304A depicts access to an array of synaptic weights ("synaptic array") that stores the connections between neurons, and 306A depicts access to an array of bias terms ("bias array") and an array of the current membrane potential of the neurons ("neural state array"). In the various embodiments depicted herein, the membrane potential is only the sum of the current membrane potential and the input to the neuron, although in other embodiments any suitable function may be used to determine the updated membrane potential.
[0044] In various embodiments, the synaptic array is stored separately from the bias array and / or the neural state array. In a particular embodiment, the bias and neural state arrays are implemented using a relatively fast memory, such as a register file (where each memory cell is a transistor, latch, or other suitable structure), while the synaptic array is stored using a relatively slow memory that is more suitable for storing large amounts of information (e.g., static random access memory (SRAM)) (due to the relatively large number of connections between neurons). However, in various embodiments, any suitable memory technology (e.g., register file, SRAM, dynamic random access memory (DRAM), flash memory, phase change memory, or other suitable memory) may be used for any of these arrays.
[0045] At time step 308A, the bias array and the neural state array are accessed, and the membrane potential of the neural unit is increased by the bias term (B) of the neural unit, and the updated membrane potential is written back to the neural unit state array. During time step 308A, other neural units may also be updated (in various embodiments, the processing logic may be shared among multiple neural units, and the neural units may be updated sequentially). At time step 308B, the bias array and the neural state array are accessed again, and the membrane potential is increased by B. At time step 308C, an input pulse 310A is received. Accordingly, the synaptic array is accessed to retrieve the weight of the connection between the neural unit being processed and the neural unit from which the pulse is received (or multiple synaptic weights if multiple pulses are received). In this example, the pulse has a negative impact on the membrane potential (although the pulse may alternatively have a positive impact on the membrane potential or no impact on the membrane potential), and the total impact on the potential at time step 308C is B - W. At time steps 308D - 308F, no input pulses are received, so only the bias array and the neural state array are accessed, and the bias term is added to the membrane potential at each time step. At time step 308G, another input pulse 310B is received, and thus the synaptic array, the bias array, and the neural state array are accessed to obtain values to update the membrane potential.
[0046] In this method, where the neural state is updated at each time step, the membrane potential can be expressed as:
[0047]
[0048] where u(t + 1) is equal to the membrane potential at the next time step, u(t) is equal to the current membrane potential, B is the bias term of the neural unit, and ( W i · I i ) is the product of the binary indication (i.e., 1 or 0) of whether a particular neural unit i coupled to the neural unit being processed is a pulse and the synaptic weight of the connection between the neural unit being processed and neural unit i. The summation can be performed over all neural units coupled to the neural unit being processed.
[0049] In this example, where the neural units are updated at each time step, the bias array and the neural state array are accessed at each time step. Such methods may use too much energy when input pulses are relatively rare (e.g., for workloads such as sparse coding of images).
[0050] Figure 3BShows an example progression of the membrane potential 302B of a neural unit of an event-driven and time-skipping neural network according to certain embodiments. The depicted progression is based on time-skipping and event-driven neural computations, where the membrane potential of the neural unit is updated only at active time steps 308C and 308G (at which one or more input pulses are received). As in Figure 3A This progression depicts an integrate-and-fire neural unit (without leakage) having the same pulse pattern and bias input as progression 302A. 304B depicts accessing the synaptic array, and 306B depicts accessing the bias array and the neural state array.
[0051] Contrary to the method shown in Figure 3A , the neural unit skips time steps 308A and 308B and does not access the bias array and the neural state array. At time step 308C, input pulse 310A is received. Similar to the progression in Figure 3A , the synaptic array is accessed to retrieve the weight of the connection between the neural unit being processed and the neural unit from which the pulse is received (or multiple synaptic weights if multiple pulses are received). The neural state array and the bias array are also accessed. In addition to the identification of the synaptic weights corresponding to any received pulses, the inputs to the neural unit for the current time step and any idle time steps not yet considered (e.g., time steps occurring between active time steps) are also determined (e.g., via bias array access or otherwise). Accordingly, the update calculation of the membrane potential at 308C is 3 * B - W, which includes three bias terms (one for the current time step and two for the idle time steps 308A and 308B that were skipped) and the weight of the incoming pulse. Then the neural unit skips time steps 308D, 308E, and 308F. At the next active time step 308G, the membrane potential is updated again based on the inputs at each idle time step and the current time step, resulting in a change of 4 * B - W to the membrane potential.
[0052] After each active time step in Figure 3B , the membrane potential 302B matches the membrane potential 302A at the same time step in Figure 3A . In this example, where the neural unit updates in response to incoming pulses instead of at each time step, the bias array and the neural state array are accessed only at active time steps, thus saving energy and improving processing time while maintaining accurate tracking of the membrane potential.
[0053] In this method, where the neural state is not updated at each time step and the bias terms remain constant from the last time step processed to the time step being processed, the membrane potential can be expressed as:
[0054]
[0055] Wherein, u(t + n) equals the membrane potential at the time step being processed, u(t) equals the membrane potential at the last time step processed, n is the number of time steps from the last time step processed to the time step being processed, B is the bias term of the neuron, and W i · I i is the product of a binary indication (i.e., 1 or 0) of whether the specific neuron i coupled to the neuron being processed is spiking and the synaptic weight of the connection between the neuron being processed and neuron i. The summation can be performed over all neurons coupled to the neuron being processed. If the bias from the last time step processed to the time step being processed is not constant, the equation can be modified to:
[0056]
[0057] where B j is the bias term of the neuron at time step j.
[0058] In various embodiments, after updating the membrane potential of a neuron, a determination can be made as to how many time steps in the future the neuron will spike in the absence of any input spikes (i.e., calculated assuming the neuron does not receive input spikes before it spikes). With a constant bias B, the number of time steps until the membrane potential exceeds a threshold θ can be determined as follows:
[0059]
[0060] where t next equals the number of time steps until the membrane potential exceeds the threshold, u equals the membrane potential calculated for the current time step, and B equals the bias term. Although the methodology is not shown here, the number of time steps until the membrane potential exceeds the threshold θ in the absence of input spikes can also be determined without keeping the bias constant by determining how many time steps will pass before the sum of the bias at each time step plus the current membrane potential exceeds the threshold.
[0061] Figure 4A shows an example progression of the membrane potential of an integrating and firing neuron according to certain embodiments. This progression depicts a time-step based method (similar to the method shown in Figure 4A ), where the membrane potential of the neuron is updated at each time step. Figure 4AThe threshold θ is also depicted. Once the membrane potential exceeds the threshold, the neuron can generate a spike and then enter a refractory period, which is configured to prevent the neuron from spiking again immediately (in some embodiments, when the neuron spikes, the potential can be reset to a specific value). As stated above, the membrane potential in the time-step method can be calculated as follows:
[0062]
[0063] Figure 4B An example progression of the membrane potential of a leaky integrate-and-fire neuron according to certain embodiments is shown. In the depicted embodiment, the membrane potential leaks between time steps, and the input is scaled based on the time constant τ. The membrane potential can be calculated according to the following equation:
[0064]
[0065] Similar to the embodiments described above, after updating the membrane potential of the leaky integrate-and-fire neuron, a determination can be made regarding how many time steps the neuron will spike in the future without any input spikes. With a constant bias B, the number of time steps until the membrane potential exceeds the threshold θ can be calculated based on the equation above. In the absence of input spikes, the equation above becomes:
[0066]
[0067] Similarly:
[0068]
[0069] Correspondingly:
[0070]
[0071] To solve for t next (the number of time steps until the neuron exceeds the threshold θ in the absence of input spikes), set u(t + n) to θ, and n (shown here as t next ) is isolated on one side of the equation:
[0072]
[0073] where u new is the most recently calculated membrane potential of the neuron. Thus, the logic implementing the above calculation can be used to determine t next . In some embodiments, the logic can be simplified by using an approximation. In a specific embodiment, the equation for u(t +n):
[0074]
[0075] It can be approximated as:
[0076]
[0077] After removing the contribution from the incoming pulse and setting u(t + n) equal to θ, t next can be calculated as:
[0078]
[0079] Accordingly, t can be solved via the logic that implements this approximation next . Although the methodology is not shown here, the number of time steps until the membrane potential exceeds the threshold θ in the absence of an input pulse can also be determined without keeping the bias constant by determining how many time steps will elapse before the sum of the bias at each time step plus the current membrane potential will exceed the threshold (and taking into account the leakage at each time step).
[0080] Figure 5 Communication of the local next pulse time across the NoC according to certain embodiments is shown. As described above, event-driven SNNs increase efficiency by determining the next time step at which an input pulse will occur (i.e., the next pulse time) for a particular group of neural units, as opposed to assuming that a pulse will occur by default at the next time step. For example, if the neural units are arranged in layers where each neuron in one layer has directed connections to neurons in a subsequent layer (e.g., a feedforward network), the next time step for the neural units to be processed for a particular layer can be the time step immediately following the time step at which any neural unit in the previous layer will spike. As another example, in a recurrent network where each neural unit has directed connections to every other neural unit, the next time step for the neural units to be processed is the next time step at which any neural unit will spike. For purposes of explanation, the following discussion will focus on embodiments involving recurrent networks, although the teachings are applicable to any suitable neural network.
[0081] In an event-driven SNN that utilizes multiple cores (e.g., each neuromorphic core can include multiple neural units of a network), the next time step at which a spike will occur can be communicated across all cores to ensure that spikes are processed in the correct order. The cores can each independently and in parallel perform spike integration and threshold calculations for their neural units. In an event-driven neural network, a core can also determine the next spike time at which any neural unit in the core will spike in the absence of input spikes before the speculative next spike time of the calculation. For example, any of the methodologies discussed above or other suitable methodologies can be used to calculate the next spike time for a neural unit.
[0082] To address spike dependencies and calculate the non-speculative spike times of a neural network (i.e., the next time step at which a spike will occur in the network), the minimum next spike time is calculated across the cores. In various embodiments, all cores process one or more spikes generated at this non-speculative next spike time. In some systems, each core uses unicast messages to communicate the next spike times of its neural units to every other core, and then each core determines the minimum next spike time of the received spike times, and then performs processing at the corresponding time step. Other systems can rely on a global event queue and a controller to coordinate the time steps of the processing. In various embodiments of the present disclosure, spike time communication is performed in a low-latency and energy-efficient manner by processing and multicasting packets in the network.
[0083] In the depicted embodiment, each router is coupled to a corresponding core. For example, router zero is coupled to core zero, router one is coupled to core one, and so on. Each of the depicted routers can have any suitable characteristics of router 104, and each core can have any suitable characteristics of core 108 or other suitable characteristics. For example, each core can be a neuromorphic core that implements any suitable number of neural units. In other embodiments, a router can be directly coupled (e.g., through a port of the router) to any number of neuromorphic cores. For example, each router can be directly coupled to four neuromorphic cores.
[0084] After processing a particular time step, a gather operation can communicate the next spike time of the network to a central entity (e.g., router in the depicted embodiment) 10 . The central entity can be any suitable processing logic, such as a router, a core, or associated logic. In a particular embodiment, the communication between the core and the router during the gather operation can follow a spanning tree with the central entity as its root. Each node of the tree (e.g., a core or a router) can send a communication with the next spike time to its parent node (e.g., a router) on the spanning tree.
[0085] The local next pulse time of a particular router is the minimum next pulse time of the next pulse times received at that router. A router can receive pulse times from each core directly connected to the router (in the depicted embodiment, each router is only directly coupled to a single core) as well as one or more next pulse times from adjacent routers. The router selects the local next pulse time as the minimum of the received next pulse times and forwards this local next pulse time to the next router. In the depicted embodiment, the local next pulse times of routers 0, 3, 4, 7, 8, 11, 12, and 15 will simply be the next pulse times of the respective cores to which the routers are coupled. Router1 will select the local next pulse time from the local next pulse time received from router0 and the next pulse time received from core1. Router5 will select the local next pulse time from the local next pulse time received from router4 and the next pulse time received from core5. Router9 will select the local next pulse time from the local next pulse time received from router8 and the next pulse time received from core9. Router13 will select the local next pulse time from the local next pulse time received from Router 12 and the next pulse time received from core13. Router2 will select the local next pulse time from the local next pulse time received from router1, the local next pulse time received from router3, and the next pulse time received from core2. Router6 will select the local next pulse time from the local next pulse time received from Router5, the local next pulse time received from router2, the local next pulse time received from router7, and the next pulse time received from core6. Router14 will select the local next pulse time from the local next pulse time received from Router13, the local next pulse time received from Router15, and the next pulse time received from core14. Finally, router10 (the root node of the spanning tree) will select the global next pulse time from the local next pulse times received from router6, router9, router11, and router14 and the next pulse time received from core10. This global next pulse time represents the next pulse time across the network at which the neuron will fire a pulse.
[0086] Thus, the leaves of the spanning tree (cores 0 through 15) send their speculative next time steps one hop towards the root of the spanning tree (e.g., in a packet). Each router collects packets from input ports, determines the minimum next pulse time among the inputs, and passes only the minimum next pulse time one hop towards the root. This process continues until the root receives the minimum pulse times of all connected cores, at which point the pulse times become non-speculative and can be passed to the cores (e.g., using a multicast message), enabling the cores to process the time steps indicated by the next pulse times (e.g., neural units of each core can be updated and new next pulse times can be determined).
[0087] Instead of sending independent unicast messages from each core to the root, this wave-based mechanism reduces network communication and improves latency and performance. Any suitable technique can be used to pre-compute or determine in real-time the topology of the tree that guides router communication. In the depicted embodiment, the routers communicate using a tree that follows a dimension-order routing scheme, specifically an X-first, Y-second routing scheme, where the local next pulse time is first transmitted in the east / west direction and then in the north / south direction. In other embodiments, any suitable routing scheme can be used.
[0088] In various embodiments, each router is programmed to know how many input ports it will receive the next pulse time from and to which output ports the local next pulse time should be sent. In various embodiments, each communication (e.g., packet) between routers that includes the local next pulse time can include a flag bit or opcode indicating that the communication includes the local next pulse time. Before determining the local next pulse time and sending it to the next hop, each router will wait to receive inputs from a specified number of input ports.
[0089] Figure 6 Shows communication of the global next pulse time across a neural network implemented on a NoC according to certain embodiments. In the depicted embodiment, a central entity (e.g., router 10 ) sends a multicast message including the global next pulse time to each core of the network. In a specific embodiment, the multicast message follows the same spanning tree (where communication moves in the opposite direction) as the local next pulse time, although in other embodiments, any suitable multicast method can be used to pass the global next pulse time to the cores. At each branch in the tree, the message can be received via one input port and copied to multiple output ports. During the multicast phase, the global next pulse time is passed to all cores, and all cores process the neuron activity that occurs during this time step, regardless of their own local speculative next time steps.
[0090] Figure 7Illustrates logic for calculating a local next pulse time according to certain embodiments. In various embodiments, the logic for calculating a local next pulse time may be included at any suitable node of the network, such as a core, a router, or a network interface between a core and a router. Similarly, the logic for calculating a global next pulse time and transmitting the global next pulse time via multicast messaging may be included at any suitable node of the network.
[0091] In various embodiments, the depicted logic may include circuitry for performing the functions described herein. In a specific embodiment, Figure 7 the logic depicted in may be located within each router and may communicate with one or more cores (or a network interface between a core and a router) and with router ports (i.e., ports coupled to other routers). When a neural network is mapped to the hardware of a NoC, the number of input ports to receive a local next pulse time from cores and / or routers and the number of output ports to send the calculated local next pulse time to the next hop may be programmed and held constant during neural network operation.
[0092] Input port 702 may include with respect to Figure 1Any suitable characteristics of the described input port. The input port can be connected to a core or another router. The depicted "data" can be a packet that includes the next pulse time (i.e., the next pulse time packet) sent by the router or core. In various embodiments, these packets can be represented by an opcode (or flag) in the packet header that differentiates them from other types of packets being passed over the NoC. Instead of forwarding these packets directly, a comparator 706 can compare the next pulse time data field of the packet with the current local next pulse time. The asynchronous merge block 704 can control which local next pulse time is provided to the comparator 706 (and can provide arbitration when multiple packets including the next pulse time are ready to be processed). The comparator 706 can compare the selected local next pulse time with the current local next pulse time stored in the buffer 708. If the selected local next pulse time is lower than the local next pulse time stored in the buffer 708, the selected local next pulse time is stored in the buffer 708 as the current local next pulse time. The asynchronous merge block 704 can also send a request signal to a counter 710 that keeps track of the number of local next pulse times that have been processed. The request signal can increment the value stored by the counter 710. The value stored by the counter can be compared with an input quantity value 712 that can be configured prior to operation of the neural network. The input quantity can be equal to the number of local next pulse times that the router is expected to receive after a processing time step and sending the local next pulse time to the central entity. Once the value of the counter 710 equals the input quantity, all local next pulse times have been processed, and the value stored by the minimum buffer 708 represents the local next pulse time for the router. The router can generate a packet that includes the local next pulse time and send the packet in a pre-programmed direction towards a central location (e.g., the root node of a spanning tree). For example, the packet can be sent to the next-hop router through an output port. If the router is a central router, the calculated local next pulse time is the global next pulse time and can be communicated as a multicast packet via multiple different output ports.
[0093] After passing the local next pulse time to the output port, the minimum buffer 708 and the counter 710 are reset. In one embodiment, the minimum buffer 708 can be set to a value high enough to ensure that any received local next pulse time will be less than the reset value and will overwrite the reset value.
[0094] Although the depicted logic is asynchronous (e.g., configured for use in an asynchronous NoC), any suitable circuit technology may be used (e.g., the logic may include synchronous circuitry suitable for use in a synchronous NoC). In a particular embodiment, the logic may utilize blocking 1-flit per-packet flow control (e.g., for request and ack signals), although any suitable flow control with guaranteed delivery may be used in various embodiments. In the depicted embodiment, request and ack signals may be utilized to provide flow control. For example, once an input (e.g., data) signal is valid and the destination for the data is ready (as indicated by an ack signal sent by the destination), the request signal may be asserted or toggled, at which point the data will be received by the destination (e.g., when the request signal is asserted, the input port may latch the data received at its input and the input port may be available to accept new data). If the downstream circuitry is not ready, the state of the ack signal may instruct the input port not to accept data. In the depicted embodiment, the ack signal sent by the output port may reset the counter 710 to zero and set the minimum buffer 708 to its maximum value after the next pulse time has been sent.
[0095] Figure 8 An example process 800 for calculating a next pulse time and receiving a global pulse time in accordance with certain embodiments is shown. The process may be performed, for example, by a network element 102 (e.g., a router and / or one or more neuromorphic cores).
[0096] At 802, a first time step is processed. For example, one or more neuromorphic cores may update the membrane potential of their neural units. At 804, one or more neuromorphic cores may determine the next time step at which any neural unit will pulse in the absence of input pulses. These next pulse times may be provided to a router connected to the one or more neuromorphic cores.
[0097] At 806, one or more next pulse times are received from one or more neighboring nodes (e.g., routers). At 808, the minimum next pulse time is selected from the next pulse times received from one or more routers and / or one or more cores. At 810, the selected minimum next pulse time is forwarded to a neighboring node (e.g., the next-hop router in a spanning tree whose root node is at the central entity).
[0098] At a later time, at 812, the router may receive the next time step (i.e., the global next pulse time) from a neighboring node. At 814, the router may forward the next time step to one or more neighboring nodes (e.g., the neuromorphic cores and / or routers from which it received the next pulse time at 806).
[0099] Where appropriate, some of the blocks shown in Figure 8 Figure 8 can be repeated, combined, modified, or deleted, and additional blocks can also be added to the flowchart. Additionally, the blocks can be executed in any suitable order without departing from the scope of the specific embodiment.
[0100] Although the above embodiments focus on passing the global time step to all cores, in some embodiments, the pulse dependencies may only need to be resolved between interconnected neural units, such as adjacent layers of neural units in a neural network. Accordingly, the global next pulse time can be passed to any suitable group of cores that are to process pulses (or otherwise have a need to receive pulse times). Thus, for example, in a specific neural network, the cores can be partitioned into separate domains, and a global time step can be calculated for each domain (in a manner similar to that described above) at a central location in the corresponding domain according to a spanning tree of the corresponding domain and only passed to the cores of that corresponding domain.
[0101] FIG. 9 shows the allowable relative time steps between two connected neuron cores for a local time step determination scheme according to certain embodiments. A neuromorphic processor can run an SNN with extremely parallel pulse processing within a time step and pulse dependencies that require ordered processing between time steps. Within a single time step, all pulses are independent. However, because the behavior of the pulses in one time step determines which neural units will pulse in subsequent time steps, there are pulse dependencies between time steps.
[0102] Coordinating time steps to resolve pulse dependencies in a multi-core neuromorphic processor to resolve pulse dependencies is a latency-critical operation. The duration of a time step is not easily predictable because the pulse neural network has variable computational amounts per core per time step. Some systems can resolve pulse dependencies globally by keeping all cores in the SNN at the same time step. Some systems can allocate the maximum possible number of hardware clock cycles to calculate each time step. In such systems, even if every neuron in the SNN pulses simultaneously, the neuromorphic processor will be able to complete all calculations before the end of the time step. The time step duration can be fixed (and can be workload-independent). Since the pulse rate of the SNN is typically low (the pulse rate may even be less than 1%), this technique may result in many wasted clock cycles and unnecessary latency penalties. When each core has completed its local processing for a time step, other systems (e.g., in conjunction with Figures 5 - 8 the embodiments described) can detect the end of the time step. Such systems benefit from shorter average time step durations (the time step duration is set by the execution time of the slowest core in each time step) but utilize global collective operations and the global time step is shared between cores.
[0103] Various embodiments of the present disclosure use local communication between cores connected in an SNN to control the time steps of neuromorphic cores on a per-core basis while maintaining proper handling of spike dependencies. Since spike dependencies only exist between connected neural units, tracking the time steps of the neurons connected to each core enables spike dependencies to be resolved without strict global synchronization. Thus, each neuromorphic core can track the time step at which an adjacent core (i.e., a core that provides input to or receives output from a particular core) is, and increment its own time step when a spike has been received from an input core (i.e., a core with fan-in neural units for the core's neural units), complete local spike processing, and any output core (i.e., a core with fan-out neural units for the core's neural units) is ready to receive new spikes. Cores closer to the SNN input (upstream cores) are allowed to compute neural unit processing for time steps ahead of downstream cores and cache future spikes and partial integration results for later use. Thus, various embodiments can utilize local communication to achieve time step control for the entire multi-core neuromorphic processor in a distributed manner.
[0104] Specific embodiments can increase hardware scalability to support larger SNNs, such as brain-scale networks. Various embodiments of the present disclosure reduce the latency of executing SNN workloads on a neuromorphic processor. For example, when each core is allowed to process one time step into the future, specific embodiments can improve the latency by approximately 24% in a 16-core fully-recursive SNN and approximately 20% for a 16-core feed-forward SNN. The latency can be further improved by increasing the number of time steps into the future that cores are allowed to process.
[0105] Figure 9AShows the allowed relative time steps between two connected neuron nuclei (“PRE nucleus” and “THIS nucleus”). The PRE nucleus can be a nucleus that includes neural units, where the neural units are fan-in neural units to one or more neural units of the THIS nucleus (so when the neural units of the PRE nucleus pulse, the pulses can be sent to one or more neural units of the THIS nucleus). The THIS nucleus can be connected to any suitable number of PRE nuclei. The depicted state assumes that the THIS nucleus is at time step t. For time step t - 1, the pulses received at the THIS nucleus from the PRE nucleus are processed in the THIS nucleus at time step t. If the PRE nucleus and the THIS nucleus are at the same time step t, the THIS nucleus can process the PRE pulses that were completed at time step t - 1, and the connection is active. If the THIS nucleus is ahead of the PRE nucleus (e.g., at time step t - 1), the PRE pulses are not complete and the THIS connection is idle because the THIS nucleus waits for the PRE nucleus to catch up. If the PRE nucleus is ahead of the THIS nucleus (e.g., at time steps t + 1, t + 2,... t + n), the THIS nucleus may be busy computing previous time steps or may be waiting for input from different connections. While the THIS nucleus is waiting for input from other PRE nuclei, the THIS nucleus may process the pulses from future time steps of the PRE nucleus, and the THIS nucleus has a look-ahead connection with the PRE nucleus. The processing results are stored in a separate buffer (e.g., a separate buffer for each time step) to ensure ordered operation. The number of available buffer resources can determine how many time steps ahead of its PRE nucleus the nucleus can process (e.g., the number of look-ahead states can vary from 1 to n, where n is the number of buffers available for storing the pulses from the PRE nucleus). When this limit is reached with respect to a specific PRE nucleus, the PRE nucleus is prevented from further incrementing its time step, which is depicted by a pre-idle connection.
[0106] Figure 9BShows the allowed relative time steps between two connected neuron cores (the THIS core and the "POST core"). The POST core can be a core that includes neural units, where the neural units are fan-out neural units of one or more neural units to the THIS core (so that when the neural units of the THIS core pulse, the pulses can be sent to one or more neural units of the POST core). The THIS core can be connected to any suitable number of POST cores. The depicted state assumes that the THIS core is at time step t. These connection states mirror the connection states between the PRE core and the THIS core. For example, when the POST core is too far behind the THIS core at t - n - 1, the connection between the THIS core and the POST core is idle (because there are not enough buffer resources in the POST core to store additional pulses from the THIS core). When the POST core is at time steps t - n to t - 1, the connection state is a look-ahead state because the POST core can buffer and process the input. When the POST core is ahead of the THIS core at time step t + 1, the connection is post-idle because the pulse at time t is not yet available to the POST core for processing at time step t + 1.
[0107] Figures 10A - 10D Shows a sequence of connection states between multiple cores according to some embodiments. This sequence shows how local time step synchronization allows look-ahead computation (i.e., for time steps before the most recent time step completed by the THIS core, allows the THIS core to process input pulses from some PRE cores), while maintaining ordered pulse execution. In these figures, the THIS core is coupled to the input cores PRE core 0 and PRE core 1. Both PRE core 0 and PRE core 1 include neural units that provide pulses to one or more neural units of the THIS core.
[0108] At Figure 10A , all cores are at time step 1, and the THIS core can process the pulses received from the two PRE cores at time step 0, so both connection states are active. At Figure 10B [[ID=ID=11]], PRE core 1 and the THIS core have completed time step 1, but PRE core 0 has not completed time step 1. The THIS core may process the pulses from PRE core 1 at time step 1, but must wait for the input pulses from PRE core 0 for time step 1 before completing time step 2, so the connection state with PRE core 0 is idle. At Figure 10CIn [the case where], the THIS core has completed the processing pulse from PRE core 0 for time step 1, but cannot complete time step 1 because it is still waiting for the pulse from PRE core 0 for time step 1. For time step 2, the THIS core can now perform look-ahead processing by receiving a pulse from Pre core 1, storing the pulse in a buffer, and performing a partial update of the membrane potential of the neural unit (for a specific time step, the update is considered complete only when all pulses have been received from all PRE cores). In Figure 10D In [the case where], PRE core 0 finally completes time step 1 and enters time step 2 and the pulse from PRE core 0 for time step 1 arrives and is processed, so the connection state between the THIS core and PRE core 0 becomes active again. Then, the THIS core can move to time step 3.
[0109] Figure 11 An example neuron core controller 1100 for tracking the time steps of neuromorphic cores according to certain embodiments is shown. In a specific embodiment, controller 1100 includes circuitry or other logic for performing the specified functions. Following the conventions of FIGS. 9 and 10, the core containing controller 1100 (or otherwise associated with controller 1100) will be referred to as the THIS core.
[0110] The neuron core controller 1100 can track the time steps of the THIS core via time step counter 1102. The neuron core controller can also track the time steps of PRE cores via time step counter 1104 and the time steps of POST cores via time step counter 1106. Counter 1102 can increment when the THIS core has completed neuron processing (e.g., of all pulses for the current time step) and the connections to all neighboring cores (PRE and POST cores) are in an active or look-ahead state. If the connection to any PRE core is in a post-idle state, then for the current time step of the THIS core, one or more additional input pulses can still be received from that PRE core, so the current time step may not increment. If the THIS core is too far ahead in time steps from a POST core, the connection can enter a pre-idle state because the POST core (or other storage space accessible to the POST core) may run out of space to store the output pulses of the THIS core at the most recent time step. Once the time step has been fully processed by the THIS core and the connection state with the neighbor cores of the THIS core allows the core to move to the next time step, the completion signal 1108 increments counter 1102.
[0111] When the time step of the THIS core increments, a completion signal can also be sent (e.g., via a multicast message) to all PRE cores and POST cores connected to the THIS core. When these cores increment their time steps, the THIS core can receive similar completion signals from its PRE and POST cores. When a completion signal is received from a PRE or POST core, the THIS core keeps track of the time steps of its PRE and POST cores by incrementing appropriate counters 1104 or 1106. For example, in the depicted embodiment, the THIS core can receive a PRE core completion signal 1110 along with a PRE core ID indicating the specific PRE core associated with the completion signal (in a specific embodiment, a packet having the PRE core ID and the PRE core completion signal can be sent from the PRE core to the THIS core). The decoder 1114 can send an increment signal to the appropriate counter 1104 based on the PRE core ID. In this way, the THIS core can keep track of the time step of each of its PRE cores. The THIS core can also keep track of the time step of each of its POST cores in a similar manner (using the POST core completion signal 1118, the POST core ID 1120, and the increment signal 1122). In other embodiments, any suitable signaling mechanism for passing completion signals between cores and incrementing the time step counters can be used.
[0112] To determine which state a connection is in, the value of the time step counter 1102 can be provided to each PRE core connection state logic block 1124 and POST core connection state logic block 1126. The difference between the value of the counter 1102 and the value of the corresponding counter 1104 or 1106 can be calculated and the corresponding connection state can be discerned based on the result. Each connection state logic block 1124 or 1126 can also include state output logic 1128 or 1130, which can output a signal that is asserted when the corresponding connection state is in an active or look-ahead state. The outputs of all state outputs can be combined and used (the output of the combined neuron processing logic 1132, which indicates whether the pulse buffer corresponding to the current time step has any pulses remaining to be processed) to determine whether the THIS core can increment its time step.
[0113] In a specific embodiment, the time step counter 1102 may maintain a counter value with more bits than the counter values maintained by the time step counters 1104 and 1106 (in some embodiments, each counter holds the same number of bits). In one example, the counter 1102 may be used for other operations of the neural network, while the time step counters 1104 and 1106 are only used to track the state of the connections of the THIS core. In embodiments where the time step counter 1102 maintains more bits than the counters 1104 and 1106, the least significant bit (LSB) group of the counter 1102 rather than the entire counter value is supplied to each connection state logic block 1124 and 1126. For example, a plurality of bits of the counter 1102 that match the number of bits stored by the counters 1104 and 1106 may be provided to the blocks 1124 and 1126. The number of bits maintained by the counters 1104 and 1106 may be sufficient to represent the number of states, e.g., an active state, all look-ahead states, and at least one idle state (in a specific embodiment, two different idle states may be conflated because they produce the same behavior). For example, a two-bit counter may be used to support two look-ahead states, an active state, and an idle state, or a three-bit counter may be used to support additional look-ahead states.
[0114] In a specific embodiment, when the THIS core increments its time step, instead of sending a completion signal to the PRE and POST cores, an event-based approach may be employed, where the THIS core sends its updated time step (or the LBS of its updated time step) to the PRE and POST cores. Accordingly, in such embodiments, the counters 1104 and 1106 may be omitted and replaced with a memory to store the received time step or replaced with other circuitry to facilitate the operation of the core state logics 1128 and 1130.
[0115] Figure 12 A neuromorphic core 1200 is shown in accordance with certain embodiments. The core 1200 may have any one or more of the characteristics of the other neuromorphic cores described herein. The core 1200 includes a neuron core controller 1100, a PRE pulse buffer 1202, a synaptic weight memory 1204, a weight summing logic 1206, a membrane potential increment buffer 1208, and a neuron processing logic 1132.
[0116] The PRE pulse buffer 1202 stores input pulses (i.e., PRE core pulses 1212) to be processed for look-ahead time steps (these pulses can be output by one or more PRE cores at the current time step or a future time step) and input pulses to be processed for the current / active time step of the core 1200 (these pulses can be output by one or more PRE cores at a previous time step). In the depicted embodiment, the PRE pulse buffer 1202 includes four entries, where one entry is dedicated to pulses received from the PRE core for the current time step, and three entries each are dedicated to pulses received from the PRE core for a specific look-ahead time step.
[0117] When a pulse 1212 is received from a neuron unit of the PRE core, it can be written to a location in the PRE pulse buffer 1202 based on the identifier of the neuron unit of the pulse (i.e., the PRE pulse address 1214) and the specified time step 1216 (at which the neuron unit pulsed). Although the buffer 1202 can be addressed in any suitable manner, in a particular embodiment, the time step 1216 can distinguish the columns of the buffer 1202, and the PRE pulse address 1214 can distinguish the rows of the buffer 1202 (thus each row of the buffer 1202 can correspond to a different neuron unit of the PRE core). In some embodiments, each column of the buffer 1202 can be used to store pulses for a specific time step.
[0118] In various embodiments, each pulse can be sent from the PRE core to the core 1200 in its own message (e.g., packet). In other embodiments, the pulses 1212 (and PRE pulse addresses 1214) can be aggregated into a message and sent to the core 1200 as a vector.
[0119] In addition to tracking the states of adjacent cores (e.g., as described above), the neuron core controller 1100 can also coordinate the processing of pulses for various time steps. When processing pulses, the neuron core controller 1100 can prioritize pulses for the earliest time steps. Thus, the controller 1100 can process any pulses for the current time step present in the buffer 1202 before processing pulses for look-ahead time steps present in the buffer 1202. The controller 1100 can also process any pulses for the first look-ahead time step present in the buffer 1202 before processing pulses for the second look-ahead time step in the buffer 1202, and so on.
[0120] In a specific embodiment, the neuron nucleus controller 1100 can read pulses from a buffer (e.g., by asserting the rows and columns of the pulses), and access the synaptic weights of the connections between the neural units of the nucleus 1200 and the pulsed neural units. For example, if the neural unit generating the pulse is connected to each neural unit of the nucleus 1200, then a row including the synaptic weights of each neural unit in the nucleus 1200 can be accessed. The synaptic weight memory 1204 includes the synaptic weights of the connections between the fan-in neural units for the PRE nucleus and the neural units of the nucleus 1200.
[0121] The weight summing logic 1206 can sum the synaptic weights of each neural unit of the nucleus 1200 individually into the membrane potential increment of that neuron. Thus, when pulses are sent to all neural units of the nucleus 1200, the weight summing logic 1206 can iterate through the neural units, adding the synaptic weights for the pulsed neural units and the neural units whose membrane potential increments are updated for the applicable time step to the membrane potential increment.
[0122] The membrane potential increment buffer 1208 can include multiple entries, each entry corresponding to a specific time step. Within each entry, a set of membrane potential increments is stored, with each increment corresponding to a specific neural unit. The membrane potential increment represents a partial processing result of the neural unit until the time step is completed (i.e., all PRE nuclei have supplied their corresponding pulses). In a specific embodiment, the same column address (e.g., time step 1218) used to access the PRE pulse buffer 1202 can also be used to access the membrane potential increment buffer 1208 during pulse processing.
[0123] Once the time step is completed, each neural unit is processed by the neuron processing logic 1132 by adding its membrane potential increment for the current time step to the membrane potential of the neural unit at the end of the previous time step (which can be stored by the neuron processing logic 1132 or stored in a memory accessible to the logic 1132). In some embodiments, if a specific neural unit is in a refractory period, the membrane potential increment is not added to the membrane potential of that neural unit. The neuron processing logic 1132 can perform any other suitable operations on the neural unit, such as applying a bias and / or leakage operation to the neural unit and determining whether the neural unit pulses at the current time step. If the neural unit pulses, the neuron processing logic can send a pulse 1220 along with a pulse address 1222 to a nucleus (i.e., a POST nucleus) having fan-out neural units for the pulsed neural unit, the pulse address 1222 including the identifier of the neural unit that pulsed.
[0124] In various embodiments, for a core having a large number of neurons, serial access to the synaptic weight memory 1204 and serial processing for weight summation and neuron processing can be performed, although any of these operations can be performed using any suitable method.
[0125] In various embodiments, the neuron core controller 1100 can facilitate the processing of input pulses 1212 by outputting a time step 1218 for accessing entries of the PRE pulse buffer 1202 and the membrane potential increment buffer 1208. If all received input pulses for the current time step have been processed (and the core 1200 is waiting for one or more PRE cores to finish generating pulses to be processed for the current time step), the neuron core controller 1100 can output an address corresponding to a look-ahead time step and process pulses from the look-ahead time step until additional input pulses are received for the current time step (or the remaining PRE cores complete the time step without sending additional pulses).
[0126] When a particular time step is completed, the corresponding entries of the PRE pulse buffer 1202 and the entries of the membrane potential increment buffer 1208 can be cleared (e.g., reset) and used for future time steps.
[0127] In a particular embodiment, the number of PRE cores and POST cores for each neuromorphic core is predetermined when mapping the SNN to hardware, and the logic of each core can be designed accordingly. For example, the neuron controller 1100 of each core can be adapted to the specific configuration of the core and can include, for example, making the number of counters 1104 and 1106 different based on the number of PRE cores and POST cores of the core. As another example, the number of rows of the PRE pulse buffer 1202 of the core 1200 can be configured based on the number of neurons of the PRE cores of the core 1200.
[0128] In the depicted embodiment, the number of allowable look-ahead states is preconfigured based on the number of entries in the PRE pulse buffer 1202 and the membrane potential increment buffer 1208 before the neural network starts operating, although in other embodiments, the number of allowable look-ahead states can be determined dynamically (i.e., the number of time steps the core can continue through adjacent cores). For example, one or more local memory pools can be shared among different time steps and / or cores, and portions of the memory can be dynamically allocated for use by the time steps and / or cores (e.g., storing outputs and / or membrane potential increments). In a particular embodiment, the central controller can dynamically allocate memory among time steps and / or cores in an intelligent manner to facilitate the efficient operation of the neural network.
[0129] Figure 13Illustrates a process for processing pulses at various time steps and incrementing the time step of a neuromorphic core according to certain embodiments. At 1302, pulses are resolved with the earliest time step. For example, the pulse buffer 1202 can be searched to determine if any pulses exist in the buffer entry corresponding to the current time step. If no pulse exists for the current time step, the buffer entry corresponding to the next time step can be searched, and so on.
[0130] At 1304, the synaptic weights of the fan-out neural units for the pulse are accessed. The synaptic weights can be the weights of the connections between the neural unit to be updated and the pulsed neural unit (i.e., the fan-out neural units). At 1306, for the time step associated with the pulse (which can actually be one time step later than the time step at which the pulse occurred), the synaptic weights are added to the membrane potential increment of the fan-out neural units.
[0131] At 1308, it is determined whether the neural unit just updated is the last fan-out neural unit of the pulsed neural unit. If not, the process returns to 1304 and additional neural units are updated. If the neural unit is the last fan-out neural unit for the pulse, a determination is made at 1310 as to whether the current time step is complete. For example, a time step can be completed when all PRE cores have provided their input pulses to the core and all pulses for that time step have been processed. If the time step is not complete, the process can return to 1302, where additional pulses (for the current time step or for future time steps) can be processed.
[0132] At 1312, after determining that the current time step is complete, neuron processing can be performed at 1312. For example, the neuron processing logic 1132 can perform any suitable operations, such as determining which neural units pulsed during the current time step, applying leakage and / or bias terms, or performing other suitable operations. Output pulses can be propagated to the appropriate cores.
[0133] At 1314, the states of adjacent cores are checked. If all adjacent cores are in a state that results in a connected state with the activity or future look-ahead of the core (e.g., a time step), then at 1316 the time step of the core can be incremented. If there are any idle connections, the core can continue to process pulses for future time steps until the connection state allows the time step of the core to be incremented.
[0134] Some of the boxes shown in Figure 13 can be repeated, combined, modified, or deleted as appropriate, and additional boxes can also be added to the flowchart. Additionally, the boxes can be executed in any suitable order without departing from the scope of the specific embodiments.
[0135] The following figures detail exemplary architectures and systems for implementing the above embodiments. For example, the neuromorphic processor described above may be included within any of the systems described below. In some embodiments, the neuromorphic processor may be communicatively coupled to any of the processors described below. In various embodiments, the neuromorphic processor may be implemented on and / or within the same chip as any of the processors described below. In some embodiments, one or more of the hardware components and / or instructions described above are simulated as detailed below, or implemented as software modules.
[0136] Processor cores may be implemented in different ways, for different purposes, and in different processors. For example, implementations of such cores may include: 1) general-purpose in-order cores intended for general-purpose computing; 2) high-performance general-purpose out-of-order cores intended for general-purpose computing; 3) specialized cores intended primarily for graphics and / or scientific (throughput) computing. Implementations of different processors may include: 1) a CPU including one or more general-purpose in-order cores intended for general-purpose computing and / or one or more general-purpose out-of-order cores intended for general-purpose computing; and 2) a coprocessor including one or more specialized cores intended primarily for graphics and / or scientific (throughput). Such different processors result in different computer system architectures, which may include: 1) a coprocessor on a separate chip from the CPU; 2) a coprocessor on a separate die within the same package as the CPU; 3) a coprocessor on the same die as the CPU (in which case, such a coprocessor is sometimes referred to as specialized logic, such as integrated graphics and / or scientific (throughput) logic, or as a specialized core); and 4) a system-on-a-chip that may include on the same die the CPU described above (sometimes referred to as one or more application cores or one or more application processors), the coprocessor described above, and additional functionality. An exemplary core architecture is next described, followed by a description of exemplary processor and computer architectures.
[0137] Figure 14A is a block diagram that illustrates both an exemplary in-order pipeline and an exemplary register renaming, out-of-order issue / execution pipeline in accordance with an embodiment of the present disclosure. Figure 14B is a block diagram that illustrates both an exemplary embodiment of an in-order architecture core to be included in a processor and an exemplary register renaming, out-of-order issue / execution architecture core in accordance with an embodiment of the present disclosure. Figure 14A The solid boxes in -B illustrate the in-order pipeline and in-order core, while the optional addition of the dashed boxes illustrates the register renaming, out-of-order issue / execution pipeline and core. Given that the in-order aspects are a subset of the out-of-order aspects, the out-of-order aspects will be described.
[0138] In Figure 14AIn [the figure], the processor pipeline 1400 includes an instruction fetch stage 1402, a length decoding stage 1404, a decoding stage 1406, an allocation stage 1408, a renaming stage 1410, a scheduling (also known as dispatch or issue) stage 1412, a register read / memory read stage 1414, an execution stage 1416, a write-back / memory write stage 1418, an exception handling stage 1422, and a commit stage 1424.
[0139] Figure 14B A processor core 1490 is shown, which includes a front-end unit 1430 coupled to an execution engine unit 1450, and both are coupled to a memory unit 1470. The core 1490 can be a reduced instruction set computing (RISC) core, a complex instruction set computing (CISC) core, a very long instruction word (VLIW) core, or a hybrid or alternative core type. As yet another option, the core 1490 can be a specialized core, such as, for example, a network or communication core, a compression and / or decompression engine, a coprocessor core, a general-purpose computing graphics processing unit (GPGPU) core, a graphics core, and so on.
[0140] The front-end unit 1430 includes a branch prediction unit 1432 coupled to an instruction cache memory unit 1434, the instruction cache memory unit 1434 being coupled to an instruction translation lookaside buffer (TLB) 1436, which is coupled to an instruction fetch unit 1438, and the instruction fetch unit 438 being coupled to a decoding unit 1440. The decoding unit 1440 (or decoder) can decode instructions and generate as output one or more micro-operations, microcode entry points, microinstructions, other instructions, or other control signals, which are decoded from, or which otherwise reflect, or are derived from the original instructions. Using a variety of different mechanisms, the decoding unit 1440 can be implemented. Examples of suitable mechanisms include, but are not limited to, lookup tables, hardware implementations, programmable logic arrays (PLAs), microcode read-only memories (ROMs), and so on. In one embodiment, the core 1490 includes a microcode ROM or another medium (e.g., in the decoding unit 1440 or otherwise within the front-end unit 1430) that stores microcode for certain macroinstructions. The decoding unit 1440 is coupled to a rename / allocator unit 1452 in the execution engine unit 450.
[0141] The execution engine unit 1450 includes a rename / allocator unit 1452 coupled to a retirement unit 1454 and a set of one or more scheduler units 1456. The one or more scheduler units 1456 represent any number of different schedulers, including reservation stations, a central instruction window, and the like. The one or more scheduler units 1456 are coupled to one or more physical register file units 1458. Each of the one or more physical register file units 1458 represents one or more physical register files, where different physical register files store one or more different data types, such as scalar integers, scalar floating point, packed integers, packed floating point, vector integers, vector floating point, status (e.g., an instruction pointer that is the address of the next instruction to be executed), and the like. In one embodiment, the one or more physical register file units 1458 include a vector register unit, a write mask register unit, and a scalar register unit. These register units may provide architected vector registers, vector mask registers, and general-purpose registers. The one or more physical register file units 1458 are overlapped via the retirement unit 1454 to illustrate various ways in which register renaming and out-of-order execution may be implemented (e.g., using one or more reorder buffers and one or more retirement register files; using one or more future files, one or more history buffers, and one or more retirement register files; using register maps and pools of registers; and the like). The retirement unit 1454 and the one or more physical register file units 1458 are coupled to one or more execution clusters 1460. The one or more execution clusters 1460 include a set of one or more execution units 1462 and a set of one or more memory access units 1464. The execution units 1462 may perform various operations (e.g., shift, add, subtract, multiply) and operate on various types of data (e.g., scalar floating point, packed integers, packed floating point, vector integers, vector floating point). While some embodiments may include multiple execution units dedicated to a particular function or set of functions, other embodiments may include multiple execution units or just one execution unit that all perform all functions. The scheduler units 1456, the one or more physical register file units 1458, and the one or more execution clusters 1460 are shown as potentially plural because certain embodiments create separate pipelines for certain types of data / operations (e.g., a scalar integer pipeline, a scalar floating point / packed integer / packed floating point / vector integer / vector floating point pipeline, and / or a memory access pipeline, each of which has its own scheduler unit, one or more physical register file units, and / or execution cluster—and in the case of a separate memory access pipeline, certain embodiments are implemented where only the execution cluster of this pipeline has one or more memory access units 1464).It should also be understood that where separate pipelines are used, one or more of these pipelines may be out-of-order issue / execution and the remainder in-order.
[0142] The set of memory access units 1464 is coupled to a memory unit 1470 which includes a data TLB unit 1472 coupled to a data cache unit 1474, the data cache unit 1474 being coupled to a level 2 (L2) cache unit 1476. In one exemplary embodiment, the memory access units 1464 may include a load unit, a store address unit, and a store data unit, each of which is coupled to the data TLB unit 1472 in the memory unit 1470. The instruction cache unit 1434 is further coupled to the level 2 (L2) cache unit 1476 in the memory unit 1470. The L2 cache unit 1476 is coupled to one or more other levels of cache and ultimately to main memory.
[0143] By way of example, an exemplary register renaming, out-of-order issue / execution core architecture may implement a pipeline 1400 as follows: 1) Instruction fetch 1438 performs a fetch and length decoding stage 1402 and 1404; 2) A decode unit 1440 performs a decode stage 1406; 3) A rename / allocator unit 1452 performs an allocation stage 1408 and a rename stage 1410; 4) One or more scheduler units 1456 perform a schedule stage 1412; 5) One or more physical register file units 1458 and the memory unit 1470 perform a register read / memory read stage 1414; An execution cluster 1460 performs an execution stage 1416; 6) The memory unit 1470 and one or more physical register file units 1458 perform a writeback / memory write stage 1418; 7) Various units may be involved in an exception handling stage 1422; and 8) A retirement unit 1454 and one or more physical register file units 1458 perform a commit stage 1424.
[0144] The core 1490 may support one or more instruction sets (e.g., the x86 instruction set (with some extensions added with more recent versions); the MIPS instruction set of MIPS Technologies of Sunnyvale, CA; the ARM instruction set of ARM Holdings of Sunnyvale, CA (with optional additional extensions such as NEON)), including one or more of the instructions described herein. In one embodiment, the core 1490 includes logic for supporting packed data instruction set extensions (e.g., AVX1, AVX2), thus allowing operations used by many multimedia applications to be performed using packed data.
[0145] It should be understood that the core supports multithreading (two or more parallel sets of operations or threads), and can do so in a variety of ways, including time-sliced multithreading, simultaneous multithreading (where in the case where a single physical core provides a logical core for each of the threads, that physical core is performing simultaneous multithreading), or a combination thereof (e.g., time-sliced fetching and decoding followed by simultaneous multithreading such as in Intel® Hyper-Threading Technology).
[0146] Although register renaming is described in the context of out-of-order execution, it should be understood that register renaming can be used in an in-order architecture. Although the illustrated embodiments of the processor also include separate instruction and data cache memory units 1434 / 1474 and a shared L2 cache memory unit 1476, alternative embodiments may have a single internal cache memory for both instructions and data, such as for example, a level 1 (L1) internal cache memory, or multiple levels of internal cache memories. In some embodiments, the system may include a combination of an internal cache memory and an external cache memory external to the core and / or the processor. Alternatively, all cache memories may be external to the core and / or the processor.
[0147] Figure 15A -B shows a block diagram of a more specific exemplary in-order core architecture where the core would be one of several logical blocks in a chip (potentially including other cores of the same type and / or different types). The logical blocks communicate via a high-bandwidth interconnect network (e.g., a ring network) with some fixed functional logic, a memory I / O interface, and other necessary I / O logic depending on the application.
[0148] Figure 15A is a block diagram of a single processor core according to various embodiments along with its connections to an on-die interconnect network 1502 and a local subset of its level 2 (L2) cache memory 1504. In one embodiment, the instruction decoder 1500 supports the x86 instruction set with packed data instruction set extensions. The L1 cache memory 1506 allows low-latency access to cache memory into scalar and vector units. Although in one embodiment (for simplicity of design), the scalar unit 1508 and the vector unit 1510 use separate register sets (correspondingly, scalar registers 1512 and vector registers 1514), and the data transferred between them is written to memory and then read back from the level 1 (L1) cache memory 1506, alternative embodiments may use different means (e.g., using a single register set or including a communication path that allows data to be transferred between the two register banks without being written and read back).
[0149] The local subset of the L2 cache 1504 is part of the global L2 cache, which is partitioned into separate local subsets (one per processor core in some embodiments). Each processor core has a direct access path to its own local subset of the L2 cache 1504. Data read by a processor core is stored in its L2 cache subset 1504 and can be accessed quickly, in parallel with other processor cores accessing their own local L2 cache subsets. Data written by a processor core is stored in its own L2 cache subset 1504 and flushed from other subsets if necessary. The ring network ensures consistency of shared data. The ring network is bi-directional to allow agents such as processor cores, L2 caches, and other logic blocks to communicate with each other within the chip. In a specific embodiment, each ring data-path is 1012-bits wide in each direction.
[0150] Figure 15B is according to an embodiment Figure 15A An expanded view of a portion of the processor core in. Figure 15B Includes a portion of the L1 data cache 1506A that includes the L1 cache 1504, and more details regarding the vector unit 1510 and vector registers 1514. Specifically, the vector unit 1510 is a 16-wide vector processing unit (VPU) (see 16-wide ALU 1528) that executes one or more of integer, single-precision floating, and double-precision floating instructions. The VPU supports shuffling register inputs via the shuffle unit 1520, numerical conversion via the numerical conversion units 1522A-B, and replication via the replication unit 1524 on memory inputs. The write mask register 1526 allows predicted result vector writes.
[0151] A processor with an integrated memory controller and graphics
[0152] Figure 16 is a block diagram of a processor 1600 that, according to an embodiment, may have more than one core, may have an integrated memory controller, and may have integrated graphics. Figure 16 The solid boxes in show the processor 1600 with a single core 1602A, a system agent 1610, and a collection of one or more bus controller units 1616, while the optional additional dashed boxes show an alternative processor 1600 with multiple cores 1602A-N, a collection of one or more integrated memory controller units 1614 in the system agent unit 1610, and dedicated logic 1608.
[0153] Accordingly, different implementations of the processor 1600 may include: 1) a CPU with dedicated logic 1608 that is integrated graphics and / or scientific (throughput) logic (which may include one or more cores), and cores 1602A-N that are one or more general-purpose cores (e.g., general-purpose in-order cores, general-purpose out-of-order cores, a combination of the two); 2) a coprocessor with cores 1602A-N that are a large number of dedicated cores intended primarily for graphics and / or scientific (throughput); and 3) a coprocessor with cores 1602A-N that are a large number of general-purpose in-order cores. Accordingly, the processor 1600 can be a general-purpose processor, a coprocessor, or a special-purpose processor, such as, for example, a network or communication processor, a compression and decompression engine, a graphics processor, a GPGPU (general-purpose graphics processing unit), a high-throughput many integrated core (MIC) coprocessor (e.g., including 30 or more cores), an embedded processor, or other fixed or configurable logic that performs logical operations. The processor can be implemented on one or more chips. Using any of a number of processing technologies (such as, for example, BiCMOS, CMOS, or NMOS), the processor 1600 can be implemented on and / or as part of one or more substrates.
[0154] In various embodiments, a processor may include any number of processing elements that may be symmetric or asymmetric. In one embodiment, a processing element refers to hardware or logic that supports a software thread. Examples of hardware processing elements include: a thread unit, a thread slot, a thread, a processing unit, a context, a context unit, a logical processor, a hardware thread, a core, and / or any other element that is capable of holding the state of a processor, such as an execution state or an architectural state. In other words, in one embodiment, a processing element refers to any hardware that is capable of independently associating with code, such as a software thread, an operating system, an application, or other code. A physical processor (or a processor socket) generally refers to an integrated circuit that potentially includes any number of other processing elements, such as cores or hardware threads.
[0155] A core may refer to logic located on an integrated circuit that is capable of maintaining an independent architectural state, where each independently maintained architectural state is associated with at least some dedicated execution resources. A hardware thread may refer to any logic located on an integrated circuit that is capable of maintaining an independent architectural state, where the independently maintained architectural states share access to execution resources. As can be seen, there is an overlap in the naming between hardware threads and cores when some resources are shared while other resources are dedicated to an architectural state. And generally, cores and hardware threads are treated as independent logical processors by an operating system, where the operating system is capable of independently scheduling operations on each logical processor.
[0156] The memory hierarchy includes one or more levels of on - core caches, a collection or one or more of shared cache memory units 1606, and external memory (not shown) coupled to a collection of integrated memory controller units 1614. The collection of shared cache memory units 1606 can include one or more mid - level caches such as level 2 (L2), level 3 (L3), level 4 (L4), or other levels of cache memory, last - level cache (LLC), and / or combinations thereof. Although in one embodiment, the ring - based interconnect unit 1612 interconnects the dedicated logic (e.g., integrated graphics logic) 1608, the collection of shared cache memory units 1606, and the system agent unit 1610 / one or more integrated memory controller units 1614, alternative embodiments can use any number of well - known techniques for interconnecting such units. In one embodiment, coherence is maintained between one or more cache memory units 1606 and the cores 1602 - A - N.
[0157] In some embodiments, one or more of the cores 1602A - N have multi - threading capabilities. The system agent 1610 includes those components that coordinate and operate the cores 1602A - N. The system agent unit 1610 can include, for example, a power control unit (PCU) and a display unit. The PCU can be or include the logic and components needed to regulate the power states of the dedicated logic 1608 and the cores 1602A - N. The display unit is used to drive one or more externally connected displays.
[0158] The cores 1602A - N can be homogeneous or heterogeneous with respect to the architectural instruction set; that is, two or more of the cores 1602A - N can have the ability to execute the same instruction set, while other cores can have the ability to execute a different instruction set or only a subset of that instruction set.
[0159] Figures 17 - 20 is a block diagram of an exemplary computer architecture. Other system designs and configurations known in the art for laptop computers, desktop computers, handheld PCs, personal digital assistants, engineering workstations, servers, network devices, network hubs, switches, embedded processors, digital signal processors (DSPs), graphics devices, video game devices, set - top boxes, microcontrollers, cellular telephones, portable media players, handheld devices, and various other electronic devices are also suitable for performing the methods described in this disclosure. In general, a vast variety of systems or electronic devices capable of incorporating a processor and / or other execution logic as disclosed herein are generally suitable.
[0160] Figure 17A block diagram depicting system 1700 according to an embodiment of the present disclosure. System 1700 may include one or more processors 1710, 1715 coupled to a controller hub 1720. In one embodiment, controller hub 1720 includes a Graphics Memory Controller Hub (GMCH) 1790 and an Input / Output Hub (IOH) 1750 (which may be on separate chips or the same chip); GMCH 1790 includes a memory and graphics controller coupled to memory 1740 and a coprocessor 1745; IOH 1750 couples input / output (I / O) devices 1760 to GMCH 1790. Alternatively, one or both of the memory and graphics controllers are integrated within a processor (as described herein), memory 1740 and coprocessor 1745 are directly coupled to processor 1710, and controller hub 1720 is a single chip including IOH 1750.
[0161] The optional nature of additional processor 1715 is denoted by a broken line in Figure 17 Each processor 1710, 1715 may include one or more of the processing cores described herein and may be a version of processor 600.
[0162] Memory 1740 may be, for example, dynamic random access memory (DRAM), phase change memory (PCM), other suitable memory, or any combination thereof. Memory 1740 may store any suitable data, such as data used by processors 1710, 1715 to provide functionality of computer system 1700. For example, data associated with a program being executed or files accessed by processors 1710, 1715 may be stored in memory 1740. In various embodiments, memory 1740 may store data and / or instruction sequences used or executed by processors 1710, 1715.
[0163] In at least one embodiment, controller hub 720 communicates with one or more processors 1710, 1715 via a front side bus (FSB) such as, a point-to-point interface such as QuickPath Interconnect (QPI), or a similar connection 1795.
[0164] In one embodiment, coprocessor 1745 is a specialized processor, such as, for example, a high throughput MIC processor, a network or communication processor, a compression and / or decompression engine, a graphics processor, a GPGPU, an embedded processor, and the like. In one embodiment, controller hub 1720 may include an integrated graphics accelerator.
[0165] There can be many differences in the spectrum of specifications between physical resources 1710, 1715 regarding metrics including architectural, microarchitectural, thermal, power consumption characteristics, and the like.
[0166] In one embodiment, the processor 1710 executes instructions that control general types of data processing operations. Coprocessor instructions may be embedded within the instructions. The processor 1710 identifies these coprocessor instructions as being of a type to be executed by an attached coprocessor 1745. Accordingly, the processor 1710 issues these coprocessor instructions (or control signals representative of the coprocessor instructions) on a coprocessor bus or other interconnect to the coprocessor 1745. One or more coprocessors 1745 receive and execute the received coprocessor instructions.
[0167] Figure 18 A block diagram depicting a first more specific exemplary system 1800 in accordance with an embodiment of the present disclosure. As Figure 18 shown, the multiprocessor system 1800 is a point-to-point interconnect system and includes a first processor 1870 and a second processor 1880 coupled via a point-to-point interconnect 1850. Each of processors 1870 and 1880 may be some version of the processor 1600. In one embodiment of the present invention, processors 1870 and 1880 are processor 1710 and 1715 respectively, and the coprocessor 1838 is the coprocessor 1745. In another embodiment, processors 1870 and 1880 are processor 1710, coprocessor 1745 respectively.
[0168] Processors 1870 and 1880 are shown to include integrated memory controller (IMC) units 1872 and 1882 respectively. Processor 1870 also includes portions of point-to-point (P-P) interfaces 1876 and 1878 as its bus controller units; similarly, second processor 1880 includes P-P interfaces 1886 and 1888. Using P-P interface circuits 1878, 1888, processors 1870, 1880 may exchange information via a point-to-point (P-P) interface 1850. As Figure 18 shown, IMCs 1872 and 1882 couple the processors to respective memories (i.e., memory 1832 and memory 1834), which may be portions of main memories locally attached to the respective processors.
[0169] Using point-to-point interface circuits 1876, 1894, 1886, 1898, processors 1870, 1880 may each exchange information with a chipset 1890 via respective P-P interfaces 1852, 1854. The chipset 1890 may optionally exchange information with a coprocessor 1838 via a high performance interface 1838. In one embodiment, the coprocessor 1838 is a special purpose processor such as, for example, a high throughput MIC processor, a network or communications processor, a compression and / or decompression engine, a graphics processor, a GPGPU, an embedded processor, and the like.
[0170] A shared cache memory (not shown) may be included in either processor or outside both processors, and is connected to the processors via a P-P interconnect such that if a processor is placed in a low power mode, local cache memory information of either or both processors can be stored in the shared cache memory.
[0171] Chipset 1890 can be coupled to first bus 1816 via interface 1896. In one embodiment, first bus 1816 can be a Peripheral Component Interconnect (PCI) bus, or a bus such as a PCI Express bus or another third generation I / O interconnect bus, although the scope of the present disclosure is not so limited.
[0172] As Figure 18 shown, various I / O devices 1814 can be coupled to first bus 1816 along with bus bridge 1818, and bus bridge 818 couples first bus 1816 to second bus 1820. In one embodiment, one or more additional processors 1815, such as a coprocessor, a high throughput MIC processor, a GPGPU, an accelerator (such as, for example, a graphics accelerator or a Digital Signal Processing (DSP) unit), a Field Programmable Gate Array, or any other processor, are coupled to first bus 1816. In one embodiment, second bus 1820 can be a Low Pin Count (LPC) bus. Various devices can be coupled to second bus 1820, including, for example, a keyboard and / or mouse 1822, a communication device 1827, and a storage unit 1828 such as a hard disk drive or other mass storage device, which can include instructions / code and data 1830 (in one embodiment). Further, audio I / O 1824 can be coupled to second bus 1820. Note that other architectures are contemplated by the present disclosure. For example, instead of Figure 18 the point-to-point architecture, the system can implement a multi-drop bus or another such architecture.
[0173] Figure 19 FIG. depicts a block diagram of a second more specific exemplary system 1900 in accordance with an embodiment of the present disclosure. Figure 18 and 19 like elements in are labeled with like reference numerals, and Figure 18 certain aspects of have been omitted from Figure 19 to avoid obscuring Figure 19 other aspects of.
[0174] Figure 19 It is shown that processors 1870, 1880 can respectively include integrated memory and I / O control logic (“CL”) 1872 and 1882. Thus, CL 1872, 1882 includes an integrated memory controller unit and includes I / O control logic.Figure 19 It is shown that not only memories 1832 and 1834 are coupled to CLs 1872 and 1882, but also I / O device 1914 is coupled to control logics 1872 and 1882. Legacy I / O device 1915 is coupled to chipset 1890.
[0175] Figure 20 A block diagram of SoC 2000 in accordance with an embodiment of the present disclosure is depicted. Figure 16 Similar elements in [figure] are labeled with like reference numerals. Also, the dashed boxes are optional features on a more advanced SoC. In Figure 20 one or more interconnect units 2002 are coupled to: an application processor 2010, which includes a set of one or more cores 1602A-N and one or more shared cache memory units 1606; a system agent unit 1610; one or more bus controller units 1616; one or more integrated memory controller units 1614; a set of co-processors 2020 or one or more, which may include integrated graphics logic, an image processor, an audio processor, and a video processor; a static random access memory (SRAM) unit 2030; a direct memory access (DMA) unit 2032; and a display unit 2040 for coupling to one or more external displays. In one embodiment, co-processor 2020 includes a dedicated processor, such as, for example, a network or communication processor, a compression and / or decompression engine, a GPGPU, a high throughput MIC processor, an embedded processor, and the like.
[0176] Embodiments of the mechanisms disclosed herein may be implemented in hardware, software, firmware, or a combination of such implementation means. Embodiments of the present invention may be implemented as program code or a computer program executed on a programmable system, the programmable system including at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device.
[0177] such as Figure 8 The program code, such as code 830 shown in [figure], may be applied to input instructions to perform the functions described herein and generate output information. The output information may be applied to one or more output devices in a known manner. For the purpose of this application, a processing system includes any system having a processor (such as, for example: a digital signal processor (DSP), a microcontroller, an application specific integrated circuit (ASIC), or a microprocessor).
[0178] The program code can be implemented in a high-level procedural or object-oriented programming language to communicate with the processing system. If desired, the program code can also be implemented in assembly or machine language. In fact, the mechanisms described herein are not limited to any specific programming language. In any case, the language can be a compiled or interpreted language.
[0179] One or more aspects of at least one embodiment can be implemented by representative instructions stored on a machine-readable medium, which represent various logics within a processor and, when read by the machine, cause the machine to fabricate the logics for performing the techniques described herein. Such representations (known as "IP cores") can be stored on a tangible, machine-readable medium and supplied to various customers or manufacturing facilities to be loaded into the fabrication machines that actually make the logics or processors.
[0180] Such machine-readable storage media can include, without limitation, non-transitory, tangible arrangements of articles manufactured or formed by a machine or device, including storage media such as hard disks, any other type of disk including floppy disks, optical disks, compact disk read-only memory (CD-ROM), rewritable compact disks (CD-RW), and magneto-optical disks, semiconductor devices such as read-only memory (ROM), random access memory (RAM) such as dynamic random access memory (DRAM), static random access memory (SRAM), erasable programmable read-only memory (EPROM), flash memory, electrically erasable programmable read-only memory (EEPROM), phase change memory (PCM), magnetic or optical cards, or any other type of medium suitable for storing electronic instructions.
[0181] Accordingly, embodiments of the present invention also include non-transitory, tangible machine-readable media that contain instructions or contain design data, such as a hardware description language (HDL), which define the structures, circuits, devices, processors, and / or system features described herein. Such embodiments may also be referred to as program products.
[0182] Emulation (including binary translation, code morphing, etc.)
[0183] In some cases, an instruction converter can be used to convert instructions from a source instruction set to a target instruction set. For example, the instruction converter can translate (e.g., using static binary translation, dynamic binary translation including dynamic compilation), morph, emulate, or otherwise convert the instructions into one or more other instructions to be processed by the core. The instruction converter is implemented in software, hardware, firmware, or a combination thereof. The instruction converter can be on the processor, off the processor, or partially on the processor and partially off the processor.
[0184] Figure 11is a block diagram that illustrates the use of a software instruction converter to convert binary instructions in a source instruction set into binary instructions in a target instruction set, according to an embodiment of the present invention. In the illustrated embodiment, the instruction converter is a software instruction converter, although alternatively the instruction converter may be implemented in software, firmware, hardware, or various combinations thereof. Figure 11 shows that using a first compiler 1104, a program in a high-level language 1102 can be compiled to generate a first binary code (e.g., x86) 1106, which can be natively executed by a processor 1116 with at least one first instruction set core. In some embodiments, the processor 1116 with at least one first instruction set core represents any processor that can generally perform the same functions as an Intel processor with at least one x86 instruction set core, either by compatibly executing or otherwise processing (1) a substantial portion of the instruction set of the Intel x86 instruction set core, or (2) an object code version of an application or other software targeted for running on an Intel processor with at least one x86 instruction set core, so as to obtain substantially the same results as an Intel processor with at least one x86 instruction set core. The first compiler 1104 represents a compiler operable to generate the binary code 1106 (e.g., object code) of a first instruction set, and the binary code 1106 of the first instruction set can be executed on the processor 1116 with at least one first instruction set core with or without additional linking processing. Similarly, Figure 11 shows the use of an alternative instruction set compiler 1108, where a program in a high-level language 1102 can be compiled to generate alternative instruction set binary code 1110, which can be natively executed by a processor 1114 without at least one first instruction set core (e.g., a processor with a core that executes the MIPS instruction set of MIPS Technologies of Sunnyvale, CA and / or the ARM instruction set of ARM Holdings of Sunnyvale, CA). An instruction converter 1112 is used to convert the first binary code 1106 into code that can be natively executed by the processor 1114 without the first instruction set core. This converted code is unlikely to be the same as the alternative instruction set binary code 1110 because it is difficult to make an instruction converter that can do so; however, the converted code will perform general operations and be composed of instructions from the alternative instruction set. Thus, the instruction converter 1112 represents software, firmware, hardware, or a combination thereof that allows a processor or another electronic device without a first instruction set processor or core to execute the first binary code 1106 through emulation, simulation, or any other process.
[0185] Designs can go through various stages from creation to simulation to fabrication. The data representing a design can be expressed in many ways to represent the design. First, as useful in simulation, hardware can be represented using a hardware description language (HDL) or another functional description language. Additionally, at certain stages of the design process, a circuit-level model with logic and / or transistor gates can be generated. Further, most designs reach a data level at some stage that represents the physical placement of various devices in the hardware model. In cases where conventional semiconductor fabrication techniques are used, the data representing the hardware model can be data specifying the presence or absence of various features on different mask layers of a mask used to fabricate an integrated circuit. In some implementations, such data can be stored in a database file format such as Graphics Data System II (GDS II), Open Artwork System Interchange Standard (OASIS), or a similar format.
[0186] In some implementations, software-based hardware models, as well as HDL and other functional description language objects, can include register transfer language (RTL) files (among other examples). Such objects can be machine-parsable such that a design tool can accept an HDL object (or model), parse the HDL object to obtain the properties of the described hardware, and determine the physical circuit and / or on-chip layout from the object. The output of the design tool can be used to fabricate a physical device. For example, in addition to other attributes of the system modeled in the HDL object that would be implemented, the design tool can also determine the configuration of various hardware and / or firmware elements from the HDL object, such as bus widths, registers (including size and type), memory blocks, physical link paths, fabric topologies. The design tool can include tools for determining the topology and fabric configuration of a system-on-chip (SoC) and other hardware devices. In some cases, the HDL object can be used as a basis for a development model and design file that can be used by a fabrication device to fabricate the described hardware. In fact, the HDL object itself can be provided as an input to the fabrication system software to cause the fabrication of the hardware.
[0187] In any representation of a design, the data representing the design can be stored in any form of machine-readable medium. A memory or a magnetic or optical storage device (such as a disk) can be a machine-readable medium to store information transmitted via modulated or otherwise generated light or radio waves to convey such information. When an electrical carrier wave indicating or carrying code or a design is transmitted, a new copy is made in terms of copying, buffering, or retransmitting the electrical signal. Thus, a communication provider or network provider can store at least temporarily on a tangible machine-readable medium an article embodying the technology of the embodiments of the present disclosure, such as information encoded into a carrier wave.
[0188] In various embodiments, a medium storing a representation of a design can be provided to a manufacturing system (e.g., a semiconductor manufacturing system capable of manufacturing integrated circuits and / or related components). The design representation can instruct the system to manufacture a device capable of performing any combination of the functions described above. For example, the design representation can instruct the system regarding which components to manufacture, how the components should be coupled together, where the components should be placed on the device, and / or regarding other suitable specifications (regarding the device to be manufactured).
[0189] Accordingly, one or more aspects of at least one embodiment can be implemented by representative instructions stored on a machine-readable medium, the representative instructions representing various logic within a processor, the logic which when read by the machine causes the machine to manufacture logic for performing the techniques described herein. Such representations (commonly referred to as “IP cores”) can be stored on a non-transitory tangible machine-readable medium and supplied to various customers or manufacturing facilities to be loaded into the manufacturing logic or the manufacturing machines of processors.
[0190] Embodiments of the mechanisms disclosed herein can be implemented in hardware, software, firmware, or a combination of such implementation approaches. Embodiments of the present disclosure can be implemented as a computer program or program code executing on a programmable system that includes at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device.
[0191] The program code (e.g., Figure 18 the code 1830 shown in
[0192] can be applied to input instructions to perform the functions described herein and generate output information. The output information can be applied to one or more output devices in a known manner. For the purposes of this application, a processing system includes any system having a processor (such as, for example, a digital signal processor (DSP), a microcontroller, an application specific integrated circuit (ASIC), or a microprocessor).
[0193] Embodiments of the methods, hardware, software, firmware, or code described above may be implemented via code or instructions executable (or otherwise accessible) by a processing element and stored on a machine-accessible, machine-readable, computer-accessible, or computer-readable medium. A non-transitory machine-accessible / readable medium includes any mechanism that provides (i.e., stores and / or transports) information in a form readable by a machine, such as a computer or an electronic system. For example, a non-transitory machine-accessible medium includes random access memory (RAM), such as static RAM (SRAM) or dynamic RAM (DRAM); ROM; magnetic or optical storage media; flash memory devices; electrical storage devices; optical storage devices; acoustic storage devices; other forms of storage devices for holding information received from a transitory (propagating) signal (e.g., a carrier wave, an infrared signal, a digital signal); etc., which is to be distinguished from a non-transitory medium from which information can be received.
[0194] Instructions for programming logic to perform embodiments of the present disclosure may be stored in a memory within the system, such as in DRAM, cache, flash memory, or other storage devices. Additionally, the instructions may be distributed via a network or through other computer-readable media. Thus, a machine-readable medium may include any mechanism for storing or transporting information in a form readable by a machine, such as a computer, but is not limited to floppy disks, optical disks, compact disks, CD-ROMs, and magneto-optical disks, ROMs, RAMs, erasable programmable ROMs (EPROMs), electrically erasable programmable ROMs (EEPROMs), magnetic or optical cards, flash memory, or tangible machine-readable storage devices used in the transmission of information via electrical, optical, acoustic, or other forms of propagated signals (e.g., carrier waves, infrared signals, digital signals, etc.) over the Internet. Accordingly, a computer-readable medium includes any type of tangible machine-readable medium suitable for storing or transporting electronic instructions or information in a form readable by a machine, such as a computer.
[0195] Logic can be used to implement any functionality of various components, such as network element 102, router 104, core 108, Figure 7Logic, neuron nucleus controller 1100, neuromorphic nucleus 1200, any processor described herein, any other components described herein, or any sub-components of any of these components. "Logic" can refer to hardware, firmware, software, and / or any combination thereof to perform one or more functions. As an example, logic can include hardware (such as a microcontroller or a processor) associated with a non-transitory medium to store code suitable for execution by the microcontroller or the processor. Thus, in one embodiment, a reference to logic refers to hardware that is specifically configured to identify and / or execute code to be held on a non-transitory medium. Additionally, in another embodiment, the use of logic refers to a non-transitory medium that includes code that is particularly suitable for execution by a microcontroller to perform a predetermined operation. And as can be inferred, in yet another embodiment, the term logic (in this example) can refer to a combination of hardware and a non-transitory medium. In various embodiments, logic can include a microprocessor or other processing element operable to execute software instructions, discrete logic such as an application-specific integrated circuit (ASIC), a programmable logic device such as a field-programmable gate array (FPGA), a memory device containing instructions, a combination of logic devices (e.g., as would be found on a printed circuit board), or other suitable hardware and / or software. Logic can include one or more gates or other circuit components, which can be implemented by, for example, transistors. In some embodiments, logic can also be implemented entirely as software. Software can be implemented as a software package, code, instructions, instruction set, and / or data recorded on a non-transitory computer-readable storage medium. Firmware can be implemented as code, instructions, or instruction set and / or data hard-coded (e.g., non-volatile) in a memory device. Generally, the logical boundaries shown as separate typically vary and potentially overlap. For example, first and second logic can share hardware, software, firmware, or a combination thereof while potentially retaining some separate hardware, software, or firmware.
[0196] In one embodiment, the use of the phrase "to" or "configured to" refers to arranging, putting together, manufacturing, offering for sale, importing, and / or designing a device, hardware, logic, or element to perform a specified or determined task. In this example, a device or its element that is not operating is still "configured to" perform the specified task (if it is designed, coupled, and / or interconnected to perform the specified task). As a purely illustrative example, a logic gate can provide a 0 or a 1 during operation. But a logic gate "configured to" provide an enable signal to a clock does not include every potential logic gate that might provide a 1 or a 0. Instead, a logic gate "configured to" provide an enable signal to a clock is a logic gate that is coupled in a certain way (such that a 1 or 0 output during operation is to enable the clock). Again, note that the use of the term "configured to" does not require operation, but rather focuses on the latent state of a device, hardware, and / or element, where in the latent state, the device, hardware, and / or element is designed to perform a specific task when the device, hardware, and / or element is operating.
[0197] In addition, in one embodiment, the use of the phrases "capable of" and / or "operable to" refers to a device, logic, hardware, and / or element designed in such a way as to enable the use of a device, logic, hardware, and / or element in a specified manner. As noted above, in one embodiment, the use of "able to", "capable of", or "operable to" refers to a latent state of a device, logic, hardware, and / or element, where the device, logic, hardware, and / or element is not operating but is designed in such a way as to enable the use of a device in a specified manner.
[0198] As used herein, a value includes any known representation of a number, state, logical state, or binary logical state. Generally, the use of the terms logic level, logic value, or logical value, also referred to as 1 and 0, simply represents a binary logical state. For example, 1 refers to a high logic level, and 0 refers to a low logic level. In one embodiment, a storage unit, such as a transistor or a flash memory cell, may be capable of holding a single logic value or multiple logic values. However, other representations of values in a computer system have been used. For example, the decimal number 10 may also be represented as the binary value 1010 and the hexadecimal letter A. Thus, a value includes any representation of information that can be held in a computer system.
[0199] In addition, a state may be represented by a value or a portion of a value. As an example, a first value, such as logic 1, may represent a default or initial state, while a second value, such as logic 0, may represent a non-default state. In addition, in one embodiment, the terms reset and set refer to a default value and an updated value or state, respectively. For example, the default value potentially includes a high logic value, i.e., reset, while the updated value potentially includes a low logic value, i.e., set. Note that any combination of values may be utilized to represent any number of states.
[0200] In at least one embodiment, a processor includes: a first neuromorphic core for implementing a plurality of neural units of a neural network, the first neuromorphic core including: a memory for storing a current time step of the first neuromorphic core; and a controller for: tracking a current time step of an adjacent neuromorphic core, the neuromorphic core receiving a pulse from or providing a pulse to the first neuromorphic core; and controlling the current time step of the first neuromorphic core based on the current time step of the adjacent neuromorphic core.
[0201] In an embodiment, the first neuromorphic core is to process pulses received from a second neuromorphic core, where the pulses occur in a first time step that is later than the current time step of the first neuromorphic core when the pulses are processed by the first neuromorphic core. In an embodiment, during a time period in which the current time step of the first neuromorphic core is the first time step, the first neuromorphic core is to receive a first pulse from the second neuromorphic core and a second pulse from a third neuromorphic core, where the first pulse occurs in a second time step and a second output pulse occurs in a time step different from the second time step. In an embodiment, during a time period in which the current time step of the first neuromorphic core is the first time step, the first neuromorphic core is to: process the first pulse by accessing a first synaptic weight associated with a first output pulse and adjusting a first membrane potential increment; and process the second pulse by accessing a second synaptic weight associated with the second output pulse and adjusting a second membrane potential increment. In an embodiment, if a second neuromorphic core that is to send a pulse to the first neuromorphic core is set to a time step earlier than the current time step of the first neuromorphic core, the controller is to prevent the first neuromorphic core from advancing to the next time step. In an embodiment, if a second neuromorphic core that is to receive a pulse from the first neuromorphic core is set to a time step earlier than the current time step of the first neuromorphic core by more than a threshold number of time steps, the controller prevents the first neuromorphic core from advancing to the next time step. In an embodiment, when the current time step of the first of the first neuromorphic cores is incremented, the controller of the first neuromorphic core is to send a message to the adjacent neuromorphic cores, the message indicating that the current time step of the first neuromorphic core has been incremented. In an embodiment, when the current time step of the first of the first neuromorphic cores changes by one or more time steps, the controller of the first neuromorphic core is to send a message including at least a portion of the current time step of the first neuromorphic core to the adjacent neuromorphic cores. In an embodiment, the first neuromorphic core includes a pulse buffer, the pulse buffer including a first entry for storing a pulse of a first time step and a second entry for storing a pulse of a second time step, where the pulse of the first time step and the pulse of the second time step are to be stored in the buffer concurrently. In an embodiment, the first neuromorphic core includes a buffer, the buffer including a first entry for storing membrane potential increment values for the plurality of neural units for a first time step and a second entry for storing membrane potential increment values for the plurality of neural units for a second time step.In an embodiment, the controller is to control the current time step of the first neuromorphic core based on the number of allowed look-ahead states, where the number of allowed look-ahead states is determined by the amount of available memory for storing pulses of the allowed look-ahead states. In an embodiment, the processor further includes a battery communicatively coupled to the processor, a display communicatively coupled to the processor, or a network interface communicatively coupled to the processor.
[0202] In at least one embodiment, a method includes: implementing a plurality of neural units of a neural network in a first neuromorphic core; storing a current time step of the first neuromorphic core; tracking current time steps of adjacent neuromorphic cores that receive pulses from or provide pulses to the first neuromorphic core; and controlling the current time step of the first neuromorphic core based on the current time steps of the adjacent neuromorphic cores.
[0203] In an embodiment, the method further includes the first neuromorphic core processing pulses received from a second neuromorphic core, where the pulses occur in a first time step that is later than the current time step of the first neuromorphic core when processing the pulses. In an embodiment, the method further includes the first neuromorphic core receiving a first pulse from the second neuromorphic core and a second pulse from a third neuromorphic core during a time period in which the current time step of the first neuromorphic core is the first time step, where the first pulse occurs in a second time step and the second output pulse occurs in a time step different from the second time step. In an embodiment, the method further includes, during a time period in which the first neuromorphic core is set to the first time step: processing the first pulse by accessing a first synaptic weight associated with the first pulse and adjusting a first membrane potential increment; and processing the second pulse by accessing a second synaptic weight associated with the second pulse and adjusting a second membrane potential increment. In an embodiment, the method further includes: preventing the first neuromorphic core from advancing to the next time step if a second neuromorphic core that is to send a pulse to the first neuromorphic core is set to a time step earlier than the current time step of the first neuromorphic core. In an embodiment, the method further includes: preventing the first neuromorphic core from advancing to the next time step if a second neuromorphic core that is to receive a pulse from the first neuromorphic core is set to a time step that is earlier than the current time step of the first neuromorphic core by more than a threshold number of time steps. In an embodiment, the method further includes: when the current time step of the first neuromorphic core is incremented, sending a message to an adjacent neuromorphic core, the message indicating that the current time step of the first neuromorphic core has been incremented. In an embodiment, the method further includes: when the current time step of the first neuromorphic core changes by one or more time steps, sending a message including at least a portion of the current time step of the first neuromorphic core to an adjacent neuromorphic core. In an embodiment, the first neuromorphic core includes a pulse buffer, the pulse buffer including a first entry for storing a pulse of a first time step and a second entry for storing a pulse of a second time step, where the pulse of the first time step and the pulse of the second time step are to be stored in the buffer concurrently. In an embodiment, the first neuromorphic core includes a buffer, the buffer including a first entry for storing membrane potential increment values of the plurality of neural units for a first time step, and a second entry for storing membrane potential increment values of the plurality of neural units for a second entry. In an embodiment, the method further includes controlling the current time step of the first neuromorphic core based on the number of allowed look-ahead states, where the number of allowed look-ahead states is determined by the amount of available memory for storing pulses of the allowed look-ahead states.
[0204] In at least one embodiment, a non-transitory machine-readable storage medium has instructions stored thereon that, when executed by a machine, cause the machine to: implement a plurality of neural units of a neural network in a first neuromorphic core; store a current time step of the first neuromorphic core; track current time steps of adjacent neuromorphic cores that receive spikes from or provide spikes to the first neuromorphic core; and control the current time step of the first neuromorphic core based on the current time steps of the adjacent neuromorphic cores.
[0205] In an embodiment, the instructions, when executed by the machine, cause the machine to: process a spike received from a second neuromorphic core in the first neuromorphic core, where the spike occurs in a first time step that is later than the current time step of the first neuromorphic core when processing the spike. In an embodiment, the instructions, when executed by the machine, cause the machine to: receive a first spike from a second neuromorphic core and a second spike from a third neuromorphic core in the first neuromorphic core during a period in which the current time step of the first neuromorphic core is a first time step, where the first spike occurs in a second time step and the second output spike occurs in a time step different from the second time step. In an embodiment, the instructions, when executed by the machine, cause the machine to: process the first spike by accessing a first synaptic weight associated with the first spike and adjusting a first membrane potential increment; and process the second spike by accessing a second synaptic weight associated with the second spike and adjusting a second membrane potential increment.
[0206] In at least one embodiment, a system includes: means for implementing a plurality of neural units of a neural network in a first neuromorphic core; means for storing a current time step of the first neuromorphic core; means for tracking current time steps of adjacent neuromorphic cores that receive spikes from or provide spikes to the first neuromorphic core; and means for controlling the current time step of the first neuromorphic core based on the current time steps of the adjacent neuromorphic cores.
[0207] In an embodiment, the system further includes components for the first neuromorphic core to process pulses received from a second neuromorphic core, where the pulses occur in a first time step that is later than the current time step of the first neuromorphic core when processing the pulses. In an embodiment, the system further includes components for the first neuromorphic core to receive a first pulse from a second neuromorphic core and a second pulse from a third neuromorphic core during a period when the current time step of the first neuromorphic core is the first time step, where the first pulse occurs in a second time step, and the second output pulse occurs in a time step different from the second time step. In an embodiment, the system further includes components for performing the following actions during a period when the first neuromorphic core is set to the first time step: processing the first pulse by accessing a first synaptic weight associated with the first pulse and adjusting a first membrane potential increment; and processing the second pulse by accessing a second synaptic weight associated with the second pulse and adjusting a second membrane potential increment.
[0208] In at least one embodiment, the system includes a processor, and the processor includes: a first neuromorphic core for implementing a plurality of neural units of a neural network; the first neuromorphic core includes a memory for storing the current time step of the first neuromorphic core; and a controller for tracking the current time steps of adjacent neuromorphic cores, where the neuromorphic cores receive pulses from the first neuromorphic core or provide pulses to the first neuromorphic core; and controlling the current time step of the first neuromorphic core based on the current time steps of the adjacent neuromorphic cores; the system further includes a memory coupled to the processor for storing the results generated by the neural network.
[0209] In an embodiment, the system further includes a network interface for transmitting the results generated by the neural network. In an embodiment, the system further includes a display for displaying the results generated by the neural network. In an embodiment, the system further includes a cellular communication interface.
[0210] The following technical solutions are also provided herein:
[0211] 1. A processor, comprising:
[0212] A first neuromorphic core, the first neuromorphic core being used to implement a plurality of neural units of a neural network, and the first neuromorphic core includes:
[0213] A memory, the memory being used to store the current time step of the first neuromorphic core; and
[0214] A controller, the controller being used for:
[0215] Tracking the current time step of adjacent neuromorphic cores that receive spikes from or provide spikes to the first neuromorphic core; and
[0216] Controlling the current time step of the first neuromorphic core based on the current time step of the adjacent neuromorphic cores.
[0217] 2. The processor according to claim 1, wherein the first neuromorphic core is to process spikes received from a second neuromorphic core, and wherein the spikes occur in a first time step that is later than the current time step of the first neuromorphic core when the first neuromorphic core processes the spikes.
[0218] 3. The processor according to claim 1, wherein during a period in which the current time step of the first neuromorphic core is a first time step, the first neuromorphic core is to receive a first spike from a second neuromorphic core and a second spike from a third neuromorphic core, wherein the first spike occurs in a second time step and a second output spike occurs in a time step different from the second time step.
[0219] 4. The processor according to claim 3, wherein during a period in which the current time step of the first neuromorphic core is the first time step, the first neuromorphic core is to:
[0220] Process the first spike by accessing a first synaptic weight associated with a first output spike and adjusting a first membrane potential increment; and
[0221] Process the second spike by accessing a second synaptic weight associated with the second output spike and adjusting a second membrane potential increment.
[0222] 5. The processor according to claim 1, wherein if a second neuromorphic core that is to send spikes to the first neuromorphic core is set to a time step earlier than the current time step of the first neuromorphic core, the controller is to prevent the first neuromorphic core from advancing to the next time step.
[0223] 6. The processor according to claim 1, wherein if a second neuromorphic core that is to receive spikes from the first neuromorphic core is set to a time step earlier than the current time step of the first neuromorphic core by more than a threshold number of time steps, the controller prevents the first neuromorphic core from advancing to the next time step.
[0224] 7. The processor according to claim 1, wherein when the current time step of the first of the first neuromorphic cores is incremented, the controller of the first neuromorphic core is to send a message to the adjacent neuromorphic core, the message indicating that the current time step of the first neuromorphic core has been incremented.
[0225] 8. The processor according to claim 1, wherein when the current time step of the first of the first neuromorphic cores changes by one or more time steps, the controller of the first neuromorphic core is to send a message including at least a portion of the current time step of the first neuromorphic core to the adjacent neuromorphic core.
[0226] 9. The processor according to claim 1, wherein the first neuromorphic core includes a pulse buffer, the pulse buffer including a first entry for storing pulses for a first time step and a second entry for storing pulses for a second time step, wherein the pulses for the first time step and the pulses for the second time step are to be stored in the buffer concurrently.
[0227] 10. The processor according to claim 1, wherein the first neuromorphic core includes a buffer, the buffer including a first entry for storing the membrane potential increment values of the plurality of neural units for a first time step and a second entry for storing the membrane potential increment values of the plurality of neural units for a second time step.
[0228] 11. The processor according to claim 1, wherein the controller is to control the current time step of the first neuromorphic core based on the number of allowed look-ahead states, wherein the number of allowed look-ahead states is determined by the amount of available memory for storing pulses for the allowed look-ahead states.
[0229] 12. The processor according to claim 1, further comprising a battery communicatively coupled to the processor, a display communicatively coupled to the processor, or a network interface communicatively coupled to the processor.
[0230] 13. A non-transitory machine-readable storage medium having instructions stored thereon that, when executed by a machine, cause the machine to:
[0231] Implement a plurality of neural units of a neural network in a first neuromorphic core;
[0232] Store the current time step of the first neuromorphic core;
[0233] Track the current time steps of adjacent neuromorphic cores that receive pulses from or provide pulses to the first neuromorphic core; and
[0234] Control the current time step of the first neuromorphic core based on the current time step of the adjacent neuromorphic cores.
[0235] 14. The medium according to claim 13, wherein when executed by the machine, the instructions cause the machine to process, in the first neuromorphic core, a pulse received from a second neuromorphic core, wherein the pulse occurs in a first time step later than the current time step of the first neuromorphic core when the pulse is being processed.
[0236] 15. The medium according to claim 13, wherein when executed by the machine, the instructions cause the machine to receive, in the first neuromorphic core, a first pulse from a second neuromorphic core and a second pulse from a third neuromorphic core during a period in which the current time step of the first neuromorphic core is a first time step, wherein the first pulse occurs in a second time step and the second output pulse occurs in a time step different from the second time step.
[0237] 16. The medium according to claim 15, wherein when executed by the machine, the instructions cause the machine, during a period in which the current time step of the first neuromorphic core is a first time step:
[0238] Process the first pulse by accessing a first synaptic weight associated with the first pulse and adjusting a first membrane potential increment; and
[0239] Process the second pulse by accessing a second synaptic weight associated with the second pulse and adjusting a second membrane potential increment.
[0240] 17. A method, comprising:
[0241] Implement a plurality of neural units of a neural network in a first neuromorphic core;
[0242] Store the current time step of the first neuromorphic core;
[0243] Track the current time steps of adjacent neuromorphic cores that receive pulses from or provide pulses to the first neuromorphic core; and
[0244] Control the current time step of the first neuromorphic core based on the current time steps of the adjacent neuromorphic cores.
[0245] 18. The method as described in Technical Solution 16 further includes the first neuromorphic core processing the pulses received from the second neuromorphic core, wherein when processing the pulses, the pulses occur in a first time step that is later than the current time step of the first neuromorphic core.
[0246] 19. The method as described in Technical Solution 16 further includes the first neuromorphic core receiving a first pulse from the second neuromorphic core and a second pulse from the third neuromorphic core during a period when the current time step of the first neuromorphic core is the first time step, wherein the first pulse occurs in a second time step, and the second output pulse occurs in a time step different from the second time step.
[0247] 20. The method as described in Technical Solution 19 further includes, during the period of setting the first neuromorphic core to the first time step:
[0248] processing the first pulse by accessing the first synaptic weight associated with the first pulse and adjusting the first membrane potential increment; and
[0249] processing the second pulse by accessing the second synaptic weight associated with the second pulse and adjusting the second membrane potential increment.
[0250] References to "an embodiment" or "embodiments" throughout this specification mean that the specific features, structures, or characteristics described in connection with the embodiment are included in at least one embodiment of the present disclosure. Thus, the phrases "in an embodiment" or "in embodiments" that appear in various places throughout this specification are not necessarily all referring to the same embodiment. Additionally, the specific features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0251] In the foregoing specification, detailed descriptions have been given with reference to specific exemplary implementations. However, it will be apparent that various modifications and changes can be made thereto without departing from the broader spirit and scope of the present disclosure as set forth in the appended claims. Accordingly, the specification and drawings are to be regarded in an illustrative rather than a limiting sense. Additionally, the foregoing use of embodiments and other exemplary language does not necessarily refer to the same embodiment or the same example, but may refer to different and distinct embodiments and potentially the same embodiment.
Claims
1. A processor, comprising: A first neuromorphic core for implementing a plurality of neural units of a neural network, the first neuromorphic core comprising: A memory for storing the current time step of the first neuromorphic core; and A controller for: Tracking the current time steps of adjacent neuromorphic cores that receive pulses from or provide pulses to the first neuromorphic core; Controlling the current time step of the first neuromorphic core based on the current time steps of the adjacent neuromorphic cores; Preventing the first neuromorphic core from advancing to the next time step if a second neuromorphic core that is to send a pulse to the first neuromorphic core is set to a time step earlier than the current time step of the first neuromorphic core; and Preventing the first neuromorphic core from advancing to the next time step if a second neuromorphic core that is to receive a pulse from the first neuromorphic core is set to more than a threshold number of time steps earlier than the current time step of the first neuromorphic core.
2. The processor according to claim 1, wherein the first neuromorphic core is to process a pulse received from a second neuromorphic core, and wherein the pulse occurs in a first time step later than the current time step of the first neuromorphic core when the pulse is processed by the first neuromorphic core.
3. The processor according to any one of claims 1-2, wherein during a period in which the current time step of the first neuromorphic core is a first time step, the first neuromorphic core is to receive a first pulse from a second neuromorphic core and a second pulse from a third neuromorphic core, wherein the first pulse occurs in a second time step and the second output pulse occurs in a time step different from the second time step.
4. The processor according to claim 3, wherein during a period in which the current time step of the first neuromorphic core is the first time step, the first neuromorphic core is to: Process the first pulse by accessing a first synaptic weight associated with a first output pulse and adjusting a first membrane potential increment; and Process the second pulse by accessing a second synaptic weight associated with the second output pulse and adjusting a second membrane potential increment.
5. The processor according to any one of claims 1-2, wherein when the current time step of the first neuromorphic core is incremented, the controller of the first neuromorphic core is to send a message to the adjacent neuromorphic core indicating that the current time step of the first neuromorphic core has been incremented.
6. The processor according to any one of claims 1-2, wherein when the current time step of the first neuromorphic core changes by one or more time steps, the controller of the first neuromorphic core is to send a message including at least a portion of the current time step of the first neuromorphic core to the adjacent neuromorphic core.
7. The processor according to any one of claims 1-2, wherein the first neuromorphic core includes a pulse buffer, the pulse buffer including a first entry for storing pulses of a first time step and a second entry for storing pulses of a second time step, wherein the pulses of the first time step and the pulses of the second time step are to be stored in the buffer concurrently.
8. The processor according to any one of claims 1-2, wherein the first neuromorphic core includes a buffer, the buffer including a first entry for storing the membrane potential increment values of the plurality of neural units for a first time step and a second entry for storing the membrane potential increment values of the plurality of neural units for a second time step.
9. The processor according to any one of claims 1-2, wherein the controller is to control the current time step of the first neuromorphic core based on the number of allowed look-ahead states, wherein the number of allowed look-ahead states is determined by the amount of available memory for storing the pulses of the allowed look-ahead states.
10. The processor according to any one of claims 1-2, further comprising a battery communicatively coupled to the processor, a display communicatively coupled to the processor, or a network interface communicatively coupled to the processor.
11. A time step processing method, comprising: Implementing a plurality of neural units of a neural network in a first neuromorphic core; Storing the current time step of the first neuromorphic core; Tracking the current time steps of adjacent neuromorphic cores that receive pulses from or provide pulses to the first neuromorphic core; Controlling the current time step of the first neuromorphic core based on the current time steps of the adjacent neuromorphic cores; Preventing the first neuromorphic core from advancing to the next time step if a second neuromorphic core that is to send a pulse to the first neuromorphic core is set to a time step earlier than the current time step of the first neuromorphic core; And Preventing the first neuromorphic core from advancing to the next time step if a second neuromorphic core that is to receive a pulse from the first neuromorphic core is set to a time step that is earlier than the current time step of the first neuromorphic core by more than a threshold number of time steps.
12. The method according to claim 11, further comprising processing, in the first neuromorphic core, a pulse received from a second neuromorphic core, wherein the pulse occurs in a first time step that is later than the current time step of the first neuromorphic core when processing the pulse.
13. The method according to claim 11, further comprising receiving, in the first neuromorphic core, a first pulse from a second neuromorphic core and a second pulse from a third neuromorphic core during a period in which the current time step of the first neuromorphic core is a first time step, wherein the first pulse occurs in a second time step and the second output pulse occurs in a time step different from the second time step.
14. The method according to claim 13, further comprising, during the period of setting the first neuromorphic core to the first time step: processing the first pulse by accessing a first synaptic weight associated with the first pulse and adjusting a first membrane potential increment; and processing the second pulse by accessing a second synaptic weight associated with the second pulse and adjusting a second membrane potential increment.
15. The method according to any one of claims 11-14, further comprising: When the current time step of the first of the first neuromorphic cores is incremented, sending a message to the adjacent neuromorphic core, the message indicating that the current time step of the first neuromorphic core has been incremented.
16. The method according to any one of claims 11-14, further comprising: When the current time step of the first of the first neuromorphic cores changes by one or more time steps, sending a message including at least a portion of the current time step of the first neuromorphic core to the adjacent neuromorphic core.
17. The method according to any one of claims 11-14, wherein the first neuromorphic core includes a pulse buffer, the pulse buffer including a first entry for storing pulses of a first time step and a second entry for storing pulses of a second time step, wherein the pulses of the first time step and the pulses of the second time step are to be stored in the buffer concurrently.
18. The method according to any one of claims 11-14, wherein the first neuromorphic core includes a buffer, the buffer including a first entry for storing membrane potential increment values of the plurality of neural units for a first time step, and a second entry for storing membrane potential increment values of the plurality of neural units for a second entry.
19. The method according to any one of claims 11-14, further comprising controlling the current time step of the first neuromorphic core based on the number of allowed look-ahead states, wherein the number of the allowed look-ahead states is determined by the amount of available memory for storing pulses of the allowed look-ahead states.
20. A time step processing system, comprising components for performing the method according to any one of claims 11-19.
21. The system according to claim 20, wherein the components include machine-readable code that, when executed, causes the machine to perform one or more steps of the method according to any one of claims 11-19.
22. A computer-readable medium having instructions stored thereon that, when executed, cause a computing device to perform the method according to any one of claims 11 to 19.
23. A computer program product including instructions that, when executed by a processor, cause the processor to perform the method according to any one of claims 11 to 19.
Citation Information
Patent Citations
Neuromorphic event-driven neural computing architecture in a scalable neural network
US20130073497A1
Globally asynchronous and locally synchronous (GALS) neuromorphic network
US20150302295A1