spiking neural unit
By introducing spike neural units into the memory device and utilizing multiplexers and comparators, the problems of long processing time and low performance of existing memory devices are solved, achieving more efficient information processing and learning capabilities.
Patent Information
- Application Number
- CN202080060400.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-08-28
- Filing Date
- 2020-08-25
- Publication Date
- 2026-01-06
- Estimated Expiration
- 2040-08-25
AI Technical Summary
Existing memory devices suffer from long processing times and low performance in information processing and machine learning, especially in artificial neural networks where they struggle to process information efficiently.
The learning process is achieved by using a spike neural unit, which receives input, accumulates weights, compares them with threshold weights, and outputs the result. Multiplexers and comparators are combined to improve processing efficiency.
By designing spike neural units, processing time is reduced and the information processing performance of memory devices is improved, especially in machine learning tasks such as image recognition, sound recognition, and natural language processing.
Smart Images

Figure CN114365223B_ABST
Abstract
Description
Technical Field
[0001] This disclosure generally relates to memory, and more specifically, to devices and methods associated with spike neural units in memory. Background Technology
[0002] Memory devices are typically provided as internal semiconductor integrated circuits in computers or other electronic devices. Many different types of memory exist, including volatile and non-volatile memory. Volatile memory requires power to maintain its data and includes Random Access Memory (RAM), Dynamic Random Access Memory (DRAM), and Synchronous Dynamic Random Access Memory (SDRAM), among others. Non-volatile memory provides persistent data by retaining stored data when no power is supplied and includes NAND flash memory, NOR flash memory, Read-Only Memory (ROM), Electrically Erasable Programmable Memory (EEPROM), Erasable Programmable Memory (EPROM), and variable resistance memory, such as Phase-Change Random Access Memory (PCRAM) and 3D XPoint. TM Resistive random access memory (RRAM) and magnetoresistive random access memory (MRAM), etc.
[0003] Memory also serves as a volatile and non-volatile data storage device for a wide range of electronic applications, including but not limited to personal computers, portable memory sticks, digital cameras, cellular phones, portable music players (e.g., MP3 players), movie players, and other electronic devices. Memory cells can be arranged in an array, wherein the array is used in a memory device.
[0004] Artificial neural networks are networks that process information by modeling a network of neurons (such as neurons in the human brain) to process information (e.g., stimuli) sensed in a specific environment. Similar to the human brain, neural networks typically contain a multi-neuron topology (e.g., which may be referred to as an artificial neuron). Attached Figure Description
[0005] Figure 1 This is a block diagram of a device including a memory array and a complementary metal-oxide-semiconductor (CMOS) according to several embodiments of the present disclosure.
[0006] Figure 2 This is a block diagram of a complementary metal-oxide-semiconductor (CMOS) comprising a plurality of spike neural units according to several embodiments of the present disclosure.
[0007] Figure 3 This is a block diagram of a system including a controller and a neural network according to several embodiments of the present disclosure.
[0008] Figure 4Example flowcharts illustrating methods for spike neural units according to several embodiments of the present disclosure.
[0009] Figure 5 Examples of artificial neurons according to several embodiments of the present disclosure are described.
[0010] Figure 6 This is a block diagram of example logic blocks of a memory device according to several embodiments of the present disclosure.
[0011] Figure 7 Example neural networks according to several embodiments of the present disclosure are described.
[0012] Figure 8 An example machine describing a computer system may execute within said computer system a set of instructions for causing said machine to perform the various methods discussed herein. Detailed Implementation
[0013] This disclosure includes devices and methods associated with spike neural units in a memory. Embodiments include: a memory array; and a complementary metal-oxide-semiconductor (CMOS) coupled to and positioned beneath the memory array, wherein the CMOS includes spike neural units comprising logic configured to: receive input to increment weights stored in memory cells of the memory array; collect the weights from the memory cells of the memory array; accumulate the weights having an increment based on the input; compare the accumulated weights with a threshold weight; and provide an output in response to the accumulated weights being greater than the threshold weight. In some instances, the spike neural units in the CMOS coupled to the memory array can reduce processing time and improve device performance.
[0014] Spike neural units can be used in information processing applications, including machine learning (e.g., deep learning). For example, spike neural units can be used for image recognition, sound recognition, and / or natural language processing.
[0015] A spiked neural unit may include a multiplexer and a comparator. The multiplexer collects weights stored in a memory cell, and the comparator compares the weights from the memory cell with a threshold weight. In some instances, it can be determined that the spiked neural unit has been spiked and learning has occurred in response to a weight from the memory cell being greater than the threshold weight. For example, spiked neural units can be used for processing computer sensing, speech recognition (e.g., from a user), machine translation, and / or social network filtering.
[0016] As used herein, “several things” can refer to one or more of such things. For example, “several memory devices” can refer to one or more memory devices. “A plurality of things” means two or more. Additionally, as used herein, particularly with respect to reference numerals in the accompanying drawings, designations such as “K,” “L,” “M,” “N,” “P,” “Q,” “R,” and “S” indicate that several specific features thus specified may be included in several embodiments of this disclosure.
[0017] The figures in this document follow a numbering convention, where the first few digits correspond to the figure number and the remaining digits identify the elements or components in the figure. Similar elements or components between different figures can be identified by using similar digits. For example, reference digit 104 may refer to... Figure 1 The element "04" in the text, and similar elements can be referenced as Figure 2 204 in the diagram. In some examples, multiple similar but functionally and / or structurally distinguishable elements or components in the same or different diagrams may be referenced sequentially using the same element number (e.g., Figure 2 (208-1, 208-2, 208-3, and 208-M in the figures). As will be understood, elements shown in the various embodiments herein may be added, swapped, and / or eliminated to provide several additional embodiments of this disclosure. Furthermore, the scale and relative dimensions of the elements provided in the figures are intended to illustrate various embodiments of this disclosure and are not intended to be limiting.
[0018] Figure 1 This is a block diagram of a device in the form of a memory device 100 including a memory array 102 and a complementary metal-oxide-semiconductor (CMOS) 104, according to several embodiments of the present disclosure. As used herein, "device" may refer to, but is not limited to, various structures or combinations thereof. For example, the memory array 102 and the CMOS 104 may also be considered individually as "devices". Figure 1 As illustrated, CMOS 104 may be positioned below (e.g., beneath) memory array 102. Memory array 102 may be formed on CMOS 104. For example, the bottom surface of memory array 102 may contact and / or couple to the top surface of CMOS 104.
[0019] The memory array 102 may include one or more memory cells 106, and the CMOS 104 may include spike neural units 108. The spike neural unit 108 may include logic. For example, the logic may be a logic component comprising multiple logic blocks configured to perform operations. The operations may include: receiving input to increment weights stored in one or more memory cells 106 of the memory array 102; collecting weights from one or more memory cells 106 of the memory array 102; accumulating weights having increments based on the input; comparing the accumulated weights with a threshold weight; and providing an output in response to the accumulated weights being greater than the threshold weight.
[0020] In some instances, the spike neuron 108 may include a multiplexer 110 and / or a comparator 112. The multiplexer 110 may be coupled to the memory array 102 and may collect weights from one or more memory cells 106 of the memory array 102. Collecting weights may include sensing signals corresponding to weighted inputs of the artificial neuron from one or more memory cells 106. One or more sensing amplifiers may sample the sensed signals and transmit them from the memory array 102 to the spike neuron 108 contained in a CMOS 104 beneath the memory array 102. Although... Figure 1 Not shown in the figure, but in some instances, the memory array 102 and the spike neural unit 108 can transmit sensing signals and are coupled to each other via one or more communication lines.
[0021] Multiplexer 110 may collect weights from one or more memory cells 106 in response to a specific time period elapsed since a previous collection of weights with an increased amount and / or in response to a specific number of signals applied to one or more memory cells 106. In some instances, multiplexer 110 may collect weights in response to a command received by spike neural unit 108.
[0022] In response to receiving input to increase weights, accumulated weights with the amount of input increment can be calculated. For example, a summation function can be executed to accumulate weights with the amount of input increment.
[0023] Comparator 112 compares the accumulated weights with a threshold weight. A threshold function (e.g., a step function) can be used to compare the accumulated weights with the threshold weight. The threshold function determines whether the accumulated weight is higher or lower than the threshold weight. For example, if the accumulated weight is greater than or equal to the threshold weight, the threshold function can produce a logic high output (e.g., logic 1) on the output, and if the accumulated weight is lower than the threshold weight, it can produce a logic low output (e.g., logic 0) on the output. The threshold weight can be a weight sufficient to indicate that learning has occurred.
[0024] In several embodiments, the spike neuron 108 may store accumulated weights in the memory array 102. The spike neuron 108 may store accumulated weights in the memory array 102 in response to accumulated weights being less than a threshold weight. For example, one or more memory cells 106 of the memory array 102 may continue to add weights and store accumulated weights before the spike neuron 108 has been spiked and before learning has occurred.
[0025] One or more memory cells 106 may be refreshed (e.g., reinforced) to store accumulated weights. After storing the accumulated weights, one or more memory cells 106 may be refreshed to prevent the accumulated weights from changing (e.g., drifting) over time. One or more memory cells 106 may be refreshed in response to receiving input to increase the weights stored in one or more memory cells and / or in response to a refresh command. In some instances, one or more memory cells 106 may be refreshed in response to a specific time period elapsed since a previous refresh of one or more memory cells 106.
[0026] In several embodiments, one or more memory cells 106 may be erased. For example, accumulated weights stored in one or more memory cells 106 may be deleted. Accumulated weights stored in one or more memory cells 106 may be deleted in response to accumulated weights exceeding a threshold weight and / or in response to providing an output. For example, one or more memory cells 106 may be erased when the spiked neuron 108 has spiked and learning has occurred. However, embodiments are not limited thereto, because in at least one embodiment, one or more memory cells 106 may continue to store accumulated weights after the spiked neuron 108 has spiked.
[0027] Figure 2This is a block diagram of a complementary metal-oxide-semiconductor (CMOS) 204 comprising a plurality of spike neural units 208-1, 208-2, 208-3, ..., 208-M according to several embodiments of the present disclosure. The plurality of spike neural units 208-1, ..., 208-M are capable of transmitting data and are coupled to each other via communication lines 220-1, 220-2, 220-3, ..., 220-4. The communication lines 220-1, ..., 220-4 interconnect the spike neural units 208-1, ..., 208-M in rows and columns. The spike neural units 208-1, ..., 208-M may also include communication lines 220-5, 220-6, 220-7, 220-8, 220-9, 220-10, 220-11, and 220-12 extending outside the spike neural units 208-1, ..., 208-M and / or the CMOS 204. Communication lines 220-5, ..., 220-12 can couple the spike neural units 208-1, ..., 208-M to spike neural units on different CMOS sensors and / or to the controller. Communication lines 220-1, ..., 220-12 enable data communication between the spike neural units 208-1, ..., 208-M, between the spike neural units 208-1, ..., 208-M and spike neural units on different CMOS sensors, and / or between the spike neural units 208-1, ..., 208-M and the controller.
[0028] Each of the plurality of spike neural units 208-1, ..., 208-M may include logic. For example, the logic may be a logic component comprising a plurality of logic blocks configured to perform operations. The operations may include: receiving input via one of the plurality of spike neural units 208-1, ..., 208-M to increase a weight stored in one or more memory cells of a memory array; collecting weights via one of the plurality of spike neural units 208-1, ..., 208-M; accumulating weights having an increment based on the input; comparing the accumulated weights with a threshold weight; and providing an output in response to the accumulated weights being greater than the threshold weights. The accumulated weights may be stored back into one or more memory cells of the memory array in response to the accumulated weights being less than the threshold weights.
[0029] Each of the plurality of spike neural units 208-1, ..., 208-M can be configured to receive input to increase the weight in a plurality of memory cells stored in a memory array. The input can be transmitted from the controller via communication lines 220-1, ..., 220-4 through one or more of the plurality of spike neural units 208-1, ..., 208-M. The input can be an electrical signal, for example, a voltage applied to one or more of the plurality of spike neural units 208-1, ..., 208-M.
[0030] Weights can be collected by one of a plurality of spike neural units 208-1, ..., 208-M. Weights can be collected in response to a specific time period elapsed since a previous weight collection, in response to a specific number of signals applied to one or more memory units, and / or in response to a command received by one of the plurality of spike neural units 208-1, ..., 208-M. In some instances, a multiplexer can be used to collect weights.
[0031] Each of the multiple peak neural units 208-1, ..., 208-M can accumulate a corresponding weight with a corresponding increment based on the input, and compare the accumulated weight with a threshold weight. The threshold weight can be a weight sufficient to cause learning to occur. The comparison between the accumulated weight and the threshold weight can be in response to the accumulation of weight and / or in response to one of the multiple peak neural units 208-1, ..., 208-M receiving a command. A specific time period can be set based on the average time taken for the accumulated weights of the multiple peak neural units 208-1, ..., 208-M to reach the threshold weight. In some instances, a comparator can be used to compare the accumulated weight with the threshold weight.
[0032] The output (e.g., a comparison result of the accumulated weights and a threshold weight) can be provided to the controller from the first peak neural unit 208-1 via the second peak neural unit 208-2 and / or directly from the first peak neural unit 208-1 to the controller. In some instances, the output can be provided in response to the accumulated weights being greater than the threshold weight. For example, the peak neural unit 208-1 can send the output to the controller to notify the controller that learning has occurred. In some instances, the controller can be external to the memory array and CMOS 204.
[0033] In several embodiments, multiple spike neural units 208-1, ..., 208-M may be included in the neural network. The neural network can execute various machine learning algorithms to process the input. Instance tasks that can be processed by the neural network may include computer vision, speech recognition, machine translation, social network filtering, and / or medical diagnosis.
[0034] A neural network may contain multiple layers, each of which contains one or more of a plurality of spike neural units 208-1, ..., 208-M, such as in combination. Figure 7Further description. The output from one of multiple peak neural units on a layer, such as the first peak neural unit 208-1, can be received by different peak neural units on different layers, such as the second peak neural unit 208-2. One or more memory cells can be refreshed. A refresh can include read and write operations on one or more memory cells. For example, a refresh can rewrite the weights stored in one or more memory cells to preserve weight data. A refresh can be performed on one or more memory cells in response to the accumulated weights being less than a threshold weight. As previously discussed, the threshold weight can be a weight sufficient to allow learning to occur.
[0035] One or more memory cells can be erased. Erasure removes data from one or more memory cells. Weights stored in one or more memory cells can be effectively erased in response to accumulated weights reaching a threshold weight. One or more memory cells can be effectively erased when reading data rather than refreshing it. In some instances using memory cells other than DRAM, one or more memory cells can be erased via an active erase mechanism.
[0036] Figure 3 This is a block diagram of a system 330 comprising a controller 332 and a neural network 334, according to several embodiments of the present disclosure. For example, system 330 may be a server system and / or a high-performance computing (HPC) system and / or a portion thereof. Controller 332 may include a state machine, a sequencer, and / or some other type of control circuitry system, and includes hardware and / or firmware (e.g., microcode instructions) in the form of an application-specific integrated circuit (ASIC), a field-programmable gate array, etc. Controller 332 may be located locally in each of a plurality of memory devices 300-1, 300-2, 300-3, ..., 300-N. In other words, although... Figure 3 The description specifies a controller 332, but the memory system 330 may include multiple controllers, each located locally in a respective memory device 300-1, ..., 300-N. In some instances, the neural network 334 may be dynamic random access memory (DRAM). The neural network 334 may include multiple memory devices 300-1, ..., 300-N.
[0037] In various embodiments, the plurality of memory devices 300-1, ..., 300-N may be three-dimensional (3D) and may include multiple layers stacked together. As an example, each of the plurality of memory devices 300-1, ..., 300-N may include: a first layer containing logic components (e.g., logic blocks, row drivers, and / or column drivers); and a second layer stacked on the first layer and containing memory components such as arrays or memory cells. The plurality of memory devices 300-1, ..., 300-N may include rows of memory cells arranged to be coupled via access lines (which may be referred to herein as word lines or select lines) and columns of memory cells coupled via sensing lines (which may be referred to herein as digital lines or data lines). For example, the memory cell array may include, but is not limited to, DRAM arrays, SRAM arrays, STT RAM arrays, PCRAM arrays, TRAM arrays, RRAM arrays, NAND flash memory arrays, and / or NOR flash memory arrays. Multiple memory devices 300-1, ..., 300-N may be in the form of multiple individual memory dies and / or dissimilar memory layers formed as integrated circuits on a chip. In some instances, each of the memory devices 300-1, ..., 300-N may include a memory array 302-1, ..., 302-P coupled to complementary metal-oxide-semiconductor (CMOS) 304-1, 304-2, 304-3, ..., 304-Q. Each CMOS 304-1, ..., 304-Q may include multiple spike neural units, such as... Figure 2 As explained in the text.
[0038] Memory devices 300-1, ..., 300-N can transmit data to and from controller 332 via communication lines 336-1, ..., 336-6. Controller 332 can send commands (e.g., inputs) to one or more of the memory devices 300-1, ..., 300-N via communication lines 336-1, ..., 336-6 and one or more of the memory devices 300-1, ..., 300-N. The commands can be electrical signals sent from controller 332 to one or more of the memory devices 300-1, ..., 300-N. In some instances, controller 332 can send commands to increase the weight of multiple memory cells stored in memory array 302-P of memory device 300-N. The commands can be sent from controller 332 to memory device 300-N via communication lines 336-3, memory device 300-3, and communication line 336-4.
[0039] In several embodiments, multiple memory devices 300-1, ..., 300-N may receive multiple different commands from controller 332. For example, the multiple memory devices 300-1, ..., 300-N may receive commands to collect weights from multiple memory cells of the memory array, accumulate weights with an increment based on the input, compare the accumulated weights with a threshold weight, and / or provide an output to controller 332 in response to the accumulated weights being greater than the threshold weight.
[0040] In some instances, one or more of the memory devices 300-1, ..., 300-N may provide output (e.g., electrical signals) to the controller 332 via communication lines 336-1, ..., 336-6 and one or more of the memory devices 300-1, ..., 300-N. For example, memory device 300-2 may send output to the controller 332 via communication line 336-2, memory device 300-1, and communication line 336-1. In response to receiving a command from the controller 332 to provide output, output may be provided by one or more of the multiple memory devices 300-1, ..., 300-N.
[0041] The controller 332 can also send refresh and / or erase commands to a plurality of memory devices 300-1, ..., 300-N. A refresh command can instruct one or more of the plurality of memory devices 300-1, ..., 300-N to refresh one or more memory cells. A refresh can include read and write operations on one or more memory cells. For example, a refresh can rewrite weights stored in one or more memory cells to preserve weight data. The controller 332 can send a refresh command in response to a weight being less than and / or equal to a threshold weight. (As in...) Figure 2 The threshold weights discussed can be weights sufficient to enable learning to occur.
[0042] An erase command can instruct one or more of a plurality of memory devices 300-1, ..., 300-N to erase one or more memory cells. Erasure can remove data from one or more memory cells. For example, erasure can dissipate weights stored in one or more memory cells. Weights stored in one or more memory cells can be removed in response to said weights being greater than and / or equal to a threshold weight.
[0043] although Figure 3 Not shown, but the memory system 330 may also include a decoder (e.g., a row / column decoder) that can be controlled by the controller 332 to decode, for example, address signals received from the host. The decoded address signals may be further provided to a row / column driver via the controller 332, which may activate rows / columns of an array of memory cells of a plurality of memory devices 300-1, ..., 300-N.
[0044] Figure 4This describes an example method 440 for a spike neural unit according to several embodiments of the present disclosure. Method 440 may, for example, be composed of... Figure 1 and 3 The described memory devices 100 and / or 300 perform this action.
[0045] At block 442, method 440 includes receiving input via a spike neural unit including logic to increase weights stored in memory cells of a memory array. The input may be an electrical signal. For example, the input may be a voltage applied to one or more memory cells coupled to the array of spike neural units. The input may be sent from different spike neural units and / or a controller. One or more data lines between the spike neural units and the different spike neural units and / or the controller may transmit the input.
[0046] At block 444, method 440 includes collecting weights via spike neural units. The weights may be collected via a multiplexer. Weights may be collected in response to a specific time period elapsed since a previous weight collection, in response to a specific number of signals applied to one or more memory units, and / or in response to a command received by one of the multiple spike neural units.
[0047] At box 446, method 440 includes accumulating weights with an increment based on the input. The weights can be accumulated by performing a summation calculation. For example, a spike neuron can perform a summation calculation to determine weights with an increment based on the input.
[0048] At box 448, method 440 includes comparing the accumulated weights with a threshold weight. The comparison of the accumulated weights with the threshold weights can be performed by a spike neural unit. The threshold weights can be weights sufficient to allow learning to occur. The comparison of the accumulated weights with the threshold weights can be performed in response to a specific time period elapsed since a previous comparison, in response to the accumulated weights, and / or in response to a command received by the spike neural unit. In some instances, a comparator can be used to compare the accumulated weights with the threshold weights.
[0049] At block 450, method 440 includes providing an output to a controller in response to an accumulated weight being greater than a threshold weight. The output (e.g., the result of a comparison between the accumulated weight and the threshold weight) may be provided from a peak neural unit to different peak neural units. For example, the peak neural unit may send the output to the different peak neural units and / or via the different peak neural units to the controller. In some instances, the output may be provided in response to an accumulated weight being greater than a threshold weight. For example, the peak neural unit may send the output to the controller to notify the controller that learning has occurred.
[0050] Figure 5Examples of artificial neurons 552 according to several embodiments of the present disclosure are described below. Artificial neurons 552 can be used to simulate (e.g., biological neurons of the human brain). Such neurons are sometimes referred to as sensory receptors. Several inputs x1 to xN, which may be referred to as stimuli, may be applied to inputs 554-1, 554-2, ..., 554-R of neuron 552, respectively. Signals corresponding to inputs x1 to xN, such as voltage, current, or specific data values (e.g., binary digits), may be generated in response to sensing some form of stimulus and may be applied to inputs 554-1, 554-2, ..., 554-R.
[0051] In various examples, inputs x1 to xN can be weighted by weights w1 to wN, which can be referred to as synaptic weights. For example, inputs x1 to xN can be multiplied by weights w1 to wN to weight inputs x1 to xN respectively. For example, each weighted input can be referred to as a synapse, and the weights can correspond to memories in human brain behavior.
[0052] Neuron 552 may include a summation function 556 that performs addition on weighted inputs to produce an output 558, such as SUM = x1w1 + x2w2 + ... + xNwN. For example, in neural network theory, "SUM" may be referred to as "NET" (e.g., from the term "NETwork"). For example, weighted signals corresponding to weighted inputs x1w1 to xNwN may be summed. In some instances, the summation function may be referred to as a transfer function. Neuron 552 further includes a function 560 configured to respond to the summation SUM and produce an output value Y at output 562, such as a function... In some instances, function 560 can be referred to as the activation function. The output of a neuron can sometimes be referred to as a class.
[0053] Various functions can be used in function 560. For example, function 560 may include a threshold function (e.g., a step function) to determine whether SUM is above or below a specific threshold level. If SUM is greater than or equal to the specific threshold amount, this threshold function may produce a logic high output (e.g., logic 1) on output 562, and if SUM is below the specific threshold amount, it may produce a logic low output (e.g., logic 0) on output 562.
[0054] In some instances, function 560 can be a sigma function, where a sigma function can be expressed as S(z) = 1 / (1+e^(z-1)). λz ), where λ is a constant and z can be SUM. For example, function 560 can be a nonlinear function. In some instances, the output value Y at output 562 can be applied to a neural network of neurons (e.g., Figure 3The neural network 334 in the model has several additional neurons, such as the inputs 554 of different neurons. Function 560 may further include a sign function and / or a linear function, etc.
[0055] Figure 6 This is a block diagram of an example logic block 664 of a memory device according to several embodiments of the present disclosure. Logic block 664 may be included in the memory device, for example, as previously described in conjunction with... Figure 1 and 3 One of the plurality of logic blocks within the described memory device 100 and / or 300.
[0056] Logic blocks can be configurable logic blocks (CLBs), which are the basic building blocks of field-programmable gate arrays (FPGAs). An FPGA is a chip capable of changing its data path and / or being reprogrammed in the field. This capability allows FPGAs to flexibly switch between, for example, a central processing unit (CPU) and a graphics processing unit (GPU). As an example, an FPGA already used as a microprocessor can be reprogrammed in the field to be used as a graphics card and / or encryption unit.
[0057] like Figure 6 As described, logic block 664 contains logic 670. Logic 670 may be a LUT-based logic. As an example, the physical location (e.g., address) of logic 670 may be mapped to a logical address, and the mapping information may be stored in a lookup table.
[0058] Logic block 664 may further include elements that can be enabled to activate as previously combined. Figure 1 and 3 The described memory arrays 102 and / or 302 include row drivers 666 and column drivers 668 for rows (or rows) and / or columns (or columns). As described herein, row drivers 666 and column drivers 668 can receive address signals decoded by corresponding row decoders and column decoders, which can be controlled by a controller, such as those previously combined with... Figure 3 The controller 332 described is used for control. Although... Figure 6 Not shown, but logic block 664 may also include (e.g., coupled to) multiple data buses that couple logic block 664 to another logic block and / or another external device (e.g., a device located outside memory device 300). The data buses of logic block 664 that couple logic block 664 to another logic block may include interconnecting optical fibers.
[0059] Figure 7 This describes an example neural network 734 according to several embodiments of the present disclosure. The neural network 734 may include nodes 774-1 to 774-S (including those previously combined...) Figure 1 and Figure 2The neural network layer 772 (e.g., input layer) of the described spike neural units 108 and / or 208, wherein the nodes receive various inputs, such as those previously combined Figure 5 The inputs x1 to xN are described. The nodes of each neural network layer (e.g., neural network layers 772, 774-1, 774-2, 774-3, and 776) may correspond to artificial neurons as described herein.
[0060] Neural network 734 may include neural network layers 774-1, 774-2, and 774-3. Neural network layer 774-1 may include nodes 778-1 to 778-L. As illustrated in interconnection region 780-1, each of the corresponding nodes 778-1 to 778-L may be coupled to receive input from nodes 774-1 to 774-L. Neural network layer 774-2 may include nodes 782-1 to 782-L. As illustrated in interconnection region 780-2, each of the corresponding nodes 782-1 to 782-L may be coupled to receive input from nodes 778-1 to 778-L. Neural network layer 774-3 may include nodes 784-1 to 784-L. As illustrated in interconnection region 780-3, each of the corresponding nodes 784-1 to 784-L may be coupled to receive input from nodes 782-1 to 782-L. The neural network 734 can be configured during training, wherein various connections in the interconnect region 780 are assigned weight values or updated with new weight values for operations or computations at nodes 778, 782, or 784. The training process may vary depending on the specific application or use of the neural network 734. For example, the neural network can be trained for image recognition, speech recognition, or any number of other processing or computational tasks.
[0061] The neural network 734 may include an output layer 776 having output nodes 786-1 to 786-K. Each of the corresponding output nodes 786-1 to 786-K may be coupled to receive input from nodes 784-1 to 784-L. The process of receiving useful outputs at the output layer 776 and output nodes 786 as a result of inputs fed to nodes 774 at the neural network layer 772 may be referred to as inference or forward propagation. That is, an input signal representing a real-world phenomenon or application can be fed into the trained neural network 734 and, through inference that occurs as a result of computations implemented by various nodes and interconnections, a result can be output. In the case of the neural network 734 trained for speech recognition, the input may be a signal representing human speech in one language, and the output may be a signal representing human speech in a different language. Or, for the neural network 734 trained for image recognition, the input may be a signal representing a photograph, and the output may be a signal representing the subject in the photograph.
[0062] As described herein, multiple neural networks can be configured within a memory device. Multiple neural networks can be trained individually (locally or remotely) and the trained neural networks can be used for inference within the memory device. Multiple neural networks can perform the same or different functions and can have the same or different weights relative to each other.
[0063] Figure 8 An example machine illustrating computer system 830 is described, within which a set of instructions can be executed to cause the machine to perform the various methods discussed herein. In various embodiments, computer system 830 (e.g., Figure 3 The computer system 330 in the middle may be coupled to or utilize the memory subsystem or may be used to execute a controller (e.g., Figure 1 The machine operates as controller 332. In an alternative embodiment, the machine may be connected (e.g., networked) to other machines in a LAN, intranet, extranet, and / or the Internet. The machine may operate as a server or client machine in a client-server network environment, as a peer machine in a peer-to-peer (or distributed) network environment, or as a server or client machine in a cloud computing infrastructure or environment.
[0064] The machine may be a personal computer (PC), tablet PC, set-top box (STB), personal digital assistant (PDA), cellular phone, network infrastructure, server, network router, switch, or bridge, or any machine capable of executing a set of instructions (sequential or otherwise) specifying the action to be taken by that machine. Furthermore, while a single machine is described, the term "machine" should also be understood to include any collection of machines that individually or jointly execute a set (or more sets) of instructions to perform any or more of the methods discussed herein.
[0065] The example computer system 830 includes a processing device 888 that communicates with each other via a bus 895, a main memory 890 (e.g., read-only memory (ROM), flash memory, dynamic random access memory (DRAM), such as synchronous DRAM (SDRAM) or Rambus DRAM (RDRAM), etc.), a static memory 892 (e.g., flash memory, static random access memory (SRAM), etc.), and a data storage system 894.
[0066] Processing device 888 represents one or more general-purpose processing devices, such as microprocessors, central processing units (CPUs), etc. More specifically, the processing device may be a Complex Instruction Set Computing (CISC) microprocessor, a Reduced Instruction Set Computing (RISC) microprocessor, a Very Long Instruction Word (VLIW) microprocessor, or a processor implementing other instruction sets, or a processor implementing combinations of instruction sets. Processing device 888 may also be one or more special-purpose processing devices, such as application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), network processors, etc. Processing device 888 is configured to execute instructions 896 to perform the operations and steps discussed herein. Computer system 830 may further include a network interface device 898 for communication via network 897.
[0067] The data storage system 894 may include a machine-readable storage medium 899 (also referred to as a computer-readable medium) thereon storing one or more sets of instructions 896 or software embodying any one or more of the methods or functions described herein. The instructions 896 may also reside wholly or at least partially in main memory 890 and / or processing device 888 during execution by computer system 830, which also constitutes a machine-readable storage medium.
[0068] In one embodiment, instruction 896 includes instructions for implementing a memory device (e.g., Figure 1 The machine-readable storage medium 899 is a single medium in the exemplary embodiment, but the term "machine-readable storage medium" should be understood as a single medium or multiple media containing one or more sets of instructions. The term "machine-readable storage medium" should also be understood as any medium capable of storing or encoding a set of instructions for machine execution and causing the machine to perform any one or more of the methods of this disclosure. Therefore, the term "machine-readable storage medium" should be understood to include, but is not limited to, solid-state memory, optical media, and magnetic media.
[0069] Although specific embodiments have been illustrated and described herein, those skilled in the art will understand that arrangements calculated to achieve the same results may replace the specific embodiments shown. This disclosure is intended to cover adaptations or variations of the various embodiments of this disclosure. It should be understood that the above description is illustrative rather than restrictive. After reviewing the above description, combinations of the above embodiments and other embodiments not specifically described herein will be apparent to those skilled in the art. The scope of the various embodiments of this disclosure includes other applications in which the above structures and methods are used. Therefore, the scope of the various embodiments of this disclosure should be determined by reference to the appended claims and the full scope of their equivalents.
[0070] In the foregoing detailed embodiments, various features are grouped in a single embodiment for the purpose of simplifying this disclosure. This disclosure method should not be construed as reflecting an intention that the disclosed embodiments of this disclosure must use more features than are expressly stated in each claim. Rather, as reflected in the appended claims, the subject matter of the invention lies in fewer than all features of a single disclosed embodiment. Therefore, the appended claims are thus incorporated into the detailed embodiments, wherein each claim is considered an independent, separate embodiment.
Claims
1. An apparatus for a spiking neural unit, comprising: a controller; a first memory array; a first complementary metal-oxide-semiconductor (CMOS) coupled to the first memory array (and positioned below the first memory array; a second memory array; and a second CMOS coupled to the second memory array and positioned below the second memory array, wherein the second CMOS is configured to: receive an input to increase a weight stored in a memory cell of the second memory array; collect the weight from the memory cell of the second memory array in response to a particular time period elapsing since a last collection of the weight, wherein the particular time period is set according to an average time required for the weight to reach a threshold weight; accumulate the weight with an amount of increase based on the input; compare the accumulated weight to the threshold weight; and provide an output to the controller in response to the accumulated weight being greater than the threshold weight, wherein the controller is configured to send a flush command to the second CMOS in response to the accumulated weight being less than or equal to the threshold weight, and wherein the accumulated weight is overwritten in response to the flush command, wherein overwriting the accumulated weight prevents the accumulated weight from changing over time.
2. The apparatus of claim 1, wherein the first memory array is coupled to the second memory array.
3. The apparatus of claim 1, wherein the controller is external to the second memory array and the second CMOS.
4. The apparatus of claim 3, wherein the second CMOS is further configured to provide the output to the controller via the first memory array.
5. The apparatus of claim 1, wherein the second memory array is formed on the second CMOS; and wherein the second CMOS is further configured to collect the weight from a plurality of memory cells of the second memory array via a multiplexer.
6. The apparatus of claim 1, wherein the second CMOS is further configured to store the accumulated weight in the second memory array in response to the accumulated weight being less than the threshold weight.
7. A method for a spiking neural unit (108, 208), comprising: receiving an input via a first complementary metal-oxide-semiconductor (CMOS) positioned below a first memory array to increase a weight stored in a memory cell of the first memory array; collecting the weight via the first CMOS in response to a particular time period elapsing since a last collection of the weight, wherein the particular time period is set according to an average time required for the weight to reach a threshold weight; accumulating the weight with an amount of increase based on the input; comparing the accumulated weight to the threshold weight; providing an output to a controller coupled with the first memory array via a second memory array positioned above a second CMOS in response to the accumulated weight being greater than the threshold weight; and sending a refresh command to the second CMOS in response to the accumulated weight being less than or equal to the threshold weight, wherein the accumulated weight is overwritten in response to the refresh command, wherein overwriting the accumulated weight prevents the accumulated weight from changing over time.
8. The method of claim 7, further comprising refreshing the weights stored in the memory cells in response to the first CMOS receiving a refresh command from the controller.
9. The method of claim 7, further comprising erasing the weights stored in the memory cells in response to the first CMOS receiving an erase command from the controller.
10. The method of claim 7, further comprising sending a result of the comparison to the controller via the second memory array.
11. The method of claim 7, further comprising collecting the weights via a multiplexer of the first CMOS in response to the first CMOS receiving a command from the controller.
12. The method of claim, further comprising collecting the weights via a multiplexer of the first CMOS in response to a particular number of signals being applied to the memory cells.
13. A system for a spiking neural unit, comprising: a controller; and a neural network coupled to the controller, wherein the neural network includes: a first memory array; and a first complementary metal-oxide-semiconductor (CMOS) coupled to the first memory array and positioned below the first memory array; a second memory array; and a second CMOS coupled to the second memory array and positioned below the second memory array, wherein the second CMOS is configured to: receive an input to increase weights stored in a plurality of memory cells of the second memory array; collect the weights from the plurality of memory cells of the second memory array in response to a particular time period elapsing since a last collection of the weights, wherein the particular time period is set according to an average time required for the weights to reach a threshold weight; accumulate the weights with an amount of increase based on the input; compare the accumulated weights to the threshold weight; and provide an output to the controller in response to the accumulated weights being greater than the threshold weight, wherein the controller is configured to send a refresh command to the second CMOS in response to the accumulated weights being less than or equal to the threshold weight, and wherein the accumulated weights are overwritten in response to the refresh command, wherein overwriting the accumulated weights prevents the accumulated weights from changing over time.
14. The system of claim 13, wherein the controller is configured to send an erase command to the second CMOS in response to the weight from the plurality of memory cells being greater than the threshold weight.
15. The system of claim 13, wherein the neural network is a dynamic random access memory (DRAM).
16. The system of claim 13, wherein the plurality of memory cells is a word of memory cells.
Citation Information
Patent Citations
Neuromorphic computing device, memory device, system, and method to maintain a spike history for neurons in a neuromorphic computing environment
US20180082176A1
Spiking neural network
US20180260696A1
Spiking neural network accelerator using external memory
US20190042920A1
Memristive nanofiber neural networks
US20190156190A1