System for converting neuronal electrical currents in deep learning artificial neural networks
By converting the neuron current output from the VMM array into time pulses and performing appropriate voltage or current conversion, the accuracy and efficiency of measuring and transmitting the output current of the VMM array is solved, and the energy efficiency and computational parallelism of the system are improved.
Patent Information
- Application Number
- CN201980089071.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-03-14
- Filing Date
- 2019-09-06
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2039-09-06
AI Technical Summary
The prior art is difficult to accurately measure and transmit the output current of a vector-matrix multiplication (VMM) array, resulting in information loss and inefficiency.
The neuron current output from the VMM array is converted into time pulses and provided as input to another VMM array in the artificial neural network, and the current is converted into a suitable voltage or current form for transmission through circuits such as an integrated analog-to-digital converter and current-voltage converter.
It realizes efficient and accurate measurement and transmission of VMM array output, reducing information loss, improving the energy efficiency and computing parallelism of the system.
Smart Images

Figure CN113302629B_ABST
Abstract
Description
[0001] Priority Declaration
[0002] This application claims priority to U.S. Provisional Application No. 62 / 814,813, filed on March 6, 2019, entitled “System for Converting Neuron Current Into Neuron Current-Based Time Pulses in an Analog Neural Memory in a Deep Learning Artificial Neural Network,” U.S. Provisional Application No. 62 / 794,492, filed on January 18, 2019, entitled “System for Converting Neuron Current Into Neuron Current-Based Time Pulses in an Analog Neural Memory in a Deep Learning Artificial Neural Network,” and U.S. Patent Application No. 16 / 353,830, filed on March 14, 2019, entitled “System for Converting Neuron Current Into Neuron Current-Based Time Pulses in an Analog Neural Memory in a Deep Learning Artiective Neural Network.” Technical Field
[0003] The present invention discloses various embodiments for converting neuronal currents output by a vector-matrix multiplication (VMM) array into temporal pulses based on the neuronal currents and providing such pulses as input to another VMM array within an artificial neural network. Background Art
[0004] Artificial neural networks simulate biological neural networks (the central nervous system of animals, especially the brain) and are used to estimate or approximate functions that may depend on a large number of inputs and are generally unknown. Artificial neural networks typically consist of layers of interconnected "neurons" that exchange messages with each other.
[0005] Figure 1An artificial neural network is shown, where circles represent inputs or layers of neurons. Connections (called synapses) are represented by arrows and have numerical weights that can be adjusted based on experience. This allows the neural network to adapt to the input and learn. Typically, a neural network includes multiple layers of inputs. There are typically one or more intermediate layers of neurons, and an output layer of neurons that provide the output of the neural network. Neurons at each level make decisions based on the data received from the synapses, either individually or collectively.
[0006] One of the main challenges in developing artificial neural networks for high-performance information processing is the lack of adequate hardware technology. In fact, practical neural networks rely on a large number of synapses to achieve high connectivity between neurons, that is, very high computational parallelism. In principle, such complexity can be achieved using digital supercomputers or clusters of dedicated graphics processing units. However, in addition to being high-cost, these approaches are also mediocre in energy efficiency compared to biological networks, which consume less energy mainly due to the low-precision analog calculations they perform. CMOS analog circuits have been used in artificial neural networks, but due to the large number of neurons and synapses required, the synapses of most CMOS implementations are too large.
[0007] Applicant previously disclosed an artificial (simulated) neural network utilizing one or more nonvolatile memory arrays as synapses in U.S. patent application Ser. No. 15 / 594,439 (published as U.S. Patent Publication No. 2017 / 0337466), which is incorporated herein by reference. The nonvolatile memory array operates as a simulated neuromorphic memory. A neural network device includes a first plurality of synapses configured to receive a first plurality of inputs and generate a first plurality of outputs therefrom, and a first plurality of neurons configured to receive the first plurality of outputs. The first plurality of synapses includes a plurality of memory cells, wherein each of the memory cells includes: spaced-apart source and drain regions formed in a semiconductor substrate, wherein a channel region extends between the source and drain regions; a floating gate disposed over and insulated from a first portion of the channel region; and a non-floating gate disposed over and insulated from a second portion of the channel region. Each of the plurality of memory cells is configured to store a weight value corresponding to a plurality of electrons on the floating gate. The plurality of memory cells is configured to multiply the first plurality of inputs by the stored weight values to generate a first plurality of outputs.
[0008] Each nonvolatile memory cell used in an analog neuromorphic memory system must be erased and programmed to maintain a very specific and precise amount of charge (i.e., number of electrons) in the floating gate. For example, each floating gate must maintain one of N different values, where N is the number of different weights that can be represented by each cell. Examples of N include 16, 32, 64, 128, and 256.
[0009] One challenge facing systems using VMM arrays is being able to accurately measure the output of a VMM array and transmit that output to another stage (such as an input block of another VMM array). Many methods are known, but each has certain drawbacks, such as current leakage leading to information loss.
[0010] Therefore, there is a need for an improved system for measuring the output current of a VMM array and converting the output current into a form more suitable for transmission to another stage of electronic devices. Summary of the Invention
[0011] The present invention discloses various embodiments for converting neuronal currents output by a vector-matrix multiplication (VMM) array into temporal pulses based on the neuronal currents and providing such pulses as input to another VMM array within an artificial neural network. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 FIG. 1 is a schematic diagram showing an artificial neural network of the prior art.
[0013] Figure 2 A prior art split-gate flash memory cell is shown.
[0014] Figure 3 Another prior art split-gate flash memory cell is shown.
[0015] Figure 4 Another prior art split-gate flash memory cell is shown.
[0016] Figure 5 Another prior art split-gate flash memory cell is shown.
[0017] Figure 6 Another prior art split-gate flash memory cell is shown.
[0018] Figure 7 A prior art stacked gate flash memory cell is shown.
[0019] Figure 8 A schematic diagram illustrating different levels of an exemplary artificial neural network using one or more non-volatile memory arrays.
[0020] Figure 9 is a block diagram showing a vector-matrix multiplication system.
[0021] Figure 10 is a block diagram illustrating an exemplary artificial neural network using one or more vector-matrix multiplication systems.
[0022] Figure 11Another embodiment of a vector-matrix multiplication system is shown.
[0023] Figure 12 Another embodiment of a vector-matrix multiplication system is shown.
[0024] Figure 13 Another embodiment of a vector-matrix multiplication system is shown.
[0025] Figure 14 Another embodiment of a vector-matrix multiplication system is shown.
[0026] Figure 15 Another embodiment of a vector-matrix multiplication system is shown.
[0027] Figure 16 A prior art long short-term memory system is shown.
[0028] Figure 17 An exemplary cell used in a long short-term memory system is shown.
[0029] Figure 18 Show Figure 17 An embodiment of an exemplary unit of .
[0030] Figure 19 Show Figure 17 Another embodiment of an exemplary unit of .
[0031] Figure 20 A prior art gate-controlled recursive cell system is shown.
[0032] Figure 21 An exemplary cell for use in a gate-controlled recursive cell system is shown.
[0033] Figure 22 Show Figure 21 An embodiment of an exemplary unit of .
[0034] Figure 23 Show Figure 21 Another embodiment of an exemplary unit of .
[0035] Figure 24 Another embodiment of a vector-matrix multiplication system is shown.
[0036] Figure 25 Another embodiment of a vector-matrix multiplication system is shown.
[0037] Figure 26 Another embodiment of a vector-matrix multiplication system is shown.
[0038] Figure 27 Another embodiment of a vector-matrix multiplication system is shown.
[0039] Figure 28 Another embodiment of a vector-matrix multiplication system is shown.
[0040] Figure 29 Another embodiment of a vector-matrix multiplication system is shown.
[0041] Figure 30 Another embodiment of a vector-matrix multiplication system is shown.
[0042] Figure 31 Another embodiment of a vector-matrix multiplication system is shown.
[0043] Figure 32 A VMM system is shown.
[0044] Figure 33 Shown is a flash memory emulated neural memory system.
[0045] Figure 34A An integrating analog-to-digital converter is shown.
[0046] Figure 34B Show Figure 34A Voltage characteristics of the integrating analog-to-digital converter.
[0047] Figure 35A An integrating analog-to-digital converter is shown.
[0048] Figure 35B Show Figure 35A Voltage characteristics of the integrating analog-to-digital converter.
[0049] Figure 36A and 36B Show Figure 34A and 35A Example waveforms of the operation of the analog-to-digital converter.
[0050] Figure 36C Shows the timing control circuit.
[0051] Figure 37 A pulse-to-voltage converter is shown.
[0052] Figure 38 A current-to-voltage converter is shown.
[0053] Figure 39 A current-to-voltage converter is shown.
[0054] Figure 40 A current-to-logarithmic voltage converter is shown.
[0055] Figure 41 A current-to-logarithmic voltage converter is shown.
[0056] Figure 42A digital data-to-voltage converter is shown.
[0057] Figure 43 A digital data-to-voltage converter is shown.
[0058] Figure 44 A reference array is shown.
[0059] Figure 45 A digital comparator is shown.
[0060] Figure 46 The converter and digital comparator are shown.
[0061] Figure 47 shows an analog comparator.
[0062] Figure 48 The converter and analog comparator are shown.
[0063] Figure 49 The output circuit is shown.
[0064] Figure 50 Shows one aspect of the output that is activated after digitization.
[0065] Figure 51 Shows one aspect of the output that is activated after digitization.
[0066] Figure 52 A charge summer circuit is shown.
[0067] Figure 53 A current summator circuit is shown.
[0068] Figure 54 A digital summer circuit is shown.
[0069] Figure 55A and 55B The digital bit-to-pulse line converter and waveform are shown respectively.
[0070] Figure 56 Shows the power management method.
[0071] Figure 57 Another power management method is shown.
[0072] Figure 58 Another power management method is shown. DETAILED DESCRIPTION
[0073] The artificial neural network of the present invention utilizes a combination of CMOS technology and non-volatile memory arrays.
[0074] Non-volatile memory cells
[0075] Digital non-volatile memory is well known. For example, U.S. Patent No. 5,029,130 (“the '130 patent”), which is incorporated herein by reference, discloses an array of split-gate non-volatile memory cells, which is a type of flash memory cell. Such a memory cell 210 is Figure 2 . Each memory cell 210 includes a source region 14 and a drain region 16 formed in a semiconductor substrate 12, with a channel region 18 therebetween. A floating gate 20 is formed over and insulated from a first portion of the channel region 18 (and controls its electrical conductivity), and is formed over a portion of the source region 14. A wordline terminal 22 (which is typically coupled to a wordline) has a first portion disposed over and insulated from a second portion of the channel region 18 (and controls its electrical conductivity), and a second portion extending upward and over the floating gate 20. The floating gate 20 and wordline terminal 22 are insulated from the substrate 12 by a gate oxide. A bitline 24 is coupled to the drain region 16.
[0076] Memory cell 210 is erased (where electrons are removed from the floating gate) by placing a high positive voltage on wordline terminal 22, which causes the electrons on floating gate 20 to tunnel through the intervening insulator via Fowler-Nordheim tunneling from floating gate 20 to wordline terminal 22.
[0077] Memory cell 210 is programmed by placing a positive voltage on wordline terminal 22 and a positive voltage on source region 14 (where electrons are placed on the floating gate). Electron current will flow from source region 14 to drain region 16. When the electrons reach the gap between wordline terminal 22 and floating gate 20, they will accelerate and become heated. Due to the electrostatic attraction from floating gate 20, some of the heated electrons will be injected through the gate oxide onto floating gate 20.
[0078] Memory cell 210 is read by placing a positive read voltage across drain region 16 and wordline terminal 22 (which turns on the portion of channel region 18 below the wordline terminal). If floating gate 20 is positively charged (i.e., electrons are erased), the portion of channel region 18 below floating gate 20 is also turned on, and current will flow through channel region 18, which is sensed as an erased state or a "1" state. If floating gate 20 is negatively charged (i.e., programmed by electrons), the portion of the channel region below floating gate 20 is mostly or completely turned off, and no current (or very little current) will flow through channel region 18, which is sensed as a programmed state or a "0" state.
[0079] Table 1 shows typical voltage ranges that may be applied to the terminals of the memory cell 110 for performing read, erase, and program operations:
[0080] Table 1: Figure 2 Operation of the flash memory unit 210
[0081] WL BL SL Read 2-3V 0.6-2V 0V Erase About 11-13V 0V 0V programming 1-2V 1-3μA 9-10V
[0082] Figure 3 Memory cell 310 is shown, which is connected to Figure 2 The memory cell 210 is similar to the memory cell 210 of FIG, but with the addition of a control gate (CG) 28. The control gate 28 is biased at a high voltage (e.g., 10V) during programming, at a low voltage or negative voltage (e.g., 0V / -8V) during erasing, and at a low voltage or medium voltage (e.g., 0V / 2.5V) during reading. The other terminals are similar to Figure 2 That's biased.
[0083] Figure 4 A quad-gate memory cell 410 is shown, comprising a source region 14, a drain region 16, a floating gate 20 over a first portion of a channel region 18, a select gate 22 (typically coupled to a word line WL) over a second portion of the channel region 18, a control gate 28 over the floating gate 20, and an erase gate 30 over the source region 14. This configuration is described in U.S. Patent 6,747,310, which is incorporated herein by reference for all purposes. Here, except for the floating gate 20, all gates are non-floating, meaning they are electrically connected or capable of being electrically connected to a voltage source. Programming is performed by heated electrons from the channel region 18 that inject themselves into the floating gate 20. Erasing is performed by electrons tunneling from the floating gate 20 to the erase gate 30.
[0084] Table 2 shows typical voltage ranges that may be applied to the terminals of the memory cell 310 for performing read, erase, and program operations:
[0085] Table 2: Figure 4 Operation of the flash memory unit 410
[0086] WL / SG BL CG EG SL Read 1.0-2V 0.6-2V 0-2.6V 0-2.6V 0V Erase -0.5V / 0V 0V 0V / -8V 8-12V 0V programming 1V 1μA 8-11V 4.5-9V 4.5-5V
[0087] Figure 5 Memory cell 510 is shown, except that it does not include an erase gate EG. Memory cell 510 is similar to Figure 4 Erasing is performed by biasing the substrate 18 to a high voltage and biasing the control gate CG 28 to a low or negative voltage. Alternatively, erasing is performed by biasing the word line 22 to a positive voltage and biasing the control gate 28 to a negative voltage. Programming and reading are similar to Figure 4 Like that.
[0088] Figure 6 A tri-gate memory cell 610 is shown, which is another type of flash memory cell. Figure 4The memory cell 410 is identical to the memory cell 610, except that the memory cell 610 does not have a separate control gate. Except that no control gate bias is applied, the erase operation (thus erasing by using the erase gate) and the read operation are the same as Figure 4 The programming operation is also completed without a control gate bias, and as a result, a higher voltage must be applied to the source line during the programming operation to compensate for the lack of control gate bias.
[0089] Table 3 shows typical voltage ranges that may be applied to the terminals of the memory cell 610 for performing read, erase, and program operations:
[0090] Table 3: Figure 6 Operation of the flash memory unit 610
[0091] WL / SG BL EG SL Read 0.7-2.2V 0.6-2V 0-2.6V 0V Erase -0.5V / 0V 0V 11.5V 0V programming 1V 2-3μA 4.5V 7-9V
[0092] Figure 7 A stacked gate memory cell 710 is shown, which is another type of flash memory cell. Figure 2 Memory cell 210 is similar to that of FIG1 , except that floating gate 20 extends over the entire channel region 18, and control gate 22 (which here will be coupled to a word line) extends over floating gate 20, separated by an insulating layer (not shown). Erase, program, and read operations operate in a manner similar to that previously described for memory cell 210.
[0093] Table 4 shows typical voltage ranges that may be applied to the terminals of the memory cell 710 and substrate 12 for performing read, erase, and program operations:
[0094] Table 4: Figure 7 Operation of the flash memory unit 710
[0095] CG BL SL substrate Read 2-5V 0.6-2V 0V 0V Erase -8 to -10V / 0V FLT FLT 8-10V / 15-20V programming 8-12V 3-5V 0V 0V
[0096] In order to utilize a memory array comprising one of the above-described types of nonvolatile memory cells in an artificial neural network, two modifications were made. First, the circuitry was configured so that each memory cell could be individually programmed, erased, and read without adversely affecting the memory states of other memory cells in the array, as explained further below. Second, continuous (analog) programming of the memory cells was provided.
[0097] Specifically, the memory state (i.e., charge on the floating gate) of each memory cell in the array can be changed continuously from a fully erased state to a fully programmed state independently and with minimal disturbance to other memory cells. In another embodiment, the memory state (i.e., charge on the floating gate) of each memory cell in the array can be changed continuously from a fully programmed state to a fully erased state, and vice versa, independently and with minimal disturbance to other memory cells. This means that the cell storage device is analog, or at least can store one of many discrete values (such as 16 or 64 different values), which allows very precise and individual tuning of all cells in the memory array, and makes the memory array ideal for storing and fine-tuning the synaptic weights of neural networks.
[0098] The methods and apparatus described herein can be applied to other non-volatile memory technologies such as SONOS (silicon-oxide-nitride-oxide-silicon, charge trapped in nitride), MONOS (metal-oxide-nitride-oxide-silicon, metal charge trapped in nitride), ReRAM (resistive RAM), PCM (phase change memory), MRAM (magnetic RAM), FeRAM (ferroelectric RAM), OTP (two-layer or multi-layer one-time programmable), and CeRAM (correlated electron RAM). The methods and apparatus described herein can be applied to volatile memory technologies used in neural networks, such as, but not limited to, SRAM, DRAM, and / or volatile synaptic cells.
[0099] Neural Networks Using Nonvolatile Memory Cell Arrays
[0100] Figure 8 A non-limiting example of a neural network using a non-volatile memory array according to the present embodiment is conceptually illustrated. This example uses the non-volatile memory array neural network for a facial recognition application, but any other suitable application may also be implemented using a non-volatile memory array-based neural network.
[0101] For this example, S0 is the input layer, which is a 32×32 pixel RGB image with 5 bits of precision (i.e., three 32×32 pixel arrays, one for each color R, G, and B, with 5 bits of precision per pixel). Synapse CB1 from input layer S0 to layer C1 applies different sets of weights in some cases and shared weights in other cases, and scans the input image with a 3×3 pixel overlapping filter (kernel), shifting the filter by 1 pixel (or more than 1 pixel as dictated by the model). Specifically, the values of 9 pixels in the 3×3 portion of the image (i.e., called the filter or kernel) are provided to synapse CB1, where these 9 input values are multiplied by the appropriate weights, and after summing the outputs of this multiplication, a single output value is determined and provided by the first synapse of CB1 for use in generating a feature map for one of the pixels of layer C1. The 3×3 filter is then shifted one pixel to the right within the input layer S0 (i.e., a column of three pixels on the right is added and a column of three pixels on the left is released), whereby the nine pixel values in this newly positioned filter are provided to the synapse CB1, where they are multiplied by the same weights and a second single output value is determined by the associated synapse. This process continues until the 3×3 filter has scanned all three colors and all bits (precision values) across the entire 32×32 pixel image of the input layer S0. This process is then repeated using different sets of weights to generate different feature maps for C1 until all feature maps for layer C1 are calculated.
[0102] At layer C1, in this example, there are 16 feature maps, each with 30×30 pixels. Each pixel is a new feature pixel extracted from the product of the input and the kernel, so each feature map is a two-dimensional array, so in this example, layer C1 is composed of a two-dimensional array of 16 layers (remember that the layers and arrays referred to in this article are logical relationships, not necessarily physical relationships, that is, arrays do not have to be oriented to physical two-dimensional arrays). Each of the 16 feature maps in layer C1 is generated by one of sixteen different sets of synaptic weights applied to the filter scan. The C1 feature maps can all relate to different aspects of the same image features, such as edge recognition. For example, a first map (generated using a first set of weights, shared for all scans used to generate the first map) can identify circular edges, a second map (generated using a second set of weights different from the first) can identify rectangular edges, or the aspect ratio of certain features, and so on.
[0103] Before passing from layer C1 to layer S1, an activation function P1 (pooling) is applied, which pools the values from consecutive non-overlapping 2×2 regions in each feature map. The purpose of the pooling function is to average the values of adjacent positions (or a max function can also be used) to, for example, reduce dependencies on edge positions and reduce the size of the data before entering the next stage. At layer S1, there are 16 15×15 feature maps (i.e., sixteen different arrays of 15×15 pixels each). The synapse CB2 from layer S1 to layer C2 scans the map in S1 using a 4×4 filter, where the filter is shifted by 1 pixel. At layer C2, there are 22 12×12 feature maps. Before passing from layer C2 to layer S2, an activation function P2 (pooling) is applied, which pools the values from consecutive non-overlapping 2×2 regions in each feature map. At layer S2, there are 22 6×6 feature maps. An activation function (pooling) is applied to the synapse CB3 from layer S2 to layer C3, where each neuron in layer C3 is connected to each map in layer S2 via a corresponding synapse on CB3. At layer C3, there are 64 neurons. Synapse CB4 from layer C3 to output layer S3 completely connects C3 to S3, that is, every neuron in layer C3 is connected to every neuron in layer S3. The output at S3 includes 10 neurons, where the highest output neuron determines the class. For example, this output can indicate the recognition or classification of the content of the original image.
[0104] The synapses at each layer are implemented using an array or a portion of an array of non-volatile memory cells.
[0105] Figure 9 A block diagram of an array that can be used for this purpose is shown in FIG. The vector-matrix multiplication (VMM) array 32 includes non-volatile memory cells and serves as a synapse between one layer and the next (such as Figure 6 CB1, CB2, CB3, and CB4 in FIG. 1 ). Specifically, the VMM array 32 includes a nonvolatile memory cell array 33, an erase gate and word line gate decoder 34, a control gate decoder 35, a bit line decoder 36, and a source line decoder 37, which decode the corresponding inputs of the nonvolatile memory cell array 33. The inputs to the VMM array 32 can come from the erase gate and word line gate decoder 34 or from the control gate decoder 35. In this example, the source line decoder 37 also decodes the output of the nonvolatile memory cell array 33. Alternatively, the bit line decoder 36 can decode the output of the nonvolatile memory cell array 33.
[0106] The non-volatile memory cell array 33 serves two purposes. First, it stores weights to be used by the VMM array 32. Second, the non-volatile memory cell array 33 effectively multiplies the inputs by the weights stored in the non-volatile memory cell array 33, and each output line (source line or bit line) adds them together to produce an output, which will serve as the input to the next layer or the final layer. By performing multiplication and addition functions, the non-volatile memory cell array 33 eliminates the need for separate multiplication and addition logic circuits and is also highly power-efficient due to its in-situ memory calculations.
[0107] The output of the non-volatile memory cell array 33 is provided to a differential summer (such as a summing operational amplifier or a summing current mirror) 38, which sums the output of the non-volatile memory cell array 33 to create a single value for the convolution. The differential summer 38 is arranged to perform the summation of positive and negative weights.
[0108] The output values of the difference summer 38 are then summed and provided to the activation function circuit 39, which corrects the output. The activation function circuit 39 can provide a sigmoid, tanh or ReLU function. The corrected output value of the activation function circuit 39 becomes the next layer (for example, Figure 8 The elements of the feature map of layer C1 in the image processing unit are then applied to the next synapse to produce the next feature map layer or the final layer. Thus, in this example, the non-volatile memory cell array 33 constitutes a plurality of synapses (which receive their inputs from existing neuron layers or from an input layer such as an image database), and the summer 38 and activation function circuit 39 constitute a plurality of neurons.
[0109] Figure 9 The inputs to the VMM array 32 (WLx, EGx, CGx, and optionally BLx and SLx) can be analog levels, binary levels, digital pulses (in which case a pulse-to-analog converter PAC may be required to convert the pulses to appropriate input analog levels), or digital bits (in which case a DAC is provided to convert the digital bits to appropriate input analog levels); the outputs can be analog levels, binary levels, digital pulses, or digital bits (in which case an output ADC is provided to convert the output analog levels into digital bits).
[0110] Figure 10 FIG. 1 is a block diagram illustrating the use of multiple layers of VMM arrays 32 (labeled here as VMM arrays 32a, 32b, 32c, 32d, and 32e). Figure 10As shown, the input (denoted as Inputx) is converted from digital to analog by a digital-to-analog converter 31 and provided to the input VMM array 32a. The converted analog input can be a voltage or a current. The first level of input D / A conversion can be accomplished by using a function or LUT (lookup table) that maps the input Inputx to the appropriate analog levels of the matrix multiplier of the input VMM array 32a. The input conversion can also be accomplished by an analog-to-analog (A / A) converter to convert the external analog input into a mapped analog input to the input VMM array 32a. The input conversion can also be accomplished by a digital-to-digital pulse (D / P) converter to convert the external digital input into one or more digital pulses that are mapped to the input VMM array 32a.
[0111] The output generated by the input VMM array 32a is provided as input to the next VMM array (hidden level 1) 32b, which in turn generates an output that is provided as input to the next VMM array (hidden level 2) 32c, and so on. The layers of the VMM array 32 serve as different layers of synapses and neurons of a convolutional neural network (CNN). Each VMM array 32a, 32b, 32c, 32d, and 32e can be a separate physical non-volatile memory array, or multiple VMM arrays can utilize different portions of the same non-volatile memory array, or multiple VMM arrays can utilize overlapping portions of the same physical non-volatile memory array. Each VMM array 32a, 32b, 32c, 32d, and 32e can also be time-division multiplexed for different portions of its array or neurons. Figure 10 The example shown includes five layers (32a, 32b, 32c, 32d, 32e): one input layer (32a), two hidden layers (32b, 32c), and two fully connected layers (32d, 32e). Those skilled in the art will appreciate that this is merely exemplary and that, on the contrary, the system may include more than two hidden layers and more than two fully connected layers.
[0112] Vector-Matrix Multiplication (VMM) Array
[0113] Figure 11 A neuron VMM array 1100 is shown, which is particularly suitable for Figure 3 The memory cells 310 shown are used as synapses and components for neurons between the input layer and the next layer. The VMM array 1100 includes a memory array 1101 of nonvolatile memory cells and a reference array 1102 of nonvolatile reference memory cells (at the top of the array). Alternatively, another reference array can be placed at the bottom.
[0114] In VMM array 1100, control gate lines (such as control gate line 1103) extend in the vertical direction (so reference array 1102 is orthogonal to control gate line 1103 in the row direction), and erase gate lines (such as erase gate line 1104) extend in the horizontal direction. Here, the inputs of VMM array 1100 are provided on control gate lines (CG0, CG1, CG2, CG3), and the outputs of VMM array 1100 appear on source lines (SL0, SL1). In one embodiment, only even-numbered rows are used, and in another embodiment, only odd-numbered rows are used. The current placed on each source line (SL0, SL1, respectively) performs a summation function of all currents from the memory cells connected to that particular source line.
[0115] As described herein for neural networks, the non-volatile memory cells of VMM array 1100 (ie, the flash memory of VMM array 1100) are preferably configured to operate in the sub-threshold region.
[0116] Biasing the nonvolatile reference memory cell and the nonvolatile memory cell described herein in weak inversion:
[0117] Ids=Io*e (Vg-Vth) / kVt =w*Io*e (Vg) / kVt ,
[0118] where w = e (-Vth) / kVt
[0119] For an I-to-V logarithmic converter that uses a memory cell (such as a reference memory cell or a peripheral memory cell) or a transistor to convert the input current to an input voltage:
[0120] Vg=k*Vt*log[Ids / wp*Io]
[0121] Here, wp is w of a reference memory cell or a peripheral memory cell.
[0122] For a memory array used as a vector matrix multiplier VMM array, the output current is:
[0123] Iout=wa*Io*e (Vg) / kVt ,Right now
[0124] Iout=(wa / wp)*Iin=W*Iin
[0125] W=e (Vthp-Vtha) / kVt
[0126] Here, wa = w for each memory cell in the memory array.
[0127] The word line or control gate may be used as the input to the memory cell for the input voltage.
[0128] Alternatively, the flash memory cells of the VMM array described herein may be configured to operate in the linear region:
[0129] Ids=β*(Vgs-Vth)*Vds;β=u*Cox*W / L
[0130] W=α(Vgs-Vth)
[0131] The word line or control gate or bit line or source line can serve as the input of the memory cell operating in the linear region. The bit line or source line can serve as the output of the memory cell.
[0132] For an IV linear converter, a memory cell (eg, a reference memory cell or a peripheral memory cell) or a transistor or resistor operating in a linear region may be used to linearly convert an input / output current into an input / output voltage.
[0133] U.S. Patent Application 15 / 826,345 describes Figure 9 Other embodiments of the VMM array 32 of the present invention are incorporated herein by reference. As described herein, the source line or the bit line can be used as the neuron output (current summing output). Alternatively, the flash memory cells of the VMM array described herein can be configured to operate in the saturation region:
[0134] Ids=α 1 / 2*β*(Vgs-Vth) 2 β=u*Cox*W / L
[0135] W=α(Vgs-Vth) 2
[0136] The word line, control gate, or erase gate can be used as the input of a memory cell operating in the saturation region. The bit line or source line can be used as the output of an output neuron.
[0137] Alternatively, the flash memory cells of the VMM arrays described herein may be used in all regions or a combination thereof (subthreshold, linear, or saturation regions).
[0138] Figure 12 A neuron VMM array 1200 is shown, which is particularly suitable for Figure 2Memory cell 210 is shown and serves as a synapse between the input layer and the next layer. VMM array 1200 includes a memory array 1203 of nonvolatile memory cells, a reference array 1201 of first nonvolatile reference memory cells, and a reference array 1202 of second nonvolatile reference memory cells. Reference arrays 1201 and 1202, arranged along the columns of the array, are used to convert current inputs flowing into terminals BLR0, BLR1, BLR2, and BLR3 into voltage inputs WL0, WL1, WL2, and WL3. In practice, the first and second nonvolatile reference memory cells are diode-connected via a multiplexer 1214 (only partially shown), with the current input flowing therein. The reference cells are tuned (e.g., programmed) to a target reference level. The target reference level is provided by a reference microarray matrix (not shown).
[0139] The memory array 1203 serves two purposes. First, it stores the weights that the VMM array 1200 will use on its corresponding memory cells. Second, the memory array 1203 effectively multiplies the inputs (i.e., the current inputs provided in terminals BLR0, BLR1, BLR2, and BLR3, which the reference arrays 1201 and 1202 convert into input voltages to be provided to the word lines WL0, WL1, WL2, and WL3) by the weights stored in the memory array 1203, and then adds all the results (memory cell currents) to produce an output on the corresponding bit lines (BL0-BLN), which will be the input to the next layer or the final layer. By performing multiplication and addition functions, the memory array 1203 eliminates the need for separate multiplication and addition logic circuits and is also highly power-efficient. Here, the voltage inputs are provided on the word lines (WL0, WL1, WL2, and WL3), and the outputs appear on the corresponding bit lines (BL0-BLN) during a read (inference) operation. The current placed on each of the bit lines BL0-BLN performs a summing function of the currents from all of the nonvolatile memory cells connected to that particular bit line.
[0140] Table 5 shows the operating voltages for VMM array 1200. The columns in the table indicate the voltages placed on the word line for a selected cell, the word line for an unselected cell, the bit line for a selected cell, the bit line for an unselected cell, the source line for a selected cell, and the source line for an unselected cell. The rows indicate read, erase, and program operations.
[0141] Table 5: Figure 12 Operation of the VMM array 1200
[0142] WL WL-Not selected BL BL-Not selected SL SL-Not selected Read 1-3.5V -0.5V / 0V 0.6-2V(Ineuron) 0.6V-2V / 0V 0V 0V Erase About 5-13V 0V 0V 0V 0V 0V programming 1-2V -0.5V / 0V 0.1-3uA Vinh about 2.5V 4-10V 0-1V / FLT
[0143] Figure 13 A neuron VMM array 1300 is shown, which is particularly suitable for Figure 2Memory cell 210 is shown and serves as a synapse and component for neurons between the input layer and the next layer. VMM array 1300 includes a memory array 1303 of nonvolatile memory cells, a reference array 1301 of first nonvolatile reference memory cells, and a reference array 1302 of second nonvolatile reference memory cells. Reference arrays 1301 and 1302 extend in the row direction of VMM array 1300. The VMM array is similar to VMM 1000, except that in VMM array 1300, the word lines extend in the vertical direction. Here, inputs are provided on word lines (WLA0, WLB0, WLA1, WLB2, WLA2, WLB2, WLA3, WLB3), and outputs appear on source lines (SL0, SL1) during a read operation. The current placed on each source line performs a summing function of all currents from the memory cells connected to that particular source line.
[0144] Table 6 shows the operating voltages for VMM array 1300. The columns in the table indicate the voltages placed on the word line for a selected cell, the word line for an unselected cell, the bit line for a selected cell, the bit line for an unselected cell, the source line for a selected cell, and the source line for an unselected cell. The rows indicate read, erase, and program operations.
[0145] Table 6: Figure 13 Operation of the VMM array 1300
[0146]
[0147] Figure 14 A neuron VMM array 1400 is shown, which is particularly suitable for Figure 3 Memory cell 310 is shown and serves as a synapse and component of neurons between the input layer and the next layer. VMM array 1400 includes a memory array 1403 of nonvolatile memory cells, a reference array 1401 of first nonvolatile reference memory cells, and a reference array 1402 of second nonvolatile reference memory cells. Reference arrays 1401 and 1402 are used to convert current inputs flowing into terminals BLR0, BLR1, BLR2, and BLR3 into voltage inputs CG0, CG1, CG2, and CG3. In practice, the first nonvolatile reference memory cell and the second nonvolatile reference memory cell are diode-connected via a multiplexer 1412 (only partially shown), with the current input flowing into them through BLR0, BLR1, BLR2, and BLR3. Multiplexers 1412 each include a corresponding multiplexer 1405 and a cascode transistor 1404 to ensure a constant voltage on a bit line (such as BLR0) of each of the first and second nonvolatile reference memory cells during a read operation. The reference cells are tuned to a target reference level.
[0148] The memory array 1403 serves two purposes. First, it stores the weights that will be used by the VMM array 1400. Second, the memory array 1403 effectively multiplies the inputs (current inputs provided to terminals BLR0, BLR1, BLR2, and BLR3, which the reference arrays 1401 and 1402 convert into input voltages to provide to the control gates CG0, CG1, CG2, and CG3) by the weights stored in the memory array and then adds all the results (cell currents) to produce an output, which appears in BL0-BLN and will be the input to the next layer or the final layer. By performing multiplication and addition functions, the memory array eliminates the need for separate multiplication and addition logic circuits and is also highly power efficient. Here, the inputs are provided on the control gate lines (CG0, CG1, CG2, and CG3) and the outputs appear on the bit lines (BL0-BLN) during a read operation. The current placed on each bit line performs a summing function of all the currents from the memory cells connected to that particular bit line.
[0149] The VMM array 1400 implements unidirectional tuning for the nonvolatile memory cells in the memory array 1403. That is, each nonvolatile memory cell is erased and then partially programmed until the desired charge on the floating gate is reached. This can be performed, for example, using the precision programming techniques described below. If too much charge is placed on the floating gate (causing an incorrect value to be stored in the cell), the cell must be erased and the sequence of partial programming operations must be restarted. As shown, two rows sharing the same erase gate (such as EG0 or EG1) need to be erased together (which is called a page erase), and thereafter, each cell is partially programmed until the desired charge on the floating gate is reached.
[0150] Table 7 shows the operating voltages for the VMM array 1400. The columns in the table indicate the voltages applied to the word line for a selected cell, the word line for an unselected cell, the bit line for a selected cell, the bit line for an unselected cell, the control gate for a selected cell, the control gate for an unselected cell in the same sector as the selected cell, the control gate for an unselected cell in a different sector from the selected cell, the erase gate for a selected cell, the erase gate for an unselected cell, the source line for a selected cell, and the source line for an unselected cell. The rows indicate read, erase, and program operations.
[0151] Table 7: Figure 14 Operation of the VMM array 1400
[0152]
[0153] Figure 15 A neuron VMM array 1500 is shown, which is particularly suitable for Figure 3Memory cells 310 are shown and serve as synapses and components for neurons between the input layer and the next layer. VMM array 1500 includes a memory array 1503 of nonvolatile memory cells, a reference array 1501 of first nonvolatile reference memory cells, and a reference array 1502 of second nonvolatile reference memory cells. EG lines EGR0, EG0, EG1, and EGR1 extend vertically, while CG lines CG0, CG1, CG2, and CG3 and SL lines WL0, WL1, WL2, and WL3 extend horizontally. VMM array 1500 is similar to VMM array 1400, except that VMM array 1500 implements bidirectional tuning, where each individual cell can be fully erased, partially programmed, and partially erased as needed to achieve a desired charge on the floating gate due to the use of separate EG lines. As shown, reference arrays 1501 and 1502 convert input currents in terminals BLR0, BLR1, BLR2, and BLR3 into control gate voltages CG0, CG1, CG2, and CG3 to be applied to the memory cells in the row direction (through the action of the reference cells connected via the diodes of multiplexer 1514). The current outputs (neurons) are in bit lines BL0-BLN, where each bit line sums all currents from the nonvolatile memory cells connected to that particular bit line.
[0154] Table 8 shows the operating voltages for VMM array 1500. The columns in the table indicate the voltages applied to the word line for a selected cell, the word line for an unselected cell, the bit line for a selected cell, the bit line for an unselected cell, the control gate for a selected cell, the control gate for an unselected cell in the same sector as the selected cell, the control gate for an unselected cell in a different sector from the selected cell, the erase gate for a selected cell, the erase gate for an unselected cell, the source line for a selected cell, and the source line for an unselected cell. The rows indicate read, erase, and program operations.
[0155] Table 8: Figure 15 Operation of the VMM array 1500
[0156]
[0157] Figure 24 A neuron VMM array 2400 is shown, which is particularly suitable for Figure 2 The memory cell 210 shown in FIG. 2 is used as a synapse and component of neurons between the input layer and the next layer. In the VMM array 2400, the inputs INPUT0 to INPUT N On bit lines BL0 to BL N outputs OUTPUT1, OUTPUT2, OUTPUT3 and OUTPUT4 are generated on source lines SL0, SL1, SL2 and SL3 respectively.
[0158] Figure 25 A neuron VMM array 2500 is shown, which is particularly suitable for Figure 2 The memory cell 210 is shown and serves as a synapse and component of the neurons between the input layer and the next layer. In this example, the inputs INPUT0, INPUT1, INPUT2 and INPUT3 are received on the source lines SL0, SL1, SL2 and SL3 respectively; the outputs OUTPUT0 to OUTPUT N On bit lines BL0 to BL N Generate on.
[0159] Figure 26 A neuron VMM array 2600 is shown, which is particularly suitable for Figure 2 The memory unit 210 shown in FIG. 2 is used as a synapse and a component of a neuron between an input layer and a next layer. In this example, inputs INPUT0 to INPUT M On word lines WL0 to WL M is received; output OUTPUT0 to OUTPUT N On bit lines BL0 to BL N Generate on.
[0160] Figure 27 A neuron VMM array 2700 is shown, which is particularly suitable for Figure 3 The memory unit 310 shown in FIG. 1 is used as a synapse and a component of a neuron between an input layer and a next layer. In this example, inputs INPUT0 to INPUT M On word lines WL0 to WL M is received; output OUTPUT0 to OUTPUT N On bit lines BL0 to BL N Generate on.
[0161] Figure 28 A neuron VMM array 2800 is shown, which is particularly suitable for Figure 4 The memory unit 410 shown in FIG. 4 is used as a synapse and a component of a neuron between an input layer and a next layer. In this example, inputs INPUT0 to INPUT n On bit lines BL0 to BL N The outputs OUTPUT1 and OUTPUT2 are generated on the erase gates EG0 and EG1.
[0162] Figure 29 A neuron VMM array 2900 is shown, which is particularly suitable for Figure 4 The memory unit 410 shown in FIG. 4 is used as a synapse and a component of a neuron between an input layer and a next layer. In this example, inputs INPUT0 to INPUT Nare received at the gates of bit line control gates 2901-1, 2901-2 to 2901-(N-1), and 2901-N, which are coupled to bit lines BL0 to BL1, respectively. N Exemplary outputs OUTPUT1 and OUTPUT2 are generated on erase gate lines SL0 and SL1.
[0163] Figure 30 A neuron VMM array 3000 is shown, which is particularly suitable for Figure 3 The memory unit 310 shown, Figure 5 The storage unit 510 and Figure 7 The storage unit 710 shown in FIG. 1 is used as a synapse and a component of a neuron between an input layer and a next layer. In this example, inputs INPUT0 to INPUT M On word lines WL0 to WL M is received; output OUTPUT0 to OUTPUT N On bit lines BL0 to BL N Generate on.
[0164] Figure 31 A neuron VMM array 3100 is shown, which is particularly suitable for Figure 3 The memory unit 310 shown, Figure 5 The storage unit 510 and Figure 7 The storage unit 710 shown in FIG. 1 is used as a synapse and a component of a neuron between an input layer and a next layer. In this example, inputs INPUT0 to INPUT M On the control gate lines CG0 to CG M OUTPUT0 is output to OUTPUT N On the source lines SL0 to SL N On the generation, each source line SL i Coupled to the source line terminals of all memory cells in column i.
[0165] Figure 32A VMM system 3200 is shown. The VMM system 3200 includes a VMM array 3201 (which may be based on any of the previously discussed VMM designs, such as VMMs 900, 1000, 1100, 1200, and 1320, or other VMM designs), a low voltage row decoder 3202, a high voltage row decoder 3203, a reference cell low voltage column decoder 3204 (shown in the column direction, meaning it provides input to output conversion in the row direction), a bit line multiplexer 3205, control logic 3206, analog circuitry 3207, a neuron output block 3208, an input VMM circuit block 3209, a pre-decoder 3210, a test circuit 3211, erase-program control logic EPCTL 3212, analog and high voltage generation circuitry 3213, a bit line PE driver 3214, redundant arrays 3215 and 3216, an NVR sector 3217, and a reference sector 3218. The input circuit block 3209 serves as an interface for external input to the input terminal of the memory array. The neuron output block 3208 serves as an interface for output from the memory array to the external interface.
[0166] The low-voltage row decoder 3202 provides bias voltages for read and program operations and provides decoding signals for the high-voltage row decoder 3203. The high-voltage row decoder 3203 provides high-voltage bias signals for program and erase operations. The reference cell low-voltage column decoder 3204 provides decoding functions for the reference cells. The bit line PE driver 3214 provides control functions for the bit lines during programming, verification, and erase operations. The analog and high-voltage generation circuit 3213 is a shared bias block that provides multiple voltages required for various programming, erasing, program verification, and read operations. Redundant arrays 3215 and 3216 provide array redundancy for replacing defective array portions. The NVR (non-volatile register, also known as the information sector) sector 3217 is an array sector used to store, but not limited to, user information, device ID, passwords, security keys, correction bits, configuration bits, and manufacturing information.
[0167] Figure 33An analog neural memory system 3300 is shown. The analog neural memory system 3300 includes macroblocks 3301a, 3301b, 3301c, 3301d, 3301e, 3301f, 3301g, and 3301h; neuron output blocks (such as summer circuits and sample and hold (S / H) circuits) 3302a, 3302b, 3302c, 3302d, 3302e, 3302f, 3302g, and 3302h; and input circuit blocks 3303a, 3303b, 3303c, 3303d, 3303e, 3303f, 3303g, and 3304h. Each of macroblocks 3301a, 3301b, 3301c, 3301d, 3301e, and 3301f is a VMM subsystem that includes a VMM array comprising rows and columns of non-volatile memory cells (such as flash memory cells). Neural memory subsystem 3333 includes macroblock 3301, input block 3303, and neuron output block 3302. Neural memory subsystem 3333 may have its own digital control block.
[0168] The analog neural memory system 3300 also includes a system control block 3304, an analog low voltage block 3305, a high voltage block 3306, and a timing control circuit 3670, which are discussed in further detail below with reference to FIG. 36.
[0169] The system control block 3304 may include one or more microcontroller cores such as ARM / MIPS / RISC_V cores to process general control functions and arithmetic operations. The system control block 3304 may also include a SIMD (single instruction multiple data) unit to utilize a single instruction to operate on multiple data. The system control block may include a DSP core. The system control block may include hardware or software for executing functions such as but not limited to pooling, averaging, minimum, maximum, softmax, addition, subtraction, multiplication, division, logarithm, antilogarithm, ReLu, sigmoid, tanh, and data compression. The system control block may include hardware or software for executing functions such as activation approximator / quantizer / normalizer. The system control block may include the ability to execute functions such as input data approximator / quantizer / normalizer. The system control block may include hardware or software for executing functions such as activation approximator / quantizer / normalizer. The control block of the neural memory subsystem 3333 may include similar elements of the system control block 3304, such as a microcontroller core, a SIMD core, a DSP core, and other functional units.
[0170] In one embodiment, neuron output blocks 3302a, 3302b, 3302c, 3302d, 3302e, 3302f, 3302g, and 3302h each include a buffered (e.g., op amp) low-impedance output type circuit that can drive long and configurable interconnects. In one embodiment, input circuit blocks 3303a, 3303b, 3303c, 3303d, 3303e, 3303f, 3303g, and 3303h each provide a summed high-impedance current output. In another embodiment, neuron output blocks 3302a, 3302b, 3302c, 3302d, 3302e, 3302f, 3302g, and 3302h each include an activation circuit, in which case an additional low-impedance buffer is required to drive the output.
[0171] In another embodiment, neuron output blocks 3302a, 3302b, 3302c, 3302d, 3302e, 3302f, 3302g, and 3302h each include an analog-to-digital conversion block that outputs digital bits rather than analog signals. In this embodiment, input circuit blocks 3303a, 3303b, 3303c, 3303d, 3303e, 3303f, 3303g, and 3303h each include a digital-to-analog conversion block that receives digital bits from the corresponding neuron output block and converts the digital bits into analog signals.
[0172] Thus, neuron output blocks 3302a, 3302b, 3302c, 3302d, 3302e, 3302f, 3302g, and 3302h receive output current from macroblocks 3301a, 3301b, 3301c, 3301d, 3301e, and 3301f and optionally convert the output current into an analog voltage, a digital bit, or one or more digital pulses, where the width of each pulse or the number of pulses varies in response to the value of the output current. Similarly, input circuit blocks 3303a, 3303b, 3303c, 3303d, 3303e, 3303f, 3303g, and 3303h optionally receive an analog current, an analog voltage, a digital bit, or a digital pulse, wherein the width or number of each pulse varies in response to the value of the output current, and provide the analog current to macro blocks 3301a, 3301b, 3301c, 3301d, 3301e, and 3301f. Input circuit blocks 3303a, 3303b, 3303c, 3303d, 3303e, 3303f, 3303g, and 3303h optionally include a voltage-to-current converter, an analog or digital counter for counting the number of digital pulses in the input signal or the length of the digital pulse width in the input signal, or a digital-to-analog converter.
[0173] Long Short-Term Memory
[0174] Prior art includes a concept known as long short-term memory (LSTM). LSTM cells are commonly used in neural networks. LSTM allows neural networks to remember information for a predetermined, arbitrary time interval and use that information in subsequent operations. A typical LSTM cell consists of a cell, an input gate, an output gate, and a forget gate. These three gates regulate the flow of information into and out of the cell, as well as the time interval over which information is remembered within the LSTM. VMMs are particularly useful in LSTM cells.
[0175] Figure 16 An exemplary LSTM 1600 is shown. The LSTM 1600 in this example includes units 1601, 1602, 1603, and 1604. Unit 1601 receives an input vector x0 and generates an output vector h0 and a unit state vector c0. Unit 1602 receives an input vector x1, an output vector (hidden state) h0 from unit 1601, and a unit state c0 from unit 1601. 0, And generates an output vector h1 and a cell state vector c1. Unit 1603 receives an input vector x2, the output vector (hidden state) h1 from unit 1602, and the cell state c1 from unit 1602, and generates an output vector h2 and a cell state vector c2. Unit 1604 receives an input vector x3, the output vector (hidden state) h2 from unit 1603, and the cell state c2 from unit 1603, and generates an output vector h3. Additional units can be used, and the LSTM with four units is only an example.
[0176] Figure 17 Shows the available Figure 16 FIG1 is an exemplary implementation of an LSTM unit 1700 for units 1601, 1602, 1603, and 1604 in FIG1. LSTM unit 1700 receives an input vector x(t), a cell state vector c(t-1) from the previous unit, and an output vector h(t-1) from the previous unit, and generates a cell state vector c(t) and an output vector h(t).
[0177] LSTM unit 1700 includes sigmoid function devices 1701, 1702, and 1703, each of which applies a number between 0 and 1 to control the amount of each component in the input vector that is allowed to pass through to the output vector. LSTM unit 1700 also includes tanh devices 1704 and 1705 for applying a hyperbolic tangent function to the input vector, multiplier devices 1706, 1707, and 1708 for multiplying two vectors together, and an addition device 1709 for adding two vectors together. The output vector h(t) can be provided to the next LSTM unit in the system, or it can be accessed for other purposes.
[0178] Figure 18 LSTM unit 1800 is shown, which is an example of a specific implementation of LSTM unit 1700. For the convenience of the reader, the same numbering is used in LSTM unit 1800 as in LSTM unit 1700. Sigmoid function devices 1701, 1702, and 1703, and tanh device 1704 each include multiple VMM arrays 1801 and activation circuit blocks 1802. Therefore, it can be seen that VMM arrays are particularly useful in LSTM units used in certain neural network systems. Multiplier devices 1706, 1707, and 1708, and adder device 1709 can be implemented digitally or in analog form. Activation function block 1802 can be implemented digitally or in analog form.
[0179] An alternative form of LSTM cell 1800 (and another example of a specific implementation of LSTM cell 1700) is Figure 19 As shown in Figure 19 In the example, sigmoid function devices 1701, 1702, and 1703 and tanh device 1704 share the same physical hardware (VMM array 1901 and activation function block 1902) in a time-division multiplexing manner. The LSTM unit 1900 also includes a multiplier device 1903 that multiplies two vectors together, an addition device 1908 that adds two vectors together, a tanh device 1705 (which includes an activation circuit block 1902), a register 1907 that stores the value i(t) when the value i(t) is output from the sigmoid function block 1902, a register 1904 that stores the value f(t)*c(t-1) when the value is output from the multiplier device 1903 through the multiplexer 1910, a register 1905 that stores the value i(t)*u(t) when the value is output from the multiplier device 1903 through the multiplexer 1910, a register 1906 that stores the value o(t)*c~(t) when the value is output from the multiplier device 1903 through the multiplexer 1910, and a multiplexer 1909.
[0180] LSTM unit 1800 includes multiple sets of VMM arrays 1801 and corresponding activation function blocks 1802, while LSTM unit 1900 includes only one set of VMM arrays 1901 and activation function blocks 1902, which are used to represent multiple layers in the implementation of LSTM unit 1900. LSTM unit 1900 will require less space than LSTM 1800 because LSTM unit 1900 only requires 1 / 4 of its space for VMM and activation function blocks compared to LSTM unit 1800.
[0181] It is also understood that an LSTM cell will typically include multiple VMM arrays, each of which requires functionality provided by certain circuit blocks outside the VMM array (such as summer and activation circuit blocks and high-voltage generation blocks). Providing a separate circuit block for each VMM array would require a significant amount of space within the semiconductor device and would be somewhat inefficient.
[0182] Gate-controlled recursive unit
[0183] The simulated VMM implementation can be used for a Gated Recurrent Unit (GRU) system. A GRU is a gate-controlled mechanism in recurrent neural networks. A GRU is similar to an LSTM, except that a GRU cell typically contains fewer components than an LSTM cell.
[0184] Figure 20 An exemplary GRU 2000 is shown. GRU 2000 in this example includes units 2001, 2002, 2003, and 2004. Unit 2001 receives an input vector x0 and generates an output vector h0. Unit 2002 receives an input vector x1, an output vector h0 from unit 2001, and generates an output vector h1. Unit 2003 receives an input vector x2 and an output vector (hidden state) h1 from unit 2002, and generates an output vector h2. Unit 2004 receives an input vector x3 and an output vector (hidden state) h2 from unit 2003, and generates an output vector h3. Additional units may be used, and a GRU with four units is merely an example.
[0185] Figure 21 Shows the available Figure 20 2001, 2002, 2003, and 2004 of a GRU unit 2100. The GRU unit 2100 receives an input vector x(t) and an output vector h(t-1) from a previous GRU unit and generates an output vector h(t). The GRU unit 2100 includes sigmoid function devices 2101 and 2102, each of which applies a number between 0 and 1 to components from the output vector h(t-1) and the input vector x(t). The GRU unit 2100 also includes a tanh device 2103 for applying a hyperbolic tangent function to the input vector, a plurality of multiplier devices 2104, 2105, and 2106 for multiplying two vectors together, an addition device 2107 for adding two vectors together, and a complement device 2108 for subtracting the input from 1 to generate an output.
[0186] Figure 222 shows a GRU unit 2200, which is an example of a specific implementation of the GRU unit 2100. For the convenience of the reader, the same numbering is used in the GRU unit 2200 as in the GRU unit 2100. Figure 22 As shown, sigmoid function devices 2101 and 2102 and tanh device 2103 each include multiple VMM arrays 2201 and activation function blocks 2202. Therefore, it can be seen that VMM arrays are particularly useful in GRU units used in certain neural network systems. Multiplier devices 2104, 2105, and 2106, addition device 2107, and complement device 2108 are implemented digitally or in analog form. Activation function block 2202 can be implemented digitally or in analog form.
[0187] An alternative form of GRU unit 2200 (and another example of a specific implementation of GRU unit 2300) is Figure 23 As shown in Figure 23 In , the GRU unit 2300 utilizes a VMM array 2301 and an activation function block 2302, which, when configured as a sigmoid function, applies a number between 0 and 1 to control the amount of each component in the input vector that is allowed to reach the output vector. Figure 23 In the example, sigmoid function devices 2101 and 2102 and tanh device 2103 share the same physical hardware (VMM array 2301 and activation function block 2302) in a time-division multiplexing manner. The GRU unit 2300 also includes a multiplier device 2303 that multiplies two vectors together, an addition device 2305 that adds two vectors together, a complement device 2309 that subtracts the input from 1 to generate an output, a multiplexer 2304, a register 2306 that holds the value h(t-1)*r(t) when it is output from the multiplier device 2303 through the multiplexer 2304, a register 2307 that holds the value h(t-1)*z(t) when it is output from the multiplier device 2303 through the multiplexer 2304, and a register 2308 that holds the value h^(t)*(1-z(t)) when it is output from the multiplier device 2303 through the multiplexer 2304.
[0188] The GRU unit 2200 includes multiple sets of VMM arrays 2201 and activation function blocks 2202, while the GRU unit 2300 includes only one set of VMM arrays 2301 and activation function blocks 2302, which are used to represent multiple layers in the implementation of the GRU unit 2300. The GRU unit 2300 will require less space than the GRU unit 2200 because the GRU unit 2300 only requires 1 / 3 of its space for VMMs and activation function blocks compared to the GRU unit 2200.
[0189] It will also be appreciated that a GRU system will typically include multiple VMM arrays, each of which requires functionality provided by certain circuit blocks outside the VMM array (such as summer and activation circuit blocks and high-voltage generation blocks). Providing a separate circuit block for each VMM array would require a significant amount of space within the semiconductor device and would be somewhat inefficient.
[0190] The inputs to the VMM array can be analog levels, binary levels, or digital bits (in which case a DAC is required to convert the digital bits to the appropriate input analog levels), and the outputs can be analog levels, binary levels, or digital bits (in which case an output ADC is required to convert the output analog levels to digital bits).
[0191] For each memory cell in the VMM array, each weight w can be implemented by a single memory cell, a differential cell, or two hybrid memory cells (the average of the two cells). In the case of a differential cell, two memory cells are required to implement the weight w as a differential weight (w = w + - w -). In the case of two hybrid memory cells, two memory cells are required to implement the weight w as the average of the two cells.
[0192] Output circuit
[0193] Figure 34A shows the application of the output neuron to the output neuron current I NEU 3406 is converted into digital pulses or digital output bits by an integrating dual hybrid slope analog-to-digital converter (ADC) 3400.
[0194] In one embodiment, ADC 3400 converts the neuron output block (such as Figure 32 The analog output current of the neuron output blocks 3302a, 3302b, 3302c, 3302d, 3302e, 3302f, 3302g, and 3302h in the memory array is converted into a digital pulse, the width of which varies in proportion to the magnitude of the analog output current of the neuron output block. The integrator including the integrating operational amplifier 3401 and the integrating capacitor 3402 is used to calculate the memory array current I NEU 3406 (which is the output neuron current) is integrated relative to the reference current IREF 3407.
[0195] Optionally, IREF 3407 may include a temperature coefficient of 0 or a temperature coefficient that tracks the neuronal current I NEU Bandgap filter of 3406. The latter temperature coefficient may optionally be obtained from a reference array containing values determined during the test phase.
[0196] Optionally, a calibration step can be performed with the circuit at or above operating temperature to cancel out any leakage currents present in the array or control circuitry, which can then be obtained from Figure 34B or Figure 35B Subtract the offset value from Ineu in.
[0197] During the initialization phase, switch 3408 is closed. Then, Vout 3403 and the input of the negative terminal of the operational amplifier 3401 will become VREF. Thereafter, as Figure 34B As shown, the switch 33408 is open, and during the fixed time period tref, the neuronal current I NEU 3406 integrates upward. During the fixed time period tref, Vout rises, and its slope changes as the neuron current changes. Thereafter, during the time period tmeas, the constant reference current IREF is integrated downward during the time period tmeas (during which Vout falls), where tmeas is the time required to integrate Vout down to VREF.
[0198] When VOUT>VREFV, the output EC 3405 will be high, otherwise it will be low. EC3405 thus generates a pulse whose width reflects the time period tmeas, which in turn is related to the current I NEU 3406 is proportional. Figure 34B In the example of tmeas=Ineu1, EC3405 is shown as waveform 3410 and as waveform 3412 in the example of tmeas=Ineu2. Therefore, the output neuron current I NEU 3406 is converted into a digital pulse EC 3405, wherein the width of the digital pulse EC 3405 is proportional to the output neuron current I NEU The value of 3406 changes proportionally.
[0199] Current I NEU 3406 = tmeas / tref * IREF. For example, for a desired output bit resolution of 10 bits, tref is equivalent to a time period of 1024 clock cycles. NEU The value of 3406 and the value of Iref, the time period tmeas varies from 0 to 1024 clock cycles. Figure 34B Shown I NEU Examples of two different values for 3406, one of which is NEU 3406=Ineu1, another I NEU 3406 = Ineu2. Therefore, the neuronal current I NEU 3406 will affect the rate and slope of charging.
[0200] Optionally, the output pulse EC 3405 can be converted into a series of pulses with a uniform period for transmission to the next stage of the circuit, such as the input block of another VMM array. At the beginning of the time period tmeas, the output EC 3405 is input into the AND gate 3440 along with the reference clock 3441. During the time period when VOUT>VREF, the output will be a pulse train 3442 (where the frequency of the pulses in the pulse train 3442 is the same as the frequency of the clock 3441). The number of pulses is proportional to the time period tmeas, which is proportional to the current I NEU Proportional to 3406.
[0201] Optionally, the pulse train 3443 can be input to a counter 3420 which will count the number of pulses in the pulse train 3442 and generate a count value 3421 which is a digital count of the number of pulses in the pulse train 3442 and which is related to the neuronal current I NEU 3406. The count value 3421 includes a set of digital bits. In another embodiment, the integrating dual slope ADC 3400 can convert the neuron current I NEU 3407 is converted into a pulse, where the width of the pulse is proportional to the neuronal current I NEU The inversion is inversely proportional to the magnitude of the 3407. This inversion can be done digitally or analogously and converted into a series of pulses or digital bits for output to the following circuit.
[0202] Figure 35A shows the cell current I applied to the output neuron to NEU 3504 into digital pulses of varying widths or a series of digital output bits. For example, the ADC 3500 can be used to convert a neuron output block (such as Figure 32 The analog output current of the neuron output blocks 3302a, 3302b, 3302c, 3302d, 3302e, 3302f, 3302g and 3302h in the embodiment of the present invention is converted into a set of digital output bits. The integrator including the integrating operational amplifier 3501 and the integrating capacitor 3502 is used to convert the neuron current I NEU 3504 is integrated with respect to a reference current IREF 3503. Switch 3505 can be closed to reset VOUT.
[0203] During the initialization phase, switch 3505 is closed and VOUT is charged to voltage V BIAS .
[0204] Afterwards, if Figure 35B As shown, the switch 3505 is open, and during the fixed time tref, the cell current I NEU3504 is integrated upwards. Thereafter, the reference current IREF 3503 is integrated downwards for a period of time tmeas until Vout drops to zero. NEU 3504 = tmeas Ineu / tref * IREF. For example, for a required 10-bit output resolution, tref is equivalent to a time period of 1024 clock cycles. NEU 3504 and the value of Iref, the time period tmeas varies from 0 to 1024 clock cycles. Figure 35B An example of two different values of Ineu is shown, one with current Ineu1 and the other with current Ineu2. Therefore, the neuron current I NEU The 3504 affects the rate and slope of charge and discharge.
[0205] When VOUT>VREF, output 3506 will be high, otherwise it will be low. Output 3506 thus generates a pulse whose width reflects the time period tmeas, which in turn is related to the current I NEU 3404 is proportional. Figure 35B , the output 3506 is shown as waveform 3512 in the example where tmeas=Ineu1 and as waveform 3515 in the example where tmeas=Ineu2. NEU 3504 is converted into a pulse, i.e. output 3506, where the width of the pulse is proportional to the output neuron current I NEU The value of 3504 changes proportionally.
[0206] Optionally, output 3506 can be converted into a series of pulses with a uniform period for transmission to the next stage of the circuit, such as the input block of another VMM array. At the beginning of time period tmeas, output 3506 is input into AND gate 3508 along with reference clock 3507. During the time period when VOUT>VREF, the output will be pulse train 3509 (where the frequency of the pulses in pulse train 3509 is the same as the frequency of reference clock 3507). The number of pulses is proportional to time period tmeas, which is proportional to the current I NEU Proportional to 3504.
[0207] Optionally, the pulse train 3509 can be input to a counter 3510, which will count the number of pulses in the pulse train 3509 and generate a count value 3511, which is a digital count of the number of pulses in the pulse train 3509, which is proportional to the neuronal current I as shown in waveforms 3514 and 3517. NEU The count value 3511 includes a set of digital bits.
[0208] In another embodiment, the integrating dual slope ADC 3500 can convert the neuronal current I NEU 3504 is converted into a pulse, wherein the width of the pulse is proportional to the neuronal current I NEU The inversion is inversely proportional to the magnitude of 3504. This inversion can be done digitally or in analog and converted into one or more pulses or digital bits for output to the follower circuit.
[0209] Figure 35B I NEU The count value 3511 (digital bits) of the two neuron current values Ineu1 and Ineu2 of 3504.
[0210] Figure 36A and 36B Waveforms associated with exemplary methods 3600 and 3650 performed in a VMM during operation are shown. In each method 3600 and 3650, word lines WL0, WL1, and WL2 receive a variety of different inputs that can optionally be converted into analog voltage waveforms to be applied to the word lines. In these examples, voltage VC represents the voltage at Figure 34A and Figure 35A The voltage across integrating capacitor 3402 or 3502 in the output block of the first VMM is measured in ADC 3400 or 3500. An OT pulse (='1') indicates a period during which the output of the neuron (which is proportional to the neuron's value) is captured using integrating dual-slope ADC 3400 or 3500. As shown in Figures 34 and 35, the output of the output block can be a pulse whose width varies proportionally to the output neuron current of the first VMM, or it can be a series of pulses of uniform width, where the number of pulses varies proportionally to the neuron current of the first VMM. These pulses can then be applied as input to the second VMM.
[0211] During method 3600, the series of pulses (such as pulse series 3442 or pulse series 3509) or an analog voltage derived from the series of pulses is applied to a word line of the second VMM array. Alternatively, the series of pulses or an analog voltage derived from the series of pulses can be applied to the control gates of cells within the second VMM array. The number of pulses (or clock cycles) directly corresponds to the magnitude of the input. In this particular example, the magnitude of the input on WL1 is four times that of WL0 (four pulses versus one pulse).
[0212] During method 3650, a single pulse of varying width (such as EC 3405 or output 3506) or an analog voltage derived from a single pulse is applied to a word line of the second VMM array, but with a variable pulse width. Alternatively, the pulse or an analog voltage derived from the pulse can be applied to a control gate. The width of the single pulse directly corresponds to the magnitude of the input. For example, the magnitude of the input on WL1 is four times that on WL0 (the pulse width of WL1 is four times the pulse width of WL0).
[0213] In addition, reference Figure 36C The timing control circuit 3670 can be used to manage the power of the VMM system by managing the output interface and input interface of the VMM array and sequentially splitting the conversion of various outputs or various inputs. Figure 56 A power management method 5600 is shown. Step 1: receiving a plurality of inputs for a vector-matrix multiplication array (step 5601); Step 2: organizing the plurality of inputs into a plurality of groups of inputs (step 5602); Step 3: providing each of the plurality of groups of inputs to the array in turn (step 5603).
[0214] An embodiment of power management method 5600 is as follows. Inputs can be applied sequentially over time to a VMM system (such as the word lines or control gates of a VMM array). For example, for a VMM array with 512 word line inputs, the word line inputs can be divided into four groups: WL0-127, WL128-255, WL256-383, and WL383-511. Each group can be enabled at a different time, and an output read operation (converting neuron currents into digital bits) can be performed on one of the four groups of word lines, such as by the output integration circuits of Figures 34 to 36. Then, after reading each of the four groups in sequence, the output digital bit results are combined. This operation can be controlled by timing control circuit 3670.
[0215] In another embodiment, the timing control circuit 3670 is used in a vector-matrix multiplication system such as Figure 33 The timing control circuit 3670 may cause inputs to be sequentially applied to the VMM subsystem 3333 over time, such as by enabling input circuit blocks 3303a, 3303b, 3303c, 3303d, 3303e, 3303f, 3303g, and 3303h at different times. Similarly, the timing control circuit 3670 may cause outputs from the VMM subsystem 3333 to be sequentially read over time, such as by enabling neuron output blocks 3302a, 3302b, 3302c, 3302d, 3302e, 3302f, 3302g, and 3302h at different times.
[0216] Figure 57 A power management method 5700 is shown. Step 1: Receive multiple outputs from a vector-matrix multiplication array (step 5701). Step 2: Organize the multiple outputs from the array into multiple groups of outputs (step 5702). Step 3: Provide each of the multiple groups of outputs sequentially to a converter circuit (step 5703).
[0217] An embodiment of the power management method 5700 is as follows. Power management can be implemented by the timing control circuit 3670 by sequentially reading groups of neuron outputs at different times, i.e., by multiplexing output circuits (such as output ADC circuits) across multiple neuron outputs (bit lines). The bit lines can be placed into different groups, and the output circuits are operated on one group at a time in sequence under the control of the timing control circuit 3670.
[0218] Figure 58 A power management method 5800 is shown. Step 1: Receive a plurality of inputs in a vector-matrix multiplication system comprising a plurality of arrays. Step 2: Sequentially enable one or more arrays of the plurality of arrays to receive some or all of the plurality of inputs (step 5802).
[0219] An embodiment of power management method 5800 is as follows. Timing control circuit 3670 can operate on one neural network layer at a time. For example, if one neural network layer is represented in a first VMM array and a second neural network layer is represented in a second VMM array, output read operations (such as where neuron outputs are converted into digital bits) can be performed sequentially on one VMM array at a time, thereby managing the power of the VMM system.
[0220] In another embodiment, the timing control circuit 3670 can be configured to: Figure 33 The multiple neural memory subsystems 3333 or multiple macros 3301 shown are operated.
[0221] In another embodiment, the timing control circuit 3670 can be configured to: Figure 33 The plurality of neural memory subsystems 3333 or the plurality of macros 3301 shown are operated without releasing the array bias (e.g., bias on the word line WL and / or bit line BL for using the control gate CG as input and the bit line BL as output, or bias on the control gate CG and / or bit line BL for using the word line WL as input and the bit line BL as output) during inactivity (i.e., the period between opening and closing sequential activations). This is to save power from unnecessary discharging and charging of the array bias that is to be used multiple times during one or more read operations (e.g., during inference or classification operations).
[0222] Figures 37 to 44Shows various circuits that can be used for a VMM input block, such as Figure 33 The input circuit blocks 3303a, 3303b, 3303c, 3303d, 3303e, 3303f, 3303g and 3303h shown or neuron output blocks such as Figure 33 Neuron output blocks 3302a, 3302b, 3302c, 3302d, 3302e, 3302f, 3302g, and 3302h are shown.
[0223] Figure 37 Pulse-to-voltage converter 3700 is shown, which can optionally be used to convert digital pulses generated by integrating dual-slope ADC 3400 or 3500 into a voltage that can be applied as an input to a VMM memory array (e.g., on the WL or CG line). Pulse-to-voltage converter 3700 includes a reference current generator 3701 that generates a reference current IREF, a capacitor 3702, and a switch 3703. The input is used to control switch 3703. When a pulse is received at the input, the switch closes and charge accumulates on capacitor 3702, so that the voltage of capacitor 3702 after the input signal is completed will indicate the number of pulses received. The capacitor can optionally be a word line or control gate capacitor.
[0224] Figure 38 A current-to-voltage converter 3800 is shown, which can optionally be used to convert the neuron output current into a voltage that can be applied, for example, as an input to a VMM memory array (e.g., on a WL or CG line). The current-to-voltage converter 3800 includes a current generator 3801, which here represents the received neuron current Ineu (or Iin), and a variable resistor 3802. The output Vout will increase as the neuron current increases. The variable resistor 3802 can be adjusted as needed to increase or decrease the maximum range of Vout.
[0225] Figure 39 A current-to-voltage converter 3900 is shown, which can optionally be used to convert the neuron output current into a voltage that can be applied, for example, as an input to a VMM memory array (e.g., on a WL or CG line). The current-to-voltage converter 3900 includes an operational amplifier 3901, a capacitor 3902, a switch 3903, a switch 3904, and a current source 3905, here representing the neuron current ICELL. During operation, switch 3903 will be open and switch 3904 will be closed. The amplitude of the output Vout will increase in direct proportion to the magnitude of the neuron current ICELL 3905.
[0226] Figure 40A current-to-log voltage converter 4000 is shown, which can optionally be used to convert a neuron output current into a logarithmic voltage that can be applied, for example, as an input to a VMM memory array (e.g., on a WL or CG line). The current-to-log voltage converter 4000 includes a memory cell 4001, a switch 4002 (which selectively connects the wordline terminal of the memory cell 4001 to a node generating Vout), and a current source 4003, here representing the neuron current Iin. During operation, the switch 4002 will be closed, and the amplitude of the output Vout will increase proportionally to the magnitude of the neuron current iIN.
[0227] Figure 41 A current-to-log voltage converter 4100 is shown, which can optionally be used to convert the neuron output current into a logarithmic voltage that can be applied, for example, as an input to a VMM memory array (e.g., on a WL or CG line). The current-to-log voltage converter 4100 includes a memory cell 4101, a switch 4102 (which selectively connects the control gate terminal of the memory cell 4101 to a node generating Vout), and a current source 4103, here representing the neuron current Iin. During operation, the switch 4102 will be closed, and the amplitude of the output Vout will increase proportionally to the magnitude of the neuron current Iin.
[0228] Figure 42 A digital data-to-voltage converter 4200 is shown, which can optionally be used to convert digital data (i.e., data of 0s and 1s) into a voltage that can be applied, for example, as an input to a VMM memory array (e.g., on a WL or CG line). Digital data-to-voltage converter 4200 includes a capacitor 4201, an adjustable current source 4202 (here, the current from a reference array of memory cells), and a switch 4203. The digital data controls switch 4203. For example, switch 4203 can be closed when the digital data is a "1" and open when the digital data is a "0." The voltage accumulated on capacitor 4201 will be the output OUT and will correspond to the value of the digital data. Optionally, the capacitor can be a word line or control gate capacitor.
[0229] Figure 43 A digital data-to-voltage converter 4300 is shown, which can optionally be used to convert digital data (i.e., data representing 0s and 1s) into a voltage that can be applied, for example, as an input to a VMM memory array (e.g., on a WL or CG line). The digital data-to-voltage converter 4300 includes a variable resistor 4301, an adjustable current source 4302 (here, the current from a reference array of memory cells), and a switch 4303. The digital data controls the switch 4303. For example, the switch 4303 can be closed when the digital data is a "1" and open when the digital data is a "0." The output voltage will correspond to the value of the digital data.
[0230] Figure 44 A reference array 4400 is shown, which may be used to provide Figure 42 and 43 The reference current of the adjustable current sources 4202 and 4302 in .
[0231] Figures 45 to 47 Components are shown for verifying, after a program operation, that a flash memory cell in a VMM contains the proper charge corresponding to the W value expected to be stored in the flash memory cell.
[0232] Figure 45 A digital comparator 4500 is shown that receives as digital inputs a set of reference W values and sensed W digital values from a plurality of programmed flash memory cells. If a mismatch exists, the digital comparator 4500 generates a flag, which would indicate that one or more flash memory cells have not been programmed with the correct value.
[0233] Figure 46 shows the cooperation with the converter 4600 Figure 45 The sensed W values are provided by multiple instantiations of converter 4600. Converter 4600 receives cell current ICELL from the flash memory cell and converts the cell current into digital data, which can be provided to digital comparator 4500 using one or more of the aforementioned converters, such as ADC 3400 or 3500.
[0234] Figure 47 An analog comparator 4700 is shown that receives as analog inputs a set of reference W values and sensed W analog values from a plurality of programmed flash memory cells. If a mismatch exists, the analog comparator 4700 generates a flag, which would indicate that one or more flash memory cells have not been programmed with the correct value.
[0235] Figure 48 shows the cooperation with the converter 4800 Figure 47 The sensed W values are provided by converter 4800. Converter 4800 receives the digital values of the sensed W values and converts them into analog signals, which can be provided to analog comparator 4700 using one or more of the previously described converters (such as pulse-to-voltage converter 3700, digital data-to-voltage converter 4200, or digital data-to-voltage converter 4300).
[0236] Figure 49 Shown is output circuit 4900. It will be appreciated that if the output of a neuron is digitized (such as by using an integrating dual slope ADC 3400 or 3500 as previously described), it may still be necessary to perform an activation function operation on the neuron output. Figure 49An embodiment is shown in which activation occurs before the neuron output is converted into a variable-width pulse or pulse train. Output circuit 4900 includes activation circuit 4901 and current-to-pulse converter 4902. The activation circuit receives Ineuron values from various flash memory cells and generates Ineuron_act, which is the sum of the received Ineuron values. Current-to-pulse converter 4902 then converts Ineuron_act into a series of digital pulses and / or digital data representing a count of the series of digital pulses. Other converters previously described (such as integrating dual-slope ADC 3400 or 3500) can be used in place of converter 4902.
[0237] In another embodiment, activation may occur after the digital pulse is generated. In this embodiment, the digital output bits are mapped to a new set of digital bits using an activation mapping table or function implemented by activation mapping unit 5010. An example of such a mapping is shown in FIG. Figure 50 and Figure 51 The activation number map can simulate sigmoid, tanh, ReLu, or any activation function. In addition, the activation number map can quantize the output neurons.
[0238] Figure 52 An example of a charge summer 5200 is shown that can be used to sum the output of a VMM during a verify operation following a program operation to obtain a single analog value that represents the output and can then optionally be converted to a digital bit value. The charge summer 5200 includes a current source 5201 and a sample-and-hold circuit that includes a switch 5202 and a sample-and-hold (S / H) capacitor 5203. As shown in the example for a 4-bit digital value, there are 4 S / H circuits to hold the values from 4 evaluation pulses, where the values are summed at the end of the process. The S / H capacitor 5203 is selected to have a ratio associated with the 2^n*DINn bit position of the S / H capacitor; for example, C_DIN3 = x8 Cu, C_DIN2 = x4 Cu, C_DIN1 = x2 Cu, DIN0 = x1 Cu. The current source 5201 is also scaled accordingly.
[0239] Figure 53A current summator 5300 is shown that can be used to sum the output of the VMM during a verify operation following a program operation. Current summator 5300 includes a current source 5301, a switch 5302, switches 5303 and 5304, and a switch 5305. As shown for the example of a 4-bit digital value, a current source circuit is present to maintain the values from the 4 evaluation pulses, where these values are summed at the end of the process. The current sources are scaled based on 2^n*DINn bit positions; for example, I_DIN3 = x8 Icell units, I_DIN2 = x4 Icell units, I_DIN1 = x2 Icell units, and I_DIN0 = x1 Icell unit.
[0240] Figure 54 A digital summer 5400 is shown that receives multiple digital values, sums them, and generates an output, DOUT, that represents the sum of the inputs. The digital summer 5400 can be used during a verify operation following a program operation. As shown in the example for a 4-bit digital value, a digital output bit exists to hold the values from the 4 evaluation pulses, where these values are summed at the end of the process. The digital outputs are digitally scaled based on 2^n*DINn bit positions; for example, DOUT3 = x8 DOUT0, I_DOUT2 = x4 DOUT1, I_DOUT1 = x2 DOUT0, and I_DOUT0 = DOUT0.
[0241] Figure 55A and 55B A digital bit to pulse width converter 5500 is shown to be used within an input block, a row decoder, or an output block. The pulse width output from the digital bit to pulse width converter 5500 is similar to the pulse width output described above with respect to Figure 36B The digital bit-to-pulse width converter includes a binary counter 5501. The state Q[N:0] of the binary counter 5501 can be loaded by serial or parallel data in a load sequence. Row control logic 5510 outputs a voltage pulse having a pulse width proportional to the value of the digital data input provided by a block such as the integrating ADC in Figures 34 and 35.
[0242] Figure 55B The waveform of the output pulse width is shown, which has a width proportional to its digital bit value. First, the data in the received digital bits is inverted and the inverted digital bits are loaded serially or in parallel into the counter 5501. The row pulse width is then generated by the row control logic 5510, which is shown as waveform 5520 by counting in binary fashion until it reaches the maximum counter value.
[0243] Optionally, a pulse train to pulse converter may be used to convert a pulse train comprising a pulse train such as Figure 34BSignals 3411 or 3413 and Figure 35B The output of the signal 3513 or 3516 in the pulse sequence is converted into a single pulse (such as Figure 36B The pulse train is used as an input to the VMM array to be applied to word lines or control gates within the VMM array. An example of a pulse train to pulse converter is a binary counter with control logic.
[0244] An example of a 4-bit digital input is shown in Table 9:
[0245] Table 9: Digital input bit to output pulse width
[0246] FROM<3:0> count Inverted DIN<3:0> loaded into the counter Output pulse width = #clks 0000 0 1111 0 0001 1 1110 1 0010 2 1101 2 0011 3 1100 3 0100 4 1011 4 0101 5 1010 5 0110 6 1001 6 0111 7 1000 7 1000 8 0111 8 1001 9 0110 9 1010 10 0101 10 1011 11 0100 11 1100 12 0011 12 1101 13 0010 13 1110 14 0001 14 1111 15 0000 15
[0247] Another embodiment uses an up binary counter and digital comparison logic. That is, the output pulse width is generated by counting up the binary counter until the digital output of the binary counter is the same as the digital input bit.
[0248] Another embodiment uses a down binary counter. First, the down binary counter is loaded with a digital data input pattern in series or in parallel. The output pulse width is then generated by counting down the down binary counter until the digital output of the binary counter reaches a minimum value (i.e., a "0" logic state).
[0249] It should be noted that, as used herein, the terms "above" and "on" both inclusively include "directly on" (no intervening material, element, or space disposed therebetween) and "indirectly on" (intervening material, element, or space disposed therebetween). Similarly, the term "adjacent" includes "directly adjacent" (no intervening material, element, or space disposed therebetween) and "indirectly adjacent" (intervening material, element, or space disposed therebetween), "mounted to" includes "directly mounted to" (no intervening material, element, or space disposed therebetween) and "indirectly mounted to" (intervening material, element, or space disposed therebetween), and "electrically coupled to" includes "directly electrically coupled to" (no intervening material or element electrically connecting the elements together) and "indirectly electrically coupled to" (intervening material or element electrically connecting the elements together). For example, forming an element "above a substrate" may include forming the element directly on the substrate without an intervening material / element therebetween, as well as forming the element indirectly on the substrate with one or more intervening materials / elements therebetween.
Claims
1. A vector-matrix multiplication system comprising: an array of non-volatile memory cells arranged in rows and columns; an input block coupled to the array for receiving one or more input pulses, converting the one or more input pulses into an analog voltage, and applying the analog voltage to a word line or a control gate line of a row of nonvolatile memory cells in the array during operation of the vector-matrix multiplier; and an output block coupled to the array for generating digital pulses of varying width in response to neuronal currents drawn by the array during operation of the vector-matrix multiplier, wherein the width of the digital pulse varies in direct proportion to the neuronal current; The output block includes: a first operational amplifier comprising an inverting input terminal, a non-inverting input terminal, and an output terminal; a reference current source for generating a reference current at the output; a first switch selectively coupling the output of the reference current source to the inverting input terminal of the first operational amplifier; a second switch that selectively couples the neuron current to the inverting input terminal of the first operational amplifier; a capacitor coupled between the inverting input terminal of the first operational amplifier and the output terminal of the first operational amplifier; a third switch selectively coupling the inverting input terminal of the first operational amplifier and the output terminal of the first operational amplifier; and A second operational amplifier includes a non-inverting input terminal coupled to the output of the first operational amplifier, an inverting input terminal, and an output terminal for generating a digital pulse in response to the neuronal current.
2. The vector-matrix multiplication system according to claim 1 further includes an AND gate arranged to receive the output of the second operational amplifier and a clock for generating a series of pulses, wherein the number of pulses generated varies in proportion to the neuronal current. 3 . The vector-matrix multiplication system according to claim 2 , further comprising a counter configured to count the series of pulses to generate a count value comprising digital bits.
4. The vector-matrix multiplication system of claim 1 , wherein the input block comprises a digital-to-analog converter.
5. The vector-matrix multiplication system of claim 4, wherein the digital-to-analog converter comprises a pulse-to-voltage converter.
6. The vector-matrix multiplication system of claim 1, wherein each nonvolatile memory cell in the array of nonvolatile memory cells is a split-gate flash memory cell.
7. The vector-matrix multiplication system of claim 1, wherein each nonvolatile memory cell in the array of nonvolatile memory cells is a stacked gate flash memory cell.
Citation Information
Patent Citations
Deep learning neural network classifier using non-volatile memory array
US11308383B2
Deep Learning Neural Network Classifier Using Non-volatile Memory Array
US20170337466A1
High Precision And Highly Efficient Tuning Mechanisms And Algorithms For Analog Neuromorphic Memory In Artificial Neural Networks
US20190164617A1
Single transistor non-valatile electrically alterable semiconductor memory device
US5029130A
Flash memory cells with separated self-aligned select and erase gates, and process of fabrication
US6747310B2