Output circuitry for analog neural memory in deep learning artificial neural network
Non-volatile memory arrays in artificial neural networks address the inefficiencies of CMOS synapses by enabling efficient, precise tuning of synaptic weights and reducing power consumption through integrated memory operations.
Patent Information
- Application Number
- JP2025064254
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-11-08
- Filing Date
- 2025-04-09
- Publication Date
- 2025-08-20
AI Technical Summary
Existing artificial neural networks face challenges in high-performance information processing due to the lack of suitable hardware technology, particularly in terms of energy efficiency and scalability, as CMOS-implemented synapses are too bulky for large numbers of neurons and synapses.
Utilizing non-volatile memory arrays as synapses in artificial neural networks, allowing for continuous programming and reading of memory states with minimal disturbance, and implementing vector-matrix multiplication arrays to perform computations efficiently.
This approach enables precise tuning of synaptic weights, reduces power consumption, and enhances computational efficiency by integrating memory operations directly within the neural network architecture.
Smart Images

Figure 2025121903000001_ABST
Abstract
Description
[Technical Field]
[0001] (Priority Claim) This application claims priority to U.S. Provisional Patent Application No. 63 / 228,529, filed August 2, 2021, entitled "Output Circuitry for Analog Neural Memory in a Deep Learning Artificial Neural Network," and U.S. Patent Application No. 17 / 521,772, filed November 8, 2021, entitled "Output Circuitry for Analog Neural Memory in a Deep Learning Artificial Neural Network."
[0002] FIELD OF THE INVENTION Numerous embodiments are disclosed for an output circuit for an analog neural memory in a deep learning artificial neural network. [Background technology]
[0003] Artificial neural networks mimic biological neural networks (the central nervous systems of animals, particularly the brain) and are used to estimate or approximate functions that may depend on multiple inputs and are generally unknown. Artificial neural networks generally contain layers of interconnected "neurons" that exchange messages between each other.
[0004] FIG. 1 shows an artificial neural network, where circles illustrate layers of inputs or neurons. Connections (called synapses) are represented by arrows and have numerical weights that can be tuned based on experience. This allows the neural network to adapt to the inputs and learn. Typically, a neural network contains multiple layers of inputs. There are typically one or more hidden layers of neurons, and an output layer of neurons that provide the neural network's output. At each level, neurons make decisions individually or collectively based on the data received from the synapses.
[0005] One of the major challenges in developing artificial neural networks for high-performance information processing is the lack of suitable hardware technology. In practice, practical neural networks rely on a very large number of synapses, which allows for high connectivity between neurons and therefore a very high degree of parallelization of computation. In principle, such complexity could be achieved using digital supercomputers or dedicated graphic processing unit clusters. However, in addition to high cost, these approaches also suffer from poor energy efficiency compared to biological networks, which primarily perform low-precision analog computations and therefore consume much less energy. While CMOS analog circuits have been used in artificial neural networks, most CMOS-implemented synapses are too bulky given the large number of neurons and synapses.
[0006] The applicant previously disclosed an artificial (analog) neural network utilizing one or more non-volatile memory arrays as synapses in U.S. Patent Application No. 15 / 594,439, which is incorporated by reference. The non-volatile memory array operates as an analog neural memory. The neural network device includes a first plurality of synapses configured to receive a first plurality of inputs and generate a first plurality of outputs therefrom, and a first plurality of neurons configured to receive the first plurality of outputs. The first plurality of synapses includes a plurality of memory cells, each including spaced apart source and drain regions formed in a semiconductor substrate with a channel region extending therebetween, a floating gate disposed insulated above a first portion of the channel region, and a non-floating gate disposed insulated above a second portion of the channel region. Each of the plurality of memory cells is configured to store a weight value corresponding to a number of electrons in the floating gate. The plurality of memory cells is configured to multiply the first plurality of inputs by the stored weight value to generate the first plurality of outputs. Non-volatile memory cell
[0007] Nonvolatile memory is well known. For example, U.S. Pat. No. 5,029,130 (the "'130 patent"), incorporated herein by reference, discloses an array of split-gate nonvolatile memory cells, which are a type of flash memory cell. Such a memory cell 210 is shown in FIG. 2. Each memory cell 210 includes a source region 14 and a drain region 16 formed in a semiconductor substrate 12, with a channel region 18 between the source region 14 and the drain region 16. A floating gate 20 is formed over and insulated from a first portion of the channel region 18 (and controls the conductivity of the first portion of the channel region 18) and over a portion of the source region 14. A word line terminal 22 (typically coupled to a word line) has a first portion disposed over and insulated from a second portion of the channel region 18 (and controls the conductivity of the second portion of the channel region 18), and a second portion extending upward above the floating gate 20. A floating gate 20 and a wordline terminal 22 are insulated from the substrate 12 by a gate oxide. A bitline 24 is coupled to the drain region 16.
[0008] The memory cell 210 is erased (electrons are removed from the floating gate) by applying a high positive voltage to the word line terminal 22, which causes electrons in the floating gate 20 to pass from the floating gate 20 to the word line terminal 22 through the insulator between them via Fowler-Nordheim (FN) tunneling.
[0009] The memory cell 210 is programmed by hot electron source side injection (SSI) (electrons are added to the floating gate) by applying a positive voltage to the word line terminal 22 and a positive voltage to the source region 14. Electrons flow from the drain region 16 toward the source region 14. The electrons accelerate and heat up when they reach the gap between the word line terminal 22 and the floating gate 20. Some of the heated electrons are injected into the floating gate 20 through the gate oxide due to electrostatic attraction from the floating gate 20.
[0010] The memory cell 210 is read by applying a positive read voltage to the drain region 16 and word line terminal 22 (turning on the portion of the channel region 18 below the word line terminal). When the floating gate 20 is positively charged (i.e., erased with electrons), the portion of the channel region 18 below the floating gate 20 is also turned on, and current flows through the channel region 18, which is sensed as an erased or "1" state. When the floating gate 20 is negatively charged (i.e., programmed with electrons), the portion of the channel region below the floating gate 20 is mostly or completely off, and no (or very little) current flows through the channel region 18, which is sensed as a programmed or "0" state.
[0011] Table 1 shows typical voltage / current ranges that may be applied to the terminals of memory cell 110 to perform read, erase, and program operations. Table 1: Operation of flash memory cell 210 of FIG. 3 [Table 1]
[0012] Other split-gate memory cell configurations, including other types of flash memory cells, are also known. For example, FIG. 3 shows a four-gate memory cell 310 including a source region 14, a drain region 16, a floating gate 20 above a first portion of a channel region 18, a select gate 22 (typically coupled to a word line, WL) above a second portion of the channel region 18, a control gate 28 above the floating gate 20, and an erase gate 30 above the source region 14. This configuration is described in U.S. Pat. No. 6,747,310, which is incorporated herein by reference for all purposes. Here, all gates, except for the floating gate 20, are non-floating gates, meaning they are electrically connected or connectable to a voltage source. Programming is performed by heated electrons injecting themselves from the channel region 18 into the floating gate 20. Erasing is performed by electrons tunneling from the floating gate 20 to the erase gate 30.
[0013] Table 2 shows typical voltage / current ranges that may be applied to the terminals of memory cell 310 to perform read, erase, and program operations. Table 2: Operation of the flash memory cell 310 of FIG. 3 [Table 2]
[0014] Figure 4 shows another type of flash memory cell, a three-gate memory cell 410. Memory cell 410 is identical to memory cell 310 of Figure 3, except that memory cell 410 does not have a separate control gate. Erase and read operations (erasure occurs through the use of an erase gate) are similar to those of Figure 3, except that no control gate bias is applied. Programming operations are also performed without a control gate bias, and as a result, a higher voltage must be applied to the source line during a program operation to compensate for the lack of control gate bias.
[0015] Table 3 shows typical voltage / current ranges that may be applied to the terminals of memory cell 410 to perform read, erase, and program operations. Table 3: Operation of flash memory cell 410 of FIG. 4 [Table 3]
[0016] 5 shows another type of flash memory cell, a stacked gate memory cell 510. Memory cell 510 is similar to memory cell 210 of FIG. 2, except that the floating gate 20 extends over the entire channel region 18, and a control gate 22 (where it is coupled to a word line) extends over the floating gate 20, separated by an insulating layer (not shown). Erasing is accomplished by FN tunneling of electrons from the FG to the substrate, and programming is accomplished by channel hot electron (CHE) injection in the region between the channel 18 and the drain region 16, by electrons flowing from the source region 14 toward the drain region 16, and by a read operation similar to that of memory cell 210, which has a higher control gate voltage.
[0017] Table 4 shows typical voltage ranges that may be applied to the terminals of memory cell 510 and substrate 12 to perform read, erase, and program operations. Table 4: Operation of flash memory cell 510 of FIG. 5 [Table 4]
[0018] The methods and means described herein may be applied to other non-volatile memory technologies such as, but not limited to, FINFET split-gate flash or stacked-gate flash memory, NAND flash, SONOS (silicon-oxide-nitride-oxide-silicon, charge traps in nitride), MONOS (metal-oxide-nitride-oxide-silicon, metal charge traps in nitride), ReRAM (resistive ram), PCM (phase change memory), MRAM (magnetic ram), FeRAM (ferroelectric ram), CT (charge trap) memory, CN (carbon-tube) memory, OTP (one time programmable), and CeRAM (correlated electron ram).
[0019] In order to utilize a memory array containing one of the non-volatile memory cell types in an artificial neural network as described above, two modifications are made. First, as explained further below, the lines are configured so that each memory cell can be individually programmed, erased, and read without adversely affecting the memory state of other memory cells in the array. Second, continuous (analog) programming of the memory cells is provided.
[0020] Specifically, the memory state (i.e., the charge on the floating gate) of each memory cell in the array can be changed continuously from a fully erased state to a fully programmed state, independently and with minimal disturbance to other memory cells. In another embodiment, the memory state (i.e., the charge on the floating gate) of each memory cell in the array can be changed continuously from a fully programmed state to a fully erased state, and vice versa, independently and with minimal disturbance to other memory cells. This means that the cell storage is analog, or at a minimum, capable of storing one of a number of discrete values (such as 16 or 64 different values), making every cell in the memory array very precisely and individually tunable and making the memory array ideal for storage and for fine tuning adjustments to the synaptic weights of neural networks. Neural networks using nonvolatile memory cell arrays
[0021] 6 conceptually illustrates a non-limiting example of a neural network utilizing the non-volatile memory array of the present embodiments. This example uses a non-volatile memory array neural network for a face recognition application, although other suitable applications can also be implemented using a non-volatile memory array-based neural network.
[0022] S0 is the input layer, which in this example is a 32x32 pixel RGB image with 5-bit precision (i.e., three 32x32 pixel arrays, one for each color R, G, and B, with each pixel having 5-bit precision). Synapse CB1 going from input layer S0 to layer C1 scans the input image with overlapping 3x3 pixel filters (kernels), applying different sets of weights to some instances and shared weights to other instances, and shifts the filters by one pixel (or two or more pixels, depending on the model). Specifically, the values of nine pixels in the 3x3 portion of the image (i.e., referred to as filters or kernels) are provided to synapse CB1, which multiplies these nine input values by the appropriate weights and, after summing the outputs of the multiplications, determines a single output value, which is applied by the first synapse of CB1 to generate one pixel of layer C1's feature map. The 3x3 filter is then shifted one pixel to the right in input layer S0 (i.e., adding a column of three pixels to the right and dropping a column of three pixels on the left), so that the nine pixel values of this newly positioned filter are provided to synapse CB1, where they are multiplied by the same weights as above to determine a second single output value by the associated synapse. This process continues until the 3x3 filter has scanned the entire 32x32 pixel image of input layer S0 for all three colors and all bits (precision values). The process is then repeated using different sets of weights to generate different feature maps for layer C1 until all of layer C1's feature maps have been calculated.
[0023] In this example, there are 16 feature maps in layer C1, each having 30x30 pixels. Each pixel is a new feature pixel extracted from the multiplication of the input and the kernel, and therefore each feature map is a two-dimensional array. Thus, in this example, layer C1 comprises 16 layers of two-dimensional arrays. (Note that the layers and arrays referred to herein are logical, not necessarily physical, relationships; i.e., the arrays are not necessarily oriented in a physical two-dimensional array.) Each of the 16 feature maps in layer C1 is generated by one of 16 different sets of synaptic weights applied to the filter scans. The C1 feature maps can all target different aspects of the same image feature, such as boundary identification. For example, a first map (generated using a first set of weights shared by all scans used to generate this first map) can identify circular edges, while a second map (generated using a second set of weights different from the first set of weights) can identify rectangular edges or the aspect ratio of a particular feature, etc.
[0024] Before going from layer C1 to layer S1, an activation function P1 (pooling) is applied, which pools values from non-overlapping, contiguous 2x2 regions in each feature map. The purpose of pooling function P1 is to average nearby locations (or a max function can be used), e.g., to reduce dependency on edge locations, and to reduce data size before going to the next stage. In layer S1, there are 16 15x15 feature maps (i.e., 16 different arrays of 15x15 pixels each). Synapse CB2 going from layer S1 to layer C2 scans the maps in layer S1 with a 4x4 filter with a filter shift of 1 pixel. In layer C2, there are 22 12x12 feature maps. Before going from layer C2 to layer S2, an activation function P2 (pooling) is applied, which pools values from non-overlapping, contiguous 2x2 regions in each feature map. In layer S2, there are 22 6x6 feature maps. At synapse CB3 going from layer S2 to layer C3, an activation function (pooling) is applied, where every neuron in layer C3 connects to every map in layer S2 through a respective synapse in CB3. There are 64 neurons in layer C3. Synapse CB4 going from layer C3 to output layer S3 fully connects C3 to S3, i.e., every neuron in layer C3 connects to every neuron in layer S3. The output at S3 includes 10 neurons, where the neuron with the highest output determines the class. This output can indicate, for example, the identification or classification of the content of the original image.
[0025] Each layer of the synapse is implemented using an array or portion of an array of non-volatile memory cells.
[0026] Figure 7 is a block diagram of an array that can be used for this purpose. A vector-by-matrix multiplication (VMM) array 32 contains nonvolatile memory cells and is utilized as a synapse between one layer and the next (such as CB1, CB2, CB3, and CB4 in Figure 6). Specifically, the VMM array 32 includes an array of nonvolatile memory cells 33, an erase gate and word line gate decoder 34, a control gate decoder 35, a bit line decoder 36, and a source line decoder 37, which decode the respective inputs to the nonvolatile memory cell array 33. Inputs to the VMM array 32 can come from the erase gate and word line gate decoder 34 or from the control gate decoder 35. The source line decoder 37 in this example also decodes the output of the nonvolatile memory cell array 33. Alternatively, the bit line decoder 36 can decode the output of the nonvolatile memory cell array 33.
[0027] The non-volatile memory cell array 33 serves two purposes. First, it stores the weights used by the VMM array 32. Second, the non-volatile memory cell array 33 effectively multiplies the inputs by the weights stored in the non-volatile memory cell array 33 and adds them for each output line (source line or bit line) to produce an output that becomes the input to the next layer or the input to the last layer. Having the non-volatile memory cell array 33 perform the multiplication and addition functions eliminates the need for separate multiplication and addition logic and is also more power efficient due to in-memory computation.
[0028] The outputs of the non-volatile memory cell array 33 are fed to a differential summer (such as a summing op-amp or summing current mirror) 38, which sums the outputs of the non-volatile memory cell array 33 to create a single value for the convolution. The differential summer 38 is arranged to perform a summation of the positive and negative weights.
[0029] The summed output values of the differential adder 38 are then provided to an activation function block 39, which rectifies the output. The activation function block 39 may provide a sigmoid, tanh, or ReLU function. The rectified output values of the activation function block 39 become elements of a feature map as the next layer (e.g., C1 in FIG. 6) and are then applied to the next synapse to generate the next feature map layer or the final layer. Thus, in this example, the non-volatile memory cell array 33 constitutes multiple synapses (receiving input from a previous layer of neurons or from an input layer such as an image database), and the summing op-amps 38 and activation function block 39 constitute multiple neurons.
[0030] The inputs to the VMM array 32 of FIG. 7 (WLx, EGx, CGx, and optionally BLx and SLx) may be analog levels, binary levels, or digital bits (in which case a DAC is provided to convert the digital bits to the appropriate input analog levels), and the outputs may be analog levels, binary levels, or digital bits (in which case an output ADC is provided to convert the output analog levels to digital bits).
[0031] FIG. 8 is a block diagram illustrating the use of multiple layers of VMM array 32, labeled in the figure as VMM arrays 32a, 32b, 32c, 32d, and 32e. As shown in FIG. 8, input (denoted Inputx) is converted from digital to analog by digital-to-analog converter 31 and provided to input VMM array 32a. The converted analog input can be a voltage or current. The first layer's input D / A conversion can be performed by using a function or LUT (look up table) that maps input Inputx to the appropriate analog level of the matrix multiplier of input VMM array 32a. The input conversion can also be performed by an analog-to-analog (A / A) converter to convert an external analog input to the mapped analog input to input VMM array 32a.
[0032] The output generated by input VMM array 32a is then provided as input to the next VMM array (hidden level 1) 32b, which then generates an output that is provided as input to input VMM array (hidden level 2) 32c, and so on. The various layers of the VMM array 32 function as layers of synapses and neurons of a convolutional neural network (CNN). Each VMM array 32a, 32b, 32c, 32d, and 32e can be a standalone physical non-volatile memory array, or multiple VMM arrays can utilize different portions of the same physical non-volatile memory array, or multiple VMM arrays can utilize overlapping portions of the same physical non-volatile memory array. The example shown in FIG. 8 includes five layers (32a, 32b, 32c, 32d, and 32e): one input layer (32a), two hidden layers (32b and 32c), and two fully connected layers (32d and 32e). Those skilled in the art will appreciate that this is merely an example, and that the system may alternatively include more than two hidden layers and more than two fully connected layers. Vector Matrix Multiplication (VMM) Array
[0033] 9 shows a neuron VMM array 900 that is particularly suited for the memory cells 310 shown in FIG. 3 and that is utilized as part of the synapses and neurons between the input layer and the next layer. The VMM array 900 includes a memory array 901 of non-volatile memory cells and a reference array 902 of non-volatile reference memory cells (located at the top of the array). Alternatively, a separate reference array can be located at the bottom.
[0034] In VMM array 900, control gate lines, such as control gate line 903, run vertically (thus, row-oriented reference array 902 is orthogonal to control gate line 903), and erase gate lines, such as erase gate line 904, run horizontally. Here, inputs to VMM array 900 are provided to control gate lines (CG0, CG1, CG2, CG3), and outputs of VMM array 900 appear on source lines (SL0, SL1). In one embodiment, only even rows are used, and in another embodiment, only odd rows are used. The current in each source line (SL0, SL1, respectively) performs a function of the sum of all currents from memory cells connected to that particular source line.
[0035] As described herein for neural networks, the non-volatile memory cells of VMM array 900, ie, memory cells 310 of VMM array 900, are preferably configured to operate in the sub-threshold region.
[0036] The nonvolatile reference memory cells and nonvolatile memory cells described herein are biased in weak inversion (sub-threshold region) as follows: Ids=Io * e (Vg-Vth) / nVt =w * Io * e (Vg) / nVt In the formula, w=e (-Vth) / nVt and Ids is the drain-source current, Vg is the gate voltage of the memory cell, Vth is the threshold voltage of the memory cell, and Vt is the thermal voltage = k * where T / q, k is Boltzmann's constant, T is temperature in Kelvin, q is electron charge, n is slope coefficient = 1 + (Cdep / Cox), where Cdep = capacitance of the depletion layer, and Cox is capacitance of the gate oxide layer, Io is memory cell current at gate voltage equal to threshold voltage, and Io is (Wt / L) * u * Cox * (n-1) * Vt 2where u is the carrier mobility, and Wt and L are the width and length of the memory cell, respectively.
[0037] When using an IV-log converter that converts input current to input voltage using a memory cell (such as a reference memory cell or peripheral memory cell) or transistor: Vg=n * Vt * log[Ids / wp * Io] where wp is the w of the reference or peripheral memory cell.
[0038] For a memory array used as a vector matrix multiplier VMM array with current inputs, the output current is: Iout=wa * Io * e (Vg) / nVt , i.e. Iout=(wa / wp) * Iin=W * Iin W=e (Vthp-Vtha) / nVt where wa=w of each memory cell in the memory array. Vthp is the effective threshold voltage of the peripheral memory cells, and Vtha is the effective threshold voltage of the main (data) memory cells. Note that the threshold voltage of a transistor is a function of the substrate body bias voltage, which is represented as Vsb, and can be modulated to compensate for various conditions at such temperature. The threshold voltage Vth can be expressed as: TIFF2025121903000006.tif22170
[0039] The word line or control gate can be used as the input of the memory cell for the input voltage.
[0040] Alternatively, the flash memory cells of the VMM arrays described herein can be configured to operate in the linear region. Ids=Beta * (Vgs-Vth) *Vds; beta = u * Cox * Wt / L W ∝ (Vgs-Vth) That is, the weight W in the linear region is proportional to (Vgs-Vth)
[0041] The word line or control gate or bit line or source line can be used as the input of a memory cell operating in the linear region, and the bit line or source line can be used as the output of the memory cell.
[0042] For the IV linear converter, memory cells (such as reference or peripheral memory cells) or transistors operating in the linear region can be used to linearly convert input and output currents to input and output voltages.
[0043] Alternatively, the memory cells of the VMM arrays described herein can be configured to operate in the saturation region. Ids=1 / 2 * beta * (Vgs-Vth) 2 ;beta=u * Cox * Wt / L W ∝ (Vgs-Vth) 2 , that is, the weight W is (Vgs-Vth) 2 is proportional to
[0044] The word line, control gate, or erase gate can be used as the input of a memory cell operating in the saturation region, and the bit line or source line can be used as the output of an output neuron.
[0045] Alternatively, the memory cells of the VMM arrays described herein may be used in all domains or combinations thereof (subthreshold, linear, or saturation) for each layer or layers of a neural network.
[0046] 7 is described in U.S. Patent No. 10,748,630, which is incorporated herein by reference. As described in that application, the source lines or bit lines can be used as neuron outputs (current sum outputs).
[0047] FIG. 10 shows a neuron VMM array 1000 that is particularly suited for the memory cells 210 shown in FIG. 2 and is utilized as a synapse between an input layer and the next layer. The VMM array 1000 includes a memory array 1003 of nonvolatile memory cells, a reference array 1001 of first nonvolatile reference memory cells, and a reference array 1002 of second nonvolatile reference memory cells. The reference arrays 1001 and 1002, arranged in columns of the array, function to convert current inputs flowing into terminals BLR0, BLR1, BLR2, and BLR3 into voltage inputs WL0, WL1, WL2, and WL3. In practice, the first and second nonvolatile reference memory cells are diode-connected through a multiplexer 1014 (only partially shown) with current inputs flowing into them. The reference cells are tuned (e.g., programmed) to a target reference level, which is provided by a reference mini-array matrix (not shown).
[0048] Memory array 1003 serves two purposes. First, memory array 1003 stores weights in each memory cell that are used by VMM array 1000. Second, memory array 1003 effectively multiplies the inputs (i.e., the current inputs provided to terminals BLR0, BLR1, BLR2, and BLR3, which reference arrays 1001 and 1002 convert to input voltages provided to word lines WL0, WL1, WL2, and WL3) by the weights stored in memory array 1003, and then adds all the results (memory cell currents) to produce outputs on the respective bit lines (BL0-BLN), which serve as inputs to the next layer or the last layer. By performing the multiplication and addition functions, memory array 1003 eliminates the need for separate multiplication and addition logic and is also power efficient. Here, voltage inputs are provided to word lines WL0, WL1, WL2, and WL3, and outputs appear on respective bit lines BL0-BLN during a read (inference) operation. The current in each of the bit lines BL0-BLN performs a function of the sum of the currents from all the non-volatile memory cells connected to that particular bit line.
[0049] Table 5 shows the operating voltages and currents for the VMM array 1000. The columns in the table indicate the voltages applied to the word line of the selected cell, the word lines of the unselected cells, the bit lines of the selected cell, the bit lines of the unselected cells, the source lines of the selected cell, and the source lines of the unselected cells. The rows indicate the read, erase, and program operations. Table 5: Operation of VMM Array 1000 in Figure 10 [Table 5]
[0050] FIG. 11 shows a neuron VMM array 1100 that is particularly suited for the memory cells 210 shown in FIG. 2 and that is utilized as part of the synapses and neurons between the input layer and the next layer. The VMM array 1100 includes a memory array 1103 of nonvolatile memory cells, a reference array 1101 of first nonvolatile reference memory cells, and a reference array 1102 of second nonvolatile reference memory cells. The reference arrays 1101 and 1102 extend in the row direction of the VMM array 1100. The VMM array is similar to the VMM 1000, except that the word lines extend vertically in the VMM array 1100. Here, inputs are provided to the word lines (WLA0, WLB0, WLA1, WLB2, WLA2, WLB2, WLA3, WLB3), and outputs appear on the source lines (SL0, SL1) during a read operation. The current in each source line performs a function of the sum of all the currents from the memory cells connected to that particular source line.
[0051] Table 6 shows the operating voltages and currents for VMM array 1100. The columns in the table indicate the voltages applied to the word line of the selected cell, the word lines of the unselected cells, the bit lines of the selected cell, the bit lines of the unselected cells, the source lines of the selected cell, and the source lines of the unselected cells. The rows indicate the read, erase, and program operations. Table 6: Operation of VMM Array 1100 in Figure 11 [Table 6]
[0052] 12 shows a neuron VMM array 1200 that is particularly suited for the memory cells 310 shown in FIG. 3 and that is utilized as part of the synapses and neurons between the input layer and the next layer. VMM array 1200 includes a memory array 1203 of nonvolatile memory cells, a reference array 1201 of first nonvolatile reference memory cells, and a reference array 1202 of second nonvolatile reference memory cells. Reference arrays 1201 and 1202 function to convert current inputs flowing into terminals BLR0, BLR1, BLR2, and BLR3 into voltage inputs CG0, CG1, CG2, and CG3. In effect, the first and second nonvolatile reference memory cells are diode-connected through multiplexer 1212 (only a portion of which is shown), with the current inputs flowing through BLR0, BLR1, BLR2, and BLR3. Multiplexer 1212 includes a corresponding multiplexer 1205 and cascoding transistor 1204 to ensure a constant voltage on the respective bit lines (e.g., BLR0) of the first and second non-volatile reference memory cells during each read operation, where the reference cells are tuned to a target reference level.
[0053] Memory array 1203 serves two purposes. First, memory array 1203 stores the weights used by VMM array 1200. Second, memory array 1203 effectively multiplies the weights stored in the memory array by the inputs (current inputs provided to terminals BLR0, BLR1, BLR2, and BLR3; reference arrays 1201 and 1202 convert these current inputs to input voltages provided to control gates (CG0, CG1, CG2, and CG3)), and then adds all the results (cell currents) to generate an output that appears on BL0-BLN and serves as the input to the next layer or the last layer. Having the memory array perform the multiplication and addition functions eliminates the need for separate multiplication and addition logic and is also power efficient. Here, the inputs are provided to the control gate lines (CG0, CG1, CG2, and CG3) and the outputs appear on the bit lines (BL0-BLN) during read operations. The current in each bit line is a function of the sum of all the currents from the memory cells connected to that particular bit line.
[0054] VMM array 1200 implements one-way tuning of the non-volatile memory cells in memory array 1203. That is, each non-volatile memory cell is erased and then partially programmed until the desired charge on the floating gate is reached. If too much charge is added to the floating gate (such as if an incorrect value is stored in the cell), the cell is erased and the series of partial programming operations starts over. As shown, two rows that share the same erase gate (such as EG0 or EG1) are erased together (known as a page erase), and then each cell is partially programmed until the desired charge on the floating gate is reached.
[0055] Table 7 shows the operating voltages and currents for VMM array 1200. The columns in the table indicate the voltages applied to the word line of the selected cell, the word lines of the unselected cells, the bit line of the selected cell, the bit lines of the unselected cells, the control gate of the selected cell, the control gates of the unselected cells in the same sector as the selected cell, the control gates of the unselected cells in a different sector from the selected cell, the erase gate of the selected cell, the erase gates of the unselected cells, the source line of the selected cell, and the source lines of the unselected cells. The rows indicate read, erase, and program operations. Table 7: Operation of VMM Array 1200 in Figure 12 [Table 7]
[0056] FIG. 13 shows a neuron VMM array 1300 that is particularly suited for the memory cells 310 shown in FIG. 3 and that is utilized as part of the synapses and neurons between the input layer and the next layer. The VMM array 1300 includes a memory array 1303 of nonvolatile memory cells, a reference array 1301 or first nonvolatile reference memory cells, and a reference array 1302 of second nonvolatile reference memory cells. EG lines EGR0, EG0, EG1, and EGR1 extend vertically, while CG lines CG0, CG1, CG2, and CG3 and SL lines WL0, WL1, WL2, and WL3 extend horizontally. The VMM array 1300 is similar to the VMM array 1400, except that the VMM array 1300 implements bidirectional tuning, meaning that each individual cell can be fully erased, partially programmed, and partially erased as needed to reach a desired amount of charge on the floating gate through the use of separate EG lines. As shown, reference arrays 1301 and 1302 convert input currents at terminals BLR0, BLR1, BLR2, and BLR3 into control gate voltages CG0, CG1, CG2, and CG3 (through the action of diode-connected reference cells via multiplexer 1314), which are applied to the memory cells in a row direction. The current outputs (neurons) are in bit lines BL0 through BLN, each bit line summing all the currents from the non-volatile memory cells connected to that particular bit line.
[0057] Table 8 shows the operating voltages and currents for VMM array 1300. The columns in the table indicate the voltages applied to the word line of the selected cell, the word lines of the unselected cells, the bit line of the selected cell, the bit lines of the unselected cells, the control gate of the selected cell, the control gates of the unselected cells in the same sector as the selected cell, the control gates of the unselected cells in a different sector from the selected cell, the erase gate of the selected cell, the erase gates of the unselected cells, the source line of the selected cell, and the source lines of the unselected cells. The rows indicate read, erase, and program operations. Table 8: Operation of VMM Array 1300 in Figure 13 [Table 8]
[0058] 22 shows a neuron VMM array 2200 that is particularly suited to the memory cells 210 shown in FIG. 2 and that is used as part of the synapses and neurons between the input layer and the next layer. In the VMM array 2200, inputs INPUT0..., INPUT N are bit lines BL0, ...BL N and outputs OUTPUT1, OUTPUT2, OUTPUT3, and OUTPUT4 are generated on source lines SL0, SL1, SL2, and SL3, respectively.
[0059] 23 shows a neuron VMM array 2300 that is particularly suited for memory cells 210 shown in FIG. 2 and that is utilized as part of the synapses and neurons between the input layer and the next layer. In this example, inputs INPUT0, INPUT1, INPUT2, and INPUT3 are received on source lines SL0, SL1, SL2, and SL3, respectively, and outputs OUTPUT0, ...OUTPUT N are the bit lines BL0, ..., BL N is generated.
[0060] 24 shows a neuron VMM array 2400 that is particularly suited for the memory cells 210 shown in FIG. 2 and that is utilized as part of the synapses and neurons between the input layer and the next layer. In this example, inputs INPUT0,..., INPUT M are the word lines WL0, ..., WL M Received and output OUTPUT0, ...OUTPUT N are the bit lines BL0, ..., BL N is generated.
[0061] 25 shows a neuron VMM array 2500 that is particularly suited for the memory cells 310 shown in FIG. 3 and that is utilized as part of the synapses and neurons between the input layer and the next layer. In this example, inputs INPUT0,..., INPUT M are the word lines WL0, ..., WL MReceived and output OUTPUT0, ...OUTPUT N are the bit lines BL0, ..., BL N is generated.
[0062] 26 shows a neuron VMM array 2600 that is particularly suited for the memory cells 410 shown in FIG. 4 and that is utilized as part of the synapses and neurons between the input layer and the next layer. In this example, the input INPUT 0、 ..., INPUT n are the vertical control gate lines CG0, ..., CG N and outputs OUTPUT1 and OUTPUT2 are generated on source lines SL0 and SL1.
[0063] 27 shows a neuron VMM array 2700 that is particularly suited for the memory cells 410 shown in FIG. 4 and that is utilized as part of the synapses and neurons between the input layer and the next layer. In this example, the inputs INPUT0, ..., INPUT N are the bit lines BL0, ..., BL N , 2701-(N-1) and 2701-N, which are coupled to the gates of the bit line control gates 2701-1, 2701-2, ..., 2701-(N-1) and 2701-N. Exemplary outputs OUTPUT1 and OUTPUT2 are generated on source lines SL0 and SL1.
[0064] 28 shows a neuron VMM array 2800 that is particularly suited for memory cells 310 shown in FIG. 3, memory cells 510 shown in FIG. 5, and memory cells 710 shown in FIG. 7, and is utilized as part of the synapses and neurons between the input layer and the next layer. In this example, inputs INPUT0, ..., INPUT M are the word lines WL0, ..., WL M Received and output OUTPUT0, ..., OUTPUT N are the bit lines BL0, ..., BL N is generated.
[0065] 29 shows a neuron VMM array 2900 that is particularly suitable for memory cells 310 shown in FIG. 3, memory cells 510 shown in FIG. 5, and memory cells 710 shown in FIG. 7, and is used as part of the synapses and neurons between the input layer and the next layer. In this example, inputs INPUT0, ..., INPUT M are the control gate lines CG0, ..., CG M Received by Output OUTPUT0, ..., OUTPUT N are the vertical source lines SL0, ..., SL N and each source line SL i is coupled to the source lines of all memory cells in column i.
[0066] 30 shows a neuron VMM array 3000 that is particularly suitable for memory cells 310 shown in FIG. 3, memory cells 510 shown in FIG. 5, and memory cells 710 shown in FIG. 7, and is used as part of the synapses and neurons between the input layer and the next layer. In this example, inputs INPUT0, ..., INPUT M are the control gate lines CG0, ..., CG M Received by Output OUTPUT0, ..., OUTPUT N are the vertical bit lines BL0, ..., BL N and each bit line BL i is coupled to the bit lines of all memory cells in column i. Long-term and short-term memory
[0067] Prior art includes a concept known as long short-term memory (LSTM). LSTM units are often used within neural networks. LSTM allows a neural network to store information for any predetermined period of time and use that information in subsequent operations. A traditional LSTM unit includes a cell, an input gate, an output gate, and a forget gate. The three gates regulate the flow of information into and out of the cell and the duration for which information is stored within the LSTM. VMMs are particularly useful in LSTM units.
[0068] Figure 14 shows an example LSTM 1400. LSTM 1400 in this example includes cells 1401, 1402, 1403, and 1404. Cell 1401 receives input vector x0 and generates output vector h0 and cell state vector c0. Cell 1402 receives input vector x1, output vector (hidden state) h0 from cell 1401, and cell state c0 from cell 1401, and generates output vector h1 and cell state vector c1. Cell 1403 receives input vector x2, output vector (hidden state) h1 from cell 1402, and cell state c1 from cell 1402, and generates output vector h2 and cell state vector c2. Cell 1404 receives input vector x3, output vector (hidden state) h2 from cell 1403, and cell state c2 from cell 1403, and generates output vector h3. Additional cells can be used; an LSTM with four cells is just an example.
[0069] Figure 15 shows an example implementation of an LSTM cell 1500 that can be used for cells 1401, 1402, 1403, and 1404 in Figure 14. LSTM cell 1500 receives an input vector x(t), a cell state vector c(t-1) from a previous cell, and an output vector h(t-1) from a previous cell, and produces a cell state vector c(t) and an output vector h(t).
[0070] LSTM cell 1500 includes sigmoid function devices 1501, 1502, and 1503, each of which applies a number between 0 and 1 to control the degree to which each component of the input vector contributes to the output vector. LSTM cell 1500 also includes tanh devices 1504 and 1505 for applying a hyperbolic tangent function to the input vector, multiplier devices 1506, 1507, and 1508 for multiplying two vectors, and adder device 1509 for adding the two vectors. The output vector h(t) can be provided to the next LSTM cell in the system or can be accessed for other purposes.
[0071] FIG. 16 shows LSTM cell 1600, an example of one implementation of LSTM cell 1500. For the convenience of the reader, the same numbering scheme from LSTM cell 1500 is used in LSTM cell 1600. Sigmoid function devices 1501, 1502, and 1503 and tanh device 1504 each include multiple VMM arrays 1601 and activation function blocks 1602. VMM arrays, therefore, prove particularly useful in LSTM cells used in certain neural network systems. Multiplier devices 1506, 1507, and 1508 and adder device 1509 are implemented in digital or analog fashion. Activation function block 1602 can be implemented in digital or analog fashion.
[0072] An alternative example of LSTM cell 1600 (and another example of one implementation of LSTM cell 1500) is shown in Figure 17. In Figure 17, sigmoid function devices 1501, 1502, and 1503 and tanh device 1504 share the same physical hardware (VMM array 1701 and activation function block 1702) in a time-multiplexed manner. LSTM cell 1700 also includes a multiplier device 1703 for multiplying two vectors, an adder device 1708 for adding two vectors, a tanh device 1505 (which includes activation function block 1702), a register 1707 for storing the value i(t) as it is output from sigmoid function block 1702, and a register 1708 for storing the value f(t). * a register 1704 for storing c(t-1) as its value is output from the multiplier device 1703 via multiplexer 1710; * a register 1705 for storing u(t) as its value is output from the multiplier device 1703 via a multiplexer 1710; * It includes a register 1706 for storing {tilde over (c)}(t) as its value is output from the multiplier device 1703 via a multiplexer 1710, and a multiplexer 1709.
[0073] While LSTM cell 1600 includes multiple sets of VMM arrays 1601 and respective activation function blocks 1602, LSTM cell 1700 includes only one set of VMM arrays 1701 and activation function blocks 1702, which are used to represent multiple layers in embodiments of LSTM cell 1700. LSTM cell 1700 requires one-quarter the space for the VMMs and activation function blocks compared to LSTM cell 1600, and therefore LSTM cell 1700 requires less space than LSTM 1600.
[0074] It can be further appreciated that an LSTM unit typically includes multiple VMM arrays, each of which requires functionality provided by specific circuit blocks outside the VMM array, such as adder and activation function blocks and high-voltage generation blocks. Providing a separate circuit block for each VMM array would require a significant amount of space within a semiconductor device and would be somewhat inefficient. Accordingly, the embodiments described below reduce the circuitry required outside the VMM array itself. Gated Recurrent Unit
[0075] Analog VMM implementations can be used for gated recurrent unit (GRU) systems. GRUs are gating mechanisms within recurrent neural networks. GRUs are similar to LSTMs, except that GRU cells generally contain fewer components than LSTM cells.
[0076] 18 shows an exemplary GRU 1800. GRU 1800 in this example includes cells 1801, 1802, 1803, and 1804. Cell 1801 receives input vector x0 and generates output vector h0. Cell 1802 receives input vector x1 and output vector h0 from cell 1801 and generates output vector h1. Cell 1803 receives input vector x2 and output vector (hidden state) h1 from cell 1802 and generates output vector h2. Cell 1804 receives input vector x3 and output vector (hidden state) h2 from cell 1803 and generates output vector h3. Additional cells can be used; a GRU with four cells is merely an example.
[0077] FIG. 19 shows an example implementation of a GRU cell 1900 that may be used for cells 1801, 1802, 1803, and 1804 of FIG. 18. GRU cell 1900 receives an input vector x(t) and an output vector h(t-1) from a preceding GRU cell and generates an output vector h(t). GRU cell 1900 includes sigmoid function devices 1901 and 1902, each of which applies a number between 0 and 1 to components from the output vector h(t-1) and the input vector x(t). GRU cell 1900 also includes a tanh device 1903 for applying a hyperbolic tangent function to the input vector, multiple multiplier devices 1904, 1905, and 1906 for multiplying two vectors, an adder device 1907 for adding the two vectors, and a complement device 1908 for subtracting the input from 1 to generate the output.
[0078] FIG. 20 shows GRU cell 2000, an example of one implementation of GRU cell 1900. For convenience of the reader, the same numbering scheme as GRU cell 1900 is used in GRU cell 2000. As can be seen from FIG. 20, sigmoid function devices 1901 and 1902 and tanh device 1903 each include multiple VMM arrays 2001 and activation function blocks 2002. Therefore, it can be seen that VMM arrays are particularly used in GRU cells used in specific neural network systems. Multiplier devices 1904, 1905, and 1906, adder device 1907, and complementary device 1908 are implemented in a digital or analog manner. Activation function block 2002 can be implemented in a digital or analog manner.
[0079] An alternative example of GRU cell 2000 (and another example of one implementation of GRU cell 1900) is shown in FIG. 21. In FIG. 21, GRU cell 2100 utilizes VMM array 2101 and activation function block 2102, which, when configured as a sigmoid function, applies a number between 0 and 1 to control the degree to which each component of the input vector contributes to the output vector. In FIG. 21, sigmoid function devices 1901 and 1902 and tanh device 1903 share the same physical hardware (VMM array 2101 and activation function block 2102) in a time-multiplexed manner. GRU cell 2100 also includes multiplier device 2103 for multiplying two vectors, adder device 2105 for adding two vectors, complementary device 2109 for subtracting the input from 1 to generate the output, multiplexer 2104, and a value h(t-1). * a register 2106 for holding r(t) as its value is output from the multiplier device 2103 via multiplexer 2104; and a register 2106 for holding the value h(t-1) * a register 2107 for holding z(t) as its value is output from the multiplier device 2103 via multiplexer 2104; and a register 2108 for holding the value ĥ(t) *and a register 2108 for holding (1-z((t)) as its value is output from the multiplier device 2103 via multiplexer 2104.
[0080] While GRU cell 2000 includes multiple sets of VMM array 2001 and activation function block 2002, GRU cell 2100 includes only one set of VMM array 2101 and activation function block 2102, which are used to represent multiple layers in embodiments of GRU cell 2100. GRU cell 2100 requires one-third the space for the VMM and activation function block compared to GRU cell 2000, so GRU cell 2100 requires less space than GRU cell 2000.
[0081] It can be further appreciated that a GRU system typically includes multiple VMM arrays, each of which requires functionality provided by specific circuit blocks outside the VMM array, such as adder and activation function blocks and high-voltage generation blocks. Providing separate circuit blocks for each VMM array would require a significant amount of space within a semiconductor device and would be somewhat inefficient. Accordingly, the embodiments described below reduce the circuitry required outside the VMM array itself.
[0082] The input to the VMM array can be an analog level, a binary level, a pulse, a time modulated pulse, or a digital bit (in which case a DAC is required to convert the digital bit to the appropriate input analog level), and the output can be an analog level, a binary level, a timing pulse, a pulse, or a digital bit (in which case an output ADC is required to convert the output analog level to a digital bit).
[0083] For each memory cell in the VMM array, each weight W can be implemented by a single memory cell, a differential cell, or two blended memory cells (the average of two cells). In the case of a differential cell, two memory cells are required to implement weight W as a differential weight (W=W+-W-). In the case of two blended memory cells, two memory cells are required to implement weight W as the average of two cells.
[0084] FIG. 31 illustrates a VMM system 3100. In some embodiments, the weights W stored in the VMM array are stored as a differential pair, W+ (positive weight) and W− (negative weight), where W=(W+)−(W−). In VMM system 3100, half of the bit lines are designated as W+ lines, i.e., bit lines connecting to memory cells that store a positive weight W+, and the other half of the bit lines are designated as W− lines, i.e., bit lines connecting to memory cells that implement a negative weight W−. W− lines are interspersed alternately among the W+ lines. Subtraction operations are performed by summing circuits, such as summing circuits 3101 and 3102, that receive current from the W+ and W− lines. The outputs of the W+ and W− lines are combined together to effectively provide W=W+−W− for each pair of (W+, W−) cells on every pair of (W+, W−) lines. Although described above with respect to W- lines interspersed alternately among W+ lines, in other embodiments, the W+ and W- lines may be arbitrarily positioned anywhere within the array.
[0085] 32 shows another embodiment. In a VMM system 3210, the positive weights W+ are implemented in a first array 3211 and the negative weights W− are implemented in a second array 3212 that is separate from the first array, and the resulting weights are suitably combined together by a summing circuit 3213.
[0086] Figure 33 shows VMM system 3300. The weights W stored in the VMM array are stored as a differential pair, W+ (positive weight) and W- (negative weight), where W = (W+) - (W-). VMM system 3300 includes array 3301 and array 3302. Half of the bit lines in each of arrays 3301 and 3302 are designated as W+ lines, i.e., bit lines connecting to memory cells that store a positive weight W+, and the other half of the bit lines in each of arrays 3301 and 3302 are designated as W- lines, i.e., bit lines connecting to memory cells that implement a negative weight W-. W- lines are interspersed alternately among the W+ lines. Subtraction operations are performed by adder circuits, such as adder circuits 3303, 3304, 3305, and 3306, that receive current from the W+ and W- lines. The outputs on the W+ and W- lines from each array 3301, 3302 are combined together, respectively, to effectively give W = W+ - W- for each pair of (W+, W-) cells on every pair of (W+, W-) lines. Additionally, the W values from each array 3301 and 3302 may be further combined via adder circuits 3307 and 3308, meaning that each W value is the result of subtracting the W value from array 3302 from the W value from array 3301, and the final result from adder circuits 3307 and 3308 is one of two difference values.
[0087] Each non-volatile memory cell used in an analog neural memory system can be erased or programmed to hold a very specific and precise amount of charge, i.e., number of electrons, in its floating gate. For example, each floating gate must hold one of N different values, where N is the number of different weights that can be represented by each cell. Examples of N include 16, 32, 64, 128, and 256.
[0088] Similarly, the read operation must be able to accurately distinguish between the N different levels.
[0089] There is a need for an improved output block in a VMM system that can quickly and accurately receive outputs from an array and identify the values represented by those outputs. Summary of the Invention
[0090] Numerous embodiments are disclosed for an output circuit for an analog neural memory in a deep learning artificial neural network.
[0091]
[0092]
[0093]
[0094]
[0095]
[0096]
[0097]
[0098]
[0099]
[0100]
[0101]
[0102]
[0103]
[0104]
[0105]
[0106]
[0107]
[0108]
[0109]
[0110]
[0111]
[0112]
[0113]
[0114]
[0115]
[0116]
[0117]
[0118]
[0119]
[0120]
[0121]
[0122]
[0123]
[0124]
[0125]
[0126]
[0127]
[0128]
[0129]
[0130]
[0131]
[0132] [Brief explanation of the drawings]
[0133] [Figure 1] FIG. 1 illustrates an artificial neural network. [Figure 2]1 shows a prior art split-gate flash memory cell. [Figure 3] 1 illustrates another prior art split-gate flash memory cell. [Figure 4] 1 illustrates another prior art split-gate flash memory cell. [Figure 5] 1 illustrates another prior art split-gate flash memory cell. [Figure 6] FIG. 1 illustrates various levels of an exemplary artificial neural network that utilizes one or more non-volatile memory arrays. [Figure 7] FIG. 1 is a block diagram illustrating a vector matrix multiplication system. [Figure 8] FIG. 1 is a block diagram illustrating an exemplary artificial neural network utilizing one or more vector-matrix multiplication systems. [Figure 9] 1 illustrates another embodiment of a vector matrix multiplication system. [Figure 10] 1 illustrates another embodiment of a vector matrix multiplication system. [Figure 11] 1 illustrates another embodiment of a vector matrix multiplication system. [Figure 12] 1 illustrates another embodiment of a vector matrix multiplication system. [Figure 13] 1 illustrates another embodiment of a vector matrix multiplication system. [Figure 14] 1 shows a prior art long-term memory system. [Figure 15] An exemplary cell for use in a long-term memory system is shown. [Figure 16] 16 illustrates one embodiment of the exemplary cell of FIG. 15. [Figure 17] 16 illustrates another embodiment of the exemplary cell of FIG. 15. [Figure 18] 1 shows a prior art gated recurrent unit system. [Figure 19] 1 shows an exemplary cell for use in a gated recurrent unit system. [Figure 20] 20 illustrates one embodiment of the exemplary cell of FIG. 19. [Figure 21] 20 illustrates another embodiment of the exemplary cell of FIG. 19. [Figure 22] 1 illustrates another embodiment of a vector matrix multiplication system. [Figure 23] 1 illustrates another embodiment of a vector matrix multiplication system. [Figure 24] 1 illustrates another embodiment of a vector matrix multiplication system. [Figure 25] 1 illustrates another embodiment of a vector matrix multiplication system. [Figure 26] 1 illustrates another embodiment of a vector matrix multiplication system. [Figure 27] 1 illustrates another embodiment of a vector matrix multiplication system. [Figure 28] 1 illustrates another embodiment of a vector matrix multiplication system. [Figure 29] 1 illustrates another embodiment of a vector matrix multiplication system. [Figure 30] 1 illustrates another embodiment of a vector matrix multiplication system. [Figure 31] 1 illustrates another embodiment of a vector matrix multiplication system. [Figure 32] 1 illustrates another embodiment of a vector matrix multiplication system. [Figure 33] 1 illustrates another embodiment of a vector matrix multiplication system. [Figure 34] 1 illustrates another embodiment of a vector matrix multiplication system. [Figure 35A] 1 illustrates an embodiment of an output block. [Figure 35B] 1 illustrates an embodiment of an output block. [Figure 35C] 1 illustrates an embodiment of an output block. [Figure 35D] 1 illustrates an embodiment of an output block. [Figure 35E] 1 illustrates an embodiment of an output block. [Figure 35F]1 illustrates an embodiment of an output block. [Figure 36] 10 illustrates another embodiment of an output circuit. [Figure 37A] 10 illustrates another embodiment of an output block. [Figure 37B] 10 illustrates another embodiment of an output block. [Figure 38A] 10 illustrates another embodiment of an output block. [Figure 38B] 10 illustrates another embodiment of an output block. [Figure 39] A replica variable resistor is shown. [Figure 40] 1 illustrates an embodiment of a current-to-voltage converter. [Figure 41] 1 shows a differential output amplifier. [Figure 42] 1 shows an offset calibration method. [Figure 43] 10 illustrates another offset calibration method. DETAILED DESCRIPTION OF THE INVENTION
[0134] The artificial neural network of the present invention utilizes a combination of CMOS technology and non-volatile memory arrays. VMM System Overview
[0135] 34 shows a block diagram of a VMM system 3400. The VMM system 3400 includes a VMM array 3401, a row decoder 3402, a high-voltage decoder 3403, a column decoder 3404, a bit line driver 3405, input circuits 3406, output circuits 3407, control logic 3408, and a bias generator 3409. The VMM system 3400 further includes a high-voltage generation block 3410, which includes a charge pump 3411, a charge pump regulator 3412, and a high-voltage analog precision level generator 3413. The VMM system 3400 further includes a (program / erase or weight adjustment) algorithm controller 3414, analog circuitry 3415, a control engine 3416 (which may include specialized functions such as, but not limited to, arithmetic functions, startup functions, embedded microcontroller logic, etc.), and test control logic 3417. The systems and methods described below may be implemented in the VMM system 3400.
[0136] The input circuit 3406 may include circuits such as a DAC (digital to analog converter), a DPC (digital to pulses converter), an AAC (analog to analog converter, such as a current-to-voltage converter or a logarithmic converter), a PAC (pulse to analog level converter), or any other type of converter. The input circuit 3406 may implement normalization, linear or nonlinear up / downscaling functions, or arithmetic functions. The input circuit 3406 may implement a temperature compensation function for the input level. The input circuit 3406 may implement an activation function such as ReLU or sigmoid. The output circuitry 3407 may include circuits such as an ADC (analog to digital converter, for converting neuron analog outputs into digital bits), an AAC (analog to analog converter, such as a current to voltage converter or a logarithmic converter), an APC (analog to pulses converter, analog to time modulated pulse converter), or any other type of converter.
[0137] The output circuit 3407 may implement activation functions such as a rectified linear activation function (ReLU) or a sigmoid. The output circuit 3407 may implement statistical normalization, regularization, up / down scaling / gain functions, statistical rounding, or arithmetic functions (e.g., addition, subtraction, division, multiplication, shift, log) of the neuron outputs. The output circuit 3407 may implement temperature compensation functions for the neuron outputs or array outputs (such as bit line outputs) to keep the power consumption of the array approximately constant or to increase the accuracy of the array (neuron) outputs, such as by keeping the IV slope approximately the same.
[0138] 35A shows output block 3500. Output block 3500 includes current-to-voltage converters (ITVs, having differential inputs and outputs) 3501-1 through 3501-i (where i is the number of pairs of bit lines W+ and W− that output block 3500 receives), multiplexer 3502, sample-and-hold circuits 3503-1 through 3503-k, channel multiplexer 3504, and differential-input analog-to-digital converter (ADC) 3505. Output block 3500 receives differential weight outputs W+ and W− from bit line pairs in the array and ultimately generates a digital output, DOUTx, that represents the output of one of the bit line pairs (e.g., the W+ and W− lines) from ADC 3505 (an ADC with a differential input).
[0139] Current-to-voltage (ITV) converters 3501-1 to 3501-i each receive analog bit line current signals BLw+ and BLw− (which are bit line outputs generated in response to the input and stored W+ and W− weights, respectively) and convert them to respective differential voltages ITVO+ and ITVO−.
[0140] The differential voltages ITVO+ and ITVO− are then received by a multiplexer 3502, which time-division multiplexes the outputs from the current-to-voltage converters 3501-1 to 3501-i to sample and hold (S / H) circuits 3503-1 to 3503k, where k may be the same as or different from i.
[0141] Each of the S / H circuits 3503-1 to 3503-k samples the differential voltages it receives and holds them as differential outputs.
[0142] The channel multiplexer 3504 receives a control signal to select one of the bit lines W+ and W− channels, i.e., one of the bit line pairs, and outputs the differential voltage held by each sample and hold circuit 3503 to the ADC 3505, which converts the analog differential voltage output by each sample and hold circuit 3503 into a set of digital bits, DOUTx. A single S / H 3503 can be shared across multiple ITV converters 3501. The ADC 3505 can operate on multiple ITV converters in a time-multiplexed manner. Each S / H 3503 can be simply a capacitor or a capacitor followed by a buffer (e.g., an operational amplifier).
[0143] The ITV converter 3501 may comprise the output current neuron circuits 3700, 3750, 3800, or 3820 from Figures 37A, 37B, 38A, and 38B, respectively, combined with the current-to-voltage converter 4000 of Figure 40. In such an instance, the inputs to the ITV converter 3501 are two current inputs (such as BLW+ and BLW- in Figures 35A-35E, 37A, 37B, 38A, or 38B), and the output of the ITV converter is a differential output (such as VOP and VON in Figure 40, or ITVO+ and ITVO- in Figures 35A-35D).
[0144] The ADC 3505 may be a hybrid ADC architecture, meaning it has more than one ADC architecture to perform the conversion. For example, if DOUTx is an 8-bit output, the ADC 3505 may include an ADC sub-architecture for generating bits B7 through B4 and another ADC sub-architecture for generating bits B3 through B0 from the differential inputs ITVSH+ and ITVSH−. That is, the ADC circuit 3505 may include multiple ADC sub-architectures.
[0145] Optionally, some ADC sub-architectures may be shared between all channels, while other ADC sub-architectures are not shared between all channels.
[0146] In another embodiment, channel multiplexer 3504 and ADC 3505 may be eliminated, and instead the output may be an analog differential voltage from S / H 3503, which may be buffered by an operational amplifier. For example, the use of an analog voltage may be implemented in an all-analog neural network (i.e., one where no digital output or digital input is required for the neural memory array).
[0147] 35B shows output block 3550. The output block includes current-to-voltage converters (ITV) 3551-1 through 3551-i (where i is the number of pairs of bit lines W+ and W− that output block 3550 receives), multiplexer 3552, differential-to-single-ended converters 3553-1 through 3553-k, sample-and-hold circuits 3554-1 through 3554-k (where k is the same as or different from i), channel multiplexer 3555, and analog-to-digital converter (ADC) 3556. Diff-to-S converter 3553 is used to convert the differential output, ITV0MX+ and ITV0MX−, from the ITV 3551 signal provided by mux 3552 to a single-ended output ITV0MX+. The single-ended output ITVSOMX+ is then input to S / H 3554, multiplexer 3555, and ADC 3556.
[0148] 35C shows output block 3560. Output block 3560 includes current-to-voltage converters (ITV) 3561-1 to 3561-i (where i is the number of pairs of bit lines W+ and W− that output block 3560 receives) and differential input analog-to-digital converters (ADC) 3566-1 to 3566-i.
[0149] 35D shows output block 3570. Output block 3570 includes current-to-voltage converters (ITV) 3571-1 to 3571-i (where i is the number of pairs of bit lines W+ and W− received by output block 3570) and single-input analog-to-digital converters (ADC) 3576-1 to 3576-i. In this case, only one output of the differential-output ITV is used, and the ITV is used with a differential input and a single output.
[0150] 35E shows output block 3580. Output block 3580 comprises current-to-voltage converters (ITV) 3581-1 to 3581-i (where i is the number of pairs of bit lines W+ and W− received by output block 3580) and differential input analog-to-digital converters (ADCs) 3586-1 to 3586-i. ITV blocks 3581-1 to 3581-i comprise common mode input circuits 3582-1 to 3582-i, respectively, and differential operational amplifiers 3583-1 to 3583-I, respectively, with feedback provided by variable resistors 3584-1 to 3584-i and 3585-1 to 3585-i, respectively.
[0151] Figure 35F shows an output block 3590 that can be used for the common mode input circuits 3582-1 to 3582-i of Figure 35E. Output block 3591 comprises two equal variable current sources Ibias+ and Ibias- connected to two current inputs BLw+ and BLw-.
[0152] 36 shows an output block 3600. The output block includes summing circuits 3601-1 to 3601-i (e.g., current mirror circuits), where i is the number of pairs of bit lines BLw+ and BLw− that output block 3600 receives, current-to-voltage converter circuits (ITV) 3602-1 to 3602-i, multiplexer 3603, sample-and-hold circuits 3604-1 to 3604-k (where k is the same as or different from i), channel multiplexer 3605, and ADC 3606. Output block 3600 receives differential weight outputs BLw+ and BLw− from the bit line pairs in the array and ultimately generates a digital output, DOUTx, from ADC 3606 that represents the output of one of the bit line pairs at a time.
[0153] Each of the current adder circuits 3601-1 to 3601-i receives a current from a bit line pair, subtracts the value of BLw- from BLw+, and outputs the result as an added current IWO.
[0154] Current-to-voltage converters 3602-1 through 3602-i receive output summed current IWO and convert each summed current into differential voltages ITVO+ and ITVO−, which are then received by multiplexer 3603 and selectively provided to sample-and-hold circuits 3604-1 through 3604-k. The differential voltages are digitized (converted to digital output bits) by a differential-input ADC (block 3606), which has various advantages such as reduced input noise (from clock feedthrough, etc.) and more accurate comparison (as in a SAR ADC).
[0155] Each sample and hold circuit 3604 receives the differential voltages ITVOMX+ and ITVOMX−, samples the received differential voltages, and holds them as differential voltage outputs OSH+ and PSH−.
[0156] The channel multiplexer 3605 receives a control signal to select one of the bit line pairs, i.e., channels BLw+ and BLw−, and outputs the voltages held by the respective sample and hold circuits 3604 to the differential input ADC 3606, which converts the voltages into a set of digital bits as DOUTx.
[0157] FIG. 37A shows an output current neuron circuit 3700 that may optionally be included in output block 3500 of FIG. 35 or output block 3600 of FIG.
[0158] The output current neuron circuit 3700 includes a first variable current source 3701, a second variable current source 3702, and a bias circuit 3703. The bias circuit 3703 generates a control voltage Vbias based on a comparison between BLW+ and VREF or between BLW- and VREF. The first variable current source 3701 generates an output current Ibias+ that varies with the control voltage Vbias (i.e., the amount of the output current Ibias+ responds to the value of Vbias) and is coupled to the first bit line BLW+. The second variable current source 3702 generates an output current Ibias- that varies with Vbias (i.e., the amount of the output current Ibias- responds to the value of Vbias) and is coupled to the second bit line BLW-. BLW+ receives a first current from a cell selected by a column decoder (not shown) to store a W+ value during a read operation, and BLW- receives a second current from a cell selected by the column decoder to store a W- value during a read operation. The W+ value and the associated W- value represent a weight value W. The outputs Ibias+ and Ibias- of current sources 3701 and 3702 are identical at any given time.
[0159] VREF is applied as an input common-mode voltage to generate the Vbias voltage, which controls variable current sources 3701 and 3702, which apply the common-mode voltage to BLW+ and BLW-, and the input common-mode voltage acts as a reference read voltage on the bit lines during a read operation. The outputs of output current neuron circuit 3700 are Iout+ and Iout-, which form a differential signal. Iout+ is the output current from bit line BLW+ after Vbias is applied to generate Ibias+, and Iout- is the output current from bit line BLW- after Vbias is applied to generate Ibias-, where Iout+ = Ibias+ - IBLW+ and Iout- = Ibias- - IBLW-.
[0160] FIG. 37B shows an output current neuron circuit 3750 illustrating one embodiment of variable current sources 3701 and 3702 using PMOS transistors 3711 and 3712.
[0161] FIG. 38A shows an output current neuron circuit 3800 that may optionally be included in output block 3500 of FIG. 35, output block 3550 of FIG. 35B, or output block 3600 of FIG.
[0162] The output current neuron circuit 3800 comprises a first variable resistor 3801 (first device) having a first end and a second end, the second end coupled to a bit line BLW+ selected during a read operation, a second variable resistor 3802 (second device) having a third end and a fourth end, the fourth end coupled to a bit line BLW− selected during a read operation, BLW+ connected to a cell in the memory array storing a W+ value and BLW− connected to a cell in the memory array storing an associated W−, a variable current source 3803, and a bias circuit operational amplifier 3804 that generates a bias voltage Vbias whose value represents the difference between BLW+ (or alternatively, BLW−) and VREF. The first end of the first variable resistor 3801 and the third end of the second variable resistor 3802 are coupled to the variable current source 3803.
[0163] VREF is used to generate the Vbias voltage that is applied to variable current source 3803 to apply an input common-mode voltage to bit lines BLW+ and BLW−, which acts as a read reference voltage on the bit lines during a read operation. The outputs of output current neuron circuit 3800 are Iout+ (first output current) from first variable resistor 3801 and Iout− (second output current) from second variable resistor 3802, which form a differential current signal. Iout+ is the output current from bit line BLW+ after Vbias is applied to generate Ibias, and Iout− is the output current from bit line BLW− after Vbias is applied to generate Ibias, according to the equations Iout+=Ibias−IBLW+ and Iout−=Ibias−IBLW−.
[0164] Figure 38B shows an output current neuron circuit 3820 that may optionally be included in output block 3500 of Figure 35, output block 3550 of Figure 35B, or output block 3600 of Figure 36. This circuit is similar to the circuit of Figure 38A, except that the output of operational amplifier 3804 is driven directly into the two terminals of two variable resistors 3801 and 3802.
[0165] 39 shows a variable resistor replica 3900 that may optionally be used in place of variable resistor 3801 and / or variable resistor 3802 of FIG. 38. Variable resistor replica 3900 comprises an NMOS transistor 3901. One terminal of NMOS transistor 3901 is coupled to bias circuit 3804. Another terminal of NMOS transistor 3901 is coupled to either BLW+ or BLW−. The gate of NMOS transistor 3901 is coupled to comparator 3902, which generates a control signal VGC that adjusts the resistance provided by NMOS transistor 3901. Thus, the resistance of NMOS 3901 is =VREF / IBIAS. By varying VREF or IBIAS, the equivalent resistance of NMOS 3901 can be changed.
[0166] FIG. 40 shows a current-to-voltage converter 4000 that can be used for current-to-voltage converter 3501 of FIG. 35A, current-to-voltage converter 3511 of FIG. 35B, or current-to-voltage converter 3602 of FIG.
[0167] Current-to-voltage converter 4000 is configured as shown and includes a differential amplifier 4001, variable integrating resistors 4002 and 4003, control switches 4004, 4005, 4006, and 4007, and variable sample and hold capacitors 4008 and 4009.
[0168] The current-to-voltage converter 4000 receives the differential currents IOUT+ and IOUT- and the output voltages VOP and VON. * R and output voltage VON=IOUT- * R, and resistors 4002 and 4003 each have a value equal to R. Scaling of the output neuron is provided by changing the values of resistors 4002 and 4003. For example, resistors 4002 and 4004 may each be provided by resistor replica circuit 3900. Capacitors 4008 and 4009 function as hold S / H capacitors, maintaining the output voltage when resistors 4002 and 4003 and the input current are interrupted. A control circuit (not shown) controls the opening and closing of switches 4004, 4005, 4006, and 4007 to provide the integration time.
[0169] In another mode of operation, variable capacitors 4008 and 4009 are used to integrate the differential output currents IOUT+ and IOUT-. In this case, resistors 4002 and 4003 are disabled (not used). Thus, the output voltage VOP is * The output voltage VON is proportional to Iout- * The value Time is proportional to Time / C. The value Time is controlled by the pulse width of pulse 4010 T. The value of C is provided by capacitors 4008 and 4009. The scaling of the output neuron value is then provided in this example by changing the pulse width T or by changing the capacitor values of capacitors 4008 and 4009.
[0170] The differential currents IOUT+ and IOUT- are derived from the first bit line current BLW+ and the second bit line current BLW-. IOUT+ and IOUT- have complementary values (one is positive and the other is negative, with the same magnitude). IOUT+ = ((current of BLW-) - (current of BLW+)) / 2, and IOUT- = ((current of BLW+) - (current of BLW-)) / 2. For example, when the current of BLW+ is 1 μA and the current of BLW- is 31 μA, Iout+ = (31 μA - 1 μA) / 2 = 15 μA, and Iout- = -15 μA.
[0171] FIG. 41 shows a differential amplifier 4100 that may optionally be included in the output block 3500 of FIG. 35A, the output block 3550 of FIG. 35B, or the output block 3600 of FIG. 36. The differential output amplifier 4100 is configured as shown and includes PMOS transistors 4101, 4102, 4103, 4104, 4105, 4106, 4107, and 4108, and NMOS transistors 4109, 4110, 4111, 4112, and 4113. The differential output amplifier 4100 receives inputs VINP and VINN and generates outputs VOUTP and VOUTN. VPBIAS is applied to the gates of PMOS transistors 4102, 4104, 4106, and 4108, and VNBIAS is applied to the gates of NMOS transistors 4111 and 4113. When VINP > VINN, VOUTP is high and VOUTN is low. When VINP < VINN, VOUTP is low and VOUTN is high. A common mode feedback circuit for the output common mode is not shown.
[0172] FIG. 42 shows an offset calibration method 4200 for an output block such as the above output blocks 3500, 3550, 3560, 3570, 3580, 3590, or 3600. This method can be executed within a sub-circuit block of the output block by an ITV block or an ADC block, etc.
[0173] First, a nominal bias is applied to the input node, which can be a midpoint offset trim setting such as a zero value or an average value (e.g., the average of the target input range for the BLw+ and BLw− inputs) (step 4202).
[0174] Second, the increased offset trim setting is applied to one of the sub-circuit blocks of the output block (such as ITV or ADC) (step 4203).
[0175] Third, the new trim output value for the entire output block is measured and compared to the expected output value to see if it is within target for the nominal output value (step 4204). If true, the method proceeds to step 4207. If not true, steps 4203 and 4204 are repeated, increasing the offset trim setting applied to the sub-circuit block each time until the new trim output value for the entire output block is within the range of the expected output value, at which point the method proceeds to step 4207.
[0176] If after a predetermined number of attempts (set by threshold T) the new trim output values for the entire output block are not within target as expected output values, the offset trim setting is returned to the nominal offset trim setting, and the offset trim setting is decreased from the nominal setting (step 4205).
[0177] The new trim output values for the entire output block are measured and compared to the expected output values for the entire output block to see if they are within the target range for the expected output values (step 4206). If true, the method proceeds to step 4207. If not true, steps 4205 and 4206 are repeated, decreasing the offset trim settings applied to the input nodes each time until the new trim output values are within the target range for the expected output values, at which point the method proceeds to step 4207.
[0178] In step 4207, the trim value that caused the output value to be within the target value for the expected output value is stored as the stored trim value. This is the trim value that results in the smallest offset by the output block.
[0179] In step 4208, optionally, the stored trim value is applied as a bias to the sub-circuit block of the output block during all operations.
[0180] Thus, the offset calibration method 4200 performs a trim operation on the entire output block by trimming sub-circuit blocks of the output block.
[0181] 43 shows an offset calibration method 4300 for an output block, such as the above-described output blocks 3500, 3550, 3560, 3570, 3580, 3590, or 3600. The method may be performed within a sub-circuit block, such as by an ITV block or an ADC block.
[0182] First, a reference bias is applied to the input nodes (eg, inputs to BLw+ and BLw-) of the sub-circuit blocks of the output block (step 4301).
[0183] Next, the output value of the output block is measured and compared with the target offset value (step 4302).
[0184] If the measured output value>target offset value, the next offset trim value in the sequence of offset trim values is applied (step 4303) and step 4302 is repeated. The offset trim is applied to one of the sub-circuit blocks of the output block (such as the ITV or ADC).
[0185] Steps 4303 and 4302 are repeated until the measured output value is less than or equal to the target offset value, at which point the offset trim value is stored (step 4304). This is the trim value that results in an acceptable level of offset.
[0186] Optionally, the stored offset trim value is applied as a bias to the sub-circuit block of the output block during all operations (step 4305).
[0187] In an alternative embodiment, the variable resistors in FIG. 35E or FIG. 40B have unequal resistances. In this case, the output voltage or current from the ITV is proportional to the resistance value. For example, in FIG. 35E, resistor 3585-1 can be made very large, in which case most of the current from the two bit lines (IBLw+-IBLw-) flows through resistor 3584-1. In another example of FIG. 35E, resistor 3585-1 is disconnected, and then all of the current from the two bit lines (IBLw+-IBLw-) flows through resistor 3584-1.
[0188] It should be noted that, as used herein, both the terms "over" and "on" are inclusive of "directly" (with no intermediate material, element, or gap disposed therebetween) and "indirectly" (with an intermediate material, element, or gap disposed therebetween). Similarly, the term "adjacent" includes "directly adjacent" (with no intermediate material, element, or gap disposed therebetween) and "indirectly adjacent" (with an intermediate material, element, or gap disposed therebetween); "attached" includes "directly attached" (with no intermediate material, element, or gap disposed therebetween) and "indirectly attached" (with an intermediate material, element, or gap disposed therebetween); and "electrically coupled" includes "directly electrically coupled" (with no intermediate material or element disposed therebetween that electrically connects the elements together) and "indirectly electrically coupled" (with an intermediate material or element disposed therebetween that electrically connects the elements together). For example, forming an element "over a substrate" can include forming the element directly on the substrate with no intermediate materials / elements therebetween, and forming the element indirectly on the substrate with one or more intermediate materials / elements therebetween.
Claims
1. An output current neuron circuit, comprising: a first bit line coupled to a W+ cell in the memory array, the first bit line drawing a first current during a read operation; a second bit line coupled to a W- cell in the memory array and sinking a second current during a read operation, the difference between the value stored in the W+ cell and the value stored in the W- cell being a weight value W; a bias circuit for generating a common mode bias voltage; a first variable current source responsive to the common mode bias voltage for applying a common mode bias current to the first bit line to generate a first output; a second variable current source responsive to the common mode bias voltage to apply the common mode bias current to the second bit line to generate a second output; an output current neuron circuit, wherein the first output is equal to the common mode bias current minus the first current, and the second output is equal to the common mode bias current minus the second current;
2. 2. The output current neuron circuit of claim 1, wherein said first variable current source comprises a first PMOS transistor.
3. 3. The output current neuron circuit of claim 2, wherein said second variable current source comprises a second PMOS transistor.
4. An output current neuron circuit, comprising: a current source; a bias circuit for applying a control voltage to the current source; a first variable resistor including a first end and a second end, the first end coupled to the current source; a second variable resistor including a third end and a fourth end, the third end coupled to the current source, the current source providing a bias current to the first variable resistor and the second variable resistor to generate a common-mode voltage; a first bit line coupled to the W+ cell during a read operation; a second bit line coupled to a W-cell during the read operation, the difference between the value stored in the W+ cell and the value stored in the W-cell being a weight value W; a first output coupled to the second end of the first variable resistor and to the first bit line to provide a first output current; a second output coupled to the fourth end of the second variable resistor and to the second bit line to provide a second output current, wherein the first output and the second output form a common-mode differential current signal.
5. 5. The circuit of claim 4, wherein the first variable resistor comprises an NMOS transistor, the voltage applied to the gate of the NMOS transistor determining the resistance of the NMOS transistor.
6. 6. The circuit of claim 5, wherein the second variable resistor comprises an NMOS transistor, and a voltage applied to a gate of the NMOS transistor determines the resistance of the NMOS transistor.
7. An output current neuron circuit, comprising: a first output node for receiving a first current from the memory array; a second output node for receiving a second current from the memory array; a bias circuit for generating a bias current; a first device that generates a first output current equal to the first current subtracted from the bias current; a second device that generates a second output current equal to the second current subtracted from the bias current.
8. 8. The output current neuron circuit of claim 7, wherein said first output current is generated from a read operation on a bit line coupled to one or more W+ cells.
9. 9. The output current neuron circuit of claim 8, wherein the first output current is generated from a read operation on a bit line coupled to one or more W-cells.
10. An output current neuron circuit, comprising: a first output node for receiving a first current from the memory array; a second output node for receiving a second current from the memory array; a bias circuit for generating a bias voltage at a bias node; a first variable resistor coupled between the bias node and the first output node; a second variable resistor coupled between the bias node and the second output node.
11. A current-to-voltage converter, a first bit line for receiving a first current generated during a read operation of the W+ cell; a second bit line for receiving a second current generated during a read operation of a W- cell, the difference between the value stored in the W+ cell and the value stored in the W- cell being a weight value W; a differential amplifier for receiving the first current and the second current and generating a differential output voltage comprising a first voltage output and a second voltage output.
12. An output block, a plurality of current-to-voltage converters, each of which receives a differential pair of bit lines and produces a differential voltage output; an output block comprising: a plurality of differential input analog-to-digital converters, each of which receives a differential voltage output from one of the plurality of current-to-voltage converters and generates a set of digital output bits;
13. An output block, a plurality of current-to-voltage converters, each receiving a differential pair of bit lines and producing a voltage output; an output block comprising: a plurality of differential input analog-to-digital converters, each receiving a voltage output from one of the plurality of current-to-voltage converters and generating a set of digital output bits;
14. An output block, a current-to-voltage converter for receiving a bit line differential pair, a differential operational amplifier having first and second inputs and first and second outputs, the first and second inputs coupled to the bit line differential pair; a first variable resistor coupled between the first input and the first output; a second variable resistor coupled between the second input and the second output; a current-to-voltage converter comprising: a common mode input circuit coupled between the first input and the second input; an output block comprising: a differential input analog-to-digital converter for receiving the first output and the second output and for generating a set of digital output bits;
15. 15. The output block of claim 14, wherein the common mode input circuit comprises a first variable current source coupled to the first input and a second variable current source coupled to the second input, the first variable current source and the second variable current source generating equal currents.
16. an output block, the output block comprising: An output current neuron circuit, comprising: a first bit line coupled to a W+ cell in the memory array, the first bit line drawing a first current during a read operation; a second bit line coupled to a W-cell in the memory array and sinking a second current; and a first bias current coupled to the first bit line; and a second bias current coupled to the second bit line, wherein the first bias current and the second bias current have the same value as the first bias current.
17. an output block, the output block comprising: An output current neuron circuit, comprising: a first bit line coupled to a W+ cell in the memory array, the first bit line drawing a first current during a read operation; a second bit line coupled to a W-cell in the memory array, the second bit line sinking a second current during the read operation; and an output current neuron circuit including: a first bias current coupled to the first bit line; a first output current proportional to a difference between the first current and the second current.
18. 18. The output block of claim 17, wherein the first output current is equal to half the difference between the first current and the second current.
19. 20. The output block of claim 17, further comprising a second output current that is complementary to the first output current.
20. 1. An offset calibration method for an output block, comprising: applying a nominal bias to an input node of a sub-circuit block of the output block; applying increased or decreased offset trim settings to the sub-circuit blocks within the output block until the output of the output block is within a threshold of a target value.
21. The method of claim 20, wherein the sub-circuit block is a current-to-voltage circuit.
22. 21. The method of claim 20, wherein the sub-circuit block is of an analog-to-digital converter circuit.
23. 21. The method of claim 20, further comprising providing, by the output block, an output from a neuron.
24. 24. The method of claim 23, wherein the neuron is part of a neural memory array within a neural network.
25. 1. An offset calibration method for an output block, comprising: measuring a new trim output of said output block in response to an increased offset trim setting; comparing the new trimmed output with a nominal biased output; repeating the applying, measuring, and comparing steps when the new trim output is equal to the nominal bias output; a comparing step, comprising: storing the new trim output as a trim value when the new trim output differs from the nominal bias output; applying the trim value to the sub-circuit block within the output block during operation.
26. 26. The method of claim 25, further comprising providing, by the output block, an output from a neuron.
27. 27. The method of claim 26, wherein the neuron is part of a neural memory array within a neural network.
28. 1. An offset calibration method for an output block, comprising: applying a nominal bias to an input node of a sub-circuit block of the output block; measuring a nominal bias output of the output block in response to the nominal bias; applying a reduced offset trim setting to the input node; measuring a new trim output of said output block in response to an increased offset trim setting; comparing the new trimmed output with the nominal biased output; repeating the applying, measuring, and comparing steps when the new trim output is equal to the nominal bias output; a comparing step, comprising: storing the new trim output as a trim value when the new trim output differs from the nominal bias output; and applying the trim value to the sub-circuit block of the output block during operation.
29. 30. The method of claim 28, further comprising providing, by the output block, an output from a neuron.
30. 30. The method of claim 29, wherein the neuron is part of a neural memory array within a neural network.
31. 1. An offset calibration method for an output block, comprising: applying input values to input nodes of sub-circuit blocks of said output block; measuring an output value in response to said input value; comparing the output value with a target offset value; when the output value exceeds the target offset value, repeating the applying, measuring, and comparing steps with a next input value; a comparing step, comprising: storing the input value as a trim value when the output value is less than or equal to the target offset value; applying the trim value to the sub-circuit block of the output block during operation of the output block.
32. 32. The method of claim 31, further comprising providing, by the output block, an output from a neuron.
33. 33. The method of claim 32, wherein the neuron is part of a neural memory array within a neural network.
Citation Information
Patent Citations
Design of digital-analog hybrid reading circuit applied to eFlash-based storage and calculation integrated circuit
CN111193511A
Neural network computing circuit, chip and system
CN112580790A
Output circuits for an analog neural memory system for deep learning neural network
US20210174185A1