Output circuits for analog neural memories in deep learning artificial neural networks
By using nonvolatile memory arrays to store weights and performing in-memory calculations in artificial neural networks, the problems of poor energy efficiency and high cost in the prior art are solved, and efficient and low-energy-consuming neural network computing is achieved.
Patent Information
- Application Number
- JP2023564551
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-08-31
- Filing Date
- 2021-09-03
- Publication Date
- 2025-05-08
- Estimated Expiration
- 2041-09-03
AI Technical Summary
The lack of appropriate hardware technology in the development of high-performance artificial neural networks leads to poor energy efficiency and high cost, especially when implementing high connectivity and complex computing processes.
A nonvolatile memory array is used as a simulated neural memory for an artificial neural network, and the weight is stored in a floating gate electrode and a vector matrix multiplication array is used to perform calculations to achieve a combination of -memory calculation and weight storage.
It improves energy efficiency and calculation density, reduces dependence on external multiplication and addition logic, realizes efficient neural network computing, and reduces the energy consumption and cost of the system.
Smart Images

Figure 0007673240000011 
Figure 0007673240000012 
Figure 0007673240000013
Abstract
Description
[Technical field]
[0001] (Priority Claim) This application claims priority to U.S. Provisional Patent Application No. 63 / 190,240, entitled "Hybrid Output Architecture for Analog Neural Memory in a Deep Learning Artificial Neural Network," filed May 19, 2021, and U.S. Patent Application No. 17 / 463,063, entitled "Output Circuit for Analog Neural Memory in a Deep Learning Artificial Neural Network," filed August 31, 2021, which are incorporated by reference herein.
[0002] FIELD OF THEINVENTION Numerous embodiments are disclosed for a hybrid output architecture for analog neural memories in deep learning artificial neural networks. [Background technology]
[0003] Artificial neural networks mimic biological neural networks (the central nervous systems of animals, particularly the brain) and are used to estimate or approximate functions that may depend on multiple inputs and are generally unknown. Artificial neural networks generally contain layers of interconnected "neurons" that exchange messages between each other.
[0004] FIG. 1 shows an artificial neural network, where the circles represent layers of inputs or neurons. The connections (called synapses) are represented by arrows and have numerical weights that can be tuned based on experience. This allows the neural network to adapt to the inputs and learn. Typically, a neural network contains multiple layers of inputs. There are typically one or more hidden layers of neurons, and one output layer of neurons that provides the output of the neural network. At each level, the neurons make decisions individually or collectively based on the data they receive from the synapses.
[0005] One of the main challenges in the development of artificial neural networks for high-performance information processing is the lack of suitable hardware technology. In practice, practical neural networks rely on a very large number of synapses, which allows high connectivity between neurons, i.e., a very high degree of parallelization of computation. In principle, such complexity can be achieved by digital supercomputers or dedicated graphic processing unit clusters. However, in addition to high costs, these approaches also suffer from poor energy efficiency compared to biological networks, which mainly perform low-precision analog computations and therefore consume much less energy. CMOS analog circuits have been used for artificial neural networks, but most CMOS-implemented synapses are too bulky given the large number of neurons and synapses.
[0006] Applicant previously disclosed an artificial (analog) neural network utilizing one or more non-volatile memory arrays as synapses in U.S. Patent Application Serial No. 15 / 594,439, which is incorporated by reference. The non-volatile memory array operates as an analog neural memory. The neural network device includes a first plurality of synapses configured to receive a first plurality of inputs and generate a first plurality of outputs therefrom, and a first plurality of neurons configured to receive the first plurality of outputs. The first plurality of synapses includes a plurality of memory cells, each of which includes spaced apart source and drain regions formed in a semiconductor substrate with a channel region extending therebetween, a floating gate insulated and disposed above a first portion of the channel region, and a non-floating gate insulated and disposed above a second portion of the channel region. Each of the plurality of memory cells is configured to store a weight value corresponding to a number of electrons in the floating gate. The plurality of memory cells is configured to multiply the first plurality of inputs by the stored weight values to generate the first plurality of outputs.
[0007] Non-volatile Memory Cell Non-volatile memories are well known. For example, U.S. Pat. No. 5,029,130 ("the '130 patent"), incorporated herein by reference, discloses an array of split-gate non-volatile memory cells, which are a type of flash memory cell. Such a memory cell 210 is shown in FIG. 2. Each memory cell 210 includes a source region 14 and a drain region 16 formed in a semiconductor substrate 12, with a channel region 18 between the source region 14 and the drain region 16. A floating gate 20 is formed over and insulated from a first portion of the channel region 18 (and controls the conductivity of the first portion of the channel region 18) and over a portion of the source region 14. A word line terminal 22 (typically coupled to a word line) has a first portion disposed over and insulated from a second portion of the channel region 18 (and controls the conductivity of the second portion of the channel region 18) and a second portion extending upwardly above the floating gate 20. A floating gate 20 and a wordline terminal 22 are insulated from the substrate 12 by a gate oxide. A bitline 24 is coupled to the drain region 16.
[0008] The memory cell 210 is erased (electrons are removed from the floating gate) by applying a high positive voltage to the wordline terminal 22, which causes electrons in the floating gate 20 to pass via Fowler-Nordheim (FN) tunneling from the floating gate 20 to the wordline terminal 22 through the insulator between them.
[0009] The memory cell 210 is programmed by source side injection (SSI) of hot electrons (electrons are injected into the floating gate) by applying a positive voltage to the word line terminal 22 and a positive voltage to the source region 14. Electrons flow from the drain region 16 towards the source region 14. The electrons accelerate and heat up when they reach the gap between the word line terminal 22 and the floating gate 20. Some of the heated electrons are injected into the floating gate 20 through the gate oxide due to electrostatic attraction from the floating gate 20.
[0010] The memory cell 210 is read by applying a positive read voltage to the drain region 16 and the word line terminal 22 (turning on the portion of the channel region 18 below the word line terminal). When the floating gate 20 is positively charged (i.e., erased with electrons), the portion of the channel region 18 below the floating gate 20 is also turned on and current flows through the channel region 18, which is sensed as the erased or "1" state. When the floating gate 20 is negatively charged (i.e., programmed with electrons), the portion of the channel region below the floating gate 20 is mostly or completely off and no (or very little) current flows through the channel region 18, which is sensed as the programmed or "0" state.
[0011] Table 1 shows typical voltage / current ranges that may be applied to the terminals of memory cell 110 to perform read, erase, and program operations. Table 1: Operation of the Flash Memory Cell 210 of FIG. [Table 1]
[0012] Other split-gate memory cell configurations are also known for other types of flash memory cells. For example, FIG. 3 shows a four-gate memory cell 310 including a source region 14, a drain region 16, a floating gate 20 above a first portion of a channel region 18, a select gate 22 (typically coupled to a word line, WL) above a second portion of the channel region 18, a control gate 28 above the floating gate 20, and an erase gate 30 above the source region 14. This configuration is described in U.S. Pat. No. 6,747,310, which is incorporated herein by reference for all purposes. Here, all gates, except for the floating gate 20, are non-floating gates, i.e., they are electrically connected or connectable to a voltage source. Programming is performed by heated electrons injecting themselves from the channel region 18 into the floating gate 20. Erasing is performed by electrons tunneling from the floating gate 20 to the erase gate 30.
[0013] Table 2 shows typical voltage / current ranges that can be applied to the terminals of memory cell 310 to perform read, erase, and program operations. Table 2: Operation of the Flash Memory Cell 310 of FIG. [Table 2]
[0014] Figure 4 shows another type of flash memory cell, a three-gate memory cell 410. Memory cell 410 is identical to memory cell 310 of Figure 3, except that memory cell 410 does not have a separate control gate. Erase and read operations (where erasure occurs through the use of an erase gate) are similar to those of Figure 3, except that no control gate bias is applied. Programming operations are also performed without a control gate bias, and as a result, a higher voltage must be applied to the source line during a program operation to compensate for the lack of control gate bias.
[0015] Table 3 shows typical voltage / current ranges that may be applied to the terminals of memory cell 410 to perform read, erase, and program operations. Table 3: Operation of Flash Memory Cell 410 of FIG. [Table 3]
[0016] Figure 5 shows another type of flash memory cell, a stacked gate memory cell 510. The memory cell 510 is similar to the memory cell 210 of Figure 2, except that the floating gate 20 extends over the entire channel region 18, and the control gate 22 (where it is coupled to a word line) extends over the floating gate 20, separated by an insulating layer (not shown). Erasing is done by FN tunneling of electrons from the FG to the substrate, and programming is done by channel hot electron (CHE) injection in the region between the channel 18 and the drain region 16, by electrons flowing from the source region 14 towards the drain region 16, and by a read operation that is similar to the read operation of the memory cell 210 with a higher control gate voltage.
[0017] Table 4 shows typical voltage ranges that may be applied to the terminals of memory cell 510 and substrate 12 to perform read, erase, and program operations. Table 4: Operation of Flash Memory Cell 510 of FIG. [Table 4]
[0018] The methods and means described herein may be applied to other non-volatile memory technologies such as, but not limited to, FINFET split-gate flash or stack-gate flash memory, NAND flash, SONOS (silicon-oxide-nitride-oxide-silicon, charge traps in nitride), MONOS (metal-oxide-nitride-oxide-silicon, metal charge traps in nitride), ReRAM (resistive random access memory), PCM (phase change memory), MRAM (magnetoresistive memory), FeRAM (ferroelectric memory), CT (charge trap) memory, CN (carbon tube) memory, OTP (bi-level or multi-level one-time programmable) and CeRAM (strongly correlated electron memory).
[0019] In order to utilize a memory array containing one of the non-volatile memory cell types in an artificial neural network as described above, two modifications are made: First, as explained further below, the lines are configured so that each memory cell can be individually programmed, erased, and read without adversely affecting the memory state of other memory cells in the array; Second, a continuous (analog) programming of the memory cells is provided.
[0020] Specifically, the memory state (i.e., the charge on the floating gate) of each memory cell in the array can be changed continuously from a fully erased state to a fully programmed state, independently and with minimal disturbance to other memory cells. In another embodiment, the memory state (i.e., the charge on the floating gate) of each memory cell in the array can be changed continuously from a fully programmed state to a fully erased state, and vice versa, independently and with minimal disturbance to other memory cells. This means that the cell storage is analog or can at a minimum store one of many discrete values (e.g., 16 or 64 different values), which allows every cell in the memory array to be very precisely and individually tunable, making the memory array ideal for storage and fine tuning adjustments to the synaptic weights of neural networks.
[0021] Neural network using non-volatile memory cell arrays 6 conceptually illustrates a non-limiting example of a neural network utilizing the non-volatile memory array of the present embodiment. This example uses a non-volatile memory array neural network for a face recognition application, although other suitable applications can also be implemented using a non-volatile memory array-based neural network.
[0022] S0 is the input layer, which in this example is a 32x32 pixel RGB image with 5-bit precision (i.e., three 32x32 pixel arrays, one for each color R, G, and B, each pixel with 5-bit precision). Synapse CB1 going from input layer S0 to layer C1 scans the input image with overlapping filters (kernels) of 3x3 pixels, applying different sets of weights to some instances and shared weights to other instances, and shifts the filters by one pixel (or two or more pixels depending on the model). Specifically, the values of the nine pixels in the 3x3 portion of the image (i.e., referred to as filters or kernels) are provided to synapse CB1, which multiplies these nine input values by the appropriate weights and, after summing the outputs of the multiplication, determines a single output value, which is given by the first synapse of CB1 to generate one pixel of the feature map of layer C1. The 3x3 filter is then shifted by one pixel to the right in the input layer S0 (i.e., adding a column of 3 pixels to the right and dropping a column of 3 pixels on the left), so that the 9 pixel values of this newly positioned filter are provided to synapse CB1, where they are multiplied by the same weights as above to determine a second single output value by the associated synapse. This process continues until the 3x3 filter scans the entire 32x32 pixel image of the input layer S0 for all three colors and all bits (precision values). The process is then repeated using different sets of weights to generate different feature maps of layer C1, until all feature maps of layer C1 have been calculated.
[0023] In this example, in layer C1, there are 16 feature maps, each with 30x30 pixels. Each pixel is a new feature pixel extracted from the multiplication of the input with a kernel, and thus each feature map is a two-dimensional array, and thus in this example, layer C1 constitutes 16 layers of two-dimensional arrays (note that layers and arrays referred to herein are logical, not necessarily physical, relationships, i.e., arrays are not necessarily oriented in a physical two-dimensional array). Each of the 16 feature maps in layer C1 is generated by one of 16 different sets of synaptic weights applied to the filter scans. The C1 feature maps can all target different aspects of the same image feature, such as boundary identification. For example, a first map (generated using a first set of weights shared by all scans used to generate this first map) can identify circular edges, and a second map (generated using a second set of weights different from the first set of weights) can identify rectangular edges or the aspect ratio of a particular feature, etc.
[0024] Before going from layer C1 to layer S1, an activation function P1 (pooling) is applied that pools values from non-overlapping, consecutive 2x2 regions in each feature map. The purpose of the pooling function P1 is to average nearby positions (or a max function can be used), e.g., to reduce edge position dependency, and to reduce data size before going to the next stage. In layer S1, there are 16 15x15 feature maps (i.e., 16 different arrays of 15x15 pixels each). The synapse CB2 going from layer S1 to layer C2 scans the maps in layer S1 with a 4x4 filter with a filter shift of 1 pixel. In layer C2, there are 22 12x12 feature maps. Before going from layer C2 to layer S2, an activation function P2 (pooling) is applied that pools values from non-overlapping, consecutive 2x2 regions in each feature map. In layer S2, there are 22 6x6 feature maps. At the synapse CB3 going from layer S2 to layer C3, an activation function (pooling) is applied, where every neuron in layer C3 connects to every map in layer S2 through a respective synapse in CB3. There are 64 neurons in layer C3. The synapse CB4 going from layer C3 to the output layer S3 fully connects C3 to S3, i.e. every neuron in layer C3 connects to every neuron in layer S3. The output at S3 contains 10 neurons, where the neuron with the highest output determines the class. This output can indicate, for example, the identity or classification (class) of the content of the original image.
[0025] Each layer of the synapse is implemented using an array or a portion of an array of non-volatile memory cells.
[0026] 7 is a block diagram of an array that can be used for that purpose. A vector matrix multiplication (VMM) array 32 contains non-volatile memory cells and is utilized as a synapse between one layer and the next (such as CB1, CB2, CB3, and CB4 in FIG. 6). Specifically, the VMM array 32 includes an array of non-volatile memory cells 33, an erase gate and word line gate decoder 34, a control gate decoder 35, a bit line decoder 36, and a source line decoder 37, which decode the respective inputs to the non-volatile memory cell array 33. Inputs to the VMM array 32 can come from the erase gate and word line gate decoder 34 or from the control gate decoder 35. The source line decoder 37 in this example also decodes the output of the non-volatile memory cell array 33. Alternatively, the bit line decoder 36 can decode the output of the non-volatile memory cell array 33.
[0027] The non-volatile memory cell array 33 serves two purposes. First, the non-volatile memory array 33 stores the weights used by the VMM array 32. Second, the non-volatile memory cell array 33 effectively multiplies the inputs by the weights stored in the non-volatile memory cell array 33 and adds them for each output line (source line or bit line) to generate an output that becomes the input to the next layer or the input to the last layer. With the non-volatile memory cell array 33 performing the multiplication and addition functions, the need for separate multiplication and addition logic is eliminated and in-memory computation is also more power efficient.
[0028] The outputs of the non-volatile memory cell array 33 are provided to a differential summer (such as a summing op-amp or summing current mirror) 38, which sums the outputs of the non-volatile memory cell array 33 to create a single value for the convolution. The differential summer 38 is arranged to perform a summation of the positive and negative weights.
[0029] The summed output value of the differential adder 38 is then fed to an activation function block 39 which rectifies the output. The activation function block 39 may provide a sigmoid, tanh, or ReLU function. The rectified output value of the activation function block 39 becomes an element of a feature map as the next layer (e.g., C1 in FIG. 6) and is then applied to the next synapse to generate the next feature map layer or the last layer. Thus, in this example, the non-volatile memory cell array 33 constitutes a number of synapses (receiving inputs from a previous layer of neurons or from an input layer such as an image database), and the summing op-amp 38 and the activation function block 39 constitute a number of neurons.
[0030] The inputs to the VMM array 32 of FIG. 7 (WLx, EGx, CGx, and optionally BLx and SLx) may be analog levels, binary levels, or digital bits (in which case a DAC is provided to convert the digital bits to appropriate input analog levels), and the outputs may be analog levels, binary levels, or digital bits (in which case an output ADC is provided to convert the output analog levels to digital bits).
[0031] FIG. 8 is a block diagram illustrating the use of multiple layers of VMM array 32, labeled in the figure as VMM arrays 32a, 32b, 32c, 32d, and 32e. As shown in FIG. 8, an input (denoted Inputx) is converted from digital to analog by a digital-to-analog converter 31 and provided to input VMM array 32a. The converted analog input can be a voltage or a current. The first layer input D / A conversion can be done by using a function or a LUT (look-up table) that maps the input Inputx to the appropriate analog level of the matrix multiplier of input VMM array 32a. The input conversion can also be done by an analog to analog (A / A) converter to convert an external analog input to the mapped analog input to input VMM array 32a.
[0032] The output generated by the input VMM array 32a is then provided as input to the next VMM array (hidden level 1) 32b, which in turn generates an output that is provided as input to the input VMM array (hidden level 2) 32c, and so on. The various layers of the VMM array 32 function as layers of synapses and neurons of a convolutional neural network (CNN). Each of the VMM arrays 32a, 32b, 32c, 32d, and 32e can be a standalone physical non-volatile memory array, or the multiple VMM arrays can utilize different portions of the same physical non-volatile memory array, or the multiple VMM arrays can utilize overlapping portions of the same physical non-volatile memory array. The example shown in FIG. 8 includes five layers (32a, 32b, 32c, 32d, 32e): one input layer (32a), two hidden layers (32b, 32c), and two fully connected layers (32d, 32e). Those skilled in the art will appreciate that this is merely an example, and that the system may instead include more than two hidden layers and more than two fully connected layers.
[0033] Vector Matrix Multiplication (VMM) Array 9 shows a neuron VMM array 900 that is particularly suited for the memory cells 310 shown in FIG. 3 and is utilized as part of the synapses and neurons between the input layer and the next layer. The VMM array 900 includes a memory array 901 of non-volatile memory cells and a reference array 902 of non-volatile reference memory cells (located at the top of the array). Alternatively, a separate reference array can be located at the bottom.
[0034] In the VMM array 900, control gate lines such as control gate line 903 run vertically (so that the row-wise reference array 902 is orthogonal to control gate line 903) and erase gate lines such as erase gate line 904 run horizontally. Here, inputs to the VMM array 900 are provided on the control gate lines (CG0, CG1, CG2, CG3) and outputs of the VMM array 900 appear on the source lines (SL0, SL1). In one embodiment, only even rows are used, and in another embodiment, only odd rows are used. The current in each source line (SL0, SL1 respectively) performs a function of the sum of all the currents from the memory cells connected to that particular source line.
[0035] As described herein for neural networks, the non-volatile memory cells of VMM array 900, namely memory cells 310 of VMM array 900, are preferably configured to operate in the sub-threshold region.
[0036] The non-volatile reference memory cells and non-volatile memory cells described herein are biased in weak inversion (sub-threshold region) as follows: Ids=Io * e (Vg-Vth) / nVt =w * Io * e (Vg) / nVt In the formula, w=e (-Vth) / nVt and where Ids is the drain-source current, Vg is the gate voltage of the memory cell, Vth is the threshold voltage of the memory cell, and Vt is the thermal voltage = k * where T / q, k is the Boltzmann constant, T is temperature in Kelvin, q is the electron charge, n is the slope coefficient = 1 + (Cdep / Cox), Cdep = capacitance of the depletion layer, and Cox is the capacitance of the gate oxide layer, Io is the memory cell current at a gate voltage equal to the threshold voltage, and Io is equal to (Wt / L). * u * Cox * (n-1) * Vt 2where u is the carrier mobility, and Wt and L are the width and length of the memory cell, respectively.
[0037] When using an IV-log converter that uses memory cells (such as reference or peripheral memory cells) or transistors to convert the input current to an input voltage: Vg=n * Vt * log[Ids / wp * Io] where wp is the w of the reference or periphery memory cell.
[0038] For a memory array used as a vector matrix multiplier VMM array with current inputs, the output current is: Iout=wa * Io * e (Vg) / nVt , i.e. Iout = (wa / wp) * Iin=W * Iin W=e (Vthp-Vtha) / nVt where wa=w of each memory cell in the memory array. Vthp is the effective threshold voltage of the peripheral memory cells and Vtha is the effective threshold voltage of the main (data) memory cells. Note that the threshold voltage of a transistor is a function of the substrate body bias voltage, which is designated Vsb, and can be modulated to compensate for various conditions at such temperature. The threshold voltage Vth can be expressed as:
number
number
[0039] The word line or control gate can be used as the input of the memory cell for an input voltage.
[0040] Alternatively, the flash memory cells of the VMM arrays described herein may be configured to operate in the linear region. Ids=β * (Vgs-Vth) * Vds; β=u * Cox * Wt / L W = α(Vgs-Vth) That is, the weight W in the linear region is proportional to (Vgs-Vth).
[0041] The word line or control gate or bit line or source line can be used as the input of a memory cell operating in the linear region. The bit line or source line can be used as the output of the memory cell.
[0042] For an IV linear converter, memory cells (such as reference or peripheral memory cells) or transistors operating in the linear region can be used to linearly convert the input and output currents to input and output voltages.
[0043] Alternatively, the memory cells of the VMM arrays described herein may be configured to operate in the saturation region. Ids=1 / 2 * β * (Vgs-Vth) 2 ; β=u * Cox * Wt / L W ∝ (Vgs-Vth) 2 , that is, the weight W is (Vgs-Vth) 2 is proportional to.
[0044] The word lines, control gates, or erase gates can be used as inputs for memory cells operating in the saturation region, and the bit lines or source lines can be used as outputs for output neurons.
[0045] Alternatively, the memory cells of the VMM arrays described herein may be used in all regions or combinations thereof (sub-threshold, linear, or saturation) for each layer or layers of a neural network.
[0046] Other embodiments for the VMM array 32 of Figure 7 are described in U.S. Patent No. 10,748,630, which is incorporated herein by reference. As described in that application, the source lines or bit lines can be used as neuron outputs (current sum outputs).
[0047] FIG. 10 shows a neuron VMM array 1000, which is particularly suitable for the memory cells 210 shown in FIG. 2 and is used as a synapse between an input layer and the next layer. The VMM array 1000 includes a memory array 1003 of non-volatile memory cells, a reference array 1001 of a first non-volatile reference memory cell, and a reference array 1002 of a second non-volatile reference memory cell. The reference arrays 1001 and 1002 arranged in the columns of the array function to convert current inputs flowing into terminals BLR0, BLR1, BLR2, and BLR3 into voltage inputs WL0, WL1, WL2, and WL3. In practice, the first and second non-volatile reference memory cells are diode-connected through a multiplexer 1014 (only a portion of which is shown) with current inputs flowing into them. The reference cells are tuned (e.g., programmed) to a target reference level. The target reference level is provided by a reference mini-array matrix (not shown).
[0048] The memory array 1003 serves two purposes. First, the memory array 1003 stores the weights in each memory cell that are used by the VMM array 1000. Second, the memory array 1003 effectively multiplies the inputs (i.e., the current inputs provided to terminals BLR0, BLR1, BLR2, and BLR3, which the reference arrays 1001 and 1002 convert to input voltages that are provided to the word lines WL0, WL1, WL2, and WL3) by the weights stored in the memory array 1003, and then adds all the results (memory cell currents) to generate outputs on the respective bit lines (BL0-BLN), which are the inputs to the next layer or the inputs to the last layer. By performing the multiplication and addition functions, the memory array 1003 eliminates the need for separate multiplication and addition logic and is also power efficient. Here, voltage inputs are provided to word lines WL0, WL1, WL2, and WL3, and outputs appear on respective bit lines BL0-BLN during a read (inference) operation. The current in each of the bit lines BL0-BLN performs a function of the sum of the currents from all the non-volatile memory cells connected to that particular bit line.
[0049] Table 5 shows the operating voltages and currents for the VMM array 1000. The columns in the table show the voltages applied to the word line of the selected cell, the word lines of the unselected cells, the bit lines of the selected cells, the bit lines of the unselected cells, the source lines of the selected cells, and the source lines of the unselected cells. The rows show the operations of read, erase, and program. Table 5: Operation of VMM Array 1000 in Figure 10 [Table 5]
[0050] FIG. 11 shows a neuron VMM array 1100 that is particularly suited for the memory cells 210 shown in FIG. 2 and that are utilized as part of synapses and neurons between an input layer and the next layer. The VMM array 1100 includes a memory array 1103 of non-volatile memory cells, a reference array 1101 of first non-volatile reference memory cells, and a reference array 1102 of second non-volatile reference memory cells. The reference arrays 1101 and 1102 run in the row direction of the VMM array 1100. The VMM array is similar to the VMM 1000, except that the word lines run vertically in the VMM array 1100. Here, the inputs are the word lines (WLA0, WLB0, WLA1, WLB2, WLB3, WLB4, WLB5, WLB6, WLB7, WLB8, WLB9, WLB10, WLB11, WLB12, WLB13, WLB14, WLB15, WLB16, WLB17, WLB18, WLB19, WLB20, WLB21, WLB22, WLB23, WLB24, WLB25, WLB26, WLB27, WLB28, WLB29, WLB30, WLB31, WLB32, WLB33, WLB34, WLB35, WLB36, WLB37, WLB38, WLB39, WLB40, WLB41, WLB42, WLB43, WLB44, WLB45, WLB46, WLB47, WLB48, WLB50, WLB51, WLB52, WLB53, WLB54, WLB55, WLB56, WLB57, WLB58, WLB69, WLB61, WLB61, WLB61, WLB71, WLB82, WLB93, 1 , WLA2, WLB2, WLA3, WLB3) and the output appears on the source lines (SL0, SL1) during a read operation. The current on each source line performs a function of the sum of all the currents from the memory cells connected to that particular source line.
[0051] Table 6 shows the operating voltages and currents for VMM array 1100. The columns in the table show the voltages applied to the word line of the selected cell, the word lines of the unselected cells, the bit lines of the selected cells, the bit lines of the unselected cells, the source lines of the selected cells, and the source lines of the unselected cells. The rows show the read, erase, and program operations. Table 6: Operation of VMM Array 1100 in FIG. 11 [Table 6]
[0052] 12 shows a neuron VMM array 1200, which is particularly suited for the memory cells 310 shown in FIG. 3 and is used as part of synapses and neurons between the input layer and the next layer. The VMM array 1200 includes a memory array 1203 of non-volatile memory cells, a reference array 1201 of a first non-volatile reference memory cell, and a reference array 1202 of a second non-volatile reference memory cell. The reference arrays 1201 and 1202 function to convert current inputs flowing into terminals BLR0, BLR1, BLR2, and BLR3 into voltage inputs CG0, CG1, CG2, and CG3. In practice, the first and second non-volatile reference memory cells are diode-connected through a multiplexer 1212 (only a portion of which is shown), with the current inputs flowing through BLR0, BLR1, BLR2, and BLR3. The multiplexer 1212 includes a corresponding multiplexer 1205 and cascoding transistor 1204 to ensure a constant voltage on each bit line (e.g., BLR0) of the first and second non-volatile reference memory cells during a read operation. The reference cells are tuned to a target reference level.
[0053] The memory array 1203 serves two purposes. First, it stores the weights used by the VMM array 1200. Second, it effectively multiplies the weights stored in the memory array by the inputs (current inputs provided to terminals BLR0, BLR1, BLR2, and BLR3, which the reference arrays 1201 and 1202 convert to input voltages and provide to the control gates (CG0, CG1, CG2, and CG3)) and then adds all the results (cell currents) to produce an output that appears on BL0-BLN and is the input to the next layer or the input to the last layer. Having the memory array perform the multiplication and addition functions eliminates the need for separate multiplication and addition logic and is also power efficient, where the inputs are provided to the control gate lines (CG0, CG1, CG2, and CG3) and the outputs appear on the bit lines (BL0-BLN) during read operations. The current in each bit line is a function of the sum of all the currents from the memory cells connected to that particular bit line.
[0054] The VMM array 1200 implements one-way tuning of the non-volatile memory cells in the memory array 1203. That is, each non-volatile memory cell is erased and then partially programmed until the desired charge on the floating gate is reached. If too much charge is added to the floating gate (such as an incorrect value being stored in the cell), the cell is erased and the series of partial programming operations starts over again. As shown, two rows that share the same erase gate (such as EG0 or EG1) are erased together (known as a page erase), and then each cell is partially programmed until the desired charge on the floating gate is reached.
[0055] Table 7 shows the operating voltages and currents for VMM array 1200. The columns in the table show the voltages applied to the word line of the selected cell, the word lines of the unselected cells, the bit line of the selected cell, the bit lines of the unselected cells, the control gate of the selected cell, the control gates of the unselected cells in the same sector as the selected cell, the control gates of the unselected cells in a different sector than the selected cell, the erase gate of the selected cell, the erase gates of the unselected cells, the source line of the selected cell, and the source lines of the unselected cells. The rows show the read, erase, and program operations. Table 7: Operation of VMM Array 1200 in FIG. 12 [Table 7]
[0056] 13 shows a neuron VMM array 1300 that is particularly suited for the memory cells 310 shown in FIG. 3 and is used as part of the synapses and neurons between the input layer and the next layer. The VMM array 1300 includes a memory array 1303 of non-volatile memory cells and , No. 1 non-volatile reference memory cell Reference Array 1301and a second reference array 1302 of non-volatile reference memory cells. EG lines EGR0, EG0, EG1, and EGR1 run vertically, and CG lines CG0, CG1, CG2, and CG3 and SL lines WL0, WL1, WL2, and WL3 run horizontally. VMM array 1300 is similar to VMM array 1400 except that VMM array 1300 implements bidirectional tuning, where each individual cell can be fully erased, partially programmed, and partially erased as needed to reach a desired amount of charge on the floating gate through the use of separate EG lines. As shown, reference arrays 1301 and 1302 convert input currents at terminals BLR0, BLR1, BLR2, and BLR3 into control gate voltages CG0, CG1, CG2, and CG3 (through the action of diode-connected reference cells via multiplexer 1314) that are applied to the memory cells in the row direction. The current outputs (neurons) are in bit lines BL0-BLN, each of which sums up all the currents from the non-volatile memory cells connected to that particular bit line.
[0057] Table 8 shows the operating voltages and currents for VMM array 1300. The columns in the table show the voltages applied to the word line of the selected cell, the word lines of the unselected cells, the bit line of the selected cell, the bit lines of the unselected cells, the control gate of the selected cell, the control gates of the unselected cells in the same sector as the selected cell, the control gates of the unselected cells in a different sector than the selected cell, the erase gate of the selected cell, the erase gates of the unselected cells, the source line of the selected cell, and the source lines of the unselected cells. The rows show the read, erase, and program operations. Table 8: Operation of VMM Array 1300 in FIG. 13 [Table 8]
[0058] 22 shows a neuron VMM array 2200 that is particularly suited for the memory cells 210 shown in FIG. 2 and is used as part of the synapses and neurons between the input layer and the next layer. In the VMM array 2200, inputs INPUT0...., INPUTN are bit lines BL0, ..., BL N and outputs OUTPUT1, OUTPUT2, OUTPUT3, and OUTPUT4 are generated on source lines SL0, SL1, SL2, and SL3, respectively.
[0059] 23 illustrates a neuron VMM array 2300 that is particularly suited for memory cells 210 shown in FIG. 2 and that is utilized as part of the synapses and neurons between an input layer and the next layer. In this example, inputs INPUT0, INPUT1, INPUT2, and INPUT3 are received on source lines SL0, SL1, SL2, and SL3, respectively, and outputs OUTPUT0, ..., OUTPUT N are bit lines BL0, ..., BL N is generated.
[0060] 24 shows a neuron VMM array 2400 that is particularly suited for the memory cells 210 shown in FIG. 2 and is used as part of the synapses and neurons between the input layer and the next layer. In this example, inputs INPUT0, ..., INPUT M are the word lines WL0, ..., WL M are received at OUTPUT0, ..., OUTPUT N are bit lines BL0, ..., BL N is generated.
[0061] 25 shows a neuron VMM array 2500 that is particularly suited for the memory cells 310 shown in FIG. 3 and is used as part of the synapses and neurons between the input layer and the next layer. In this example, inputs INPUT0, ..., INPUT M are the word lines WL0, ..., WL M are received at OUTPUT0, ..., OUTPUT N are bit lines BL0, ..., BL N is generated.
[0062] 26 shows a neuron VMM array 2600 that is particularly suited for the memory cells 410 shown in FIG. 4 and is used as part of the synapses and neurons between the input layer and the next layer. In this example, the input INPUT 0、 ..., INPUT n are the vertical control gate lines CG0, ..., CG N and outputs OUTPUT1 and OUTPUT2 are generated on source lines SL0 and SL1.
[0063] 27 shows a neuron VMM array 2700 that is particularly suited for the memory cells 410 shown in FIG. 4 and is used as part of the synapses and neurons between the input layer and the next layer. In this example, inputs INPUT0, ..., INPUT N are bit lines BL0, ..., BL N , 2701-(N-1) and 2701-N, which are respectively coupled to bit line control gates 2701-1, 2701-2, ..., 2701-(N-1) and 2701-N. Exemplary outputs OUTPUT1 and OUTPUT2 are generated on source lines SL0 and SL1.
[0064] 28 shows a neuron VMM array 2800 that is particularly suitable for memory cells 310 shown in FIG. 3, memory cells 510 shown in FIG. 5, and memory cells 710 shown in FIG. 7, and is used as part of synapses and neurons between the input layer and the next layer. In this example, inputs INPUT0, ..., INPUT M are the word lines WL0, ..., WL M Received at OUTPUT0, ..., OUTPUT N are bit lines BL0, ..., BL N are generated respectively.
[0065] 29 shows a neuron VMM array 2900 that is particularly suitable for memory cells 310 shown in FIG. 3, memory cells 510 shown in FIG. 5, and memory cells 710 shown in FIG. 7, and is used as part of synapses and neurons between the input layer and the next layer. In this example, inputs INPUT0, ..., INPUTM are the control gate lines CG0, ..., CG M Received at OUTPUT0, ..., OUTPUT N are the vertical source lines SL0, ..., SL N and each source line SL i is coupled to the source lines of all memory cells in column i.
[0066] 30 shows a neuron VMM array 3000 that is particularly suitable for memory cells 310 shown in FIG. 3, memory cells 510 shown in FIG. 5, and memory cells 710 shown in FIG. 7, and is used as part of synapses and neurons between the input layer and the next layer. In this example, inputs INPUT0, ..., INPUT M are the control gate lines CG0, ..., CG M Received at OUTPUT0, ..., OUTPUT N are the vertical bit lines BL0, ..., BL N and each bit line BL i is coupled to the bit lines of all memory cells in column i.
[0067] Long and short term memory Prior art includes a concept known as long short-term memory (LSTM). LSTM units are often used within neural networks. LSTM allows a neural network to store information for any given period of time and use that information in subsequent operations. A traditional LSTM unit includes a cell, an input gate, an output gate, and a forget gate. The three gates regulate the flow of information into and out of the cell, and the period of time that information is stored within the LSTM. VMMs are particularly useful in LSTM units.
[0068] FIG. 14 illustrates an exemplary LSTM 1400. The LSTM 1400 in this example includes cells 1401, 1402, 1403, and 1404. Cell 1401 receives an input vector x0 and generates an output vector h0 and a cell state vector c0. Cell 1402 receives an input vector x1, an output vector (hidden state) h0 from cell 1401, and a cell state c0 from cell 1401, and generates an output vector h1 and a cell state vector c1. Cell 1403 receives an input vector x2, an output vector (hidden state) h1 from cell 1402, and a cell state c1 from cell 1402, and generates an output vector h2 and a cell state vector c2. Cell 1404 receives an input vector x3, an output vector (hidden state) h2 from cell 1403, and a cell state c2 from cell 1403, and generates an output vector h3. Additional cells can be used, the LSTM with four cells is just an example.
[0069] Figure 15 shows an example implementation of an LSTM cell 1500 that can be used for cells 1401, 1402, 1403, and 1404 in Figure 14. The LSTM cell 1500 receives an input vector x(t), a cell state vector c(t-1) from a previous cell, and an output vector h(t-1) from a previous cell, and produces a cell state vector c(t) and an output vector h(t).
[0070] LSTM cell 1500 includes sigmoid function devices 1501, 1502, and 1503, each of which applies a number between 0 and 1 to control the degree to which each component of the input vector contributes to the output vector. LSTM cell 1500 also includes tanh devices 1504 and 1505 for applying a hyperbolic tangent function to the input vector, multiplier devices 1506, 1507, and 1508 for multiplying two vectors, and adder device 1509 for adding the two vectors. The output vector h(t) can be provided to the next LSTM cell in the system or can be accessed for other purposes.
[0071] FIG. 16 shows an example of an implementation of the LSTM cell 1500, the LSTM cell 1600. For the convenience of the reader, the same numbering scheme from the LSTM cell 1500 is used in the LSTM cell 1600. The sigmoid function devices 1501, 1502, and 1503, and the tanh device 1504 each include a plurality of VMM arrays 1601 and activation function blocks 1602. Thus, the VMM array proves to be particularly useful in the LSTM cell used in certain neural network systems. The multiplier devices 1506, 1507, and 1508, and the adder device 1509 are implemented in a digital or analog manner. The activation function block 1602 can be implemented in a digital or analog manner.
[0072] An alternative example of LSTM cell 1600 (and another example of one implementation of LSTM cell 1500) is shown in Figure 17. In Figure 17, sigmoid function devices 1501, 1502, and 1503, and tanh device 1504 share the same physical hardware (VMM array 1701 and activation function block 1702) in a time-multiplexed manner. LSTM cell 1700 also includes a multiplier device 1703 for multiplying two vectors, an adder device 1708 for adding two vectors, a tanh device 1505 (which includes activation function block 1702), a register 1707 for storing the value i(t) as it is output from sigmoid function block 1702, and a register 1708 for storing the value f(t). * a register 1704 for storing c(t-1) as that value is output from the multiplier device 1703 via a multiplexer 1710; * A register 1705 for storing u(t) as its value is output from the multiplier device 1703 via a multiplexer 1710; * It includes a register 1706 for storing {tilde over (c)}(t) as its value is output from the multiplier device 1703 via a multiplexer 1710, and a multiplexer 1709.
[0073] Whereas LSTM cell 1600 includes multiple sets of VMM arrays 1601 and respective activation function blocks 1602, LSTM cell 1700 includes only one set of VMM arrays 1701 and activation function blocks 1702, which are used to represent multiple layers in embodiments of LSTM cell 1700. LSTM cell 1700 requires less space than LSTM cell 1600 because LSTM cell 1700 requires ¼ the space for the VMMs and activation function blocks compared to LSTM cell 1600.
[0074] It can be further appreciated that an LSTM unit typically includes multiple VMM arrays, each of which requires functionality provided by specific circuit blocks outside the VMM array, such as adder and activation function blocks and high voltage generation blocks. Providing a separate circuit block for each VMM array would require a significant amount of space within a semiconductor device and would be somewhat inefficient. Thus, the embodiments described below reduce the circuitry required outside the VMM array itself.
[0075] Gated Recurrent Unit Analog VMM implementations can be used for gated recurrent unit (GRU) systems. GRUs are a gating mechanism in recurrent neural networks. GRUs are similar to LSTMs, except that GRU cells generally contain fewer components than LSTM cells.
[0076] 18 illustrates an exemplary GRU 1800. The GRU 1800 in this example includes cells 1801, 1802, 1803, and 1804. Cell 1801 receives input vector x0 and generates output vector h0. Cell 1802 receives input vector x1 and output vector h0 from cell 1801 and generates output vector h1. Cell 1803 receives input vector x2 and output vector (hidden state) h1 from cell 1802 and generates output vector h2. Cell 1804 receives input vector x3 and output vector (hidden state) h2 from cell 1803 and generates output vector h3. Additional cells may be used, and a GRU with four cells is merely an example.
[0077] FIG. 19 illustrates an exemplary implementation of a GRU cell 1900 that may be used for cells 1801, 1802, 1803, and 1804 of FIG. 18. GRU cell 1900 receives an input vector x(t) and an output vector h(t-1) from a preceding GRU cell and generates an output vector h(t). GRU cell 1900 includes sigmoid function devices 1901 and 1902, each of which applies a number between 0 and 1 to the output vector h(t-1) and components from the input vector x(t). GRU cell 1900 also includes a tanh device 1903 for applying a hyperbolic tangent function to the input vector, a number of multiplier devices 1904, 1905, and 1906 for multiplying two vectors, an adder device 1907 for adding the two vectors, and a complement device 1908 for subtracting the input from 1 to generate an output.
[0078] FIG. 20 shows a GRU cell 2000, which is an example of one implementation of the GRU cell 1900. For the convenience of the reader, the same numbering method from the GRU cell 1900 is used in the GRU cell 2000. As can be seen from FIG. 20, the sigmoid function devices 1901 and 1902 and the tanh device 1903 each include a plurality of VMM arrays 2001 and activation function blocks 2002. Therefore, it can be seen that the VMM array is particularly used in the GRU cell used in a specific neural network system. The multiplier devices 1904, 1905, 1906, the adder device 1907, and the complementary device 1908 are implemented in a digital or analog manner. The activation function block 2002 can be implemented in a digital or analog manner.
[0079] An alternative example of GRU cell 2000 (and another example of one implementation of GRU cell 1900) is shown in FIG. 21. In FIG. 21, GRU cell 2100 utilizes a VMM array 2101 and an activation function block 2102, which, when configured as a sigmoid function, applies a number between 0 and 1 to control how much each component of the input vector contributes to the output vector. In FIG. 21, sigmoid function devices 1901 and 1902, and tanh device 1903 share the same physical hardware (VMM array 2101 and activation function block 2102) in a time-multiplexed manner. GRU cell 2100 also includes a multiplier device 2103 for multiplying two vectors together, an adder device 2105 for adding two vectors together, a complement device 2109 for subtracting the input from 1 to generate the output, a multiplexer 2104, and a value h(t-1) * A register 2106 for holding r(t) as its value is output from the multiplier device 2103 via multiplexer 2104, and a value h(t-1) * A register 2107 for holding z(t) as its value is output from the multiplier device 2103 via multiplexer 2104, and a register 2108 for holding the value ĥ(t) *and a register 2108 for holding (1-z((t)) as its value is output from the multiplier device 2103 via multiplexer 2104.
[0080] Whereas GRU cell 2000 includes multiple sets of VMM array 2001 and activation function blocks 2002, GRU cell 2100 includes only one set of VMM array 2101 and activation function blocks 2102, which are used to represent multiple tiers in an embodiment of GRU cell 2100. GRU cell 2100 requires less space than GRU cell 2000 because GRU cell 2100 requires ⅓ of the space for the VMM and activation function blocks compared to GRU cell 2000.
[0081] It can be further appreciated that a GRU system typically includes multiple VMM arrays, each of which requires functionality provided by specific circuit blocks outside the VMM array, such as adder and activation function blocks and high voltage generation blocks. Providing separate circuit blocks for each VMM array would require a significant amount of space within a semiconductor device and would be somewhat inefficient. Thus, the embodiments described below reduce the circuitry required outside the VMM array itself.
[0082] The input to the VMM array can be an analog level, a binary level, a pulse, a time modulated pulse, or a digital bit (in this case a DAC is required to convert the digital bit to the appropriate input analog level), and the output can be an analog level, a binary level, a timing pulse, a pulse, or a digital bit (in this case an output ADC is required to convert the output analog level to a digital bit).
[0083] For each memory cell in the VMM array, each weight W can be implemented by a single memory cell, or by a differential cell, or by two blended memory cells (average of two cells). In the case of differential cells, two memory cells are required to implement the weight W as a differential weight (W=W+-W-). In the case of two blended memory cells, two memory cells are required to implement the weight W as the average of two cells.
[0084] FIG. 31 illustrates a VMM system 3100. In some embodiments, the weights W stored in the VMM array are stored as a differential pair, W+ (positive weight) and W- (negative weight), where W=(W+)-(W-). In the VMM system 3100, half of the bit lines are designated as W+ lines, i.e., bit lines that connect to memory cells that store a positive weight W+, and the other half of the bit lines are designated as W- lines, i.e., bit lines that connect to memory cells that implement a negative weight W-. The W- lines are interspersed alternately among the W+ lines. The subtraction operation is performed by summing circuits, such as summing circuits 3101 and 3102, that receive currents from the W+ and W- lines. The output of the W+ lines and the output of the W- lines are combined together to effectively give W=W+-W- for each pair of (W+, W-) cells of every pair of (W+, W-) lines. Although described thus far with respect to W- lines interspersed alternatingly among W+ lines, in other embodiments the W+ and W- lines may be arbitrarily positioned anywhere within the array.
[0085] 32 shows another embodiment. In a VMM system 3210, the positive weights W+ are implemented in a first array 3211 and the negative weights W− are implemented in a second array 3212 that is separate from the first array, and the resulting weights are appropriately combined together by a summing circuit 3213.
[0086] FIG. 33 illustrates a VMM system 3300. The weights W stored in the VMM array are stored as a differential pair, W+ (positive weight) and W- (negative weight), where W=(W+)-(W-). The VMM system 3300 includes an array 3301 and an array 3302. Half of the bit lines in each of the arrays 3301 and 3302 are designated as W+ lines, i.e., bit lines that connect to memory cells that store a positive weight W+, and the other half of the bit lines in each of the arrays 3301 and 3302 are designated as W- lines, i.e., bit lines that connect to memory cells that implement a negative weight W-. The W- lines are interspersed alternately among the W+ lines. Subtraction operations are performed by adder circuits, such as adder circuits 3303, 3304, 3305, and 3306, that receive current from the W+ and W- lines. The output on the W+ line and the output on the W- line from each array 3301, 3302 respectively are combined together to effectively give W=W+-W- for each pair of (W+,W-) cells of every pair of (W+,W-) lines. Furthermore, the W values from each array 3301 and 3302 may be further combined via summing circuits 3307 and 3308, meaning that each W value is the result of subtracting the W value from array 3302 from the W value from array 3301, and the final result from summing circuits 3307 and 3308 is one of two difference values.
[0087] Each non-volatile memory cell used in an analog neural memory system can be erased or programmed to hold a very specific and precise amount of charge, i.e., number of electrons, in its floating gate. For example, each floating gate must hold one of N different values, where N is the number of different weights that can be exhibited by each cell. Examples of N include 16, 32, 64, 128, and 256.
[0088] Similarly, the read operation must be able to accurately distinguish between the N different levels.
[0089] What is needed in a VMM system is an improved output block that can quickly and accurately receive outputs from arrays and identify the values represented by those outputs. Summary of the Invention
[0090] Numerous embodiments are disclosed for an output circuit for an analog neural memory in a deep learning artificial neural network.
[0091]
[0092]
[0093]
[0094]
[0095]
[0096]
[0097]
[0098]
[0099]
[0100]
[0101]
[0102]
[0103]
[0104]
[0105]
[0106]
[0107]
[0108]
[0109]
[0110]
[0111]
[0112]
[0113]
[0114]
[0115]
[0116]
[0117]
[0118]
[0119]
[0120]
[0121]
[0122]
[0123]
[0124]
[0125]
[0126]
[0127]
[0128]
[0129]
[0130]
[0131]
[0132]
[0133]
[0134]
[0135]
[0136]
[0137]
[0138]
[0139]
[0140]
[0141] [Brief description of the drawings]
[0142] [Figure 1] FIG. 1 illustrates an artificial neural network. [Diagram 2] 1 shows a prior art split-gate flash memory cell. [Diagram 3] 1 illustrates another prior art split-gate flash memory cell. [Figure 4] 1 illustrates another prior art split-gate flash memory cell. [Diagram 5] 1 illustrates another prior art split-gate flash memory cell. [Figure 6] FIG. 1 illustrates various levels of an exemplary artificial neural network that utilizes one or more non-volatile memory arrays. [Figure 7] FIG. 1 is a block diagram illustrating a vector matrix multiplication system. [Figure 8] FIG. 1 is a block diagram illustrating an example artificial neural network utilizing one or more vector matrix multiplication systems. [Figure 9] 4 illustrates another embodiment of a vector matrix multiplication system. [Figure 10] 4 illustrates another embodiment of a vector matrix multiplication system. [Figure 11] 4 illustrates another embodiment of a vector matrix multiplication system. [Figure 12] 4 illustrates another embodiment of a vector matrix multiplication system. [Figure 13]4 illustrates another embodiment of a vector matrix multiplication system. [Figure 14] 1 shows a prior art long-term and short-term memory system. [Figure 15] An exemplary cell for use in a long- and short-term memory system is shown. [Figure 16] 16 illustrates an embodiment of the exemplary cell of FIG. 15. [Figure 17] 16 illustrates an alternative embodiment of the exemplary cell of FIG. [Figure 18] 1 shows a prior art gated recurrent unit system. [Figure 19] 1 shows an exemplary cell for use in a gated recurrent unit system. [Figure 20] 20 illustrates an embodiment of the exemplary cell of FIG. 19. [Figure 21] 20 illustrates an alternative embodiment of the exemplary cell of FIG. 19. [Figure 22] 4 illustrates another embodiment of a vector matrix multiplication system. [Figure 23] 4 illustrates another embodiment of a vector matrix multiplication system. [Figure 24] 4 illustrates another embodiment of a vector matrix multiplication system. [Diagram 25] 4 illustrates another embodiment of a vector matrix multiplication system. [Figure 26] 4 illustrates another embodiment of a vector matrix multiplication system. [Figure 27] 4 illustrates another embodiment of a vector matrix multiplication system. [Figure 28] 4 illustrates another embodiment of a vector matrix multiplication system. [Figure 29] 4 illustrates another embodiment of a vector matrix multiplication system. [Diagram 30] 4 illustrates another embodiment of a vector matrix multiplication system. [Diagram 31] 4 illustrates another embodiment of a vector matrix multiplication system. [Diagram 32]4 illustrates another embodiment of a vector matrix multiplication system. [Diagram 33] 4 illustrates another embodiment of a vector matrix multiplication system. [Diagram 34] 4 illustrates another embodiment of a vector matrix multiplication system. [Figure 35A] 1 illustrates an embodiment of an output block. [Figure 35B] 1 illustrates an embodiment of an output block. [Diagram 36] 13 illustrates another embodiment of an output block. [Figure 37A] 13 illustrates another embodiment of an output block. [Figure 37B] 13 illustrates another embodiment of an output block. [Figure 38A] 13 illustrates another embodiment of an output block. [Figure 38B] 13 illustrates another embodiment of an output block. [Figure 39A] 13 illustrates another embodiment of an output block. [Figure 39B] 13 illustrates another embodiment of an output block. [Figure 40A] 13 illustrates another embodiment of an output block. [Figure 40B] 13 illustrates another embodiment of an output block. [Figure 40C] 13 illustrates another embodiment of an output block. [Diagram 41] 1 shows a serial analog-to-digital converter circuit. [Diagram 42] 1 shows a successive approximation register analog-to-digital converter circuit. [Diagram 43] 1 illustrates a pipelined successive approximation register analog-to-digital converter circuit. [Figure 44A] 1 illustrates a hybrid successive approximation register and serial analog-to-digital converter circuit. [Figure 44B] 1 illustrates a hybrid successive approximation register and serial analog-to-digital converter circuit. [Diagram 45] 1 shows an algorithmic analog-to-digital converter block. [Diagram 46] 1 shows the tracking reference generator used in the output block. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0143] The artificial neural network of the present invention utilizes a combination of CMOS technology and non-volatile memory arrays.
[0144] VMM System Overview 34 shows a block diagram of a VMM system 3400. The VMM system 3400 includes a VMM array 3401, a row decoder 3402, a high voltage decoder 3403, a column decoder 3404, a bit line driver 3405, an input circuit 3406, an output circuit 3407, a control logic 3408, and a bias generator 3409. The VMM system 3400 further includes a high voltage generation block 3410 including a charge pump 3411, a charge pump regulator 3412, and a high voltage analog precision level generator 3413. The VMM system 3400 further includes a (program / erase, or weight adjustment) algorithm controller 3414, an analog circuit 3415, a control engine 3416 (which may include special functions such as, but not limited to, arithmetic functions, startup functions, embedded microcontroller logic, etc.), and a test control logic 3417. The systems and methods described below may be implemented in the VMM system 3400.
[0145] The input circuit 3406 may include circuits such as a DAC (digital-to-analog converter), a DPC (digital-to-pulse converter, digital-to-time modulated pulse converter), an AAC (analog-to-analog converter such as a current-to-voltage converter, a logarithmic converter), a PAC (pulse-to-analog level converter), or any other type of converter. The input circuit 3406 may implement normalization, linear or nonlinear up / downscaling functions, or arithmetic functions. The input circuit 3406 may implement a temperature compensation function for the input level. The input circuit 3406 may implement an activation function such as ReLU or sigmoid. The output circuit 3407 may include circuits such as an ADC (analog-to-digital converter for converting neuron analog output to digital bits), an AAC (analog-to-analog converter such as a current-to-voltage converter, a logarithmic converter), an APC (analog-to-pulse converter, analog-to-time modulated pulse converter), or any other type of converter.
[0146] The output circuit 3407 may implement activation functions such as a rectified linear activation function (ReLU) or a sigmoid. The output circuit 3407 may implement statistical normalization, regularization, up / down scaling / gain functions, statistical rounding, or arithmetic functions (e.g., add, subtract, divide, multiply, shift, log) of the neuron outputs. The output circuit 3407 may implement temperature compensation functions for the neuron outputs or array outputs (such as bit line outputs) to keep the power consumption of the array approximately constant, or to increase the accuracy of the array (neuron) outputs, such as by keeping the IV slope approximately the same.
[0147] 35A shows an output block 3500. The output block includes current-to-voltage converters (ITV) 3501-1 through 3501-i (where i is the number of pairs of bit lines W+ and W− that the output block 3500 receives), a multiplexer 3502, sample and hold circuits 3503-1 through 3503-i, a channel multiplexer 3504, and an analog-to-digital converter (ADC) 3505. The output block 3500 receives differential weighted outputs W+ and W− from the bit line pairs in the array, and ultimately generates a digital output, DOUTx, that represents the output of one of the bit line pairs (e.g., the W+ and W− lines) from the ADC 3505.
[0148] Each of the current-to-voltage converters 3501-1 to 3501-i receives analog bit line current signals BLw+ and BLw- (which are bit line outputs generated in response to the input and the stored W+ and W- weights, respectively) and converts them to differential voltages ITVO+ and ITVO-.
[0149] ITVO+ and ITVO- are then received by a multiplexer 3502, which time division multiplexes the outputs from the current-to-voltage converters 3501-1 to 3501-I to S / H circuits 3503-1 to 3503k, where k may be the same as or different from i.
[0150] Each of the S / H circuits 3503-1 to 3503-k samples the received differential voltages and holds them as differential outputs.
[0151] The channel multiplexer 3504 then receives a control signal to select one of the bit lines W+ and W- channels, i.e., one of the bit line pairs, and outputs the differential voltage held by the respective sample and hold circuit 3503 to the ADC 3505, which converts the analog differential voltage output by the respective sample and hold circuit 3503 to a set of digital bits, DOUTx. As shown, the S / H 3503 may be shared across multiple ITV circuits 3501, and the ADC 3505 may operate with multiple ITV circuits in a time division multiplexed manner. Each S / H 3503 may be simply a capacitor or a capacitor followed by a buffer (e.g., an operational amplifier).
[0152] The ADC 3505 may be a hybrid ADC architecture, meaning that it has more than one ADC architecture to perform the conversion. For example, if DOUTx is an 8-bit output, the ADC 3505 may include an ADC sub-architecture for generating bits B7-B4 and another ADC sub-architecture for generating bits B3-B0 from the differential inputs ITVSH+ and ITVSH-. That is, the ADC circuit 3505 may include multiple ADC sub-architectures. Bua It may include an architecture.
[0153] Optionally, certain ADC sub-architectures may be shared among all channels, while other ADC sub-architectures are not shared among all channels.
[0154] In another embodiment, channel mux 3504 and ADC 3505 may be removed and instead the output may be an analog differential voltage from S / H 3503, which may be buffered by an operational amplifier. For example, the use of an analog voltage may be implemented in an all-analog neural network (i.e., one where no digital output or digital input is required for the neural memory array).
[0155] 35B shows an output block 3550. The output block includes current-to-voltage converters (ITV) 3551-1 to 3551-i (where i is the number of pairs of bit lines W+ and W- that the output block 3550 receives), a multiplexer 3552, a differential-to-single-ended converter Diff-to-S converter 3553, sample-and-hold circuits 3554-1 to 3554-k (where k is the same as or different from i), and a channel multiplexer 3554-k. 5 and an analog-to-digital converter (ADC) 355 6 A Diff-to-S converter 3553 is used to convert the differential output from the ITV 3551 signal provided by mux 3552 to a single-ended output. The single-ended output is then input to S / H 3554, mux 3555, and ADC 3556.
[0156] 36 shows an output block 3600. The output block includes summing circuits 3601-1 to 3601-i (such as current mirror circuits, where i is the number of pairs of bit lines BLw+ and BLw- that the output block 3600 receives), current-to-voltage converter circuits (ITV) 3602-1 to 3602-i, a multiplexer 3603, sample and hold circuits 3604-1 to 3604-k (where k is the same as or different from i), a channel multiplexer 3605, and an ADC 3606. The output block 3600 receives differential weighted outputs BLw+ and BLw- from the bit line pairs in the array, and finally produces a digital output, DOUTx, from the ADC 3606 that represents the output of one of the bit line pairs at a time.
[0157] Each of the current addition circuits 3601-1 to 3601-i receives a current from a bit line pair, subtracts the BLw- value from BLw-, and outputs the result as an added current.
[0158] The current-to-voltage converters 3602-1 to 3602-i receive the output summed currents and convert the respective summed currents into differential voltages ITVO+ and ITVO-, which are then received by the multiplexer 3603 and selectively provided to the sample and hold circuits 3604-1 to 3604-k.
[0159] Respective sample and hold circuits 3604 receive the differential voltages ITVOMX+ and ITVOMX−, sample the received differential voltages, and hold them as differential voltage outputs OSH+ and PSH−.
[0160] The channel multiplexer 3605 receives a control signal to select one of the bit line pairs, i.e., channels BLw+ and BLw-, and outputs the voltages held by the respective sample and hold circuits 3604 to the ADC 3606, which converts the voltages into a set of digital bits as DOUTx.
[0161] 37A shows a current-to-voltage converter 3700. The current-to-voltage converter 3700 includes operational amplifiers 3701 and 3702 configured as shown, and variable resistors 3703, 3704, and 3705. The current-to-voltage converter 3700 receives differential output currents shown as variable current sources, BLw+ from the W+ bit line and BLw- from the W- bit line, and generates a single-ended output Vout. The output voltage Vout is = (BLw+ - Blw-) * R, and resistors 3703, 3704, and 3705 each have a value equal to R. The variable resistors in FIG. 37A can be used to scale the output.
[0162] 37B shows current-to-voltage converter 3710. Current-to-voltage converter 3710 includes operational amplifiers 3711, 3712, and 3713, and variable resistors 3714, 3715, 3716, and 3717, configured as shown. Current-to-voltage converter 3710 receives an output current BLw+ from the W+ bit line, shown as a variable current source, and generates an output Vout+ for that line, and receives an output current Blw- from the W- bit line, shown as a variable current source, and generates an output Vout- for that line. Thus, unlike output block 3700, output block 3710 is not a single-ended output, but rather generates a differential voltage representing the differential values BLw+ and BLw-, respectively. The output voltage is Vout+ = Iw+ * R and Vout- = -Rw- * R, and resistors 3714, 3715, 3716, and 3717 each have a value equal to R. The variable resistors in FIG. 37B can be used to scale the output.
[0163] Optionally, the differential output voltages Vout+ and Vout- may be input to an ADC3718, which converts them into a set of digital output bits, Doutx.
[0164] 38A shows a current-to-voltage converter 3800. The current-to-voltage converter 3800 includes operational amplifiers 3801 and 3802, variable capacitors 3803, 3805, and 3806, and controlled switches 3804 and 3807, configured as shown. The current-to-voltage converter 3800 receives a differential output current BLw+ from the W+ bit line, shown as a variable current source, and BLw- from the W- bit line, shown as a variable current source, and generates a single-ended output Vout. The output voltage Vout is given by:= (Iw+ - Iw-) * The integration time t_integration is t_integration / C, and capacitors 3803, 3805, and 3806 each have a capacitance value equal to C. A control circuit (not shown) controls the opening and closing of switches 3804, 3807 to provide the integration time t_integration.
[0165] FIG. 38B illustrates a current-to-voltage converter 3810. The current-to-voltage converter 3810 includes operational amplifiers 3811, 3812, and 3813, variable capacitors 3815, 3816, 3817, and 3819, and switches 3814, 3818, and 3820. The current-to-voltage converter 3810 receives an output current BLw+ from the W+ bit line, shown as a variable current source, and generates an output Vout+ for that line, and receives an output current BLw- from the W- bit line, shown as a variable current source, and generates a Vout- for that line. Thus, unlike the output block 3800, the output block 3810 generates two voltages that represent the respective difference values BLw+ and BLw-. The output voltage is Vout+ = BLw+ * t_integration / C and Vout- = BLw- * The integration time t_integration is t_integration / C, and capacitors 3815, 3816, 3817, and 3819 each have a capacitance value equal to C. A control circuit (not shown) controls the opening and closing of switches 3814, 3818, and 3820 to provide the integration time t_integration.
[0166] Optionally, the differential output voltages Vout+ and Vout- may be input to an ADC3821, which converts them into a set of digital output bits, Doutx.
[0167] 39A shows a current-to-voltage converter 3900. The current-to-voltage converter 3900 includes an operational amplifier 3901 configured as shown, variable integrating resistors 3902 and 3903, controlled switches 3904, 3905, 3906, and 3907, and sample-and-hold capacitors 3908 and 3909. The current-to-voltage converter 3900 receives differential currents BLw+ from the W+ bit line and BLw- from the W- bit line, and outputs voltages Vout+ and Vout-, respectively. The output voltage is Vout+ = (Blw+) * R and Vout- = (BLw-) *R, and resistors 3902 and 3903 each have a value equal to R. Capacitors 3908 and 3909 function as resistors 3902 and 3903, respectively, and as holding S / H capacitors to hold up the output voltage when the input current is interrupted. A control circuit (not shown) controls the opening and closing of switches 3904, 3905, 3906, and 3907 to provide the integration time.
[0168] Optionally, the differential output voltages Vout+ and Vout- may be input to an ADC3910, which converts them into a set of digital output bits, Doutx.
[0169] 39B shows a differential voltage to single-ended voltage converter (Diff-to-S) 3950. The Diff-to-S converter 3950 includes an operational amplifier 3951 and variable integration resistors 3952 and 3953. Output voltage Vout-(Vin1-Vin2) * (R_3852 / R_3953). This is used, for example, as block 3553 in FIG. 35B.
[0170] FIG. 40A shows an output block 4000, which is a hybrid output conversion block. The output block 400 includes multiple sub-architectures, such as SAR and serial ADC sub-architectures, as shown. The output block 4000 receives differential signals Iw+ and Iw-. A successive approximation register analog-to-digital converter SAR 4001 converts the differential signals Iw+ and Iw- to upper digital bits, and then a serial block ADC 4002 converts the signal remaining after the upper bit conversion to lower bits, and outputs all output digital bits together. In one example, for an 8-bit ADC conversion, the SAR ADC 4001 converts a portion of the received differential voltage to MSB bits B7-B4, and the serial ADC 4002 converts a portion of the received differential voltage to LSB bits B3-B0.
[0171] FIG. 40B illustrates an output block 4010. The output block 4010 includes multiple sub-architectures, such as algorithmic ADC and serial ADC sub-architectures, as shown. The output block 4010 receives differential signals Iw+ and Iw-. The algorithmic analog-to-digital converter 4003 converts the differential signals Iw+ and Iw- to more significant digital bits, and then the serial ADC block 4004 converts the signal remaining after the more significant bit conversion to less significant bits, and outputs all output digital bits together. In one example, for an 8-bit ADC conversion, the algorithmic ADC 4003 converts a portion of the received differential voltage to MSB bits B7-B4, and the serial ADC 4004 converts a portion of the received differential voltage to LSB bits B3-B0.
[0172] Figure 40C shows an output block 4020. The output block 4020 receives the differential signals Iw+ and Iw-. The output block 4020 includes a hybrid analog-to-digital converter that converts the differential signals Iw+ and Iw- into digital bits by combining different conversion schemes (such as those shown in Figures 40A and 40B) in one block.
[0173] FIG. 41 shows a configurable serial analog-to-digital converter 4100. The serial analog-to-digital converter 4100 outputs a neuron output current I, which is shown as a variable current source. NEU onto an integration capacitor 4102 (Cint). The integrator 4170 includes a differential amplifier 4101, controlled switches 4108 and 4110, and a control circuit (not shown) that controls the opening and closing of the switches 4108 and 4110 to provide the integration time.
[0174] In one embodiment, VRAMP 4150 is provided to an inverting input of comparator 4104. Digital output (count value) 4121 is generated by ramping VRAMP 4150 until the output of comparator 4104, shown as EC 4105, switches polarity, and counter 4120 counts clock pulses from the start of the ramp of VRAMP 4150 and stops when the output of comparator 4104 switches polarity in response to AND gate 4140 blocking the passage of clock 4141 from reaching counter 4120 as pulse series 4142.
[0175] In another embodiment, VREF4155 is provided to the inverting input of the comparator 4104. VC4110 is ramped down by the ramp current 4151 (IREF) until VOUT4103 reaches VREF4155, at which point the output EC4105 of the comparator 4104 switches polarity and disables the count of the counter 4120. Thus, the counter 4120 is enabled with the closure of switch S2 (which is disabled after the opening of S2, when the output of the comparator 4104, EC4105, switches polarity). S3 is used for initialization (equalization) at the start of operation. The (n-bit) ADC 4100 is configurable to have lower precision (less than n-bit) or higher precision (more than n-bit) depending on the target application. Configurability of the accuracy is achieved by configuring, without limitation, the capacitance of capacitor 4102, the current 4151 (IREF), the ramping rate of VRAMP 4150, or the clock frequency of clock 4141.
[0176] In another embodiment, the ADC circuitry of a VMM array is configured to have less than n bits of precision, and the ADC circuitry of another VMM array is configured to have more than n bits of precision.
[0177] In another embodiment, one instance of the serial ADC circuit 4100 of one neuron (array output) circuit is configured in combination with another instance of the serial ADC circuit 4100 of an adjacent neuron circuit to generate an ADC circuit having an accuracy greater than n bits, such as by combining the integrating capacitors 4102 of the two instances of the serial ADC circuit 4100.
[0178] FIG. 42 shows a configurable SAR (successive approximation register) analog-to-digital converter 4200 used in a neuron output circuit (array output circuit). This circuit is a successive approximation converter based on charge redistribution using binary capacitors. The successive approximation converter includes a binary CDAC (capacitor digital-to-analog converter) 4201, a comparator 4202, and a SAR logic and register 4203. As shown, GndV 4204 is a low voltage reference level, e.g., ground level. The SAR logic and register 4203 provides a digital output 4206. Other non-binary capacitor structures can be implemented using weighted reference voltages or compensation with the output.
[0179] FIG. 43 shows a pipelined SAR ADC circuit 4300 that can be used in combination with a subsequent SAR ADC to increase the number of bits in a pipelined manner. The SAR ADC circuit 4300 includes a binary CDAC 4301, a comparator 4302 (operating as an opamp or comparator), an opamp / comparator 4303, and a SAR logic and registers 4304. As shown, GndV 3104 is a low voltage reference level, e.g., ground level. The SAR logic and registers 4304 provide a digital output 4306. Vin is at the input voltage, VREF is a reference voltage, and GndV is a low voltage, such as ground voltage. The V residual is generated by a capacitor 4305 and provided as an input to the next stage of the SAR ADC conversion sequence.
[0180] FIG. 44A shows a hybrid SAR+serial ADC circuit 4400 that can be used to increase the number of bits in a hybrid manner. The SAR ADC circuit 4400 includes a binary CDAC 4401 and a comparator 4402. 2 and , SAR logic and registers 4403. As shown, GndV is a low voltage reference level, e.g., ground level during SAR ADC operation. SAR logic and registers 4403 provides the digital output. Vin is the input voltage. VREFRAMP is used as a reference ramp voltage during serial ADC operation with appropriate control circuitry and signal multiplexing (not shown).
[0181] Other hybrid ADC architectures that may be used include SAR ADC+sigma delta ADC, flash ADC+serial ADC, pipelined ADC+serial ADC, serial ADC+SAR ADC, and other architectures.
[0182] FIG. 44B shows a hybrid differential SAR+serial ADC circuit 4400 that can be used to increase the number of bits in a hybrid manner.
[0183] 45 shows an algorithmic ADC output block 4500. Output block 4500 includes a sample and hold circuit 4501, a 1-bit analog to digital converter 4502, a 1-bit digital to analog converter 4503, a summer 4504, an operational amplifier 4505, and controlled switches 4506 and 4507, configured as shown. The operational amplifier 4505 is shown configured to provide a gain of 2.
[0184] FIG. 46 shows a tracking voltage reference generator 4600 used to generate reference voltages that may be used by the output circuits described herein and components of such output circuits, such as in FIGS. 37A, 37B, 38A, 38B, 41, 42, 43, 44A, and 44B.
[0185] The tracking voltage reference generator 4600 includes a bias current 4601 and a variable resistor 4602, and outputs VREFx 4603=i * R, where i is the current from bias current 4601 and R is the resistance of variable resistor 4602.
[0186] It should be noted that, as used herein, both the terms "over" and "on" are inclusive of "directly" (with no intermediate material, element, or gap disposed between them) and "indirectly" (with an intermediate material, element, or gap disposed between them). Similarly, the term "adjacent" includes "directly adjacent" (with no intermediate material, element, or gap disposed between them) and "indirectly adjacent" (with an intermediate material, element, or gap disposed between them), "attached" includes "directly attached" (with no intermediate material, element, or gap disposed between them) and "indirectly attached" (with an intermediate material, element, or gap disposed between them), and "electrically coupled" includes "directly electrically coupled" (without an intermediate material or element disposed between them that electrically connects the elements together) and "indirectly electrically coupled" (with an intermediate material or element disposed between them that electrically connects the elements together). For example, forming an element "above the substrate" can include forming the element directly on the substrate, with no intermediate materials / elements between them, and forming the element indirectly on the substrate, with one or more intermediate materials / elements between them.
Claims
1. 1. An output circuit for generating an output from one or more arrays of non-volatile memory cells, comprising: a plurality of current-to-voltage converters, each of the plurality of current-to-voltage converters receiving current from a respective bit line coupled to one or more non-volatile memory cells of the one or more arrays that store a W+ value and from a respective bit line coupled to one or more non-volatile memory cells of the one or more arrays that store a W- value; a multiplexer for receiving respective voltage outputs from the plurality of current-to-voltage converters; a plurality of sample and hold circuits, each sample and hold circuit selectively coupled to one of the plurality of current to voltage converters by the multiplexer to generate a held voltage output; a channel multiplexer that receives the held voltage outputs from the plurality of sample and hold circuits and that generates as an output a pair of the held voltage outputs in response to a control signal; an analog-to-digital converter that converts the output from the channel multiplexer into a digital output;
2. 2. The output circuit of claim 1, wherein the bit lines coupled to the one or more non-volatile memory cells storing a W+ value and the bit lines coupled to the one or more non-volatile memory cells storing a W- value are located within the same array of the one or more arrays.
3. 2. The output circuit of claim 1, wherein the bit lines coupled to the one or more non-volatile memory cells storing a W+ value and the bit lines coupled to the one or more non-volatile memory cells storing a W- value are located in different arrays of the one or more arrays.
4. 2. The output circuit of claim 1, wherein each of the one or more arrays of non-volatile memory cells is a neural network memory array.
5. 10. The output circuit of claim 1, wherein each non-volatile memory cell in the one or more arrays is capable of storing one of three or more possible values.
6. 2. The output circuit of claim 1, wherein each of the non-volatile memory cells in the one or more arrays is a split-gate flash memory cell.
7. An output circuit for generating an output from one or more arrays of non-volatile memory cells, the output circuit comprising: a plurality of summing circuits, each of the plurality of summing circuits receiving current from a respective bit line coupled to one or more non-volatile memory cells of the one or more arrays of non-volatile memory cells storing a W+ value and from a respective bit line coupled to one or more non-volatile memory cells of the one or more arrays of non-volatile memory cells storing a W− value, each of the plurality of summing circuits generating a respective summing current; a plurality of current-to-voltage converters, each of the current-to-voltage converters receiving a respective summed current from one of the plurality of summing circuits and generating a respective summed voltage; a multiplexer receiving respective summed voltages from the plurality of current-to-voltage converters; a plurality of sample and hold circuits, each sample and hold circuit selectively coupled to one of the plurality of current-to-voltage converters by the multiplexer to generate a respective held voltage; a channel multiplexer that receives the respective held voltages from the plurality of sample and hold circuits and generates an output of the held voltages as a channel output in response to a control signal; an analog-to-digital converter that converts the channel outputs received from the channel multiplexer into digital outputs;
8. 8. The output circuit of claim 7, wherein the bit lines coupled to the one or more non-volatile memory cells storing a W+ value and the bit lines coupled to the one or more non-volatile memory cells storing a W- value are located within the same array of one or more arrays of the non-volatile memory cells.
9. 8. The output circuit of claim 7, wherein the bit lines coupled to the one or more non-volatile memory cells storing a W+ value and the bit lines coupled to the one or more non-volatile memory cells storing a W- value are located in different arrays of the one or more arrays of non-volatile memory cells.
10. 8. The output circuit of claim 7, wherein each of the one or more arrays of non-volatile memory cells is a neural network memory array.
11. 8. The output circuit of claim 7, wherein each non-volatile memory cell in the one or more arrays is capable of storing one of three or more possible values.
12. 8. The output circuit of claim 7, wherein each of the non-volatile memory cells in the one or more arrays is a split-gate flash memory cell.
Citation Information
Patent Citations
Design of digital-analog hybrid reading circuit applied to eFlash-based storage and calculation integrated circuit
CN111193511A
Read-out unit for memory cell array and storage and calculation integrated chip comprising same
CN112349316A