Analog Neural Memory Array in an Artificial Neural Network with a Source Line Pull-Down Mechanism
The improved source line pull-down mechanism in non-volatile memory arrays addresses the inefficiencies of existing neural networks by minimizing voltage drops and enhancing energy efficiency and computational speed in neural network operations.
Patent Information
- Application Number
- JP2024065620
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-11-05
- Filing Date
- 2024-04-15
- Publication Date
- 2025-07-10
- Estimated Expiration
- 2040-11-16
AI Technical Summary
Existing artificial neural networks face challenges in high-performance information processing due to the lack of appropriate hardware technologies, particularly in efficiently implementing large numbers of synapses with low energy consumption and minimal voltage drops during operations.
An improved mechanism for quickly pulling down the source line to ground in non-volatile memory arrays, minimizing voltage drops during read, program, or erase operations, and enabling efficient in-memory computing using CMOS technology and non-volatile memory arrays.
The solution enhances the efficiency and accuracy of neural network operations by reducing parasitic impedance variations and minimizing voltage drops, thereby improving energy efficiency and computational speed.
Smart Images

Figure 0007705978000022 
Figure 0007705978000023 
Figure 0007705978000024
Abstract
Description
Technical Field
[0001] (Claims of Priority) This application claims the benefit of priority of U.S. Provisional Patent Application No. 63 / 022,551, filed May 10, 2020, and entitled "Analog Neural Memory Array in Artificial Neural Network with Source Line Pulldown Mechanism", and U.S. Patent Application No. 17 / 090,481, filed Nov. 5, 2020, and entitled "Analog Neural Memory Array in Artificial Neural Network with Source Line Pulldown Mechanism".
[0002] (Field of the Invention) Numerous embodiments of analog neural memory arrays are disclosed. Certain embodiments include an improved mechanism for accurately pulldown a source line to ground. This is useful, for example, to minimize voltage drops during read, program, or erase operations.
Background Art
[0003] An artificial neural network closely resembles a biological neural network (e.g., the central nervous system of an animal, particularly the brain), and is used to estimate or approximate a function that may depend on a large number of inputs and is generally unknown. An artificial neural network generally includes layers of interconnected "neurons" that exchange messages.
[0004] Figure 1 shows an artificial neural network, in which the circles represent input or neuron layers. Connections (referred to as synapses) are represented by arrows and have numerical weights that can be tuned based on experience. This enables the artificial neural network to adapt to the input and become learnable. Typically, an artificial neural network includes multiple input layers. Typically, there is one or more intermediate layers of neurons and an output layer of neurons that provides the output of the neural network. At each level, the neurons make decisions individually or collectively based on the data received from the synapses.
[0005] One of the main challenges in the development of artificial neural networks for high-performance information processing is the lack of appropriate hardware technologies. In practice, practical artificial neural networks rely on a very large number of synapses, which enables a high degree of connectivity between neurons, that is, a very high degree of parallelization of computational processing. In principle, such complexity can be realized by digital supercomputers or dedicated graphics processing unit clusters. However, in addition to high costs, these approaches also suffer from poor energy efficiency compared to biological networks, which mainly perform low-precision analog calculations and consume far less energy. CMOS analog circuits have been used in artificial neural networks, but most CMOS implementation synapses have been too bulky assuming a large number of neurons and synapses.
[0006] The applicant previously disclosed in U.S. Patent Application No. 15 / 594,439, published as U.S. Patent Publication 2017 / 0337466, incorporated by reference, an artificial (analog) neural network that utilizes one or more non-volatile memory arrays as synapses. The non-volatile memory arrays operate as analog neuromorphic memories. As used herein, the term neuromorphic means a circuit that implements a model of a nervous system. An analog neuromorphic memory includes a first plurality of synapses configured to receive a first plurality of inputs and then generate a first plurality of outputs, and a first plurality of neurons configured to receive the first plurality of outputs. The first plurality of synapses includes a plurality of memory cells, each memory cell including a spaced-apart source region and drain region formed in a semiconductor substrate with a channel region extending therebetween, a floating gate insulated and disposed above a first portion of the channel region, and a non-floating gate insulated and disposed above a second portion of the channel region. Each of the plurality of memory cells is configured to store a weight value corresponding to a plurality of electrons in the floating gate. The plurality of memory cells is configured to multiply the stored weight values by the first plurality of inputs to generate the first plurality of outputs. An array of memory cells arranged in this manner may be referred to as a vector matrix multiplication (VMM) array.
[0007] Here, examples of different non-volatile memory cells that can be used in a VMM are discussed.
[0008] <<Non-volatile memory cell>> Various types of known non-volatile memory cells can be used in a VMM array. For example, U.S. Patent No. 5,029,130 (the “’130 Patent”), which is incorporated herein by reference, discloses an array of split-gate non-volatile memory cells, which are a type of flash memory cell. Such a memory cell 210 is shown in FIG. 2. Each memory cell 210 includes a source region 14 and a drain region 16 formed in a semiconductor substrate 12, with a channel region 18 therebetween. A floating gate 20 is formed insulated above a first portion of the channel region 18 (and controls the conductivity of the first portion of the channel region 18) and extends over a portion of the source region 14. A word line terminal 22 (typically coupled to a word line) is disposed insulated above a second portion of the channel region 18 and has a first portion (which controls the conductivity of the second portion of the channel region 18) and a second portion that extends upward above the floating gate 20. The floating gate 20 and the word line terminal 22 are insulated from the substrate 12 by a gate oxide. A bit line terminal 24 is coupled to the drain region 16.
[0009] By applying a high positive voltage to the word line terminal 22, the memory cell 210 is erased (electrons are removed from the floating gate), whereby electrons in the floating gate 20 pass through the insulator therebetween from the floating gate 20 to the word line terminal 22 via Fowler-Nordheim tunneling.
[0010] The memory cell 210 is programmed (electrons are applied to the floating gate) by applying a positive voltage to the word line terminal 22 and a positive voltage to the source region 14. An electron current flows from the drain region 16 toward the source region 14 (source line terminal). The electrons are accelerated and are excited (heat up) when they reach the gap between the word line terminal 22 and the floating gate 20. A portion of the heated electrons is injected into the floating gate 20 through the gate oxide due to the electrostatic attraction from the floating gate 20.
[0011] The memory cell 210 is read by applying a positive read voltage to the drain region 16 and the word line terminal 22 (turning on the portion of the channel region 18 below the word line terminal). When the floating gate 20 is positively charged (i.e., electrons are erased), the portion of the channel region 18 below the floating gate 20 is also turned on, and current flows through the channel region 18, which is detected as the erased state, i.e., the "1" state. When the floating gate 20 is negatively charged (i.e., programmed with electrons), the portion of the channel region below the floating gate 20 becomes almost or completely off, and current does not (or hardly) flow through the channel region 18, which is detected as the programmed state, i.e., the "0" state.
[0012] Table 1 shows typical voltage ranges that can be applied to the terminals of the memory cell 110 to perform read, erase, and program operations. Table 1: Operation of the flash memory cell 210 in FIG. 2
[0013] [Table 1] "Read 1" is a read mode in which the cell current is output to the bit line. "Read 2" is a read mode in which the cell current is output to the source line terminal.
[0014] FIG. 3 shows a memory cell 310 similar to the memory cell 210 in FIG. 2 with an additional control gate (CG) terminal 28. The control gate terminal 28 is biased at a high voltage (e.g., 10V) during programming, a low or negative voltage (e.g., 0V / -8V) during erase, and a low or medium voltage (e.g., 0V / 2.5V) during read. The other terminals are biased in the same manner as the terminals in FIG. 2.
[0015] FIG. 4 shows a four-gate memory cell 410 comprising a source region 14, a drain region 16, a floating gate 20 above a first portion of the channel region 18, a select gate 22 (typically coupled to a word line, WL) above a second portion of the channel region 18, a control gate 28 above the floating gate 20, and an erase gate 30 above the source region 14. This configuration is described in U.S. Patent No. 6,747,310, which is incorporated herein by reference for all purposes. Here, all gates except the floating gate 20 are non-floating gates, i.e., they are electrically connected or connectable to a voltage source. Programming is performed by injecting hot electrons from the channel region 18 into the floating gate 20 by the electrons themselves. Erasure is performed by tunneling electrons from the floating gate 20 to the erase gate 30.
[0016] Table 2 shows typical voltage ranges that can be applied to the terminals of the memory cell 410 to perform read, erase, and program operations. Table 2: Operation of the Flash Memory Cell 410 of FIG. 4 [Table 2] "Read 1" is a read mode in which the cell current is output to the bit line. "Read 2" is a read mode in which the cell current is output to the source line terminal.
[0017] FIG. 5 shows a memory cell 510 similar to the memory cell 410 of FIG. 4 except that the memory cell 510 does not include an erase gate (EG) terminal. Erasure is performed by biasing the substrate 18 to a high voltage and biasing the control gate CG terminal 28 to a low voltage or a negative voltage. Alternatively, erasure is performed by biasing the word line terminal 22 to a positive voltage and biasing the control gate terminal 28 to a negative voltage. Programming and reading are the same as those in FIG. 4.
[0018] Figure 6 shows a three-gate memory cell 610, which is another type of flash memory cell. Memory cell 610 is identical to memory cell 410 of FIG. 4, except that memory cell 610 does not have a separate control gate terminal. (Erasure occurs through the use of an erase gate terminal) The erase operation and the read operation are the same as those of FIG. 4, except that no control gate bias is applied. The programming operation is also performed without a control gate bias. As a result, during the program operation, a higher voltage must be applied to the source line terminal to compensate for the lack of control gate bias.
[0019] Table 3 shows the typical voltage ranges that can be applied to the terminals of memory cell 610 to perform read, erase, and program operations. Table 3: Operation of Flash Memory Cell 610 of FIG. 6
Table 3
[0020] Figure 7 shows a stacked gate memory cell 710, which is another type of flash memory cell. Memory cell 710 is the same as memory cell 210 of FIG. 2, except that the floating gate 20 extends over the entire channel region 18 and the control gate terminal 22 (coupled to the word line) is separated by an insulating layer (not shown) and extends above the floating gate 20. Programming is performed using hot electron injection from channel 18 to floating gate 20 in the channel region adjacent to drain region 16, and erasure is performed using Fowler-Nordheim electron tunneling from floating gate 20 to substrate 12. The read operation operates in the same manner as described above for memory cell 210.
[0021] Table 4 shows the typical voltage ranges that can be applied to the terminals of memory cell 710 and substrate 12 to perform read, erase, and program operations. Table 4: Operations of Flash Memory Cell 710 in FIG. 7 [Table 4]
[0022] "Read 1" is a read mode in which the cell current is output to the bit line. "Read 2" is a read mode in which the cell current is output to the source line terminal. Optionally, in an array including rows and columns of memory cells 210, 310, 410, 510, 610, or 710, the source line may be coupled to one row of memory cells or two adjacent rows of memory cells. That is, the source line terminal may be shared by adjacent rows of memory cells.
[0023] FIG. 8 shows a twin split gate memory cell 810. The memory cell 810 includes a floating gate (FG) 20 that is disposed above the substrate 12 and insulated, a control gate 28 (CG) that is disposed above the floating gate 20 and insulated, and an erase gate 30 (EG) that is disposed above the floating gate 20 and the control gate 28 and insulated and is disposed above the substrate 12. The erase gate 30 is formed in a T-shape such that the upper corner of the control gate CG28 faces the inner corner of the T-shaped erase gate to improve the erase efficiency. The erase gate 30 (EG) and a drain region 16 (DR) in the substrate adjacent to the floating gate 20 where a bit line contact 24 (BL) is connected to the drain diffusion region 16 (DR). The memory cell is formed as a memory cell pair (A on the left and B on the right) and shares a common erase gate 30. This cell design is different from the memory cells described above with reference to FIGS. 2-7 in that at least it lacks a source region under the erase gate EG30, lacks a select gate (also called a word line), and lacks a channel region for each memory cell 810. Instead, a single continuous channel region 18 extends under both memory cells 810 (i.e., extends from the drain region 16 of one memory cell 810 to the drain region 16 of the other memory cell 810). To read or program one memory cell 810, the control gate 28 of the other memory cell 810 is raised to a sufficient voltage, and the voltage coupling to the floating gate 20 therebetween activates the underlying channel region portion (e.g., to read or program cell A, the voltage of FGB is raised by voltage coupling from CGB to activate the channel region under FGB). Erase is performed using Fowler Nordheim electron tunneling from the floating gate 20 to the erase gate 30. Programming is performed using hot electron injection from the channel region 18 to the floating gate 20, which is shown as Program 1 in Table 5 below. Alternatively, programming is performed using Fowler Nordheim electron tunneling from the erase gate 30 to the floating gate 20, which is shown as Program 2 in Table 5.Alternatively, programming is performed using Fowler-Nordheim electron tunneling from channel 18 to floating gate 20, where the conditions are the same as program 2 except that the substrate 12 is biased at a low voltage or a negative voltage while the erase gate 30 is biased at a low positive voltage.
[0024] Table 5 shows typical voltage ranges that can be applied to the terminals of memory cell 810 to perform read, erase, and program operations. Cell A (FG, CGA, BLA) is selected for read, program, and erase operations. Table 5: Operation of Flash Memory Cell 810 of FIG. 8
[0025]
Table 5
[0026] To utilize a memory array that includes one of the types of non-volatile memory cells in the artificial neural network described above, two modifications are made. First, the lines are configured so that each memory cell can be individually programmed, erased, and read without adversely affecting the memory states of other memory cells in the array, as further described below. Second, continuous (analog) programming of the memory cells is provided.
[0027] Specifically, the memory state of each memory cell in the array (i.e., the charge in the floating gate) can be continuously changed independently, with minimal interference from other memory cells, from a completely erased state to a completely programmed state. In another embodiment, the memory state of each memory cell in the array (i.e., the charge in the floating gate) can be continuously changed independently, with minimal interference from other memory cells, from a completely programmed state to a completely erased state and vice versa. This means that the cell memory is either analog or can store at least one of a number of discrete values (such as 16 or 64 different values), which allows all cells in the memory array to be very accurately and individually tunable, and makes the memory array ideal for fine-tuning adjustments to memory and the synaptic weights of neural networks.
[0028] The methods and means described herein can be applied, without limitation, to other non-volatile memory technologies such as FINFET split gate flash or stacked gate flash memory, NAND flash, SONOS (silicon-oxide-nitride-oxide-silicon, charge trapping in nitride), MONOS (metal-oxide-nitride-oxide-silicon, metal charge trapping in nitride), ReRAM (resistive change memory), PCM (phase change memory), MRAM (magnetoresistive memory), FeRAM (ferroelectric memory), OTP (one-time programmable with two or multi-levels only), and CeRAM (strongly correlated electron memory). The methods and means described herein can be applied, without limitation, to volatile memory technologies used in neural networks such as SRAM, DRAM, and other volatile synaptic cells.
[0029] <<Neural Networks Using Non-Volatile Memory Cell Arrays>> Figure 9 conceptually shows a non-limiting example of a neural network that utilizes the non-volatile memory array of this embodiment. This example uses a non-volatile memory array neural network for a face recognition application, but it is also possible to implement other suitable applications using a non-volatile memory array-based neural network.
[0030] S0 is the input layer, and in this example, it is a 32×32 pixel RGB image with 5-bit precision (i.e., three 32×32 pixel arrays, one for each of the colors R, G, and B, and each pixel has 5-bit precision). The synapses CB1 from the input layer S0 to the layer C1 apply different sets of weights to some instances and shared weights to other instances, scan the input image with an overlapping 3×3 pixel filter (kernel), and shift the filter by one pixel (or more than two pixels depending on the model) at a time. Specifically, the nine pixel values in the 3×3 portion of the image (i.e., what is referred to as the filter or kernel) are provided to the synapses CB1, where these nine input values are multiplied by appropriate weights, and after summing the outputs of the multiplications, a single output value is determined and given by the first synapse of CB1 to generate one pixel of the layer of the feature map C1. The 3×3 filter is then shifted one pixel to the right within the input layer S0 (i.e., a 3-pixel column is added on the right and a 3-pixel column is dropped on the left), and thus the nine pixel values of this newly positioned filter are provided to the synapses CB1, where they are multiplied by the same weights as above, and a second single output value is determined by the associated synapses. This process is continued until the 3×3 filter has scanned the entire 32×32 pixel image of the input layer S0 for all three colors and all bits (precision values). The process is then repeated using different sets of weights until all the feature maps of layer C1 are calculated, generating different feature maps of C1.
[0031] In this example, in layer C1, there are 16 feature maps each having 30×30 pixels. Each pixel is a new feature pixel extracted from the multiplication of the input and the kernel. Thus, each feature map is a two-dimensional array. Therefore, in this example, layer C1 consists of 16 layers of two-dimensional arrays (note that the layers and arrays referred to in this specification are not necessarily physical relationships but logical relationships, that is, the arrays are not necessarily oriented in a physical two-dimensional array). Each of the 16 feature maps in layer C1 is generated by one of 16 different sets of synaptic weights applied to the filter scan. All of the C1 feature maps can target different aspects of the same image feature, such as edge identification. For example, the first map (generated using a first set of weights shared by all scans used to generate this first map) can identify circular edges, and the second map (generated using a second set of weights different from the first set of weights) can identify rectangular edges or the aspect ratio of a specific feature, etc.
[0032] Before going from layer C1 to layer S1, an activation function P1 (pooling) that pools values from non-overlapping and consecutive 2×2 regions within each feature map is applied. The purpose of the pooling function is to average neighboring positions (or it is also possible to use the max function), for example, to reduce the dependence on edge positions, and to reduce the data size before going to the next stage. In layer S1, there are 16 15×15 feature maps (i.e., 16 different arrays of 15×15 pixels each). The synapse CB2 from layer S1 to layer C2 scans the maps in S1 with a 4×4 filter with a 1-pixel filter shift. In layer C2, there are 22 12×12 feature maps. Before going from layer C2 to layer S2, an activation function P2 (pooling) that pools values from non-overlapping and consecutive 2×2 regions within each feature map is applied. In layer S2, there are 22 6×6 feature maps. In the synapse CB3 from layer S2 to layer C3, an activation function (pooling) is applied, where all neurons in layer C3 are connected to all maps in layer S2 via each synapse of CB3. In layer C3, there are 64 neurons. The synapse CB4 from layer C3 to the output layer S3 fully connects C3 to S3, i.e., all neurons in layer C3 are connected to all neurons in layer S3. The output in S3 contains 10 neurons, and the neuron with the highest output determines the class (classification). This output can indicate, for example, the identification or classification of the content of the original image.
[0033] Each layer of the synapse is implemented using an array or a part of an array of non-volatile memory cells.
[0034] Figure 10 is a block diagram of a system that can be used for that purpose. The VMM system 32 includes non-volatile memory cells and is used as synapses (such as CB1, CB2, CB3, and CB4 in FIG. 6) between one layer and the next layer. Specifically, the VMM system 32 includes a VMM array 33 including non-volatile memory cells arranged in rows and columns, an erase gate and word line gate decoder 34, a control gate decoder 35, a bit line decoder 36, and a source line decoder 37, and these decoders decode respective inputs to the non-volatile memory cell array 33. Inputs to the VMM array 33 can be made from the erase gate and word line gate decoder 34 or from the control gate decoder 35. The source line decoder 37 in this example also decodes the output of the VMM array 33. Alternatively, the bit line decoder 36 can decode the output of the VMM array 33.
[0035] The VMM array 33 serves two purposes. First, it stores the weights used by the VMM system 32. Second, the VMM array 33 effectively multiplies the inputs by the weights stored in the VMM array 33 and sums them for each output line (source line or bit line) to generate an output, which becomes the input to the next layer or the input to the last layer. By performing the functions of multiplication and addition, the VMM array 33 eliminates the need for separate multiplication and addition logic circuits and is also power-efficient due to in-memory computing on the fly.
[0036] The output of the VMM array 33 is supplied to a differential adder (such as an addition operation amplifier or an addition current mirror) 38 that sums the outputs of the VMM array 33 to create a single value for convolution. The differential adder 38 is arranged to perform the sum of both the positive weight input and the negative weight input and output a single value.
[0037] The summed output value of the differential adder 38 is then supplied to an activation function circuit 39 that rectifies the output. The activation function circuit 39 can provide a sigmoid function, a tanh function, a ReLU function, or any other non-linear function. The rectified output value of the activation function circuit 39 becomes an element of the feature map of the next layer (e.g., C1 in FIG. 9), and is then applied to the next synapse to generate the next feature map layer or the last layer. Thus, in this example, the VMM array 33 constitutes a plurality of synapses (which receive inputs from the previous layer of neurons or from an input layer such as an image database), and the adder 38 and the activation function circuit 39 constitute a plurality of neurons.
[0038] The inputs (WLx, EGx, CGx, and optionally BLx and SLx) to the VMM system 32 of FIG. 10 can be at an analog level, a binary level, a digital pulse (in which case a pulse - analog converter PAC may be required to convert the pulse to an appropriate input analog level), or a digital bit (in which case a DAC is provided to convert the digital bit to an appropriate input analog level), and the output can be at an analog level (e.g., current, voltage, or charge), a binary level, a digital pulse, or a digital bit (in which case an output ADC is provided to convert the output analog level to a digital bit).
[0039] FIG. 11 is a block diagram showing the use of multiple layers of the VMM system 32, labeled as VMM systems 32a, 32b, 32c, 32d, and 32e in the figure. As shown in FIG. 11, an input (indicated as Inputx) is converted from digital to analog by a digital - analog converter 31 and provided to the input VMM system 32a. The converted analog input can be a voltage or a current. The input D / A conversion of the first layer can be performed by using a function or a LUT (look - up table) that maps the input Inputx to an appropriate analog level of the matrix multiplier of the input VMM system 32a. The input conversion can also be performed by an analog - analog (A / A) converter to convert an external analog input to the mapped analog input to the input VMM system 32a. The input conversion can also be performed by a digital - digital pulse (D / P) converter to convert an external digital input to the mapped digital pulse to the input VMM system 32a.
[0040] The output generated by the input VMM system 32a is then provided as input to the next VMM system (hidden level 1) 32b, which then generates an output that is in turn provided as input to the next input VMM system (hidden level 2) 32c, and so on. The various layers of the VMM system 32 function as the respective layers of synapses and neurons of a convolutional neural network (CNN). Each of the VMM systems 32a, 32b, 32c, 32d, and 32e can be a stand-alone physical system with a corresponding non-volatile memory array, or multiple VMM systems can utilize different portions of the same physical non-volatile memory array, or multiple VMM systems can utilize overlapping portions of the same physical non-volatile memory array. Each VMM system 32a, 32b, 32c, 32d, and 32e can also be time-multiplexed with respect to the various portions of its array or neurons. The example shown in FIG. 11 includes five layers (32a, 32b, 32c, 32d, 32e), namely, one input layer (32a), two hidden layers (32b, 32c), and two fully-connected layers (32d, 32e). One of ordinary skill in the art will understand that this is merely exemplary and that the system can alternatively include more than two hidden layers and more than two fully-connected layers.
[0041] <<VMM Array>> FIG. 12 shows a neuron VMM array 1200 that is particularly suitable for the memory cell 310 shown in FIG. 3 and that is utilized as part of the synapses and neurons between the input layer and the next layer. The VMM array 1200 includes a memory array 1201 of non-volatile memory cells and a reference array 1202 of non-volatile reference memory cells (located at the top of the array). Alternatively, another reference array can be located at the bottom.
[0042] In the VMM array 1200, control gate lines such as the control gate line 1203 extend in the vertical direction (thus, the reference array 1202 in the row direction is orthogonal to the control gate line 1203), and erase gate lines such as the erase gate line 1204 extend in the horizontal direction. Here, the input to the VMM array 1200 is provided to the control gate lines (CG0, CG1, CG2, CG3), and the output of the VMM array 1200 appears on the source lines (SL0, SL1). In one embodiment, only even rows are used, and in another embodiment, only odd rows are used. The current applied to each source line (SL0 and SL1 respectively) performs a summation function of all the currents from the memory cells connected to that particular source line.
[0043] As described herein for neural networks, the non-volatile memory cells of the VMM array 1200, i.e., the flash memory of the VMM array 1200, are preferably configured to operate in the subthreshold region.
[0044] The non-volatile reference memory cells and non-volatile memory cells described herein are biased with weak inversion as follows: Ids = Io * e (Vg-Vth) / nVt = w * Io * e (Vg) / nVt where w = e (-Vth) / nVt and where Ids is the drain-source current, Vg is the gate voltage of the memory cell, Vth is the threshold voltage of the memory cell, Vt is the thermal voltage = k * T / q, where k is the Boltzmann constant, T is the Kelvin temperature, q is the electron charge, n is the slope factor = 1+(Cdep / Cox), Cdep is the capacitance of the depletion layer, and Cox is the capacitance of the gate oxide layer, Io is the memory cell current at a gate voltage equal to the threshold voltage, and Io is (Wt / L) * u * Cox * (n - 1) * Vt 2Proportional to, where u is the carrier mobility, and Wt and L are the width and length of the memory cell, respectively.
[0045] When using an I-V logarithmic converter that converts the input current Ids to the input voltage Vg using a memory cell (such as a reference memory cell or a peripheral memory cell) or a transistor, Vg is as follows: Vg = n * Vt * log[Ids / wp * Io] Where wp is the w of the reference or peripheral memory cell.
[0046] When using an I-V logarithmic converter that converts the input current Ids to the input voltage Vg using a memory cell (such as a reference memory cell or a peripheral memory cell) or a transistor, Vg is as follows:
[0047] Vg = n * Vt * log[Ids / wp * Io] Where wp is the w of the reference or peripheral memory cell.
[0048] For the memory array used as the vector matrix multiplier VMM array, the output current is as follows: Iout = wa * Io * e (Vg) / nVt That is Iout = (wa / wp) * Iin = W * Iin W = e (Vthp-Vtha) / nVt Iin = wp * Io * e (Vg) / nVt Where wa = w for each memory cell of the memory array. Vthp is the effective threshold voltage of the peripheral memory cell, and Vtha is the effective threshold voltage of the main (data) memory cell. Note that the threshold voltage of the transistor is a function of the substrate body bias voltage, and the substrate body bias can be modulated for various compensations, such as overheating or modulation of the cell current. Vth=Vth0+gamma(SQRT(Vsb+|2 * φF|)-SQRT|2 * φF|) Vth0 is the threshold voltage with zero substrate bias, φF is the surface potential, and gamma is the body effect parameter.
[0049] The word line or control gate can be used as the input of the memory cell for the input voltage.
[0050] Alternatively, the non-volatile memory cells of the VMM array described herein can be configured to operate in the linear region. Ids = beta * (Vgs-Vth) * Vds; beta = u * Cox * Wt / L W α (Vgs-Vth) That is, the weight W in the linear region is proportional to (Vgs-Vth).
[0051] The word line or control gate or bit line or source line can be used as the input of the memory cell operating in the linear region. The bit line or source line can be used as the output of the memory cell.
[0052] For an I-V linear converter, a memory cell (such as a reference memory cell or a peripheral memory cell) or a transistor or a resistor operating in the linear region can be used to linearly convert the input / output current into the input / output voltage.
[0053] Alternatively, the memory cells of the VMM array described herein can be configured to operate in the saturation region. Ids = 1 / 2 * beta * (Vgs-Vth) 2; Beta = u * Cox * Wt / L W α (Vgs - Vth) 2 i.e., the weight W is 2 proportional to (Vgs - Vth).
[0054] The word line, control gate, or erase gate can be used as an input to a memory cell operating in the saturation region. The bit line or source line can be used as an output of the output neuron.
[0055] Alternatively, the memory cells of the VMM array described herein can be used in all regions or combinations thereof (subthreshold, linear, or saturated) for each layer or multiple layers of a neural network.
[0056] FIG. 13 shows a neuron VMM array 1300 particularly suitable for the memory cell 210 shown in FIG. 2 and is utilized as a synapse between the input layer and the next layer. The VMM array 1300 includes a memory array 1303 of non - volatile memory cells, a reference array 1301 of first non - volatile reference memory cells, and a reference array 1302 of second non - volatile reference memory cells. The reference arrays 1301 and 1302 arranged in the column direction of the array function to convert the current inputs flowing into the terminals BLR0, BLR1, BLR2, and BLR3 into voltage inputs WL0, WL1, WL2, and WL3. In practice, the first and second non - volatile reference memory cells are diode - connected through a multiplexer 1314 (only part shown) with the current inputs flowing in. The reference cells are tuned (e.g., programmed) to a target reference level. The target reference level is provided by a reference mini - array matrix (not shown).
[0057] The memory array 1303 serves two purposes. First, it stores the weights used by the VMM array 1300 in each memory cell. Second, the memory array 1303 effectively multiplies the input (i.e., the current inputs provided to terminals BLR0, BLR1, BLR2, and BLR3, which are converted by the reference arrays 1301 and 1302 into input voltages and supplied to the word lines WL0, WL1, WL2, and WL3) by the weights stored in the memory cell array 1303, and then adds up all the results (memory cell currents) to generate the output of each bit line (BL0 to BLN), and this output serves as the input to the next layer or the input to the last layer. By the memory array 1303 performing the functions of multiplication and addition, the need for separate multiplication and addition logic circuits is eliminated, and the power efficiency is also good. Here, the voltage input is provided to the word lines WL0, WL1, WL2, and WL3, and the output appears on each bit line BL0 to BLN during the read (inference) operation. The current arranged on each bit line BL0 to BLN performs the total function of the currents from all the non-volatile memory cells connected to that specific bit line.
[0058] Table 6 shows the operating voltages of the VMM array 1300. The columns in the table show the voltages applied to the word lines of the selected cells, the word lines of the non-selected cells, the bit lines of the selected cells, the bit lines of the non-selected cells, the source lines of the selected cells, and the source lines of the non-selected cells, and FLT indicates floating, i.e., no voltage is applied. The rows indicate the read, erase, and program operations. Table 6: Operation of the VMM Array 1300 in FIG. 13 [Table 6]
[0059] FIG. 14 shows a neuron VMM array 1400 that is particularly suitable for the memory cell 210 shown in FIG. 2 and is used as part of synapses and neurons between the input layer and the next layer. The VMM array 1400 includes a memory array 1403 of non-volatile memory cells, a reference array 1401 of first non-volatile reference memory cells, and a reference array 1402 of second non-volatile reference memory cells. The reference arrays 1401 and 1402 extend in the row direction of the VMM array 1400. The VMM array is similar to the VMM 1300 except that the word lines extend vertically in the VMM array 1400. Here, the inputs are provided to the word lines (WLA0, WLB0, WLA1, WLB2, WLA2, WLB2, WLA3, WLB3), and the outputs appear on the source lines (SL0, SL1) during a read operation. The current applied to each source line performs a sum function of all the currents from the memory cells connected to that particular source line.
[0060] Table 7 shows the operating voltages of the VMM array 1400. The columns in the table show the voltages applied to the word line of the selected cell, the word line of the non-selected cell, the bit line of the selected cell, the bit line of the non-selected cell, the source line of the selected cell, and the source line of the non-selected cell. The rows show the operations of read, erase, and program. Table 7: Operation of the VMM Array 1400 in FIG. 14
Table 7
[0061] FIG. 15 shows a neuron VMM array 1500 that is particularly suitable for the memory cell 310 shown in FIG. 3 and is used as part of synapses and neurons between the input layer and the next layer. The VMM array 1500 includes a memory array 1503 of non-volatile memory cells, a reference array 1501 of first non-volatile reference memory cells, and a reference array 1502 of second non-volatile reference memory cells. The reference arrays 1501 and 1502 function to convert the current inputs flowing into the terminals BLR0, BLR1, BLR2, and BLR3 into voltage inputs CG0, CG1, CG2, and CG3. In practice, the first and second non-volatile reference memory cells are diode-connected through a multiplexer 1512 (only part shown) in a state where the current inputs flow through BLR0, BLR1, BLR2, and BLR3. The multiplexer 1512 includes a corresponding multiplexer 1505 and a cascode transistor 1504 respectively to ensure a constant voltage for each bit line (such as BLR0) of the first and second non-volatile reference memory cells during the read operation. The reference cells are tuned to a target reference level.
[0062] The memory array 1503 serves two purposes. First, it stores the weights used by the VMM array 1500. Second, the memory array 1503 multiplies the input (the current inputs provided to the terminals BLR0, BLR1, BLR2, and BLR3, and the reference arrays 1501 and 1502 convert these current inputs into input voltages and supply them to the control gates (CG0, CG1, CG2, and CG3)) by the weights stored in the memory cell array, then adds up all the results (cell currents) to generate an output, which appears on BL0 to BLN and becomes the input to the next layer or the input to the last layer. By the memory array performing the functions of multiplication and addition, the need for separate multiplication and addition logic circuits is eliminated, and the power efficiency is also good. Here, the input is provided to the control gate lines (CG0, CG1, CG2, and CG3), and the output appears on the bit lines (BL0 to BLN) during the read operation. The current applied to each bit line performs the total function of all the currents from the memory cells connected to that particular bit line.
[0063] The VMM array 1500 implements unidirectional tuning of non-volatile memory cells within the memory array 1503. That is, each non-volatile memory cell is erased and then partially programmed until the desired charge on the floating gate is reached. This can be performed, for example, using the precision programming techniques described below. If too much charge is applied to the floating gate (in which case an incorrect value is stored in the cell), the cell must be erased and a series of partial programming operations must be repeated. As shown, two rows sharing the same erase gate (such as EG0 or EG1) need to be erased together (known as page erase), and then each cell is partially programmed until the desired charge on the floating gate is reached.
[0064] Table 8 shows the operating voltages of the VMM array 1500. The columns in the table show the voltages applied to the word line of the selected cell, the word line of the non-selected cell, the bit line of the selected cell, the bit line of the non-selected cell, the control gate of the selected cell, the control gate of the non-selected cell within the same sector as the selected cell, the control gate of the non-selected cell in a different sector from the selected cell, the erase gate of the selected cell, the erase gate of the non-selected cell, the source line of the selected cell, and the source line of the non-selected cell. The rows show the read, erase, and program operations. Table 8: Operation of the VMM Array 1500 of FIG. 15 [Table 8]
[0065] FIG. 16 shows a neuron VMM array 1600 that is particularly suitable for the memory cell 310 shown in FIG. 3 and is used as part of the synapses and neurons between the input layer and the next layer. The VMM array 1600 includes a memory array 1603 of non-volatile memory cells, a reference array 1601 or a first non-volatile reference memory cell, and a reference array 1602 of second non-volatile reference memory cells. The EG lines EGR0, EG0, EG1, and EGR1 extend vertically, and the CG lines CG0, CG1, CG2, and CG3 and the SL lines WL0, WL1, WL2, and WL3 extend horizontally. The VMM array 1600 is similar to the VMM array 1600 except that the VMM array 1600 implements bidirectional tuning, and each individual cell can be completely erased, partially programmed, and partially erased as needed to reach the desired charge level of the floating gate by using an individual EG line. As shown, the reference arrays 1601 and 1602 convert the input current in the terminals BLR0, BLR1, BLR2, and BLR3 into control gate voltages CG0, CG1, CG2, and CG3 (through the action of diode-connected reference cells via the multiplexer 1614), and these voltages are applied to the memory cells in the row direction. The current output (neuron) is in the bit lines BL0 to BLN, and each bit line sums all the currents from the non-volatile memory cells connected to that particular bit line.
[0066] Table 9 shows the operating voltages of the VMM array 1600. The columns in the table show the voltages applied to the word line of the selected cell, the word line of the non-selected cell, the bit line of the selected cell, the bit line of the non-selected cell, the control gate of the selected cell, the control gate of the non-selected cell in the same sector as the selected cell, the control gate of the non-selected cell in a different sector from the selected cell, the erase gate of the selected cell, the erase gate of the non-selected cell, the source line of the selected cell, and the source line of the non-selected cell. The rows show the operations of read, erase, and program. Table 9: Operation of the VMM Array 1600 in FIG. 16
Table 9
[0067] The input to the VMM array can be an analog level, binary level, timing pulse, or digital bit, and the output can be an analog level, binary level, timing pulse, or digital bit (in which case an output ADC is required to convert the output analog level current or voltage to a digital bit).
[0068] For each memory cell within the VMM array, each weight w can be implemented by a single memory cell, or by differential cells, or by two blended memory cells (the average of two or more cells). In the case of differential cells, two memory cells are required to implement the weight w as a differential weight (w = w+ - w-). In the case of two blended memory cells, two memory cells are required to implement the weight w as the average of two cells.
[0069] One drawback of prior art arrays of non-volatile memory cells is that a relatively large amount of time is required to pull down the source line to ground in order to perform a read or erase operation.
[0070] What is needed is an improved VMM system that includes a source line pull-down mechanism that can pull down the source line to ground more quickly than prior art systems. SUMMARY OF THE INVENTION
[0071] Numerous embodiments of an analog neural memory array are disclosed. Certain embodiments include an improved mechanism for quickly pulling down the source line to ground. This is useful, for example, to minimize the voltage drop during read, program, or erase operations. Other embodiments include the implementation of negative and positive inputs with negative and positive weights.
[0072] In one embodiment, a non-volatile memory system includes an array of non-volatile memory cells arranged in rows and columns, a plurality of bit lines, each of which is coupled to a column of non-volatile memory cells, and a plurality of pull-down bit lines, each of which is coupled to a row of non-volatile memory cells.
[0073] In another embodiment, a non-volatile memory system includes an array of non-volatile memory cells arranged in rows and columns, a plurality of bit lines, each of which is coupled to a column of non-volatile memory cells, and a plurality of pull-down bit lines, each of which is coupled to a column of non-volatile memory cells. During a read or verify operation of a selected cell, current enters the selected cell through one of the plurality of bit lines, enters a plurality of rows of pull-down cells, and flows into two or more of the plurality of pull-down bit lines.
[0074] In another embodiment, a non-volatile memory system includes an array of non-volatile memory cells arranged in rows and columns, a plurality of bit lines, each of which is coupled to a column of non-volatile memory cells, and a plurality of pull-down bit lines, each of which is coupled to a column of non-volatile memory cells. During a read or verify operation of a selected cell, current enters the selected cell through one of the plurality of bit lines, enters a pull-down cell adjacent to the selected cell, and flows into one of the plurality of pull-down bit lines.
[0075] In another embodiment, the non-volatile memory system includes an array of non-volatile memory cells arranged in rows and columns, a plurality of bit lines, each of the plurality of bit lines being coupled to a column of non-volatile memory cells, and a plurality of pull-down cells, each of the plurality of pull-down cells being coupled to a source line of a non-volatile memory cell.
[0076] In another embodiment, the non-volatile memory system includes an array of non-volatile memory cells arranged in rows and columns, a plurality of bit lines, each of the plurality of bit lines being coupled to a column of non-volatile memory cells, a plurality of rows, each of the plurality of row lines being coupled to a row of non-volatile memory cells, and a row that receives a negative value input.
[0077]
[0078]
[0079]
[0080]
[0081]
[0082]
[0083]
[0084]
[0085]
[0086]
[0087]
[0088]
[0089]
[0090]
[0091]
[0092]
[0093]
[0094]
[0095]
[0096]
[0097]
[0098]
[0099]
[0100]
[0101]
[0102]
[0103]
[0104]
[0105]
[0106]
[0107]
[0108]
[0109]
[0110]
[0111]
[0112]
[0113]
[0114]
Brief Description of the Drawings
[0115]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18A
Figure 18B
Figure 18C
Figure 19A
Figure 19B
Figure 19C
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
Figure 26
Figure 27A
Figure 27B
Figure 27C
Figure 28
Figure 29
Figure 30
Figure 31
Figure 32
Figure 33A
Figure 33B
Figure 34A
Figure 34B
Figure 34C
Figure 35
Figure 36
Figure 37
Figure 38
Figure 39
Embodiments for Carrying Out the Invention
[0116] The artificial neural network of the present invention utilizes a combination of CMOS technology and a non-volatile memory array.
[0117] <<Overview of the VMM System>> FIG. 17 shows a block diagram of a VMM system 1700. The VMM system 1700 includes a VMM array 1701, a row decoder 1702, a high-voltage decoder 1703, a column decoder 1704, a bit-line driver 1705, an input circuit 1706, an output circuit 1707, control logic 1708, and a bias generator 1709. The VMM system 1700 further includes a high-voltage generation block 1710 that includes a charge pump 1711, a charge pump regulator 1712, and a high-voltage level generator 1713. The VMM system 1700 further includes an algorithm controller 1714, an analog circuit 1715, control logic 1716, and test control logic 1717. The systems and methods described below may be implemented in the VMM system 1700.
[0118] The input circuit 1706 may include circuits such as a DAC (Digital - Analog Converter), DPC (Digital - Pulse Converter), AAC (Analog - Analog Converter such as a current - voltage converter), PAC (Pulse - Analog Level Converter), or any other type of converter. The input circuit 1706 may implement a normalization function, a scaling function, or an arithmetic function. The input circuit 1706 may implement a temperature compensation function for the input. The input circuit 1706 may implement an activation function such as ReLU or sigmoid. The output circuit 1707 may include circuits such as an ADC (Analog - Digital Converter for converting neuron analog output to digital bits), AAC (Analog converter such as a current - voltage converter), APC (Analog - Pulse Converter), or any other type of converter. The output circuit 1707 may implement an activation function such as ReLU or sigmoid. The output circuit 1707 may implement a statistical normalization, regularization, up / down scaling function, statistical rounding, or arithmetic function (e.g., addition, subtraction, division, multiplication, shift, log) of the neuron output. The output circuit 1707 may implement a temperature compensation function for the neuron output or the array output (such as bit - line output) in order to keep the power consumption of the array approximately constant or to improve the accuracy of the array (neuron) output by keeping the slope of the IV approximately the same, etc.
[0119] FIG. 18A shows a prior art VMM system 1800. The VMM system 1800 includes exemplary cells 1801 and 1802, an exemplary bit line switch 1803 (connecting the bit line to a sense circuit), an exemplary dummy bit line switch 1804 (coupled to a low level such as ground level in a read), exemplary dummy cells 1805 and 1806 (source line pull-down cells). The bit line switch 1803 is coupled to a column of cells including cells 1801 and 1802 that are used to store data in the VMM system 1800. The dummy bit line switch 1804 is coupled to a column of cells (bit lines) that are dummy cells not used to store data in the VMM system 1800. This dummy bit line (also known as a source line pull-down bit line) is used as a source line pull-down in a read, which means it is used to pull down the source lines SL (e.g., SL0 and SL1) to a low level such as ground level through memory cells in the dummy bit line. The dummy bit line switch 1804 is at the end of the same array as the bit line switch 1803.
[0120] One drawback of the VMM system 1800 is that the input impedance of each cell varies due to the length of the electrical path through the associated bit line switch, the cell itself, and the associated dummy bit line switch. For example, FIG. 18B shows the electrical path through bit line switch 1803, cell 1801, source line SL0, dummy cell 1805, and dummy bit line switch 1804. Similarly, FIG. 18C shows the electrical path through bit line switch 1803, vertical metal bit line 1807, cell 1802, source line SL1, dummy cell 1806, vertical metal bit line 1808, and dummy bit line switch 1804. As can be seen, the path of FIG. 18C through cell 1802 crosses bit lines and dummy bit lines of significantly greater length than the path of FIG. 18B through cell 1801, and thus the path through cell 1802 is associated with a higher capacitance and a higher resistance than the path through cell 1801. This results in a cell 1802 having a greater parasitic impedance in the bit line or source line than cell 1801. This variability is a drawback because it results in variations in the accuracy of the cell output applied to the readout or verification of the cell (for the program / erase tuning cycle), for example, depending on the position of the cell within the array.
[0121] FIG. 19A shows an improved VMM system 1900. The VMM system 1900 includes exemplary cells 1901 and 1902, an exemplary bit line switch 1903 (connecting the bit line to the sense circuit), exemplary dummy cells 1905 and 1906 (each being a pull-down cell), and an exemplary dummy bit line switch 1904 (coupling to a low level such as ground level in a read). The dummy bit line switch is connected to a dummy bit line that connects to the dummy cells 1905, 1906 used as pull-down cells. As can be seen, the exemplary dummy bit line switch 1904 and other dummy bit line switches are located at opposite ends of the array from the bit line switch 1903 and other bit line switches. Four columns of cells and associated dummy cells are shown, and the bit line switches are labeled 1903a, 1903b, 1903c, 1903d respectively, and each dummy bit line switch is labeled 1904a, 1904b, 1904c, and 1904d.
[0122] The advantages of this design can be seen in FIGS. 19B and 19C. FIG. 19B shows an electrical path through the bit line switch 1903, cell 1901, source line SL0, dummy cell 1905 (source line pull-down cell), vertical metal bit line 1908, and dummy bit line switch 1904 (coupling to a low level such as ground in a read). FIG. 19C shows an electrical path through the bit line switch 1903, vertical metal line 1907, cell 1902, source line SL1, dummy cell 1906 (source line pull-down cell), and dummy bit line switch 1904. The paths are substantially the same (cell, interconnect length), which holds for all cells of the VMM system 1900. As a result, the bit line impedance + source line impedance of each cell is substantially the same, which means that the variation in the amount of parasitic voltage drop in the read or verification operations of various cells in the array is relatively the same.
[0123] FIG. 20 shows a VMM system 2000 having global source line pull-down bit lines. The VMM system 2000 is the same as the VMM system 1900, except that dummy bit lines 2005a - 2005n and 2007a - 2007n are connected together (to act as global source line pull-down lines for pulling the memory cell source lines to the ground level during read or verify), dummy bit line switches such as dummy bit line switches 2001 and 2002 are connected or coupled to a common ground indicated as ARYGND, and the source lines are coupled together to a source line switch 2008 that selectively pulls down the source lines to ground. These changes further reduce the variation in (array) parasitic impedance between cells during read or verify operations.
[0124] FIG. 21 shows a VMM system 2100. The VMM system 2100 includes a bit line switch 2101, a dummy bit line (also known as a pull-down bit line) switch 2102, a dummy bit line switch 2103, a bit line switch 2104, a data cell 2105 (in this specification, a "data cell" is a memory cell used to store the weight values of a neural network), a dummy cell (also known as a pull-down cell) 2106, a dummy cell 2107, and a data cell 2018. Note that the dummy cells 2106 and 2107 are adjacent to each other. Thereby, the vertical metal lines BLpdx of the two dummy cells 2106 and 2107 are connected together (line 2111), and it becomes possible to reduce the parasitic resistance by the resulting wider metal line. During the read or verification (for the program / erase tuning cycle) operation of the data cell 2105, current enters the bit line terminal of the cell 2105 through the bit line switch 2101, exits to the source line terminal of the cell 2015, then enters the source line 2110, enters the source line terminals of the dummy cells 2106 and 2107, passes through the vertical dummy bit line 2111s, and flows through the dummy bit line switches 2102 and 2103. During the read or verification (for the program / erase tuning cycle) operation of the data cell 2108, current enters the bit line terminal of the data cell 2108 through the bit line switch 2104, exits to the source line terminal of the cell 2108, then enters the source line 2110, enters the source line terminals of the dummy cells 2106 and 2107, passes through the vertical dummy bit line 2111s, and flows through the dummy bit line switches 2102 and 2103. The pattern of this column is repeated across the entire array, and all four columns include two columns of data cells and two adjacent array columns used for the pull-down operation located between the two columns of data cells. In another embodiment, the diffusion of the two pull-down cells in two adjacent columns can be merged into one larger diffusion to enhance the pull-down ability. In another embodiment, the diffusion of the pull-down cell can be made larger than the data cell diffusion to enhance the pull-down ability. In another embodiment, each pull-down cell has a bias condition different from the bias condition of the selected data cell.
[0125] In one embodiment, the pull-down cell has the same physical structure as a normal data memory cell. In another embodiment, the pull-down cell has a physical structure different from that of a normal data memory cell. For example, the pull-down cell can be a modified version of a normal data memory cell by modifying one or more physical dimensions (such as width, length without limitation) or electrical parameters (such as layer thickness, implant without limitation). In another embodiment, the pull-down cell is a normal transistor (without a floating gate) such as an I / O or high-voltage transistor.
[0126] FIG. 22 shows a VMM system 2200. The VMM system 2200 includes bit lines 2201, pull-down bit lines 2202, data cells 2203 and 2206, pull-down cells 2204 and 2205, and source lines 2210, 2211. During the read or verify operation of cell 2203, current enters the bit line terminal of cell 2203 through bit line switch 2201, exits the source line terminal of cell 2203, then enters source line 2210, enters the source line terminal of pull-down cell 2204, and flows through pull-down bit line BLpd 2202 connected to the drain of pull-down cell 2204 (through a contact or the like). This design is repeated for all columns, and finally, the row including pull-down cell 2204 becomes the row of pull-down cells.
[0127] During the read or verify (for program / erase tuning cycle) operation of cell 2206, current enters the bit line terminal of cell 2206 through bit line switch 2201, exits the source line terminal of cell 2206, then enters source line 2211, enters the source line terminal of pull-down cell 2205, and flows through pull-down bit line BLpd 2202 connected to the drain of pull-down cell 2205 (through a contact or the like). This design is repeated for all columns, and finally, the row including pull-down cell 2205 becomes the row of pull-down cells. As shown in FIG. 22, there are four rows, the two central adjacent rows are used for pull-down cells, and the top and bottom rows are data cells.
[0128] Table 10 shows the operating voltages of the VMM system 2200. The columns in the table indicate the bit lines of the selected cell, the pull-down bit lines, the word line WL of the selected cell, the control gate of the selected cell, the word line WLS of the selected pull-down cell, the control gate CGS of the selected pull-down cell, the erase gate EG of all cells, and the source line SL of all cells. The rows indicate the read, erase, and program operations. Note that the voltage biases of CGS and WLS in read are higher than the voltage biases of the normal WL and CG biases to enhance the driving ability of the pull-down cell. The voltages biased for WLS and CGS can be negative during programming to reduce interference. Table 10: Operation of the VMM array 2200 of FIG. 22 [Table 10]
[0129] FIG. 23 shows a VMM system 2300. The VMM system 2300 includes bit lines 2301, pull-down bit lines 2302, data cells 2303 and 2306, and pull-down cells 2304 and 2305. During the read or verification (for program / erase tuning cycles) operation of cell 2303, current enters the bit line terminal of cell 2303 through bit line 2301, exits to the source line terminal of cell 2303, and then enters the source line terminal of pull-down cell 2304 (connected to source line SL0), and flows through pull-down bit line 2302 connected to the drain terminal of pull-down cell 2304. This design is repeated for all columns, and finally, the row including pull-down cell 2304 in the first mode becomes the row of pull-down cells. During the read or verification (for program / erase tuning cycles) operation of data cell 2306, current enters the bit line terminal of cell 2306 through bit line 2301, exits to the source line terminal of cell 2306 (connected to source line SL1), and then enters the source line terminal of pull-down cell 2305, and flows through pull-down bit line 2302 connected to the drain terminal of pull-down cell 2305. This design is repeated for all columns, and finally, the row including pull-down cell 2305 in the second mode becomes the row of pull-down cells. As shown in FIG. 23, there are four rows, and alternative odd (or even) rows are used for pull-down cells, and alternative even (or odd) rows are data cells.
[0130] In particular, during the second mode, cells 2304 and 2306 are active in read or verification, cells 2303 and 2305 are used for the pull-down process, and the roles of bit lines 2301 and 2302 are reversed compared to the first mode. In other words, the role of each row of cells can change between the first mode and the second mode, and the role of each bit line can change between the first mode and the second mode.
[0131] Table 11 shows the operating voltages of the VMM system 2300. The columns in the table are the bit lines of the selected data cell, the bit lines BL0A, BL0B of the selected pull-down cell, the word line WL of the selected data cell, the control gate CG of the selected data cell, the word line WLS of the selected pull-down cell, the control gate CGS of the selected pull-down cell, the erase gate EG of all cells, and the source line SL of all cells. The rows indicate the operations of read, erase, and program. Table 11: Operations of the VMM System 2300 in FIG. 23 [Table 11]
[0132] FIG. 24 shows the VMM system 2400. The VMM system 2400 includes bit lines 2401, pull-down bit lines 2402, (data) cells 2403, source lines 2411, and pull-down cells 2404, 2405, and 2406. During the read or verification operation of cell 2403, current enters the bit line terminal of cell 2403 through bit line 2401, exits the source line terminal of cell 2403, then enters source line 2411, and then enters the source line terminals of pull-down cells 2404, 2405, and 2406, and flows from there through pull-down bit line 2402. This design is repeated for all columns, and finally, the rows including pull-down cells 2404, 2405, and 2406 become the rows of pull-down cells respectively. Thereby, when current is drawn into pull-down bit line 2402 through three cells, the pull-down applied to the source line terminal of cell 2403 is maximized. Note that the source lines of four rows are connected together by vertical lines 4020s.
[0133] Table 12 shows the operating voltages of the VMM system 2400. The columns in the table represent the bit line BL of the selected cell, the pull-down bit line BPpd, the word line WL of the selected cell, the control gate CG of the selected cell, the erase gate EG of the selected cell, the word line WLS of the selected pull-down cell, the control gate CGS of the selected pull-down cell, the erase gate EGS of the selected pull-down cell, and the source line SL of all cells. The rows represent the operations of read, erase, and program. Table 12: Operations of the VMM System 2400 in FIG. 24 [Table 12]
[0134] FIG. 25 shows an exemplary layout 2500 of the VMM system 2200 of FIG. 22. The bright squares represent metal contacts (metal-to-diffusion contacts) that connect the memory cell diffusion part (drain region of the memory cell) to metal bit lines such as bit line 2201 (e.g., connecting the drain diffusion of memory cell 2203 to bit line 2201) and metal pull-down bit lines such as pull-down bit line 2202 (e.g., connecting the drain diffusion part of memory cell 2204 to pull-down bit line 2202).
[0135] FIG. 26 shows an alternative layout 2600 of a VMM system similar to the VMM system 2200 of FIG. 22, but with the difference that the vertical pull-down bit line 2602 is very wide and connects across two columns of pull-down cells. This means that there is one vertical pull-down line for every two bit lines of the data memory cells. That is, the area of the pull-down bit line 2602 is wider than the area of the bit line 2202 in FIG. 25. Layout 2600 further shows data cells 2603 and pull-down cells 2604, as well as source line 2610. In another embodiment, the diffusions of the two pull-down cells (left and right) can be merged into a larger single diffusion region.
[0136] FIG. 27A shows a VMM system 2700. To implement the negative and positive weights of a neural network, half of the bit lines are designated as w+ lines (bit lines connected to memory cells implementing positive weights), and the other half of the bit lines are designated as w- lines (bit lines connected to memory cells implementing negative weights), and are interspersed alternately among the w+ lines. The application of negative operations is performed on the outputs of the w- bit lines (neuron outputs) by summing circuits such as summing circuits 2701 and 2702. The outputs of the w+ lines and the outputs of the w- lines are combined together to effectively give w = w+ - w- for each pair of all pairs of (w+, w-) cells of the (w+, w-) lines. The dummy bit lines or pull-down bit lines used respectively to avoid the floating gate (FG-FG coupling) of adjacent cells in readout and / or to reduce the IR voltage drop in the source line are not shown in the figure. Inputs to the system 2700 (such as CG or WL) can have positive or negative values. When the input has a negative value, since the actual input to the array is positive (such as the voltage level of CG or WL), the array output (bit line output) is made negative before output to realize the equivalent function of the negative value input.
[0137] Alternatively, referring to FIG. 27B, the positive weights can be implemented in a first array 2711, the negative weights can be implemented in a second array 2712 separate from the first array, and the resulting weights are appropriately combined by a summing circuit 2713. Similarly, dummy bit lines (not shown) or pull-down bit lines (not shown) are used respectively to avoid the FG-FG coupling of the source line in readout and / or to reduce the IR voltage drop.
[0138] Alternatively, FIG. 27C shows a VMM system 2750 for implementing the negative and positive weights of a neural network with positive or negative inputs. A first array 2751 implements positive value inputs with negative and positive weights in alternating columns, and a second array 2752 implements negative value inputs with negative and positive weights in corresponding alternating columns. The output of the second array is made negative before being added to the output of the first array by an adder 2755.
[0139] Table 10A shows an exemplary layout of the physical array arrangement of the (w+, w-) pairs of bit lines BL0 / 1 and BL2 / 3, with four rows coupled to the pull-down bit line BLPWDN. The pairs of bit lines (BL0, BL1; BL2, BL3) are used to implement the (w+, w-) lines. Between the pairs of (w+, w-) lines, there is a pull-down bit line (BLPWDN). This is used to prevent the coupling of current (w+, w-) lines from adjacent (w+, w-) lines (e.g., FG-FG coupling). Basically, the pull-down bit line (BLPWDN) functions as a physical barrier between the pairs of (w+, w-) lines.
[0140] Additional details regarding the FG-FG coupling phenomenon and the mechanism to counteract that phenomenon can be found in U.S. Patent Provisional Application No. 62 / 981,757, filed on February 26, 2020, entitled "Ultra-Precise Tuning of Analog Neural Memory Cells in a Deep Learning Artificial Neural Network," which is incorporated herein by reference.
[0141] Table 10B shows different exemplary combinations of weights. "1" means the cell is used and has an actual output value, and "0" means the cell is not used and does not have a large output value.
[0142] In another embodiment, dummy bit lines can be used instead of pull-down bit lines.
[0143] In another embodiment, dummy rows can also be used as a physical barrier to avoid coupling between rows. Table 10A: Exemplary Layout [Table 10A] Table 10B: Exemplary Combinations of Weights [Table 10B]
[0144] Table 11A shows another embodiment of an array of the physical layout of bit line pairs BL0 / 1 and BL2 / 3 having redundant lines BL01, BL23, and pull-down bit line BLPWDN (w+, w-). BL01 is used for remapping the weights of pair BL0 / 1 (meaning redistributing the weights among the cells), and BL23 is used for remapping the weights of pair BL2 / 3.
[0145] Table 11B shows the case of distributed weights that do not require remapping. Basically, there are no adjacent '1's between BL1 and BL3, and adjacent '1's cause the coupling of adjacent bit lines. Table 11C shows the case of distributed weights that require remapping due to adjacent '1's between BL1 and BL2 that cause the coupling of adjacent bit lines. This remapping is shown in Table 11D, and as a result, there will be no '1' values between adjacent bit lines. Further, by remapping the '1' actual values of the weights between the bit lines, the total current along the bit lines decreases at this point, and the values within the bit lines (output neurons) become more precise. In this case, additional columns (bit lines) are required to act as redundant columns (BL01, BL23). Tables 11E and 11F show another embodiment of remapping noisy cells (or defective cells) to redundant (spare) columns such as BL01, BL23 of Table 10E or BL0B and BL1B of Table 11F. The adder is used to sum the appropriately mapped bit line outputs. Table 11A: Exemplary Layout
Table 11A
Table 11B
Table 11C
[0146] Table 11G shows an embodiment of the physical layout of an array suitable for FIG. 27B. Since each array has either a positive or negative weight, a dummy bit line that acts as a pull-down, and a physical barrier to avoid FG-FG coupling are required for each bit line. Table 11G: Exemplary Layout [Table 11G]
[0147] In another embodiment, the tuning bit line is used as a bit line adjacent to the target bit line to tune the target bit line to its final target by FG-FG capacitive coupling. In this case, the pull-down bit line (BLPWDN) is inserted on the side of the target bit line that is not bounded by the tuning bit line.
[0148] In another embodiment, noisy or defective cells are designated as cells that are not used (after they are identified as noisy or defective by the sensing circuit), which means that they are (deeply) programmed so that they do not contribute any value to the neuron output.
[0149] In another embodiment, more precise programming algorithms are applied to these cells, such as a programming sequence that identifies fast cells and uses smaller voltage increment pulses or does not use voltage increment pulses or uses a floating gate coupling algorithm.
[0150] FIG. 28 shows an optional redundant array 2801 that can be included in any of the VMM arrays discussed so far. The redundant array 2801 can be used as redundancy to replace a defective column if any of the columns attached to the bit line switches are considered defective. The redundant array can have its own redundant neuron output (e.g., bit line) and an ADC circuit for redundancy purposes. If redundancy is required, the output of the redundant ADC is used to replace the output of the ADC of the defective bit line. The redundant array 2801 can also be used for weight mapping as described in Tables 11D, 11E, and 11F for power distribution between bit lines.
[0151] FIG. 29 shows a VMM system 2900 including an array 2901, an array 2902, a column multiplexer 2903, top local bit lines LBL2905a - d, bottom local bit lines LBL2905e - h, global bit lines GBL2908 and 2909, and dummy bit line switches 2905. The column multiplexer 2903 is used to respectively select the top local bit lines 2905 of array 2901 or the bottom local bit lines 2905 of array 2902 to the global bit lines 2908, 2909. In one embodiment, the (metal) global bit line 2908 has the same number of lines as the number of local bit lines, for example, 8 or 16 lines. In another embodiment, the global bit line 2908 has only one (metal) line per number N of local bit lines, such as one global bit line per 8 or 16 local bit lines. The column multiplexer 2903 further includes multiplexing adjacent global bit lines (such as GBL2909) to the current global bit line (such as GBL2908) to effectively increase the width of the current global bit line. This reduces the voltage drop across the global bit line.
[0152] FIG. 30 shows a VMM system 3000. The VMM system 3000 includes an array 3010, a (shift register) SR3001, a digital - to - analog converter 3002 (receiving an input from SR3001 and outputting an equivalent (analog or pseudo - analog) level or information), an adder circuit 3003, an analog - to - digital converter (ADC) 3004, bit line switches (not shown), dummy bit lines (not shown), and dummy bit line switches (not shown). As shown, the ADCs 3004 can be combined together to fabricate a single ADC having higher precision (i.e., a larger number of bits).
[0153] The adder circuit 3003 can include the circuits shown in FIGS. 31 - 33. This can include, without limitation, circuits for normalization, scaling, arithmetic operations, activation, or statistical rounding.
[0154] FIG. 31 shows a current-voltage summing circuit 3100 adjustable by a variable resistor, including current sources 3101-1, ..., 3101-n that respectively draw currents Ineu1, ..., Ineun (only Ineu1 and Ineu2 are shown here, which are currents received from the bit line(s) of the VMM array), an operational amplifier 3102, a variable holding capacitor 3104, a variable resistor 3103, and a switch 3106. The operational amplifier 3102 outputs a voltage Vneuout = R3103 * (Ineu1 + Ineu2), which is proportional to the current Ineux (the current from a column in the VMM array). The variable holding capacitor 3104 is used to hold the output voltage when the switch 3106 is open. This held output voltage is used, for example, to be converted into digital bits by an ADC circuit. The variability of the capacitor 3103 and the resistor 3103 is achieved, for example, by a trimming circuit (trimming the value of the capacitor or resistor) to adjust the dynamic range of the output of the operational amplifier 3102 according to the input current range of the neuron currents 3101s.
[0155] FIG. 32 shows a current-voltage summing circuit 3200 adjustable by a variable capacitor (basically an integrator), including current sources 3201-1, ..., 3201-n that respectively draw currents Ineu1, ..., Ineun (currents received from the bit line(s) of the VMM array), an operational amplifier 3202, a variable capacitor 3203, and a switch 3204. The operational amplifier 3202 outputs a voltage Vneuout = Ineu * Integral time / C3203, which is proportional to the current(s) Ineu.
[0156] FIG. 33A shows a voltage adder 3300 adjustable by a variable capacitor (i.e., a switched capacitor (SC) circuit), which includes switches 3301 and 3302, variable capacitors 3303 and 3304, an operational amplifier 3305, a variable capacitor 3306, and a switch 3307. When switch 3301 is closed, input Vin0 is provided to operational amplifier 3305. When switch 3302 is closed, input Vin1 is provided to operational amplifier 3305. Preferably, switches 3301 and 3302 are not closed simultaneously. Operational amplifier 3305 generates an output Vout that is an amplified version of the input (either Vin0 and / or Vin1 depending on which of switches 3301 and 3302 is closed). That is, Vout = Cin / Cout * (Vin), where Cin is C3303 or C3304, and Cout is C3306. For example, Vout = Cin / Cout * Σ(Vinx), where Cin = C3303 or C3304. In one embodiment, Vin0 is a w+ voltage, Vin1 is a w− voltage, and voltage adder 3300 sums them to generate an output voltage Vout.
[0157] FIG. 33B shows a voltage adder 3350 that includes switches 3351, 3352, 3353, and 3354, a variable input capacitor 3358, an operational amplifier 3355, a variable feedback capacitor 3356, and a switch 3357. In one embodiment, Vin0 is a w+ voltage, Vin1 is a w− voltage, and voltage adder 3300 sums them to generate an output voltage Vout.
[0158] When input = Vin0: When switches 3354 and 3351 are closed, input Vin0 is provided to the upper terminal of capacitor 3358. Then, switch 3351 is opened and switch 3353 is closed to transfer charge from capacitor 3358 to feedback capacitor 3356. Basically then, output VOUT = (C3358 / C3356) * Vin0 (e.g., when VREF = 0).
[0159] When the input is Vin1: When switches 3353 and 3354 are closed, both terminals of capacitor 3358 are discharged to VREF. Then, switch 3354 is opened and switch 3352 is closed to charge the bottom terminal of capacitor 3358 to Vin1, and then the feedback capacitor 3356 is charged to VOUT = -(C3358 / C3356) * Vin1 (when VREF = 0).
[0160] Therefore, for example, when VREF = 0 and the Vin1 input is enabled after the Vin0 input is enabled, VOUT = (C3358 / C3356) * (Vin0 - Vin1). This is used, for example, to implement w = w+ - w-.
[0161] The method of input / output operation applied to the above VMM array may be in digital format or analog format. The method includes the following: · Sequential input IN[0:q] to the DAC: · Operate IN0, then IN1,..., then INq sequentially. All input bits have the same VCGin. All bit line (neuron) outputs are summed using a scaled binary index multiplier. This is done either before or after the ADC. ·Adjusted Neuron (Bit Line) Binary Index Multiplication Method: As shown in FIG. 30, an exemplary adder has two bit lines BL0 and Bln. The weights are distributed across a plurality of bit lines from BL0 to BLn. For example, there are four bit lines BL0, BL1, BL2, BL3. The output from bit line BL0 is multiplied by 2^0 = 1. The output from bit line BLn representing the nth binary bit position is multiplied by 2^n, for example, when n = 3, 2^3 = 8. Then, the outputs from all bit lines after being appropriately multiplied by the binary bit position 2^n are added together. This is then digitized by an ADC. This method means that all cells need to have only a binary range, and the multi-level range (n bits) is achieved by peripheral circuitry (i.e., by the adder circuit). Thus, the voltage drop across all bit lines is approximately the same. ·Operate IN0, IN1,..., and then INq sequentially. Each input bit has a corresponding analog value VCG. All neuron outputs are added together for all input bit evaluations. This is done either before or after the ADC. ·Parallel Input to DAC ·Each input IN[0:q] has a corresponding analog value VCGin. All neuron outputs are added together using the adjusted binary index multiplication method as described above. This is done either before or after the ADC.
[0162] FIGS. 34A, 34B, and 34C show output circuits that can be used in the adder circuit 3003 and analog-to-digital converter 3004 of FIG. 30.
[0163] FIG. 34A shows an output circuit 3400 that includes an analog-to-digital converter 3402 that receives a neuron output 3401 and outputs an output digital bit 3403.
[0164] Figure 34B shows an output circuit 3410 including a neuron output circuit 3411 and an analog-to-digital converter 3412, which together receive a neuron output 3401 and generate an output 3413.
[0165] Figure 34C shows an output circuit 3420 including a neuron output circuit 3421 and a converter 3422, which together receive a neuron output 3401 and generate an output 3423.
[0166] The neuron output circuit 3411 or 3412 may perform, without limitation, for example, addition, scaling, normalization, or arithmetic operations. The converter 3422 may perform, without limitation, for example, ADC (analog-to-digital converter), PDC (pulse-to-digital converter), AAC (analog-to-analog converter), or APC (analog-to-pulse converter) operations.
[0167] Figure 35 shows a neuron output circuit 3500 including an adjustable (scaling) current source 3501 and an adjustable (scaling) current source 3502, which together generate an output i OUT which is the neuron output. This circuit may simultaneously perform addition of positive and negative weights, i.e., w = w+ - w-, i.e., Iw+ and Iw-, and up or down scaling of the output neuron current.
[0168] Figure 36 shows a configurable neuron serial analog-to-digital converter 3600. The converter includes an integrator 3670 that integrates the neuron output current into an integration capacitor 3602. In one embodiment, the digital output (count output) 3621 is generated by clocking a ramping VRAMP 3650 until the comparator 3604 switches polarity. In another embodiment, it is generated by ramping down node VC 3610 with a ramp current 3651 until VOUT 3603 reaches the value of VREF 3650, at which point the output EC 3605 signal of the comparator 3064 disables the counter 3620. The (n-bit) ADC can be configured to have a lower bit accuracy <n bits or a higher bit accuracy >n bits depending on the target application. The configurability can be achieved, without limitation, by configuring the capacitor 3602, the current 3651, or the ramping speed of VRAMP 3650, or the clock speed of the clock 3641. In another embodiment, the ADC circuit of one VMM array is configured to have a lower accuracy <n bits, and the ADC circuit of another VMM array is configured to have a higher accuracy >n bits. Further, the ADC circuit of one neuron circuit can be configured to generate a higher n-bit ADC accuracy, such as by combining the integration capacitors 3602 of two ADC circuits in combination with the next ADC of the next neuron circuit.
[0169] Figure 37 shows a configurable neuron SAR (successive approximation register) analog-to-digital converter 3700. This circuit is a successive approximation converter based on charge redistribution using binary capacitors. It includes a binary CDAC (capacitor-based DAC) 3701, an operational amplifier / comparator 3702, and SAR logic 3703. As shown, GndV 3704 is a low voltage reference level, for example, ground level.
[0170] Figure 38 shows a configurable neuron combo SAR analog-to-digital converter 3800. This circuit combines two ADCs from two neuron circuits into one to achieve higher n-bit accuracy. For example, in the case of a 4-bit ADC of one neuron circuit, this circuit can achieve an accuracy of >4 bits, such as 8-bit ADC accuracy, by combining two 4-bit ADCs. The combo SAR ADC circuit topology is equivalent to a split-cap (bridge capacitor (cap) or attenuation cap) SAR ADC circuit. For example, an 8-bit 4C-4C SAR ADC is brought about by combining two adjacent 4-bit 4C SAR ADC circuits. A bridge circuit 3804 is required to achieve this, and the capacitance of this circuit = (total number of CDAC cap units / total number of CDAC cap units - 1).
[0171] Figure 39 shows a configurable neuron pipelined SAR CDAC ADC circuit 3900 that can be used to increase the number of bits in a pipelined manner in combination with the following SAR ADC. The residual voltage 3906 is generated by capacitor 3930Cf to be provided as an input to the next stage of the pipelined ADC (for example, to provide a gain of 2 (the Cf-to-C ratio of all caps of DAC 3901) as an input to the next SAR CDAC ADC).
[0172] Additional implementation details regarding circuits of configurable output neurons (such as configurable neuron ADCs) can be found in U.S. Patent Application No. 16 / 449,201, filed on June 21, 2019, by the same assignee and titled "Configurable Input Blocks and Output Blocks and Physical Layout for Analog Neural Memory in a Deep Learning Artificial Neural Network", which is incorporated herein by reference.
[0173] As used herein, it should be noted that both the terms "over" and "on" include both "directly over" (with no intervening material, element, or gap therebetween) and "indirectly over" (with an intervening material, element, or gap therebetween). Similarly, the term "adjacent" includes "directly adjacent" (with no intervening material, element, or gap therebetween) and "indirectly adjacent" (with an intervening material, element, or gap therebetween), "attached to" includes "directly attached to" (with no intervening material, element, or gap therebetween) and "indirectly attached to" (with an intervening material, element, or gap therebetween), and "electrically coupled" includes "directly electrically coupled" (with no intervening material or element electrically connecting the elements together therebetween) and "indirectly electrically coupled" (with an intervening material or element electrically connecting the elements together therebetween). For example, forming an element "over a substrate" can include forming the element directly on the substrate without an intervening material / element therebetween, and forming the element indirectly over the substrate with one or more intervening material / elements therebetween.
Claims
1. A non-volatile memory system comprising: An array of non-volatile memory cells arranged in rows and columns; A plurality of bit lines, each of the plurality of bit lines being coupled to a column of non-volatile memory cells; A plurality of word lines, each of the plurality of word lines being coupled to a row of non-volatile memory cells; A plurality of pull-down cells, each pull-down cell having the same physical structure as a respective non-volatile memory cell; At least one of the plurality of rows arranged to receive a negative value input.
2. The non-volatile memory system according to claim 1, further comprising at least one of the plurality of rows arranged to receive a positive value input.
3. The non-volatile memory system according to claim 1, wherein the output of the negative value input is made negative at the output of the array.
4. The memory system according to claim 1, wherein the memory system is part of a neural network.
5. The memory system according to claim 4, wherein the neural network is an analog neural network.
6. The non-volatile memory system according to claim 4, wherein the array implements negative weights of the neural network.
7. The non-volatile memory system according to claim 6, wherein the output of the negative weight is made negative at the array output.
8. The non-volatile memory system according to claim 6, wherein the array implements positive weights of the neural network.
9. The non-volatile memory system according to claim 8, wherein the outputs of the negative value input and the positive value input are combined at the output of the array.
10. During a read or verification operation of a selected non-volatile memory cell, current flows sequentially through a bit line, the selected non-volatile cell, and a pull-down cell, according to claim 1 of the non-volatile memory system.
11. During a read or verification operation of a selected non-volatile memory cell, current flows sequentially through one of the plurality of bit lines, the selected non-volatile memory cell, and a pull-down cell, according to claim 1 of the non-volatile memory system.
12. The non-volatile memory system according to claim 1, wherein the non-volatile memory cells in the array of the non-volatile memory cells are split-gate non-volatile memory cells.
13. The non-volatile memory system according to claim 1, wherein the non-volatile memory cells in the array of the non-volatile memory cells are stacked-gate non-volatile memory cells.
Citation Information
Patent Citations
Semiconductor memory
JP2000251489A
Vector-by-matrix multiplier modules based on non-volatile 2d and 3D memory arrays
US20190213234A1
Programmable neuron for analog non-volatile memory in deep learning artificial neural network
WO2019135839A1