Analog Neural Memory Array in an Artificial Neural Network with Logic Cells and an Improved Programming Mechanism
By grouping physical memory cells into logical cells and employing different programming mechanisms, the analog neural memory array addresses the challenges of energy efficiency and scalability in artificial neural networks, achieving high accuracy and speed.
Patent Information
- Application Number
- JP2024065576
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-10-28
- Filing Date
- 2024-04-15
- Publication Date
- 2025-06-30
- Estimated Expiration
- 2040-10-30
AI Technical Summary
Existing artificial neural networks face challenges in achieving high-performance information processing due to the lack of appropriate hardware technologies, particularly in terms of energy efficiency and scalability for large numbers of synapses.
The development of an analog neural memory array using non-volatile memory cells, where two or more physical memory cells are grouped to form a logical cell capable of storing one of N possible levels, and programmed using different mechanisms such as coarse, fine, and ultra-fine programming.
This approach achieves high programming accuracy and speed while optimizing area, thereby enhancing the energy efficiency and scalability of artificial neural networks.
Smart Images

Figure 0007700312000024 
Figure 0007700312000025 
Figure 0007700312000026
Abstract
Description
Technical Field
[0001] (Claim of Priority) This application claims priority to U.S. Provisional Patent Application No. 63 / 024,351, filed May 13, 2020, entitled "Analog Neural Memory Array in Artificial Neural Network Comprising Logical Cells and Improved Programming Mechanism", and U.S. Patent Application No. 17 / 082,956, filed October 28, 2020, entitled "Analog Neural Memory Array in Artificial Neural Network Comprising Logical Cells and Improved Programming Mechanism".
[0002] (Field of the Invention) Numerous embodiments of analog neural memory arrays are disclosed. Two or more physical memory cells are grouped together to form a logical cell that stores one of N possible levels. Within each logical cell, the memory cells can be programmed using different mechanisms. For example, one or more of the memory cells within a logical cell can be programmed using a coarse programming mechanism, one or more of the memory cells can be programmed using a fine mechanism, and one or more of the memory cells can be programmed using an ultra-fine mechanism. This achieves very high programming accuracy and programming speed with optimal area.
Background Art
[0003] An artificial neural network mimics a biological neural network (the central nervous system of an animal, particularly the brain), can depend on a large number of inputs, and is used to estimate or approximate a generally unknown function. An artificial neural network generally includes layers of interconnected "neurons" that exchange messages.
[0004] Figure 1 shows an artificial neural network, in which the circles represent layers of inputs or neurons. The connections (referred to as synapses) are represented by arrows and have numerical weights that can be tuned based on experience. This enables the artificial neural network to adapt to the inputs and become learnable. Typically, an artificial neural network includes multiple input layers. Typically, there is one or more intermediate layers of neurons and an output layer of neurons that provides the output of the neural network. At each level, the neurons make decisions individually or collectively based on the data received from the synapses.
[0005] One of the main challenges in the development of artificial neural networks for high-performance information processing is the lack of appropriate hardware technologies. In practice, practical artificial neural networks rely on a very large number of synapses, which enables high connectivity between neurons, i.e., a very high degree of parallelization of computational processing. In principle, such complexity can be realized by digital supercomputers or dedicated graphics processing unit clusters. However, in addition to high costs, these approaches also suffer from poor energy efficiency compared to biological networks, which mainly perform low-precision analog calculations and consume far less energy. CMOS analog circuits have been used in artificial neural networks, but most CMOS-implemented synapses have been too bulky assuming a large number of neurons and synapses.
[0006] The applicant has previously disclosed in U.S. Patent Application No. 15 / 594,439, published as U.S. Patent Publication No. 2017 / 0337466, which is incorporated by reference, an artificial (analog) neural network that utilizes one or more non-volatile memory arrays as synapses. The non-volatile memory arrays operate as analog neuromorphic memories. As used herein, the term neuromorphic means a circuit that implements a model of the nervous system. An analog neuromorphic memory includes a first plurality of synapses configured to receive a first plurality of inputs and then generate a first plurality of outputs, and a first plurality of neurons configured to receive the first plurality of outputs. The first plurality of synapses includes a plurality of memory cells, each memory cell including a spaced-apart source region and drain region formed in a semiconductor substrate with a channel region extending therebetween, a floating gate insulated and disposed above a first portion of the channel region, and a non-floating gate insulated and disposed above a second portion of the channel region. Each of the plurality of memory cells is configured to store a weight value corresponding to a plurality of electrons on the floating gate. The plurality of memory cells is configured to multiply the stored weight values by the first plurality of inputs to generate the first plurality of outputs. An array of memory cells arranged in this manner may be referred to as a vector matrix multiplication (VMM) array.
[0007] Here, examples of different non-volatile memory cells that can be used in VMM are discussed. <<Non-volatile memory cell>>
[0008] Various types of known non-volatile memory cells can be used in a VMM array. For example, U.S. Patent No. 5,029,130 (the “’130 patent”), which is incorporated herein by reference, discloses an array of split-gate non-volatile memory cells, which are a type of flash memory cell. Such a memory cell 210 is shown in FIG. 2. Each memory cell 210 includes a source region 14 and a drain region 16 formed in a semiconductor substrate 12, and there is a channel region 18 between the source region 14 and the drain region 16. A floating gate 20 is formed insulated above a first portion of the channel region 18 (and controls the conductivity of the first portion of the channel region 18), and is formed over a portion of the source region 14. A word line terminal 22 (typically coupled to a word line) is disposed insulated above a second portion of the channel region 18 (and controls the conductivity of the second portion of the channel region 18), and has a first portion that extends upward above the floating gate 20 and a second portion that extends upward above the floating gate 20. The floating gate 20 and the word line terminal 22 are insulated from the substrate 12 by a gate oxide. A bit line terminal 24 is coupled to the drain region 16.
[0009] By applying a high positive voltage to the word line terminal 22, the memory cell 210 is erased (electrons are removed from the floating gate), whereby the electrons in the floating gate 20 pass through the insulator therebetween from the floating gate 20 to the word line terminal 22 via Fowler-Nordheim tunneling.
[0010] The memory cell 210 is programmed (electrons are applied to the floating gate) by applying a positive voltage to the word line terminal 22 and a positive voltage to the source region 14. An electron current flows from the drain region 16 toward the source region 14 (source line terminal). The electrons are accelerated and excited (heated) when they reach the gap between the word line terminal 22 and the floating gate 20. A portion of the heated electrons is injected into the floating gate 20 through the gate oxide due to the electrostatic attraction from the floating gate 20.
[0011] The memory cell 210 is read by applying a positive read voltage to the drain region 16 and the word line terminal 22 (turning on the portion of the channel region 18 below the word line terminal). When the floating gate 20 is positively charged (i.e., electrons are erased), the portion of the channel region 18 below the floating gate 20 also turns on similarly, and current flows through the channel region 18, which is detected as the erased state, i.e., the "1" state. When the floating gate 20 is negatively charged (i.e., programmed with electrons), the portion of the channel region below the floating gate 20 becomes almost or completely off, and current does not (or hardly) flow through the channel region 18, which is detected as the programmed state, i.e., the "0" state.
[0012] Table 1 shows the typical voltage ranges that can be applied to the terminals of the memory cell 110 to perform read, erase, and program operations. Table 1: Operation of the flash memory cell 210 in FIG. 2
Table 1
[0013] FIG. 3 shows a memory cell 310 similar to the memory cell 210 in FIG. 2 with an additional control gate (CG) terminal 28. The control gate terminal 28 is biased at a high voltage (e.g., 10V) during programming, a low or negative voltage (e.g., 0V / -8V) during erase, and a low or medium voltage (e.g., 0V / 2.5V) during read. The other terminals are biased in the same manner as the terminals in FIG. 2.
[0014] FIG. 4 shows a four-gate memory cell 410 comprising a source region 14, a drain region 16, a floating gate 20 above a first portion of the channel region 18, a select gate 22 (typically coupled to a word line, WL) above a second portion of the channel region 18, a control gate 28 above the floating gate 20, and an erase gate 30 above the source region 14. This configuration is described in U.S. Patent No. 6,747,310, which is incorporated herein by reference for all purposes. Here, all gates except the floating gate 20 are non-floating gates, i.e., they are electrically connected or connectable to a voltage source. Programming is performed by injecting hot electrons themselves from the channel region 18 into the floating gate 20. Erasure is performed by tunneling electrons from the floating gate 20 to the erase gate 30.
[0015] Table 2 shows typical voltage ranges that can be applied to the terminals of the memory cell 410 to perform read, erase, and program operations. Table 2: Operation of the Flash Memory Cell 410 of FIG. 4
Table 2
[0016] FIG. 5 shows a memory cell 510 similar to the memory cell 410 of FIG. 4, except that the memory cell 510 does not include an erase gate (EG) terminal. Erasure is performed by biasing the substrate 18 to a high voltage and biasing the control gate CG terminal 28 to a low voltage or a negative voltage. Alternatively, erasure is performed by biasing the word line terminal 22 to a positive voltage and biasing the control gate terminal 28 to a negative voltage. Programming and reading are the same as those of FIG. 4.
[0017] FIG. 6 shows a three-gate memory cell 610, which is another type of flash memory cell. Memory cell 610 is identical to memory cell 410 of FIG. 4, except that memory cell 610 does not have a separate control gate terminal. (Erasure occurs through the use of an erase gate terminal) The erase operation and the read operation are the same as those of FIG. 4, except that no control gate bias is applied. The programming operation is also performed without a control gate bias. As a result, during the program operation, a higher voltage must be applied to the source line terminal to compensate for the lack of control gate bias.
[0018] Table 3 shows the typical voltage ranges that can be applied to the terminals of memory cell 610 to perform read, erase, and program operations. Table 3: Operation of Flash Memory Cell 610 of FIG. 6
Table 3
[0019] FIG. 7 shows a stacked gate memory cell 710, which is another type of flash memory cell. Memory cell 710 is the same as memory cell 210 of FIG. 2, except that the floating gate 20 extends over the entire channel region 18 and the control gate terminal 22 (coupled to the word line) is separated by an insulating layer (not shown) and extends above the floating gate 20. Programming is performed using hot electron injection from the channel 18 to the floating gate 20 in the channel region adjacent to the drain region 16, and erasure is performed using Fowler-Nordheim electron tunneling from the floating gate 20 to the substrate 12. The read operation operates in the same manner as described above for memory cell 210.
[0020] Table 4 shows the typical voltage ranges that can be applied to the terminals of memory cell 710 and substrate 12 to perform read, erase, and program operations. Table 4: Operations of Flash Memory Cell 710 in FIG. 7 [Table 4]
[0021] "Read 1" is a read mode in which the cell current is output to the bit line. "Read 2" is a read mode in which the cell current is output to the source line terminal. Optionally, in an array including rows and columns of memory cells 210, 310, 410, 510, 610, or 710, the source line may be coupled to one row of memory cells or two adjacent rows of memory cells. That is, the source line terminal may be shared by adjacent rows of memory cells.
[0022] FIG. 8 shows a twin split gate memory cell 810. The memory cell 810 includes a floating gate (FG) 20 that is disposed insulated above a substrate 12, a control gate 28 (CG) that is disposed insulated above the floating gate 20, and an erase gate 30 (EG) that is disposed insulated above the floating gate 20 and the control gate 28 and is disposed insulated above the substrate 12. The erase gate is formed in a T-shape such that the upper corner of the control gate CG faces the inner corner of the T-shaped erase gate to improve the erase efficiency. The erase gate 30 (EG) and a drain region 16 (DR) in the substrate adjacent to the floating gate 20 (where the bit line contact 24 (BL) is connected to the drain diffusion region 16 (DR)). The memory cell is formed as a memory cell pair (A on the left and B on the right) and shares a common erase gate 30. This cell design differs from the memory cells described above with reference to FIGS. 2-7 in that it lacks at least the source region under the erase gate EG, lacks a select gate (also called a word line), and lacks the channel region of each memory cell. Instead, a single continuous channel region 18 extends under both memory cells (i.e., from the drain region 16 of one memory cell to the drain region 16 of the other memory cell). To read or program one memory cell, the control gate 28 of the other memory cell is raised to a sufficient voltage, and the channel region portion below is activated by voltage coupling to the floating gate 20 therebetween (e.g., to read or program cell A, the voltage on FGB is raised by voltage coupling from CGB to activate the channel region under FGB). Erase is performed using Fowler Nordheim electron tunneling from the floating gate 20 to the erase gate 30. Programming is performed using hot electron injection from the channel 18 to the floating gate 20, which is shown as Program 1 in Table 5. Alternatively, programming is performed using Fowler Nordheim electron tunneling from the erase gate 30 to the floating gate 20, which is shown as Program 2 in Table 5.Alternatively, programming is performed using Fowler-Nordheim electron tunneling from channel 18 to floating gate 20, where the conditions are the same as Program 2, except that the substrate is biased at a low or negative voltage while the erase gate is biased at a low positive voltage.
[0023] Table 5 shows the typical voltage ranges that can be applied to the terminals of memory cell 810 to perform read, erase, and program operations. Cell A (FG, CGA, BLA) is selected for read, program, and erase operations.
[0024] Table 5: Operation of Flash Memory Cell 810 in FIG. 8
Table 5
[0025] To utilize a memory array including one of the types of non-volatile memory cells in the artificial neural network described above, two modifications are made. First, the lines are configured so that each memory cell can be individually programmed, erased, and read without adversely affecting the memory states of other memory cells in the array, as further described below. Second, continuous (analog) programming of the memory cells is provided.
[0026] Specifically, the memory state of each memory cell in the array (i.e., the charge on the floating gate) can be continuously changed independently, with minimal interference from other memory cells, from a completely erased state to a completely programmed state. In another embodiment, the memory state of each memory cell in the array (i.e., the charge on the floating gate) can be continuously changed independently, with minimal interference from other memory cells, from a completely programmed state to a completely erased state and vice versa. This means that the cell memory can be analog or can store at least one of a number of discrete values (such as 16 or 64 different values), which allows all cells in the memory array to be very precisely and individually tunable, and makes the memory array ideal for memory and fine-tuning adjustments to the synaptic weights of neural networks.
[0027] The methods and means described herein can be applied, without limitation, to other non-volatile memory technologies such as FINFET split-gate flash or stacked-gate flash memory, NAND flash, SONOS (silicon-oxide-nitride-oxide-silicon, charge trapping in nitride), MONOS (metal-oxide-nitride-oxide-silicon, metal charge trapping in nitride), ReRAM (resistive change memory), PCM (phase change memory), MRAM (magnetoresistive memory), FeRAM (ferroelectric memory), OTP (one-time programmable, bi-level or multi-level), and CeRAM (strongly correlated electron memory). The methods and means described herein can be applied, without limitation, to volatile memory technologies used in neural networks such as SRAM, DRAM, and other volatile synaptic cells. <<Neural Networks Using Non-Volatile Memory Cell Arrays>>
[0028] FIG. 9 conceptually shows a non-limiting example of a neural network that utilizes the non-volatile memory array of the present embodiment. This example uses a non-volatile memory array neural network for a face recognition application, but it is also possible to implement other suitable applications using a non-volatile memory array-based neural network.
[0029] S0 is the input layer, and in this example, it is a 32×32 pixel RGB image with 5-bit precision (i.e., three 32×32 pixel arrays, one for each of the colors R, G, and B, and each pixel has 5-bit precision). The synapses CB1 going from the input layer S0 to the layer C1 apply different sets of weights to some instances and shared weights to other instances, scan the input image with a 3×3 pixel overlapping filter (kernel), and shift the filter one pixel (or more than two pixels depending on the model) at a time. Specifically, the values of the 9 pixels in the 3×3 portion of the image (i.e., what is referred to as the filter or kernel) are provided to the synapses CB1, where these 9 input values are multiplied by appropriate weights, and after summing the outputs of the multiplications, a single output value is determined and given by the first synapse of CB1 to generate one pixel of the layer of the feature map C1. The 3×3 filter is then shifted one pixel to the right within the input layer S0 (i.e., a 3-pixel column is added on the right and a 3-pixel column is dropped on the left), and thereby the 9 pixel values of this newly positioned filter are provided to the synapses CB1, where they are multiplied by the same weights as above, and a second single output value is determined by the relevant synapses. This process is continued until the 3×3 filter has scanned the entire 32×32 pixel image of the input layer S0 for all three colors and all bits (precision values). The process is then repeated using different sets of weights until all the feature maps of layer C1 are calculated, generating different feature maps of C1.
[0030] In this example, in layer C1, there are 16 feature maps each having 30×30 pixels. Each pixel is a new feature pixel extracted from the multiplication of the input and the kernel. Thus, each feature map is a two-dimensional array. Therefore, in this example, layer C1 consists of 16 layers of two-dimensional arrays (note that the layers and arrays referred to in this specification are logical relationships rather than necessarily physical relationships, that is, the array is not necessarily oriented as a physical two-dimensional array). Each of the 16 feature maps within layer C1 is generated by one of 16 different sets of synaptic weights applied to the filter scan. All of the C1 feature maps can target different aspects of the same image feature, such as edge identification. For example, the first map (generated using a first set of weights shared by all scans used to generate this first map) can identify circular edges, and the second map (generated using a second set of weights different from the first set of weights) can identify rectangular edges or the aspect ratio of a specific feature, etc.
[0031] Before going from layer C1 to layer S1, an activation function P1 (pooling) that pools values from non - overlapping consecutive 2×2 regions within each feature map is applied. The purpose of the pooling function is to average neighboring positions (or it is also possible to use the max function), for example, to reduce the dependence on edge positions, and to reduce the data size before going to the next stage. In layer S1, there are 16 15×15 feature maps (i.e., 16 different arrays of 15×15 pixels each). The synapse CB2 from layer S1 to layer C2 scans the maps in S1 with a 4×4 filter with a 1 - pixel filter shift. In layer C2, there are 22 12×12 feature maps. Before going from layer C2 to layer S2, an activation function P2 (pooling) that pools values from non - overlapping consecutive 2×2 regions within each feature map is applied. In layer S2, there are 22 6×6 feature maps. In the synapse CB3 from layer S2 to layer C3, an activation function (pooling) is applied, where all neurons in layer C3 are connected to all maps in layer S2 via each synapse of CB3. In layer C3, there are 64 neurons. The synapse CB4 from layer C3 to the output layer S3 fully connects C3 to S3, that is, all neurons in layer C3 are connected to all neurons in layer S3. The output in S3 contains 10 neurons, and the neuron with the highest output determines the class (classification). This output can indicate, for example, the identification or classification of the content of the original image.
[0032] Each layer of synapses is implemented using an array or a part of an array of non - volatile memory cells.
[0033] Figure 10 is a block diagram of a system that can be used for that purpose. The VMM system 32 includes non-volatile memory cells and is utilized as synapses (such as CB1, CB2, CB3, and CB4 in FIG. 6) between one layer and the next. Specifically, the VMM system 32 includes a VMM array 33 having non-volatile memory cells arranged in rows and columns, an erase gate and word line gate decoder 34, a control gate decoder 35, a bit line decoder 36, and a source line decoder 37, and these decoders decode respective inputs to the non-volatile memory cell array 33. Inputs to the VMM array 33 can be made from the erase gate and word line gate decoder 34 or from the control gate decoder 35. The source line decoder 37 in this example also decodes the output of the VMM array 33. Alternatively, the bit line decoder 36 can decode the output of the VMM array 33.
[0034] The VMM array 33 serves two purposes. First, it stores the weights used by the VMM system 32. Second, the VMM array 33 effectively multiplies the inputs by the weights stored in the VMM array 33 and sums them for each output line (source line or bit line) to generate an output, which becomes the input to the next layer or the input to the last layer. By performing the functions of multiplication and addition, the VMM array 33 eliminates the need for separate multiplication and addition logic circuits and is also power-efficient due to in-memory computing.
[0035] The output of the VMM array 33 is supplied to a differential adder (such as an addition operation amplifier or an addition current mirror) 38 that sums the outputs of the VMM array 33 to create a single value for convolution. The differential adder 38 is arranged to perform the sum of both the positive weight input and the negative weight input and output a single value.
[0036] The summed output value of the differential adder 38 is then supplied to an activation function circuit 39 that rectifies the output. The activation function circuit 39 can provide a sigmoid function, a tanh function, a ReLU function, or any other non-linear function. The rectified output value of the activation function circuit 39 becomes an element of the feature map of the next layer (e.g., C1 in FIG. 9), and is then applied to the next synapse to generate the next feature map layer or the last layer. Thus, in this example, the VMM array 33 constitutes a plurality of synapses (which receive inputs from the previous layer of neurons or from an input layer such as an image database), and the adder 38 and the activation function circuit 39 constitute a plurality of neurons.
[0037] The inputs (WLx, EGx, CGx, and optionally BLx and SLx) to the VMM system 32 of FIG. 10 can be at an analog level, a binary level, a digital pulse (in which case a pulse - analog converter PAC may be required to convert the pulse to an appropriate input analog level), or a digital bit (in which case a DAC is provided to convert the digital bit to an appropriate input analog level), and the output can be at an analog level (e.g., current, voltage, or charge), a binary level, a digital pulse, or a digital bit (in which case an output ADC is provided to convert the output analog level to a digital bit).
[0038] FIG. 11 is a block diagram showing the use of multiple layers of the VMM system 32, labeled as VMM systems 32a, 32b, 32c, 32d, and 32e in the figure. As shown in FIG. 11, an input (denoted as Inputx) is converted from digital to analog by a digital-to-analog converter 31 and provided to the input VMM system 32a. The converted analog input can be a voltage or a current. The input D / A conversion of the first layer can be performed by using a function or a LUT (look-up table) that maps the input Inputx to an appropriate analog level of the matrix multiplier of the input VMM system 32a. The input conversion can also be performed by an analog-to-analog (A / A) converter so as to convert an external analog input to the mapped analog input to the input VMM system 32a. The input conversion can also be performed by a digital-to-digital pulse (D / P) converter so as to convert an external digital input to the mapped digital pulse to the input VMM system 32a.
[0039] The output generated by the input VMM system 32a is then provided as input to the next VMM system (hidden level 1) 32b, which then generates an output that is in turn provided as input to the next input VMM system (hidden level 2) 32c, and so on. The various layers of the VMM system 32 function as the layers of synapses and neurons of a convolutional neural network (CNN). Each of the VMM systems 32a, 32b, 32c, 32d, and 32e can be a stand-alone physical system with its corresponding non-volatile memory array, or multiple VMM systems can utilize different portions of the same physical non-volatile memory array, or multiple VMM systems can utilize overlapping portions of the same physical non-volatile memory array. Each of the VMM systems 32a, 32b, 32c, 32d, and 32e can also be time-multiplexed across the various portions of its array or neurons. The example shown in FIG. 11 includes five layers (32a, 32b, 32c, 32d, 32e), namely, one input layer (32a), two hidden layers (32b, 32c), and two fully connected layers (32d, 32e). One skilled in the art will understand that this is merely exemplary and that the system can alternatively include more than two hidden layers and more than two fully connected layers. <<VMM Array>>
[0040] FIG. 12 shows a neuron VMM array 1200 that is particularly suitable for the memory cell 310 shown in FIG. 3 and that is utilized as part of the synapses and neurons between the input layer and the next layer. The VMM array 1200 includes a memory array 1201 of non-volatile memory cells and a reference array 1202 of non-volatile reference memory cells (located at the top of the array). Alternatively, another reference array can be located at the bottom.
[0041] In the VMM array 1200, control gate lines such as the control gate line 1203 extend vertically (thus, the reference array 1202 in the row direction is orthogonal to the control gate line 1203), and erase gate lines such as the erase gate line 1204 extend horizontally. Here, the input to the VMM array 1200 is provided to the control gate lines (CG0, CG1, CG2, CG3), and the output of the VMM array 1200 appears on the source lines (SL0, SL1). In one embodiment, only even rows are used, and in another embodiment, only odd rows are used. The current applied to each source line (SL0 and SL1 respectively) performs a summation function of all the currents from the memory cells connected to that particular source line.
[0042] As described herein for neural networks, the non-volatile memory cells of the VMM array 1200, i.e., the flash memory of the VMM array 1200, are preferably configured to operate in the subthreshold region.
[0043] The non-volatile reference memory cells and non-volatile memory cells described herein are biased with weak inversion as follows: Ids = Io * e (Vg-Vth) / nVt = w * Io * e (Vg) / nVt where w = e (-Vth) / nVt and where Ids is the drain-source current, Vg is the gate voltage of the memory cell, Vth is the threshold voltage of the memory cell, Vt is the thermal voltage = k * T / q, where k is the Boltzmann constant, T is the Kelvin temperature, q is the electron charge, n is the slope factor = 1+(Cdep / Cox), Cdep is the capacitance of the depletion layer, and Cox is the capacitance of the gate oxide layer, Io is the memory cell current at a gate voltage equal to the threshold voltage, and Io is (Wt / L) * u * Cox * (n - 1) * Vt 2Proportional to, where u is the carrier mobility, and Wt and L are the width and length of the memory cell, respectively.
[0044] When an I-V log converter that converts the input current Ids to the input voltage Vg using a memory cell (such as a reference memory cell or a peripheral memory cell) or a transistor is used, Vg is as follows: Vg = n * Vt * log[Ids / wp * Io] Where wp is the w of the reference or peripheral memory cell.
[0045] When an I-V log converter that converts the input current Ids to the input voltage Vg using a memory cell (such as a reference memory cell or a peripheral memory cell) or a transistor is used, Vg is as follows: Vg = n * Vt * log[Ids / wp * Io]
[0046] Where wp is the w of the reference or peripheral memory cell.
[0047] For the memory array used as the vector matrix multiplier VMM array, the output current is as follows: Iout = wa * Io * e (Vg) / nVt That is Iout = (wa / wp) * Iin = W * Iin W = e (Vthp-Vtha) / nVt Iin = wp * Io * e (Vg) / nVt Where wa = w for each memory cell of the memory array. Vthp is the effective threshold voltage of the peripheral memory cell, and Vtha is the effective threshold voltage of the main (data) memory cell. Note that the threshold voltage of the transistor is a function of the substrate body bias voltage, and the substrate body bias can be modulated for various compensations such as overheating or modulation of the cell current. Vth = Vth0 + gamma(SQRT(Vsb + |2 * φF|) - SQRT|2 * φF|) Vth0 is the threshold voltage with zero substrate bias, φF is the surface potential, and gamma is the body effect parameter.
[0048] The word line or control gate can be used as an input to the memory cell for the input voltage.
[0049] Alternatively, the non-volatile memory cells of the VMM array described herein can be configured to operate in the linear region. Ids = beta * (Vgs - Vth) * Vds; beta = u * Cox * Wt / L W α (Vgs - Vth) That is, the weight W in the linear region is proportional to (Vgs - Vth).
[0050] The word line or control gate or bit line or source line can be used as an input to the memory cell operating in the linear region. The bit line or source line can be used as an output of the memory cell.
[0051] For an I-V linear converter, a memory cell (such as a reference memory cell or a peripheral memory cell) or a transistor or a resistor operating in the linear region can be used to linearly convert the input / output current to the input / output voltage.
[0052] Alternatively, the memory cells of the VMM array described herein can be configured to operate in the saturation region. Ids = 1 / 2 * beta * (Vgs - Vth) 2; Beta = u * Cox * Wt / L W α (Vgs - Vth) 2 That is, the weight W is proportional to (Vgs - Vth). 2 is proportional to
[0053] The word line, control gate, or erase gate can be used as an input to a memory cell operating in the saturation region. The bit line or source line can be used as the output of an output neuron.
[0054] Alternatively, the memory cells of the VMM array described herein can be used in all regions or combinations thereof (subthreshold, linear, or saturation) for each layer or multiple layers of a neural network.
[0055] FIG. 13 shows a neuron VMM array 1300 particularly suitable for the memory cell 210 shown in FIG. 2 and is utilized as a synapse between the input layer and the next layer. The VMM array 1300 includes a memory array 1303 of non - volatile memory cells, a reference array 1301 of first non - volatile reference memory cells, and a reference array 1302 of second non - volatile reference memory cells. The reference arrays 1301 and 1302 arranged in the column direction of the array function to convert the current inputs flowing into the terminals BLR0, BLR1, BLR2, and BLR3 into voltage inputs WL0, WL1, WL2, and WL3. In practice, the first and second non - volatile reference memory cells are diode - connected through a multiplexer 1314 (only a part shown) in a state where the current inputs flow in. The reference cells are tuned (e.g., programmed) to a target reference level. The target reference level is provided by a reference mini - array matrix (not shown).
[0056] The memory array 1303 serves two purposes. First, it stores the weights used by the VMM array 1300 in respective memory cells. Second, the memory array 1303 effectively multiplies the input (i.e., the current inputs provided to terminals BLR0, BLR1, BLR2, and BLR3, which are converted by the reference arrays 1301 and 1302 into input voltages and supplied to the word lines WL0, WL1, WL2, and WL3) by the weights stored in the memory cell array 1303, and then adds up all the results (memory cell currents) to generate the outputs of respective bit lines (BL0 to BLN), and this output becomes the input to the next layer or the input to the last layer. By the memory array 1303 performing the functions of multiplication and addition, the need for separate multiplication and addition logic circuits is eliminated, and the power efficiency is also good. Here, the voltage input is provided to the word lines WL0, WL1, WL2, and WL3, and the output appears on respective bit lines BL0 to BLN during the read (inference) operation. The current arranged in each bit line BL0 to BLN performs the total function of the currents from all the non-volatile memory cells connected to that specific bit line.
[0057] Table 6 shows the operating voltages of the VMM array 1300. The columns in the table indicate the voltages applied to the word lines of the selected cells, the word lines of the non-selected cells, the bit lines of the selected cells, the bit lines of the non-selected cells, the source lines of the selected cells, and the source lines of the non-selected cells, where FLT indicates floating, i.e., no voltage is applied. The rows indicate the operations of read, erase, and program. Table 6: Operation of the VMM array 1300 in FIG. 13:
Table 6
[0058] FIG. 14 shows a neuron VMM array 1400 that is particularly suitable for the memory cell 210 shown in FIG. 2 and is used as part of synapses and neurons between the input layer and the next layer. The VMM array 1400 includes a memory array 1403 of non-volatile memory cells, a reference array 1401 of first non-volatile reference memory cells, and a reference array 1402 of second non-volatile reference memory cells. The reference arrays 1401 and 1402 extend in the row direction of the VMM array 1400. The VMM array is similar to the VMM 1300 except that the word lines extend vertically in the VMM array 1400. Here, the inputs are provided to the word lines (WLA0, WLB0, WLA1, WLB2, WLA2, WLB2, WLA3, WLB3), and the outputs appear on the source lines (SL0, SL1) during a read operation. The current applied to each source line performs a sum function of all the currents from the memory cells connected to that particular source line.
[0059] Table 7 shows the operating voltages of the VMM array 1400. The columns in the table show the voltages applied to the word line of the selected cell, the word line of the non-selected cell, the bit line of the selected cell, the bit line of the non-selected cell, the source line of the selected cell, and the source line of the non-selected cell. The rows show the read, erase, and program operations. Table 7: Operation of the VMM Array 1400 in FIG. 14
Table 7
[0060] FIG. 15 shows a neuron VMM array 1500 that is particularly suitable for the memory cell 310 shown in FIG. 3 and is used as part of synapses and neurons between the input layer and the next layer. The VMM array 1500 includes a memory array 1503 of non-volatile memory cells, a reference array 1501 of first non-volatile reference memory cells, and a reference array 1502 of second non-volatile reference memory cells. The reference arrays 1501 and 1502 function to convert the current inputs flowing into the terminals BLR0, BLR1, BLR2, and BLR3 into voltage inputs CG0, CG1, CG2, and CG3. In practice, the first and second non-volatile reference memory cells are diode-connected through a multiplexer 1512 (only part shown) in a state where the current inputs flow through BLR0, BLR1, BLR2, and BLR3. The multiplexer 1512 includes corresponding multiplexers 1505 and cascode transistors 1504, respectively, to ensure a constant voltage of each bit line (such as BLR0) of the first and second non-volatile reference memory cells during the read operation. The reference cells are tuned to a target reference level.
[0061] The memory array 1503 serves two purposes. First, it stores the weights used by the VMM array 1500. Second, the memory array 1503 multiplies the input (the current inputs provided to the terminals BLR0, BLR1, BLR2, and BLR3, which are converted by the reference arrays 1501 and 1502 into input voltages and supplied to the control gates (CG0, CG1, CG2, and CG3)) by the weights stored in the memory cell array, and then adds up all the results (cell currents) to generate an output, which appears on BL0 to BLN and serves as an input to the next layer or the last layer. By the memory array performing the multiplication and addition functions, the need for separate multiplication and addition logic circuits is eliminated, and the power efficiency is also good. Here, the input is provided to the control gate lines (CG0, CG1, CG2, and CG3), and the output appears on the bit lines (BL0 to BLN) during the read operation. The current applied to each bit line performs the summation function of all the currents from the memory cells connected to that specific bit line.
[0062] The VMM array 1500 implements unidirectional tuning of non-volatile memory cells within the memory array 1503. That is, each non-volatile memory cell is erased and then partially programmed until the desired charge on the floating gate is reached. This can be performed, for example, using the precision programming techniques described below. If too much charge is applied to the floating gate (in which case an incorrect value is stored in the cell), the cell must be erased and a series of partial programming operations must be repeated. As shown, two rows sharing the same erase gate (such as EG0 or EG1) need to be erased together (known as page erase), and then each cell is partially programmed until the desired charge on the floating gate is reached.
[0063] Table 8 shows the operating voltages of the VMM array 1500. The columns in the table show the voltages applied to the word line of the selected cell, the word line of the non-selected cell, the bit line of the selected cell, the bit line of the non-selected cell, the control gate of the selected cell, the control gate of the non-selected cell within the same sector as the selected cell, the control gate of the non-selected cell in a different sector from the selected cell, the erase gate of the selected cell, the erase gate of the non-selected cell, the source line of the selected cell, and the source line of the non-selected cell. The rows show the operations of read, erase, and program. Table 8: Operation of the VMM Array 1500 of FIG. 15 [Table 8]
[0064] FIG. 16 shows a neuron VMM array 1600 that is particularly suitable for the memory cell 310 shown in FIG. 3 and is used as part of the synapses and neurons between the input layer and the next layer. The VMM array 1600 includes a memory array 1603 of non-volatile memory cells, a reference array 1601 or a first non-volatile reference memory cell, and a reference array 1602 of second non-volatile reference memory cells. The EG lines EGR0, EG0, EG1, and EGR1 extend vertically, and the CG lines CG0, CG1, CG2, and CG3 and the SL lines WL0, WL1, WL2, and WL3 extend horizontally. The VMM array 1600 is similar to the VMM array 1600 except that the VMM array 1600 implements bidirectional tuning, and each individual cell can be completely erased, partially programmed, and partially erased as needed to reach the desired charge level of the floating gate by using the individual EG lines. As shown, the reference arrays 1601 and 1602 convert the input current in the terminals BLR0, BLR1, BLR2, and BLR3 into the control gate voltages CG0, CG1, CG2, and CG3 (through the action of the diode-connected reference cells via the multiplexer 1614), and these voltages are applied to the memory cells in the row direction. The current outputs (neurons) are in the bit lines BL0 to BLN, and each bit line sums all the currents from the non-volatile memory cells connected to that particular bit line.
[0065] Table 9 shows the operating voltages of the VMM array 1600. The columns in the table show the voltages applied to the word line of the selected cell, the word line of the non-selected cell, the bit line of the selected cell, the bit line of the non-selected cell, the control gate of the selected cell, the control gate of the non-selected cell in the same sector as the selected cell, the control gate of the non-selected cell in a different sector from the selected cell, the erase gate of the selected cell, the erase gate of the non-selected cell, the source line of the selected cell, and the source line of the non-selected cell. The rows show the read, erase, and program operations. Table 9: Operation of the VMM Array 1600 in FIG. 16
Table 9
[0066] The inputs to the VMM array can be analog levels, binary levels, timing pulses, or digital bits, and the outputs can be analog levels, binary levels, timing pulses, or digital bits (in which case an output ADC is required to convert the output analog level current or voltage to digital bits).
[0067] For each memory cell in the VMM array, each weight w can be implemented by a single memory cell, or by differential cells, or by two blended memory cells (the average of two or more cells). In the case of differential cells, two memory cells are required to implement the weight w as a differential weight (w = w+ - w-). In the case of two blended memory cells, two memory cells are required to implement the weight w as the average of two cells.
[0068] One challenge in the VMM array is that it requires extremely high accuracy during the programming process. For example, if each cell in the VMM array can store one of N different values (e.g., N = 64 or 128), the system must deposit additional charge little by little on the floating gate of the selected cell to achieve the desired level change. On the other hand, it is still important that the programming be as fast as possible, and that there is an inherent trade-off between programming accuracy and programming speed.
[0069] What is needed is an improved VMM system that can complete programming at a relatively fast pace while still achieving precise programming. Summary of the Invention
[0070] Numerous embodiments of an analog neural memory array are disclosed. Two or more memory cells are grouped together to form a logic cell that stores one of N possible levels. Within each logic cell, the memory cells can be programmed using different mechanisms. For example, one or more of the memory cells within a logic cell can be programmed using a coarse programming mechanism, and one or more of the memory cells can be programmed using a fine mechanism. This achieves extremely high programming accuracy and programming speed.
[0071]
[0072]
[0073]
[0074]
[0075]
[0076]
[0077]
[0078]
[0079]
[0080]
[0081]
[0082]
[0083]
[0084]
[0085]
[0086]
[0087]
[0088]
[0089]
[0090]
[0091]
[0092]
[0093]
[0094]
[0095]
[0096]
[0097]
[0098]
[0099]
[0100]
[0101]
[0102]
[0103]
[0104]
[0105]
[0106]
[0107]
[0108]
[0109]
[0110]
[0111]
[0112]
[0113]
[0114]
[0115]
[0116]
[0117]
Brief Description of the Drawings
[0118]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18A
Figure 18B
Figure 18C
Figure 19A
Figure 19B
Figure 19C
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
Figure 26
Figure 27A
Figure 27B
Figure 27C
Figure 28
Figure 29
Figure 30
Figure 31
Figure 32
Figure 33A
Figure 33B
Figure 34A
Figure 34B
Figure 34C
Figure 35
Figure 36
Figure 37
Figure 38
Figure 39
Figure 40
Figure 41
Figure 42
Figure 43
Figure 44
Figure 45
Figure 46
Figure 47
Figure 48
Embodiments for Carrying Out the Invention
[0119] The artificial neural network of the present invention utilizes a combination of CMOS technology and a non-volatile memory array. <<Embodiment of an improved VMM system>>
[0120] FIG. 17 shows a block diagram of a VMM system 1700. The VMM system 1700 includes a VMM array 1701, a row decoder 1702, a high-voltage decoder 1703, a column decoder 1704, a bit-line driver 1705, an input circuit 1706, an output circuit 1707, control logic 1708, and a bias generator 1709. The VMM system 1700 further includes a high-voltage generation block 1710 including a charge pump 1711, a charge pump regulator 1712, and a high-voltage level generator 1713. The VMM system 1700 further includes an algorithm controller 1714, an analog circuit 1715, control logic 1716, and test control logic 1717. The systems and methods described below may be implemented in the VMM system 1700.
[0121] The input circuit 1706 may include circuits such as a DAC (Digital-to-Analog Converter), DPC (Digital-to-Pulse Converter), AAC (Analog-to-Analog Converter such as a Current-to-Voltage Converter), PAC (Pulse-to-Analog Level Converter), or any other type of converter. The input circuit 1706 may implement a normalization function, a scaling function, or an arithmetic function. The input circuit 1706 may implement a temperature compensation function for the input. The input circuit 1706 may implement an activation function such as ReLU or sigmoid. The output circuit 1707 may include circuits such as an ADC (Analog-to-Digital Converter for converting neuron analog output to digital bits), AAC (Analog Converter such as a Current-to-Voltage Converter), APC (Analog-to-Pulse Converter), or any other type of converter. The output circuit 1707 may implement an activation function such as ReLU or sigmoid. The output circuit 1707 may implement statistical normalization, regularization, up / down scaling function, statistical rounding, or arithmetic functions (e.g., addition, subtraction, division, multiplication, shift, log) of the neuron output. The output circuit 1707 may implement a temperature compensation function for the neuron output or the array output (such as bit line output) in order to keep the power consumption of the array substantially constant or to improve the accuracy of the array (neuron) output by keeping the slope of the IV substantially the same, etc.
[0122] FIG. 18A shows a prior art VMM system 1800. The VMM system 1800 includes exemplary cells 1801 and 1802, an exemplary bit line switch 1803 (connecting the bit line to a sense circuit), an exemplary dummy bit line switch 1804 (coupled to a low level such as ground level in a read), and exemplary dummy cells 1805 and 1806 (source line pull-down cells). The bit line switch 1803 is coupled to a column of cells including cells 1801 and 1802 that are used to store data in the VMM system 1800. The dummy bit line switch 1804 is coupled to a column of cells (bit lines) that are dummy cells not used to store data in the VMM system 1800. This dummy bit line (also known as a source line pull-down bit line) is used as a source line pull-down in a read, which means it is used to pull the source line SL to a low level such as ground level through a memory cell in the dummy bit line.
[0123] One drawback of the VMM system 1800 is that the input impedance of each cell varies due to the length of the electrical path through the associated bit line switch, the cell itself, and the associated dummy bit line switch. For example, FIG. 18B shows the electrical path through the bit line switch 1803, the cell 1801, the dummy cell 1805, and the dummy bit line switch 1804. Similarly, FIG. 18C shows the electrical path through the bit line switch 1803, the vertical metal bit line 1807, the cell 1802, the dummy cell 1808, the vertical metal bit line 1808, and the dummy bit line switch 1804. As can be seen, the path through the cell 1802 crosses significantly longer bit lines and dummy bit lines, which is associated with higher capacitance and higher resistance. This results in a cell 1802 having a larger parasitic impedance in the bit line or source line than the cell 1801. This variability is a drawback because it causes variations in the accuracy of the cell output applied to the read or verification (for program / erase tuning cycles) of the cells, depending on, for example, the position of the cells within the array.
[0124] FIG. 19A shows a VMM system 1900. The VMM system 1900 includes exemplary cells 1901 and 1902, an exemplary bit line switch 1903 (connecting a bit line to a sense circuit), exemplary dummy cells 1905 and 1906 (source line pull-down cells), and an exemplary dummy bit line switch 1904 (coupled to a low level such as ground level during read. This switch connects to a dummy bit line that connects to a dummy cell used as a source line pull-down). As can be seen, the exemplary dummy bit line switch 1904 and other dummy bit line switches are located at opposite ends of the array from the bit line switch 1903 and other bit line switches.
[0125] The advantages of this design can be seen in FIGS. 19B and 19C. FIG. 19B shows an electrical path through bit line switch 1903, cell 1901, dummy cell 1905 (source line pull-down cell), vertical metal bit line 1908, and dummy bit line switch 1904 (coupled to a low level such as ground level during read). FIG. 19C shows an electrical path through bit line switch 1903, vertical metal line 1907, cell 1902, dummy cell 1906 (source line pull-down cell), and dummy bit line switch 1904. The paths are substantially the same (cell, interconnect length), which is true for all cells of the VMM system 1900. As a result, the impedance of the bit line impedance + source line impedance of each cell is substantially the same, which means that the variation in the amount of parasitic voltage drop drawn to read or verify the operation of various cells in the array is relatively the same.
[0126] FIG. 20 shows a VMM system 2000 having global source line pull-down bit lines. The VMM system 2000 is the same as the VMM system 1900, except that dummy bit lines 2005a - 2005n or 2007a - 2007n are connected together (to act as global source line pull-down lines for pulling the memory cell source lines to the ground level during read or verification), dummy bit line switches such as dummy bit line switches 2001 and 2002 are connected or coupled to a common ground, and the source lines are coupled together to a source line switch 2003 that selectively pulls the source lines to the ground. These changes further reduce the variation in (array) parasitic impedance between cells during read or verification operations.
[0127] FIG. 21 shows a VMM system 2100. The VMM system 2100 includes bit line switches 2101, pull-down bit line switches 2102, pull-down bit line switches 2103, bit line switches 2104, data cells 2105 (in this specification, a "data cell" is a memory cell used to store weight values of a neural network), pull-down cells 2106, pull-down cells 2107, and data cell 2018. Note that pull-down cells 2106 and 2107 are adjacent to each other. Thereby, the vertical metal lines BLpdx of the two pull-down cells 2106 and 2107 are connected together (line 2111), and it becomes possible to reduce parasitic resistance by the resulting wider metal line. During the read or verification (for the program / erase tuning cycle) operation of data cell 2105, current enters the bit line terminal of cell 2105 through bit line switch 2101, exits to the source line terminal of cell 2015, then enters source line 2110, where it enters the source line terminals of pull-down cells 2106 and 2107 and passes through pull-down bit line switches 2102 and 2103. During the read or verification (for the program / erase tuning cycle) operation of cell 2104, current enters the bit line terminal of data cell 2108 through bit line switch 2104, exits to the source line terminal of cell 2108, then enters source line 2110, where it enters the source line terminals of pull-down cells 2106 and 2107 and passes through pull-down bit line switches 2102 and 2103. This column pattern is repeated across the entire array, and all four columns include two columns of data cells and two adjacent array columns used for the pull-down operation. In another embodiment, the diffusion of two pull-down cells in two adjacent columns can be merged into one larger diffusion to enhance the pull-down ability. In another embodiment, the diffusion of the pull-down cells can be made larger than the data cell diffusion to enhance the pull-down ability. In another embodiment, each pull-down cell has a bias condition different from the bias condition of the selected data cell.
[0128] In one embodiment, the pull-down cell has the same physical structure as a normal data memory cell. In another embodiment, the pull-down cell has a physical structure different from that of a normal data memory cell. For example, the pull-down cell can be a modified version of a normal data memory cell by modifying one or more physical dimensions (such as width, length) of electrical parameters (layer thickness, implant, etc.). In another embodiment, the pull-down cell is a normal transistor (without a floating gate) such as an IO or high-voltage transistor.
[0129] FIG. 22 shows a VMM system 2200. The VMM system 2200 includes bit lines 2201, pull-down bit lines 2202, data cells 2203 and 2206, pull-down cells 2204 and 2205, and a source line 2210. During the read or verification operation of cell 2203, current enters the bit line terminal of cell 2203 through the bit line switch 2201, exits to the source line terminal of cell 2203, then enters the source line 2210 and the source line terminal of pull-down cell 2204, and passes through the pull-down bit line BLpd 2202. This design is repeated for all columns, and finally, the row containing pull-down cell 2204 becomes the row of pull-down cells.
[0130] During the read or verification (for program / erase tuning cycles) operation of cell 2206, current enters the bit line terminal of cell 2206 through the bit line switch 2201, exits to the source line terminal of cell 2206, then enters the source line 2211 and the source line terminal of pull-down cell 2205, and passes through the pull-down bit line 2202. This design is repeated for all columns, and finally, the row containing pull-down cell 2205 becomes the row of pull-down cells. As shown in FIG. 22, there are four rows, the two central adjacent rows are used for pull-down cells, and the top and bottom rows are data cells.
[0131] Table 10 shows the operating voltages of the VMM system 2200. The columns in the table show the voltages on the bit line of the selected cell, the bit line pull-down, the word line of the selected cell, the control gate of the selected cell, the word line WLS of the selected pull-down cell, the control gate CGS of the selected pull-down cell, the erase gate of all cells, and the source line of all cells. The rows show the read, erase, and program operations. Note that the voltage biases of CGS and WLS in a read are higher than the voltage biases of the normal WL and CG biases to enhance the driving ability of the pull-down cell. The voltages biased for WLS and CGS can be negative in programming to reduce interference. Table 10: Operation of the VMM array 2200 of FIG. 22 [Table 10]
[0132] Figure 23 shows a VMM system 2300. The VMM system 2300 includes bit lines 2301, 2302, data cells 2303 and 2306, and pull - down cells 2304 and 2305. During the read or verify (for program / erase tuning cycles) operation of cell 2303, current enters the bit - line terminal of cell 2303 through bit line 2301, exits through the source - line terminal of cell 2303, then enters the source - line terminal of pull - down cell 2304 and passes through bit line 2302 (which acts as a pull - down bit line in this case). This design is repeated for all columns, and finally, the row containing pull - down cell 2304 in the first mode becomes the row of pull - down cells. During the read or verify (for program / erase tuning cycles) operation of data cell 2306, current enters the bit - line terminal of cell 2306 through bit line 2301, exits through the source - line terminal of cell 2306, then enters the source - line terminal of pull - down cell 2305 and passes through bit line 2302 (which acts as a pull - down bit line in this case). This design is repeated for all columns, and finally, the row containing pull - down cell 2305 in the second mode becomes the row of pull - down cells. As shown in Figure 23, there are four rows, and alternative odd (or even) rows are used for pull - down cells, and alternative even (or odd) rows are data cells.
[0133] In particular, during the second mode, cells 2305 and 2306 are active in read or verify, cells 2303 and 2305 are used in the pull - down process, and the roles of bit lines 2301 and 2302 are reversed.
[0134] Table 11 shows the operating voltages of the VMM system 2300. The columns in the table show the bit line of the selected data cell, the bit line of the selected pull - down cell, the word line of the selected cell, the control gate of the selected data cell, the word line WLS of the selected pull - down cell, the control gate CGS of the selected pull - down cell, the erase gate of all cells, and the source line of all cells. The rows show the read, erase, and program operations. Table 11: Operation of the VMM System 2300 in Figure 23
Table 11
[0135] Figure 24 shows the VMM system 2400. The VMM system 2400 includes bit lines 2401, pull-down bit lines 2402, (data) cells 2403, source lines 2411, and pull-down cells 2404, 2405, and 2406. During the read or verification operation of cell 2403, current enters the bit line terminal of cell 2403 through bit line 2401, exits to the source line terminal of cell 2403, then enters source line 2411, and then enters the source line terminals of pull-down cells 2404, 2405, and 2406, from where it flows through pull-down bit line 2402. This design is repeated for all columns, and finally, the rows containing pull-down cells 2404, 2405, and 2406 each become rows of pull-down cells. Thereby, when current is drawn into pull-down bit line 2402 through three cells, the pull-down applied to the source line terminal of cell 2403 is maximized. Note that the source lines of four rows are connected together.
[0136] Table 12 shows the operating voltages of the VMM system 2400. The columns in the table represent the bit line of the selected cell, the bit line pull-down, the word line of the selected cell, the control gate of the selected cell, the erase gate of the selected cell, the word line WLS of the selected pull-down cell, the control gate CGS of the selected pull-down cell, the erase gate of the selected pull-down cell, and the source line of all cells. The rows represent the read, erase, and program operations. Table 12: Operation of the VMM System 2400 in Figure 24
Table 12
[0137] FIG. 25 shows an exemplary layout 2500 of the VMM system 2200 of FIG. 22. The bright squares indicate metal contacts between bit lines such as bit line 2201 and pull-down bit lines such as pull-down bit line 2202.
[0138] FIG. 26 shows an alternative layout 2600 of a VMM system similar to the VMM system 2200 of FIG. 22, but with the difference that the pull-down bit line 2602 is very wide and crosses two columns of pull-down cells. That is, the diffusion region of the pull-down bit line 2602 is wider than the diffusion region of the bit line 2601. Layout 2600 further shows cells 2603 and 2604 (pull-down cells), source line 2610, and bit line 2601. In another embodiment, the diffusion of the two pull-down cells (left and right) can be merged into a larger single diffusion.
[0139] FIG. 27A shows a VMM system 2700. To implement the negative and positive weights of a neural network, half of the bit lines are designated as w+ lines (bit lines connected to memory cells implementing positive weights), and the other half of the bit lines are designated as w- lines (bit lines connected to memory cells implementing negative weights), and they are interspersed alternately between the w+ lines. Negative operations are performed on the outputs (neuron outputs) of the w- bit lines by summing circuits such as summing circuits 2701 and 2702. The outputs of the w+ lines and the outputs of the w- lines are combined together to effectively give w = w+ - w- for each pair of (w+, w-) cells for all pairs of (w+, w-) lines. The dummy bit lines or source line pull-down bit lines used to avoid FG-FG coupling of the source lines during readout and / or to reduce IR voltage drop are not shown in the figure. Inputs to the system 2700 (such as CG or WL) can have positive or negative values. When the input has a negative value, since the actual input to the array is still positive (such as the voltage level of CG or WL), the array output (bit line output) is invalidated before output to implement the equivalent function of the negative value input.
[0140] Alternatively, referring to FIG. 27B, positive weights may be implemented in the first array 2711, negative weights may be implemented in a second array 2712 separate from the first array, and the resulting weights are appropriately combined by the adder circuit 2713. Similarly, dummy bit lines (not shown) or source line pull-down bit lines (not shown) are used to avoid source line FG-FG coupling during readout and / or to reduce IR voltage drop.
[0141] Alternatively, FIG. 27C shows a VMM system 2750 for implementing the negative and positive weights of a neural network having positive or negative inputs. The first array 2751 implements positive value inputs having negative and positive weights, and the second array 2752 implements negative value inputs having negative and positive weights. Since any input to either array has only positive values (such as the analog voltage levels of CG or WL), the output of the second array is invalidated before being added to the output of the first array by the adder 2755.
[0142] Table 10A shows an exemplary layout of the physical array arrangement of the (w+, w-) pairs of bit lines BL0 / 1 and BL2 / 3, with four rows coupled to the source line pull-down bit line BLPWDN. The (BL0, BL1) pair of bit lines is used to implement the (w+, w-) lines. Between the (w+, w-) lines is a source line pull-down bit line (BLPWDN). This is used to prevent coupling (e.g., FG-FG coupling) of current from adjacent (w+, w-) lines to the (w+, w-) lines. Basically, the source line pull-down bit line (BLPWDN) functions as a physical barrier between the (w+, w-) line pairs.
[0143] Additional details regarding the FG-FG binding phenomenon and the mechanism for counteracting that phenomenon can be found in U.S. Patent Provisional Application No. 62 / 981,757, entitled "Ultra-Precise Tuning of Analog Neural Memory Cells in a Deep Learning Artificial Neural Network," filed on February 26, 2020, by the same assignee and incorporated herein by reference.
[0144] Table 10B shows different exemplary combinations of weights. "1" means that the cell is used and has an actual output value, and "0" means that the cell is not used and has no value or does not have a large output value.
[0145] In another embodiment, dummy bitlines may be used instead of source line pull-down bitlines.
[0146] In another embodiment, the dummy row can also be used as a physical barrier to avoid coupling between rows. Table 10A: Exemplary Layout [Table 10A] Table 10B: Exemplary Combinations of Weights [Table 10B]
[0147] Table 11A shows another array embodiment of the physical arrangement of the redundant lines BL01, BL23, and the source line pull-down bitline BLPWDN for the (w+, w-) pair of lines BL0 / 1 and BL2 / 3. BL01 is used to weight the remapping of the pair BL0 / 1, and BL23 is used to weight the remapping of the pair BL2 / 3.
[0148] Table 11B shows the case of distributed weights that do not require remapping. Basically, there is no adjacent "1" between BL1 and BL3, which causes the combination of adjacent bit lines.
[0149] In one embodiment, the weight mapping is such that the total current along the bit lines is substantially constant so as to maintain a substantially constant bit line voltage drop. In another embodiment, the weight mapping is such that the total current along the source lines is substantially constant so as to maintain a substantially constant source line voltage drop.
[0150] Table 11C shows the case of distributed weights that require remapping. Basically, there is an adjacent "1" between BL1 and BL3, which causes the combination of adjacent bit lines. This remapping is shown in Table 11D, and as a result, there will be no "1" value between adjacent bit lines. Further, by remapping the "1" actual value weights between bit lines, that is, redistributing the weights, the total current along the bit lines decreases at this point, and the values within the bit lines (output neurons) become more precise. In this case, additional columns (bit lines) are required to act as redundant columns (BL01, BL23). Tables 11E and 11F show another embodiment of remapping noisy cells (or defective cells) to redundant (spare) columns such as BL01 and BL23 in Table 10E or BL0B and BL1B in Table 11F. The adder is used to sum the properly mapped bit line outputs. Table 11A: Exemplary Layout
Table 11A
Table 11B
Table 11C
[0151] Table 11G shows an embodiment of the physical layout of an array suitable for FIG. 27B. Since each array has either a positive or negative weight, a dummy bit line acting as a source line pull-down and a physical barrier for avoiding FG-FG coupling are required for each bit line. Table 11G: Exemplary Layout [Table 11G]
[0152] Another embodiment has a tuning bit line as a bit line adjacent to the target bit line to tune the target bit line to the final target by FG-FG coupling. In this case, the source line pull-down bit line (BLPWDN) is inserted on one side of the target bit line that does not border the tuning bit line.
[0153] An alternative embodiment for mapping noisy or defective cells is to designate these cells as unused cells (after they are identified as noisy or defective by a detection circuit), which means they are (deeply) programmed so that they do not contribute any value to the neuron output.
[0154] Embodiments for processing high-speed cells first identify these cells and then apply more precise algorithms to these cells, such as reducing the voltage increment pulse, or having no voltage increment pulse, or using a floating gate coupling algorithm.
[0155] FIG. 28 shows an optional redundant array 2801 that can be included in any of the VMM arrays considered so far. The redundant array 2801 can be used as redundancy to replace a defective column if any column attached to the bit line switch is considered defective. The redundant array can have its own redundant neuron output (e.g., bit line) and an ADC circuit for redundancy purposes. If redundancy is required, the output of the redundant ADC replaces the output of the ADC of the defective bit line. The redundant array 2801 can also be used for weight mapping as described in Table 10x for power distribution between bit lines.
[0156] FIG. 29 shows a VMM system 2900 comprising an array 2901, an array 2902, a column multiplexer 2903, local bit lines LBL2905a - d, global bit lines GBL2908 and 2909, and dummy bit line switches 2905. The column multiplexer 2903 is used to select the top local bit line 2905 of the array 2901 or the bottom local bit line 2905 of the array 2902 to the global bit line 2908. In one embodiment, the (metal) global bit line 2908 has the same number of lines as the number of local bit lines, e.g., 8 or 16. In another embodiment, the global bit line 2908 has only one (metal) line per N local bit lines, such as one global bit line per 8 or 16 local bit lines. The column multiplexer 2903 further includes multiplexing an adjacent global bit line (such as GBL2909) to the current global bit line (such as GBL2908) to effectively increase the width of the current global bit line. This reduces the voltage drop across the global bit line.
[0157] Figure 30 shows a VMM system 3000. The VMM system 3000 includes an array 3010, a (shift register) SR3001, a digital-to-analog converter 3002 (which receives an input from SR3001 and outputs an equivalent (analog or pseudo-analog) level or information), an adder circuit 3003, an analog-to-digital converter 3004, and a bit line switch 3005. There are dummy bit lines and dummy bit line switches, but they are not shown. As shown, the ADC circuits can be combined together to create a single ADC with higher accuracy (i.e., a larger number of bits).
[0158] The adder circuit 3003 may include the circuits shown in FIGS. 31-33. This may include, without limitation, circuits for normalization, scaling, arithmetic operations, activation, statistical rounding, etc.
[0159] Figure 31 shows a variable resistor-adjustable current-voltage adder circuit 3100 including current sources 3101-1,..., 3101-n that respectively draw currents Ineu(1),..., Ineu(n) (which are currents received from the bit line(s) of the VMM array), an operational amplifier 3102, a variable holding capacitor 3104, and a variable resistor 3103. The operational amplifier 3102 outputs a voltage Vneuro = R3103 * (Ineu1 + Ineu0), which is proportional to the current Ineux. The holding capacitor 3104 is used to hold the output voltage when the switch 3106 is open. This held output voltage is used, for example, to be converted to digital bits by an ADC circuit.
[0160] Figure 32 shows a variable capacitor (basically an integrator)-adjustable current-voltage adder circuit 3200 including current sources 3201-1,..., 3201-n that respectively draw currents Ineu(1),..., Ineu(n) (which are currents received from the bit line(s) of the VMM array), an operational amplifier 3202, a variable capacitor 3203, and a switch 3204. The operational amplifier 3202 outputs a voltage Vneuout = Ineu *Output the integration time / C3203, which is proportional to the current Ineu (plural available).
[0161] FIG. 33A shows a voltage adder 3300 adjustable by a variable capacitor (i.e., a switched capacitor SC circuit), which includes switches 3301 and 3302, variable capacitors 3303 and 3304, an operational amplifier 3305, a variable capacitor 3306, and a switch 3306. When switch 3301 is closed, input Vin0 is provided to operational amplifier 3305. When switch 3302 is closed, input Vin1 is provided to operational amplifier 3305. Optionally, switches 3301 and 3302 are not closed simultaneously. Operational amplifier 3305 generates an output Vout that is an amplified version of the input (either Vin0 and / or Vin1 depending on which switch among switches 3301 and 3302 is closed). That is, Vout = Cin / Cout * (Vin), where Cin is C3303 or C3304 and Cout is C3306. For example, Vout = Cin / Cout * Σ(Vinx), Cin = C3303 = C3304. In one embodiment, Vin0 is the W+ voltage, Vin1 is the W- voltage, and voltage adder 3300 sums them to generate an output voltage Vout.
[0162] FIG. 33B shows a voltage adder 3350 that includes switches 3351, 3352, 3353, and 3354, a variable input capacitor 3358, an operational amplifier 3355, a variable feedback capacitor 3356, and a switch 3357. In one embodiment, Vin0 is the W+ voltage, Vin1 is the W- voltage, and voltage adder 3300 sums them to generate an output voltage Vout.
[0163] When the input = Vin0: When switches 3354 and 3351 are closed, the input Vin0 is provided to the upper terminal of capacitor 3358. Then, switch 3351 is opened and switch 3353 is closed to transfer charge from capacitor 3358 to feedback capacitor 3356. Basically, thereafter, the output VOUT = (C3358 / C3356) * Vin0 (for example, when VREF = 0).
[0164] When the input = Vin1: When switches 3353 and 3354 are closed, both terminals of capacitor 3358 are discharged to VREF. Then, switch 3354 is opened and switch 3352 is closed to charge the bottom terminal of capacitor 3358 to Vin1, and then charge feedback capacitor 3356 to VOUT = -(C3358 / C3356) * Vin1 (when VREF = 0).
[0165] Therefore, for example, when VREF = 0, if the Vin1 input is enabled after the Vin0 input is enabled, VOUT = (C3358 / C3356) * (Vin0 - Vin1). This is used, for example, to implement w = w+ - w-.
[0166] The method of input / output operation to FIG. 2 applied to the VMM array described above may be in digital format or analog format. The method includes the following: · Sequential input IN[0:q] to the DAC: · Operate IN0, then IN1,..., then INq sequentially. All input bits have the same VCGin. All bit line (neuron) outputs are summed after adjusting by binary index multipliers. This is done either before or after the ADC. · Adjustment of the Neuron (Bit Line) Binary Index Multiplication Method: As shown in FIG. 20, an exemplary adder has two bit lines BL0 and Bln. The weights are distributed across a plurality of bit lines BL0 to BLn. For example, there are four bit lines BL0, BL1, BL2, BL3. The output from bit line BL0 is multiplied by 2^0 = 1. The output from bit line BLn, which represents the nth binary bit position, is multiplied by 2^n. For example, when n = 3, 2^3 = 8. Then, the outputs from all bit lines after being appropriately multiplied by the binary bit position 2^n are added together. This is then digitized by an ADC. This method means that all cells have only a binary range, and the multi-level range (n bits) is achieved by the peripheral circuit (which means by the adder circuit). Therefore, the voltage drop across all bit lines is approximately the same for the highest bias level of the memory cell. · Operate IN0, IN1,..., and then INq sequentially. Each input bit has a corresponding analog value VCGin. All neuron outputs are added together for all input bit evaluations. This is done either before or after the ADC. · Parallel Input to the DAC · Each input IN[0:q] has a corresponding analog value VCGin. All neuron outputs are added together with the binary index multiplication method adjusted. This is done either before or after the ADC.
[0167] By operating sequentially in the array, the power is more evenly distributed. The Neuron (Bit Line) Binary Index Method also reduces the power in the array because each cell in the bit line has only a binary level, and the 2^n levels are achieved by the adder circuit 2603.
[0168] Each ADC shown in FIG. 33 can be configured to be combined with the next ADC for a higher bit implementation using an appropriate design of the ADC.
[0169] Figures 34A, 34B, and 34C show output circuits that can be used in the adder circuit 3003 and the analog-to-digital converter 3004 of FIG. 30.
[0170] FIG. 34A shows an output circuit 3400 that receives a neuron output 3401 and outputs an output digital bit 3403, including an analog-to-digital converter 3402.
[0171] FIG. 34B shows an output circuit 3410 that includes a neuron output circuit 3411 and an analog-to-digital converter 3412, which together receive the neuron output 3401 and generate an output 3413.
[0172] FIG. 34C shows an output circuit 3420 that includes a neuron output circuit 3421 and a converter 3422, which together receive the neuron output 3401 and generate an output 3423.
[0173] The neuron output circuit 3411 or 3411 can perform, for example, addition, scaling, normalization, arithmetic operations, etc. The converter 3422 can perform, for example, ADC, PDC, AAC, APC operations, etc.
[0174] FIG. 35 shows a neuron output circuit 3500 that includes an adjustable (scaling) current source 3501 and an adjustable (scaling) current source 3502, which together generate an output i that is a neuron output. OUT This circuit can perform the addition of positive and negative weights, i.e., w = w+ - w-, and can simultaneously perform up or down scaling of the output neuron current.
[0175] FIG. 36 shows a configurable neuron serial analog-to-digital converter 3600. The converter includes an integrator 3670 that integrates the neuron output current into an integration capacitor 3602. In one embodiment, the digital output (count output) 3621 is generated by clocking a ramp VRAMP 3650 until the comparator 3604 switches polarity. In another embodiment, it is generated by ramping down the node VC 3610 with a ramp current 3651 until VOUT 3603 reaches VREF 3650 and at that point the EC 3605 signal disables the counter 3620. The (n-bit) ADC can be configured to have a lower bit accuracy <n bits or a higher bit accuracy >n bits depending on the target application. The configurability is done, for example, by configuring the capacitor 3602, the current 3651, or the ramping rate of VRAMP 3650, the clocking 3641, etc. In another embodiment, the ADC circuit of one VMM array is configured to have a lower accuracy <n bits, and the ADC circuit of another VMM array is configured to have a higher accuracy >n bits. Further, the ADC circuit of one neuron circuit can be configured to generate a higher n-bit ADC accuracy, for example, by combining the integration capacitors 3602 of two ADC circuits in combination with the next ADC of the next neuron circuit.
[0176] FIG. 37 shows a configurable neuron SAR (successive approximation register) analog-to-digital converter 3700. This circuit is a successive approximation converter based on charge redistribution using binary capacitors. It includes a binary CDAC (capacitor-based DAC) 3701, an operational amplifier / comparator 3702, and SAR logic 3703. As shown, GndV 3704 is a low voltage reference level, for example, ground level.
[0177] FIG. 38 shows a configurable neuron combo SAR analog-to-digital converter 3800. This circuit combines two ADCs from two neuron circuits into one to achieve higher n-bit accuracy. For example, in the case of a 4-bit ADC of one neuron circuit, this circuit can achieve an accuracy greater than 4 bits, such as 8-bit ADC accuracy, by combining two 4-bit ADCs. The combo circuit topology is equivalent to a split-cap (bridge capacitor (cap) or attenuation cap) SAR ADC circuit. For example, an 8-bit 4C-4C SAR ADC is brought about by combining two adjacent 4-bit 4C SAR ADC circuits. A bridge circuit 3804 is required to accomplish this, and the capacitance of this circuit = (total number of CDAC cap units / total number of CDAC cap units - 1).
[0178] FIG. 39 shows a configurable neuron pipelined SAR CDAC ADC circuit 3900 that can be used to increase the number of bits in a pipelined manner in combination with the following SAR ADC. The residual voltage 3906 is generated by capacitor 3930Cf to be provided as an input to the next stage of the pipelined ADC (for example, to provide a gain of 2 (Cf to C ratio of all caps of DAC 3901) as an input to the next SAR CDAC ADC).
[0179] Additional implementation details regarding circuits of configurable output neurons (such as configurable neuron ADCs) can be found in U.S. Patent Application No. 16 / 449,201, filed on June 21, 2019, by the same assignee and entitled "Configurable Input Blocks and Output Blocks and Physical Layout for Analog Neural Memory in a Deep Learning Artificial Neural Network", which is incorporated herein by reference.
[0180] The applicant has previously invented a mechanism for achieving precise data tuning in analog neural memory within an artificial neural network described in U.S. Patent Application No. 16 / 985,147, entitled "Ultra-Precise Tuning of Analog Neural Memory Cells in a Deep Learning Artificial Neural Network," filed on August 4, 2020, which is incorporated herein by reference. The prior application discloses embodiments for performing coarse programming, fine programming, and ultra-fine programming of selected cells within the VMM. Thus, the prior application contemplates performing up to three types of programming for each selected cell. This approach can achieve very precise programming, but it is also quite time-consuming because each selected cell within the array must undergo all three types of programming processes.
[0181] Figures 40-48 illustrate embodiments that improve upon the programming mechanisms of the prior art and prior applications.
[0182] Figure 40 shows a logic cell 4000. In this example, the logic cell includes three memory cells labeled as fine cell 4001-1, coarse cell 4001-2, and coarse cell 4001-3. The fine cell 4001-1 is coupled to a fine bit line 4002-1, the coarse cell 4001-2 is coupled to a coarse bit line 4002-2, and the coarse cell 4001-3 is coupled to a coarse bit line 4002-3. The logic cell 4000 includes data that in the prior art was stored in a (physical) single cell. For example, the logic cell 4000 can hold one of N different values, where N is the total number of different values that can be stored in the logic cell 4000 (e.g., N = 64 or 128). Unlike the prior art, rather than performing multiple types of programming (e.g., coarse programming and fine programming) on each physical cell, the coarse programming method is performed on the coarse cell 4001-3 with the coarse bit line 4002-3, the coarse programming method is performed on the coarse cell 4001-2 with the coarse bit line 4002-2, and the coarse and / or fine programming methods are performed on the fine cell 4001-1 with the fine bit line 4002-1. That is, the fine programming method is not performed on the coarse cells 4001-2 and 4001-3. By implementing this approach, the programming time is much shorter compared to an approach that performs coarse programming and fine programming on all three cells, especially since fine programming requires a relatively large amount of time.
[0183] Figure 41 shows a program and verification method 4100 executed on the logic cell 4000 of FIG. 40.
[0184] The first step is to erase the fine cell 4001-1, the coarse cell 4001-2, and the coarse cell 4001-3 (step 4101). Optionally, the first step further includes, after erasure, performing a coarse programming method to an intermediate value on all three cells.
[0185] The second step is to program the coarse cells 4001-3 using the coarse programming method, and after its operation, to verify the logic cell 4000 to confirm that the coarse cells 4001-3 are correctly programmed to the coarse values intended for the coarse cells 4001-3 (step 4102). An alternative method is to verify the coarse cells themselves.
[0186] The third step is to program the coarse cell 4001-2 using the coarse programming method, and after its operation, to verify the logic cell 4000 to confirm that the coarse cells 4001-2 and 4001-3 together are correctly programmed to the coarse values intended for the coarse cells 4001-2 and 4001-3 that are reflected as the value of the logic cell 4000 (step 4103).
[0187] The fourth step is to program the fine cell 4001-1 using the fine programming method, and after its operation, to verify the logic cell 4000 to confirm that the fine cell 4001-1, the coarse cell 4001-2, and the coarse cell 4003 together are correctly programmed to the value intended for the logic cell 4000 (step 4104).
[0188] Table 12A shows examples of the target values for the logic cell 4000, the fine cell 4001-1, the coarse cell 4001-2, and the coarse cell 4001-3. Table 12A: Exemplary Target Values for Logic Cell 4000
Table 12A
[0189] As can be understood with reference to Table 12A, only the fine cell 4001-1 needs to have a precise and accurate value within the allowable percentage (e.g., without limitation, + / -0.5%, + / -0.25%) of the final target value of the logic cell 4000. For example, the applicant has determined that the coarse cell can have, by way of example, + / -20% of the target value of the coarse cell (since any inaccuracies can be corrected by the fine cell 4001-1), while the fine cell can have + / -0.5% of the target value of the logic cell. Thus, the coarse cells can be programmed with coarser voltage steps and can reach their targets much faster. One way to assign charge levels to each of the N levels of the memory cell is as follows. First, determine the current range of the maximum current Imax in the subthreshold or any other region from the data characteristic evaluation. Typically, the current range Imax is within the range of the floating voltage of about Vtfg - 0.2V. Second, determine the leakage current Ileak from the physical memory cell when the cell is in the off state (e.g., WL = 0V, CG = 0V). The lowest charge for the lowest level of the N levels is some coefficient a * Ileak, for example, a = 128 in the case of a 128-row array. The highest charge for the highest level of the N levels is the charge associated with the maximum current Imax of the current range. Third, use a coarse / fine or coarse / fine / ultrafine algorithm to determine the programming resolution when programming to Imax. Typically, the variation of the standard deviation (sigma) of the Imax target is the single electron programming resolution for the coarse / fine algorithm and the sub-electron programming resolution for the coarse / fine / ultrafine (bit line tuning or floating gate - floating gate coupling tuning method). For example, the target delta level of the neural network is = (IdeltaL = I(Ln) - I(Ln-1) = b * 1 sigma variation, and b can typically be 1 or 2 or 3 depending on the network accuracy desired for a particular application. In an exemplary embodiment, the number of levels NL is = (Imax - a * Ileak) / IdeltaL, and a is a predetermined number based on the data characteristic evaluation.
[0190] Figure 42 shows a logic cell 4200. In this example, the logic cell 4200 includes four memory cells labeled as tuning cell 4201-1, fine cell 4201-2, coarse cell 4201-3, and coarse cell 4201-4. The tuning cell 4201-1 is coupled to a tuning bit line 4202-1, the fine cell 4201-2 is coupled to a fine bit line 4202-2, the coarse cell 4201-3 is coupled to a coarse bit line 4202-3, and the coarse cell 4201-4 is coupled to a coarse bit line 4202-4. The logic cell 4200 includes data that in the prior art was stored in a single physical cell. For example, the logic cell 4200 may hold one of N different values, where N is the total number of different values that can be stored in the logic cell (e.g., N = 64 or 128). Unlike the prior art, rather than performing multiple types of programming (e.g., coarse programming, fine programming, ultra-fine (FG-FG tuning) or tuning programming) on each physical cell, coarse programming is performed on the coarse cell 4201-4 with the coarse bit line 4202-4, coarse programming is performed on the coarse cell 4201-3 with the coarse bit line 4202-3, coarse and / or fine programming is performed on the fine cell 4201-2 with the fine bit line 4202-2, and FG-FG tuning (ultra-fine) programming is performed on the tuning cell 4201-1 with the tuning bit line 4202-1. That is, ultra-fine and fine programming are not performed on the coarse cells 4201-3 and 4201-4, and ultra-fine programming is not performed on the fine cell 4201-2.
[0191] Through ultra-fine programming, the logic cell can reach within the target percentage of the final target value, for example, without limitation, within a target percentage of + / - 0.5% or + / - 0.25%. Ultra-fine programming is performed by programming the tuning cell 4201-1. The tuning cell 4201-1 tunes the fine cell 4201-2 through an FG-FG connection (the FG of the tuning cell 4201-1 is connected to the FG of the fine cell 4201-2). For example, if the percentage of the FG-FG connection from the tuning cell to the fine cell is about 3%, this means that a 4mV change in the FG of the tuning cell (from, for example, a 10mV CG program increment per cell) results in a 0.12mV change in the FG of the fine cell (from the FG-FG connection of two adjacent cells). The respective coarse cell targets of the coarse cells 4201-3 and 4201-5 can be, for example, within + / - 20%, the fine cell target can be within 15%, and the tuning target of the tuning cell can be + / - 0.2%. It should be noted that only one tuning cell is required for three (or more) physical cells to realize one logic cell.
[0192] Figure 43 shows a program and verification method 4300 executed for the logic cell 4200 of Figure 42.
[0193] The first step is to erase the tuning cell 4201-1, the fine cell 4201-2, the coarse cell 4201-3, and the coarse cell 4201-4 (step 4301). Optionally, the first step further includes performing a coarse programming method to an intermediate value for the fine cell 4201-2, the coarse cell 4201-3, and the coarse cell 4201-4.
[0194] The second step is to program the coarse cell 4201-4 using the coarse programming method and, after its operation, to verify the logic cell 4200 to confirm that the coarse cell 4201-4 is correctly programmed to the intended coarse value in the coarse cell 4002-4 (step 4302).
[0195] The third step is to program the coarse cell 4201-3 using a coarse programming method and, after its operation, verify the logic cell 4200 to confirm that the coarse cells 4201-3 and 4201-4 together are correctly programmed to the intended coarse values for the coarse cells 4201-3 and 4201-4 (step 4303).
[0196] The fourth step is to program the fine cell 4201-2 using a fine programming method and, after its operation, verify the logic cell 4200 to confirm that the fine cell 4201-2, the coarse cell 4201-3, and the coarse cell 4201-4 together are correctly programmed to the intended value for the logic cell 4200 (step 4304).
[0197] The fifth step is to program the tuning cell 4201-1 using a tuning method and, after its operation, verify the logic cell 4200 to confirm that the tuning cell 4201-1, the fine cell 4201-2, the coarse cell 4201-3, and the coarse cell 4201-4 together are correctly programmed to the intended value for the logic cell 4200 (step 4305).
[0198] Table 13 shows examples of the target values for the logic cell 4200, the tuning cell 4201-1, the fine cell 4001-2, the coarse cell 4201-3, and the coarse cell 4201-4. Table 13: Exemplary Target Values for Logic Cell 4200
Table 13
[0199] Note that since the value of the logic cell 4200 is the sum of the two coarse cells 4201-3, 4201-4 and the fine cell 4201-2, the absolute value of the tuning cell 4201-1 is not important. The purpose of the tuning cell 4201-1 is to tune the value of the fine cell 4201-2.
[0200] FIG. 44 shows an array 4400. The array 4400 includes a plurality of logic cells such as an exemplary logic cell 4451, which follows the structure of the logic cell 4200 here. Thus, the logic cell 4451 includes a tuning cell 4411-1 coupled to a tuning bit line 4401-1, a fine cell 4411-2 coupled to a fine bit line 4401-2, a coarse cell 4411-3 coupled to a coarse bit line 44010-3, and a coarse cell 4411-4 coupled to a coarse bit line 4401-4. Here, each row includes a plurality of logic cells having the same structure as the logic cell 4200, and the array 4400 includes a plurality of rows such as exemplary rows 4410 and 4420. In the same row, the next logic cell has its tuning cell next to the coarse cell of the previous logic cell, which is used to minimize the FG-FG coupling between two logic cells, noting that the capacitive effect is relatively small in the coarse cell 4411-4 compared to, for example, the fine cell 4411-6.
[0201] FIG. 45 shows an array 4500. The array 4500 is similar to the array 4400, except that the array 4500 also includes columns of isolation cells coupled to isolation bit lines. For example, the array 4500 includes a logic cell 4551 that includes a tuning cell 4511-2 coupled to a tuning bit line 4501-2, a fine cell 4511-3 coupled to a fine bit line 4501-3, a coarse cell 4511-4 coupled to a coarse bit line 4501-4, and a coarse cell 4511-5 coupled to a coarse bit line 4501-5. The array 4500 further includes an isolation cell 4511-1 coupled to an isolation bit line 4501-1 and an isolation cell 4511-6 coupled to an isolation bit line 4501-6, and the isolation cells 4511-1 and 4511-6 are adjacent to the logic cells 4551 on both sides thereof. The isolation cells 4511-1 and 4511-6 are not used for storing data, but rather are used to provide a buffer between the logic cells to reduce an undesirable interference effect between the logic cells. Preferably, the isolation cells are deeply programmed such that the FG voltage of the isolation cells is at as low a value as possible. Alternatively, the isolation cells are partially erased, fully erased, or in a native state (no erase or program). Alternatively, the isolation cells are partially programmed. Alternatively, the isolation cells are dummy cells. It should be noted that for the same row, the isolation cells between the logic cells are adjacent to one coarse cell of the previous logic cell and adjacent to the tuning cell of the subsequent logic cell.
[0202] FIG. 46 shows an array 4600. The array 4600 is similar to the array 4400, except that the array 4600 also includes columns of strap cells coupled to any bit lines. For example, the array 4600 includes a tuning cell 4611-2 coupled to a tuning bit line 4601-2, a fine cell 4611-3 coupled to a fine bit line 4601-3, a coarse cell 4611-4 coupled to a coarse bit line 4601-4, and a coarse cell 4611-5 coupled to a coarse bit line 4601-5, in a logic cell 4651. The array 4600 further includes strap cells 4611-1 and 4611-6 positioned adjacent to the logic cells 4651 on both sides thereof. The strap cells 4611-1 and 4611-6 are not used for storing data; rather, they are used as areas where conductive connections (such as metal interconnects) can be made between various lines (poly lines) within the array 4600 (WL straps for word lines, EG straps for erase gates, CG straps for control gates, SL straps for source lines, or combinations of straps such as SLWL straps, SLCG straps, SLEG straps; the strap cells may still have dummy floating gate structures in their construction) and devices and connections outside the array 4600 (such as driver circuits). Alternatively, isolation cells are arranged between the strap cells and the tuning cells. Alternatively, the isolation cells are arranged next to the strap cells.
[0203] FIG. 47 shows an array 4700. The array 4700 is similar to the array 4400, except that the array 4700 also includes a column of pull-down cells where the array 4700 is coupled to the pull-down source line of the array. For example, the array 4700 includes logic cells 4751 including tuning cells 4711-2 coupled to tuning bit lines 4701-2, fine cells 4711-3 coupled to fine bit lines 4701-3, coarse cells 4711-4 coupled to coarse bit lines 4701-4, and coarse cells 4711-5 coupled to coarse bit lines 4701-5. The array 4700 further includes pull-down cells 4711-1 coupled to pull-down bit lines 4701-1 and pull-down cells 4711-6 coupled to pull-down bit lines 4701-6, and the pull-down cells 4711-1 and 4711-6 are adjacent to the logic cells 4751 on both sides thereof. The pull-down cells 4711-1 and 4711-6 are not used for storing data, but rather are used to pull down the source line terminal to ground as needed, as described above with reference to FIGS. 21-27.
[0204] FIG. 48 shows an array 4800. The array 4800 includes logic cells 4851 including tuning cells 4811-1 and 4821-1 coupled to tuning bit lines 4801-1, coarse cells 4811-2 and fine cells 4821-2 coupled to mixed bit lines 4801-2, and coarse cells 4811-3 and fine cells 4821-3 coupled to coarse bit lines 4810-3. Thus, in the array 4800, each logic cell includes three cells in one row (such as an even row) and three cells in an adjacent row (such as an odd row). Three of the cells are coarse cells, one cell is a fine cell, and two cells are tuning cells. Consistent with the programming method described above, when the logic cell 4851 is programmed, the order in which the cells are programmed is: coarse cell 4811-3, coarse cell 4821-3, coarse cell 4811-2, fine cell 4821-2, tuning cell 4811-1 (expected to have a minimal effect on coarse cell 4811-2), and tuning cell 4821-1. During a read or verification operation, all six cells are read as one logic cell. Through this approach, the mismatch between odd and even rows is averaged together to minimize the I-V gradient mismatch.
[0205] As used herein, it should be noted that both the terms "over" and "on" include both "directly over" (with no intervening material, element, or gap therebetween) and "indirectly over" (with an intervening material, element, or gap therebetween). Similarly, the term "adjacent" includes "directly adjacent" (with no intervening material, element, or gap therebetween) and "indirectly adjacent" (with an intervening material, element, or gap therebetween), "attached to" includes "directly attached to" (with no intervening material, element, or gap therebetween) and "indirectly attached to" (with an intervening material, element, or gap therebetween), and "electrically coupled" includes "directly electrically coupled" (with no intervening material or element electrically connecting the elements together therebetween) and "indirectly electrically coupled" (with an intervening material or element electrically connecting the elements together therebetween). For example, forming an element "over a substrate" can include forming the element directly on the substrate without an intervening material / element therebetween and forming the element indirectly over the substrate with one or more intervening materials / elements therebetween.
Claims
1. 1. A method of level assignment for an array of non-volatile memory cells, the method comprising: determining a program resolution current value; and setting levels for a programming operation of a plurality of non-volatile memory cells in the array such that a delta current between levels of each pair of adjacent cells among the plurality of non-volatile memory cells is a multiple of the program resolution current value.
2. The method of claim 1 , wherein the array is an analog memory.
3. The method of claim 1 , wherein the lowest level among said levels is a multiple of a leakage value of an OFF cell.
4. 2. The method of claim 1, wherein the delta current is at least one sigma relative to the program resolution current value.
5. 2. The method of claim 1, wherein the delta current is at least two sigma relative to the program resolution current value.
6. The method of claim 1 , wherein the delta current is at least 3 sigma with respect to the program resolution current value.
7. 10. The method of claim 1, wherein the non-volatile memory cells are organized into logic cells, each logic cell including one or more coarse cells and one or more fine cells.
8. The method of claim 1 , wherein the non-volatile memory cells are split-gate flash memory cells.
9. The method of claim 1 , wherein the non-volatile memory cells are stacked gate flash memory cells.
10. The method of claim 1 , wherein the array of non-volatile memory cells is part of a neural network.
11. The method of claim 10 , wherein the neural network is an analog neural network.
Citation Information
Patent Citations
Deep Learning Neural Network Classifier Using Non-Volatile Memory Arrays
JP2019517138A
Memory device and method for varying program state separation based on frequency of use
JP2022514111A