Precision programming method and apparatus for analog neural memories in artificial neural networks - Patents.com

Through precise programming algorithms and equipment, the control gate voltage is gradually adjusted to control floating gate charge, which solves the problem of insufficient charge accuracy of memory cells in the prior art, and realizes high-precision storage of weights in simulated neural networks.

JP7676489B2Active Publication Date: 2025-05-14SILICON STORAGE TECHNOLOGY INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2023136978
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-01-23
Filing Date
2023-08-25
Publication Date
2025-05-14
Estimated Expiration
2040-05-22

AI Technical Summary

Technical Problem

The prior art is difficult to accurately program the floating gate charge of nonvolatile memory cells, resulting in insufficient accuracy when storing weights in simulated neural networks.

Method used

Using precision programming algorithms and equipment, precise control of floating gate charges is ensured that each memory cell can store specific charge values ​​by gradually adjusting the control gate voltage.

Benefits of technology

High-precision programming of nonvolatile memory units is realized to ensure accurate storage and calculation of weights in simulated neural networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007676489000021
    Figure 0007676489000021
  • Figure 0007676489000022
    Figure 0007676489000022
  • Figure 0007676489000023
    Figure 0007676489000023
Patent Text Reader

Abstract

To provide a precise programming algorithm and device for precisely and rapidly depositing precise amounts of charge on a floating gate of a non-volatile memory cell in a vector-by-matrix multiplication (VMM) array in an artificial neural network.SOLUTION: A method of programming a non-volatile memory cell includes: programming all cells to a "0" state; erasing all cells to an intermediate weak erase level; executing a hard program for adding electrons to a floating gate of the cell to a very deep programmed state on all non-selected cells; executing a coarse programming method on the selected cell; executing the precision programming method on the selected cell; and programming each selected cell with the exact value desired.SELECTED DRAWING: Figure 33A
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] (Priority Claim) This application claims priority to U.S. Provisional Patent Application No. 62 / 933,809, filed November 11, 2019, entitled "PRECISE PROGRAMMING METHOD AND APPARATUS FOR ANALOG NEURAL MEMORY IN A DEEP LEARNING ARTIFICIAL NEURAL NETWORK," and to U.S. Patent Application No. 16 / 751,202, filed January 23, 2020, entitled "PRECISE PROGRAMMING METHOD AND APPARATUS FOR ANALOG NEURAL MEMORY IN A DEEP LEARNING ARTIFICIAL NEURAL NETWORK."

[0002] FIELD OF THEINVENTION Numerous embodiments are disclosed for a precision programming algorithm and apparatus for precisely and rapidly depositing precise amounts of charge onto the floating gates of non-volatile memory cells in vector matrix multiplication (VMM) arrays in artificial neural networks. [Background technology]

[0003] Artificial neural networks mimic biological neural networks (the central nervous systems of animals, particularly the brain), and are used to estimate or approximate functions that may depend on multiple inputs and are generally unknown. Artificial neural networks generally contain layers of interconnected "neurons" that exchange messages.

[0004] FIG. 1 shows an artificial neural network, where the circles represent layers of inputs or neurons. The connections (called synapses) are represented by arrows and have numerical weights that can be adjusted based on experience. This allows the artificial neural network to adapt to the inputs and learn. Typically, an artificial neural network contains multiple layers of inputs. There is typically one or more intermediate layers of neurons, and an output layer of neurons that provides the output of the neural network. At each level, the neurons make decisions, either individually or together, based on the data they receive from the synapses.

[0005] One of the main challenges in the development of artificial neural networks for high-performance information processing is the lack of suitable hardware technology. In practice, practical artificial neural networks rely on a very large number of synapses, which allows high connectivity between neurons, i.e., a very high degree of parallelization of computation. In principle, such complexity can be achieved by digital supercomputers or dedicated graphic processing unit clusters. However, in addition to high costs, these approaches also suffer from poor energy efficiency compared to biological networks, which mainly perform low-precision analog computations and therefore consume much less energy. CMOS analog circuits have been used for artificial neural networks, but most CMOS-implemented synapses are too bulky given the large number of neurons and synapses.

[0006] Applicant previously disclosed an artificial (analog) neural network utilizing one or more non-volatile memory arrays as synapses in U.S. Patent Application Serial No. 15 / 594,439, published as U.S. Patent Publication 2017 / 0337466, which is incorporated by reference. The non-volatile memory array operates as an analog neuromorphic memory. As used herein, the term neuromorphic refers to a circuit that implements a model of a neural system. The analog neuromorphic memory includes a first plurality of synapses configured to receive a first plurality of inputs and generate a first plurality of outputs therefrom, and a first plurality of neurons configured to receive the first plurality of outputs. The first plurality of synapses includes a plurality of memory cells, each memory cell including spaced apart source and drain regions formed in a semiconductor substrate with a channel region extending therebetween, a floating gate disposed above and insulated from a first portion of the channel region, and a non-floating gate disposed above and insulated from a second portion of the channel region. Each of the plurality of memory cells is configured to store a weight value corresponding to a number of electrons on the floating gate. The plurality of memory cells are configured to multiply a first plurality of inputs by the stored weight values ​​to generate a first plurality of outputs. An array of memory cells arranged in this manner may be referred to as a vector matrix multiplication (VMM) array.

[0007] Each non-volatile memory cell used in an analog neuromorphic memory array must store a very specific and precise amount of charge, i.e., number of electrons, in its floating gate in response to erasure and programming. For example, each floating gate must store one of N different values, where N is the number of different weights that can be exhibited by each cell. Examples of N include 16, 32, 64, 128, and 256. One challenge in analog neuromorphic memory systems is the ability to program selected cells to the different values ​​of N with the precision and granularity required.

[0008] What is needed is an improved programming system and method suitable for use with VMM arrays in analog neuromorphic memories. Summary of the Invention

[0009] Numerous embodiments are disclosed of precision programming algorithms and apparatus for precisely and rapidly depositing precise amounts of charge onto the floating gates of non-volatile memory cells in vector matrix multiplication (VMM) arrays in analog neuromorphic memories, such that selected cells can be programmed with extremely high precision to hold one of N distinct values.

[0010]

[0011]

[0012]

[0013]

[0014]

[0015]

[0016]

[0017]

[0018]

[0019]

[0020]

[0021]

[0022]

[0023]

[0024]

[0025]

[0026]

[0027]

[0028]

[0029]

[0030]

[0031]

[0032]

[0033]

[0034]

[0035]

[0036]

[0037]

[0038]

[0039]

[0040]

[0041]

[0042]

[0043]

[0044]

[0045]

[0046]

[0047]

[0048]

[0049]

[0050]

[0051]

[0052]

[0053]

[0054]

[0055]

[0056]

[0057]

[0058]

[0059]

[0060]

[0061]

[0062]

[0063]

[0064]

[0065]

[0066]

[0067]

[0068]

[0069]

[0070] [Brief description of the drawings]

[0071] [Figure 1] FIG. 1 illustrates a prior art artificial neural network. [Diagram 2] 1 shows a prior art split-gate flash memory cell. [Diagram 3] 1 illustrates another prior art split-gate flash memory cell. [Figure 4] 1 illustrates another prior art split-gate flash memory cell. [Diagram 5]1 illustrates another prior art split-gate flash memory cell. [Figure 6] 1 illustrates another prior art split-gate flash memory cell. [Figure 7] 1 shows a prior art stacked gate flash memory cell. [Figure 8] FIG. 1 illustrates various levels of an exemplary artificial neural network that utilizes one or more non-volatile memory arrays. [Figure 9] FIG. 1 is a block diagram illustrating a vector matrix multiplication system. [Figure 10] FIG. 1 is a block diagram illustrating an example artificial neural network utilizing one or more vector-matrix multiplication systems. [Figure 11] 4 illustrates another embodiment of a vector matrix multiplication system. [Figure 12] 4 illustrates another embodiment of a vector matrix multiplication system. [Figure 13] 4 illustrates another embodiment of a vector matrix multiplication system. [Figure 14] 4 illustrates another embodiment of a vector matrix multiplication system. [Figure 15] 4 illustrates another embodiment of a vector matrix multiplication system. [Figure 16] 4 illustrates another embodiment of a vector matrix multiplication system. [Figure 17] 4 illustrates another embodiment of a vector matrix multiplication system. [Figure 18] 4 illustrates another embodiment of a vector matrix multiplication system. [Figure 19] 4 illustrates another embodiment of a vector matrix multiplication system. [Figure 20] 4 illustrates another embodiment of a vector matrix multiplication system. [Figure 21] 4 illustrates another embodiment of a vector matrix multiplication system. [Figure 22] 4 illustrates another embodiment of a vector matrix multiplication system. [Figure 23] 4 illustrates another embodiment of a vector matrix multiplication system. [Figure 24]4 illustrates another embodiment of a vector matrix multiplication system. [Diagram 25] 1 shows a prior art long-term and short-term memory system. [Figure 26] An exemplary cell for use in a long- and short-term memory system is shown. [Figure 27] 27 illustrates an embodiment of the exemplary cell of FIG. 26. [Figure 28] 27 illustrates another embodiment of the exemplary cell of FIG. 26. [Figure 29] 1 shows a prior art gated recurrent unit system. [Diagram 30] 1 shows an exemplary cell for use in a gated recurrent unit system. [Diagram 31] 31 illustrates an embodiment of the exemplary cell of FIG. 30. [Diagram 32] 31 illustrates another embodiment of the exemplary cell of FIG. 30. [Figure 33A] 1 illustrates one embodiment of a method for programming a non-volatile memory cell. [Figure 33B] 1 illustrates another embodiment of a method for programming a non-volatile memory cell. [Diagram 34] 1 illustrates one embodiment of a coarse programming method. [Diagram 35] 3 illustrates exemplary pulses used in programming non-volatile memory cells. [Figure 36A] 3 illustrates exemplary pulses used in programming non-volatile memory cells. [Figure 36B] 1 illustrates exemplary complementary increment and decrement pulses used in programming non-volatile memory cells. [Figure 37] 1 illustrates a calibration algorithm for programming non-volatile memory cells that adjusts programming parameters based on the tilt characteristics of the cell. [Figure 38] 38 shows the circuit used in the calibration algorithm of FIG. 37. [Figure 39] 1 illustrates a calibration algorithm for programming non-volatile memory cells. [Diagram 40]1 illustrates a calibration algorithm for programming non-volatile memory cells. [Diagram 41] 41 shows the circuit used in the calibration algorithm of FIG. 40. [Diagram 42] 4 illustrates an exemplary progression of voltages applied to the control gates of non-volatile memory cells during a programming operation. [Diagram 43] 4 illustrates an exemplary progression of voltages applied to the control gates of non-volatile memory cells during a programming operation. [Diagram 44] 1 illustrates a system for applying programming voltages during programming of non-volatile memory cells in a vector multiplication matrix system. [Diagram 45] 1 shows a vector multiplication matrix system having a modulator, an analog-to-digital converter, and an output block including a summer. [Diagram 46] 1 shows a charge adder circuit. [Figure 47] 2 shows a current adder circuit. [Figure 48] 1 shows a digital adder circuit. [Figure 49A] 1 illustrates one embodiment of an integrated analog-to-digital converter for neuron outputs. [Figure 49B] 49B shows a graph illustrating the voltage output over time of the integrating analog-to-digital converter of FIG. 49A. [Figure 49C] 13 shows another embodiment of an integrating analog-to-digital converter for neuron outputs. [Figure 49D] 49D shows a graph illustrating the voltage output over time of the integrating analog-to-digital converter of FIG. 49C. [Figure 49E] 13 shows another embodiment of an integrating analog-to-digital converter for neuron outputs. [Fig.49F] 13 shows another embodiment of an integrating analog-to-digital converter for neuron outputs. [Figure 50A] 1 shows a successive approximation analog-to-digital converter for neuron outputs. [Figure 50B] 1 shows a successive approximation analog-to-digital converter for neuron outputs. [Figure 51] 1 illustrates an embodiment of a sigma-delta analog-to-digital converter. [Figure 52A] 1 illustrates an embodiment of a ramp-type analog-to-digital converter. [Figure 52B] 1 illustrates an embodiment of a ramp-type analog-to-digital converter. [Figure 52C] 1 illustrates an embodiment of a ramp-type analog-to-digital converter. [Figure 53] 1 illustrates an embodiment of an algorithmic analog-to-digital converter. [Figure 54] 1 shows a logical multi-bit cell. [Figure 55] 55 illustrates a method for programming the logical multi-bit cell of FIG. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0072] The artificial neural network of the present invention utilizes a combination of CMOS technology and non-volatile memory arrays. [Non-volatile memory cell]

[0073] Digital non-volatile memories are well known. For example, U.S. Pat. No. 5,029,130 ​​("the '130 patent"), incorporated herein by reference, discloses an array of split-gate non-volatile memory cells, which are a type of flash memory cell. Such a memory cell 210 is shown in FIG. 2. Each memory cell 210 includes a source region 14 and a drain region 16 formed in a semiconductor substrate 12, with a channel region 18 between the source region 14 and the drain region 16. A floating gate 20 is formed over and insulated from (and controls the conductivity of) a first portion of the channel region 18, and over a portion of the source region 14. A word line terminal 22 (typically coupled to a word line) is disposed over and has a first portion insulated from (and controls the conductivity of) the second portion of the channel region 18, and a second portion extending upwardly over the floating gate 20. A floating gate 20 and a wordline terminal 22 are insulated from the substrate 12 by a gate oxide. A bitline terminal 24 is coupled to the drain region 16.

[0074] The memory cell 210 is erased (electrons are removed from the floating gate) by applying a high positive voltage to the wordline terminal 22, which causes electrons in the floating gate 20 to pass via Fowler-Nordheim tunneling from the floating gate 20 to the wordline terminal 22 through the insulator between them.

[0075] The memory cell 210 is programmed (electrons are applied to the floating gate) by applying a positive voltage to the word line terminal 22 and a positive voltage to the source region 14. An electron current flows from the source region 14 (source line terminal) towards the drain region 16. The electrons accelerate and are heated when they reach the gap between the word line terminal 22 and the floating gate 20. Some of the heated electrons are injected into the floating gate 20 through the gate oxide due to electrostatic attraction from the floating gate 20.

[0076] The memory cell 210 is read by applying a positive read voltage to the drain region 16 and the word line terminal 22 (turning on the portion of the channel region 18 below the word line terminal). When the floating gate 20 is positively charged (i.e., erased with electrons), the portion of the channel region 18 below the floating gate 20 is also turned on and current flows through the channel region 18, which is sensed as the erased or "1" state. When the floating gate 20 is negatively charged (i.e., programmed with electrons), the portion of the channel region below the floating gate 20 is mostly or completely off and no (or very little) current flows through the channel region 18, which is sensed as the programmed or "0" state.

[0077] Table 1 shows typical voltage ranges that may be applied to the terminals of memory cell 110 to perform read, erase, and program operations. Table 1: Operation of the Flash Memory Cell 210 of FIG. 2 [Table 1] "Read 1" is a read mode in which the cell current is output to the bit line. "Read 2" is a read mode in which the cell current is output to the source line terminal.

[0078] Figure 3 shows a memory cell 310 similar to memory cell 210 of Figure 2 with the addition of a control gate (CG) terminal 28. The control gate terminal 28 is biased at a high voltage (e.g., 10V) during programming, a low or negative voltage (e.g., 0v / -8V) during erasure, and a low or medium voltage (e.g., 0v / 2.5V) during reading. The other terminals are biased similarly to those of Figure 2.

[0079] FIG. 4 shows a four-gate memory cell 410 including a source region 14, a drain region 16, a floating gate 20 above a first portion of a channel region 18, a select gate 22 (typically coupled to a word line, WL) above a second portion of the channel region 18, a control gate 28 above the floating gate 20, and an erase gate 30 above the source region 14. This configuration is described in U.S. Pat. No. 6,747,310, which is incorporated herein by reference for all purposes. Here, all gates, except for the floating gate 20, are non-floating gates, i.e., they are electrically connected or connectable to a voltage source. Programming is performed by heated electrons injecting themselves from the channel region 18 into the floating gate 20. Erasing is performed by electrons tunneling from the floating gate 20 to the erase gate 30.

[0080] Table 2 shows typical voltage ranges that may be applied to the terminals of memory cell 410 to perform read, erase, and program operations. Table 2: Operation of Flash Memory Cell 410 of FIG. [Table 2] "Read 1" is a read mode in which the cell current is output to the bit line. "Read 2" is a read mode in which the cell current is output to the source line terminal.

[0081] Figure 5 shows a memory cell 510 similar to memory cell 410 of Figure 4, except that memory cell 510 does not include an erase gate (EG) terminal. Erasing is accomplished by biasing the substrate 18 to a high voltage and the control gate CG terminal 28 to a low or negative voltage. Alternatively, erasure is accomplished by biasing the word line terminal 22 to a positive voltage and the control gate terminal 28 to a negative voltage. Programming and reading are similar to those of Figure 4.

[0082] Figure 6 shows another type of flash memory cell, a three-gate memory cell 610. Memory cell 610 is identical to memory cell 410 of Figure 4, except that memory cell 610 does not have a separate control gate terminal. Erase and read operations (where erasure occurs through the use of the erase gate terminal) are similar to those of Figure 4, except that no control gate bias is applied. Programming operations are also performed without a control gate bias, and as a result, a higher voltage must be applied to the source line terminal during a program operation to compensate for the lack of control gate bias.

[0083] Table 3 shows typical voltage ranges that may be applied to the terminals of memory cell 610 to perform read, erase, and program operations. Table 3: Operation of Flash Memory Cell 610 of FIG. [Table 3] "Read 1" is a read mode in which the cell current is output to the bit line. "Read 2" is a read mode in which the cell current is output to the source line terminal.

[0084] Figure 7 shows another type of flash memory cell, a stacked gate memory cell 710. Memory cell 710 is similar to memory cell 210 of Figure 2, except that the floating gate 20 extends over the entire channel region 18, and a control gate terminal 22 (coupled to a wordline) extends over the floating gate 20, separated by an insulating layer (not shown). Erase, programming, and read operations operate in a similar manner as previously described for memory cell 210.

[0085] Table 4 shows typical voltage ranges that may be applied to the terminals of memory cell 710 and substrate 12 to perform read, erase, and program operations. Table 4: Operation of Flash Memory Cell 710 of FIG. [Table 4]

[0086] "Read 1" is a read mode in which the cell current is output to the bit line. "Read 2" is a read mode in which the cell current is output at the source line terminal. Optionally, in an array including rows and columns of memory cells 210, 310, 410, 510, 610, or 710, a source line may be coupled to one row of memory cells or to two adjacent rows of memory cells. That is, a source line terminal may be shared by adjacent rows of memory cells.

[0087] In order to utilize a memory array containing one of the non-volatile memory cell types in an artificial neural network as described above, two modifications are made: First, as explained further below, the lines are configured so that each memory cell can be individually programmed, erased, and read without adversely affecting the memory state of other memory cells in the array; Second, a continuous (analog) programming of the memory cells is provided.

[0088] Specifically, the memory state (i.e., the charge on the floating gate) of each memory cell in the array can be changed continuously from a fully erased state to a fully programmed state, independently and with minimal disturbance to other memory cells. In another embodiment, the memory state (i.e., the charge on the floating gate) of each memory cell in the array can be changed continuously from a fully programmed state to a fully erased state, and vice versa, independently and with minimal disturbance to other memory cells. This means that the cell storage is analog, or can at a minimum store one of many discrete values ​​(such as 16 or 64 different values), making every cell in the memory array very precisely and individually adjustable, and making the memory array ideal for storage and fine tuning the synaptic weights of a neural network.

[0089] The methods and means described herein can be applied to other non-volatile memory technologies, such as, but not limited to, SONOS (silicon-oxide-nitride-oxide-silicon, charge traps in nitride), MONOS (metal-oxide-nitride-oxide-silicon, metal charge traps in nitride), ReRAM (resistive RAM), PCM (phase change RAM), MRAM (magnetoresistive RAM), FeRAM (ferroelectric RAM), OTP (bi-level or multi-level one-time programmable) and CeRAM (strongly correlated electron RAM). The methods and means described herein can be applied to, but not limited to, volatile memory technologies used in neural networks, such as SRAM, DRAM, and other volatile synapse cells. [Neural network using non-volatile memory cell array]

[0090] 8 conceptually illustrates a non-limiting example of a neural network utilizing the non-volatile memory array of the present embodiments. This example uses the non-volatile memory array neural network for a face recognition application, although other suitable applications can also be implemented using the non-volatile memory array-based neural network.

[0091] S0 is the input layer, which in this example is a 32x32 pixel RGB image with 5-bit precision (i.e., three 32x32 pixel arrays, one for each color R, G, and B, with each pixel being 5-bit precision). Synapse CB1 going from input layer S0 to layer C1 scans the input image with overlapping filters (kernels) of 3x3 pixels, applying different sets of weights to some instances and shared weights to other instances, and shifts the filters by one pixel (or two or more pixels, depending on the model). Specifically, the values ​​of the nine pixels in the 3x3 portion of the image (i.e., called filters or kernels) are provided to synapse CB1, which multiplies these nine input values ​​by the appropriate weights, and after summing the outputs of the multiplications, a single output value is determined and given by the first synapse of CB1 to generate one pixel of the layer of feature map C1. The 3x3 filter is then shifted one pixel to the right in the input layer S0 (i.e., add a column of 3 pixels to the right and drop a column of 3 pixels on the left), so that the 9 pixel values ​​of this newly positioned filter are provided to synapse CB1, where they are multiplied by the same weights as above to determine a second single output value by the associated synapse. This process continues until the 3x3 filter scans the entire 32x32 pixel image of input layer S0 for all three colors and all bits (precision values). The process is then repeated using different sets of weights to generate different feature maps of C1, until all feature maps of layer C1 have been calculated.

[0092] In this example, in layer C1, there are 16 feature maps, each with 30x30 pixels. Each pixel is a new feature pixel extracted from the multiplication of the input with a kernel, and thus each feature map is a two-dimensional array, and thus in this example layer C1 constitutes 16 layers of two-dimensional arrays (note that layers and arrays referred to herein are logical, not necessarily physical, i.e., arrays are not necessarily oriented in a physical two-dimensional array). Each of the 16 feature maps in layer C1 is generated by one of 16 different sets of synaptic weights applied to the filter scans. The C1 feature maps can all target different aspects of the same image feature, such as boundary identification. For example, a first map (generated using a first set of weights shared by all scans used to generate this first map) can identify circular edges, a second map (generated using a second set of weights different from the first set of weights) can identify rectangular edges or the aspect ratio of a particular feature, etc.

[0093] From layer C1, an activation function P1 (pooling) is applied before going to layer S1, which pools values ​​from non-overlapping, consecutive 2x2 regions in each feature map. The purpose of the pooling function is to average nearby positions (or a max function can be used), e.g. to reduce edge position dependency, and to reduce data size before going to the next stage. In layer S1, there are 16 15x15 feature maps (i.e. 16 different arrays of 15x15 pixels each). Synapse CB2 going from layer S1 to layer C2 scans the maps in S1 with a 4x4 filter with a filter shift of 1 pixel. In layer C2, there are 22 12x12 feature maps. From layer C2, an activation function P2 (pooling) is applied before going to layer S2, which pools values ​​from non-overlapping, consecutive 2x2 regions in each feature map. In layer S2, there are 22 6x6 feature maps. At synapse CB3 going from layer S2 to layer C3, an activation function (pooling) is applied, where all neurons in layer C3 connect to all maps in layer S2 through respective synapses in CB3. There are 64 neurons in layer C3. Synapse CB4 going from layer C3 to output layer S3 fully connects C3 to S3, i.e. all neurons in layer C3 connect to all neurons in layer S3. The output at S3 contains 10 neurons, where the neuron with the highest output determines the class. This output can indicate, for example, the identification or classification of the content of the original image.

[0094] Each layer of synapses is implemented using an array or part of an array of non-volatile memory cells.

[0095] Figure 9 is a block diagram of a system that can be used for this purpose. A vector matrix multiplication (VMM) system 32 includes non-volatile memory cells and is used to control the synapses (Fig. 8In particular, the VMM system 32 includes a VMM array 33 including non-volatile memory cells arranged in rows and columns, an erase gate and word line gate decoder 34, a control gate decoder 35, a bit line decoder 36, and a source line decoder 37, which decode respective inputs to the non-volatile memory cell array 33. Inputs to the VMM array 33 can come from the erase gate and word line gate decoder 34 or from the control gate decoder 35. The source line decoder 37 in this example also decodes the output of the VMM array 33. Alternatively, the bit line decoder 36 can decode the output of the VMM array 33.

[0096] The VMM array 33 serves two purposes. First, it stores the weights used by the VMM system 32. Second, the VMM array 33 effectively multiplies the inputs by the weights stored in the VMM array 33 and sums them for each output line (source line or bit line) to generate an output, which becomes the input to the next layer or the input to the last layer. By performing multiplication and addition functions, the VMM array 33 eliminates the need for separate multiplication and addition logic and is also power efficient with in-place memory computation.

[0097] The outputs of the VMM array 33 are fed to a differential summer (such as a summing op-amp or summing current mirror) 38, which sums the outputs of the VMM array 33 to create a single value for the convolution. The differential summer 38 is arranged to perform a summation of both the positive and negative weight inputs to output a single value.

[0098] The summed output values ​​of the differential summer 38 are then fed to an activation function circuit 39 which rectifies the output. The activation function circuit 39 may provide a sigmoid function, a tanh function, a ReLU function, or any other non-linear function. The rectified output values ​​of the activation function circuit 39 become elements of the feature map of the next layer (e.g., C1 in FIG. 8) and are then applied to the next synapse to generate the next feature map layer or the last layer. Thus, in this example, the VMM array 33 constitutes a number of synapses (which receive input from a previous layer of neurons or from an input layer such as an image database), and the summer 38 and the activation function circuit 39 constitute a number of neurons.

[0099] The inputs to the VMM system 32 of FIG. 9 (WLx, EGx, CGx, and optionally BLx and SLx) may be analog levels, binary levels, digital pulses (in which case a pulse-to-analog converter PAC may be required to convert the pulses to appropriate input analog levels) or digital bits (in which case a DAC is provided to convert the digital bits to appropriate input analog levels), and the outputs may be analog levels, binary levels, digital pulses, or digital bits (in which case an output ADC is provided to convert the output analog levels to digital bits).

[0100] FIG. 10 is a block diagram illustrating the use of multiple layers of the VMM system 32, labeled in the figure as VMM systems 32a, 32b, 32c, 32d, and 32e. As shown in FIG. 10, an input (denoted Inputx) is converted from digital to analog by a digital-to-analog converter 31 and provided to the input VMM system 32a. The converted analog input can be a voltage or a current. The first layer input D / A conversion can be done by using a function or a LUT (look-up table) that maps the input Inputx to the appropriate analog level of the matrix multiplier of the input VMM system 32a. The input conversion can also be done by an analog-to-analog (A / A) converter to convert an external analog input to a mapped analog input to the input VMM system 32a. The input conversion can also be done by a digital-to-digital pulse (D / P) converter to convert an external digital input to a mapped digital pulse(s) to the input VMM system 32a.

[0101] The output generated by the input VMM system 32a is then provided as an input to the next VMM system (hidden level 1) 32b, which in turn generates an output that is provided as an input to the input VMM system (hidden level 2) 32c, and so on. The various layers of the VMM systems 32 function as layers of synapses and neurons of a convolutional neural network (CNN). Each of the VMM systems 32a, 32b, 32c, 32d, and 32e can be a standalone physical system that includes a corresponding non-volatile memory array, or the multiple VMM systems can utilize different portions of the same physical non-volatile memory array, or the multiple VMM systems can utilize overlapping portions of the same physical non-volatile memory array. Each of the VMM systems 32a, 32b, 32c, 32d, and 32e can also be time multiplexed to various portions of its array or neurons. The example shown in Figure 10 includes five layers (32a, 32b, 32c, 32d, 32e): one input layer (32a), two hidden layers (32b, 32c), and two fully connected layers (32d, 32e). Those skilled in the art will appreciate that this is merely an example and that the system may instead include more than two hidden layers and more than two fully connected layers. [VMM Array]

[0102] 11 shows a neuron VMM array 1100 that is particularly suited for the memory cells 310 shown in FIG. 3 and is utilized as part of the synapses and neurons between the input layer and the next layer. The VMM array 1100 includes a memory array 1101 of non-volatile memory cells and a reference array 1102 of non-volatile reference memory cells (located at the top of the array). Alternatively, a separate reference array could be located at the bottom.

[0103] In the VMM array 1100, control gate lines such as control gate line 1103 run vertically (so that row-wise reference array 1102 is orthogonal to control gate line 1103) and erase gate lines such as erase gate line 1104 run horizontally. Here, inputs to the VMM array 1100 are provided to the control gate lines (CG0, CG1, CG2, CG3) and outputs of the VMM array 1100 appear on the source lines (SL0, SL1). In one embodiment, only even rows are used, and in another embodiment, only odd rows are used. The current applied to each source line (SL0, SL1, respectively) performs a function of the sum of all the currents from the memory cells connected to that particular source line.

[0104] As described herein for neural networks, the non-volatile memory cells of VMM array 1100, ie, the flash memory of VMM array 1100, are preferably configured to operate in the sub-threshold region.

[0105] The non-volatile reference memory cells and non-volatile memory cells described herein are biased in weak inversion as follows: Ids=Io * e (Vg-Vth) / nVt =w * Io * e (Vg) / nVt In the formula, w=e (-Vth) / nVt and where Ids is the drain-source current, Vg is the gate voltage of the memory cell, Vth is the threshold voltage of the memory cell, and Vt is the thermal voltage = k * T / q, k is the Boltzmann constant, T is temperature in Kelvin, q is the electron charge, n is the slope coefficient = 1 + (Cdep / Cox), Cdep = capacitance of the depletion layer, and Cox is the capacitance of the gate oxide layer, Io is the memory cell current at a gate voltage equal to the threshold voltage, and Io is equal to (Wt / L) * u * Cox * (n-1) * Vt 2where u is the carrier mobility, and Wt and L are the width and length of the memory cell, respectively.

[0106] When using an IV log converter that converts the input current Ids to an input voltage Vg using a memory cell (such as a reference memory cell or a peripheral memory cell) or transistor: Vg=n * Vt * log[Ids / wp * Io] where wp is the w of the reference or periphery memory cell.

[0107] When using an IV log converter that converts the input current Ids to an input voltage Vg using a memory cell (such as a reference memory cell or a peripheral memory cell) or transistor: Vg=n * Vt * log[Ids / wp * Io]

[0108] where wp is the w of the reference or periphery memory cell.

[0109] For a memory array used as a vector matrix multiplier VMM array, the output current is: Iout=wa * Io * e (Vg) / nVt , i.e. Iout = (wa / wp) * Iin=W * Iin W=e (Vthp-Vtha) / nVt Iin=wp * Io * e (Vg) / nVt where wa=w for each memory cell in the memory array.

[0110] The word line or control gate can be used as the input of the memory cell for an input voltage.

[0111] Alternatively, the non-volatile memory cells of the VMM arrays described herein can be configured to operate in the linear region. Ids=β * (Vgs-Vth) * Vds; β=u * Cox * Wt / L W α (Vgs-Vth) That is, the weight W in the linear region is proportional to (Vgs-Vth).

[0112] The word line or control gate or bit line or source line can be used as the input of a memory cell operating in the linear region. The bit line or source line can be used as the output of the memory cell.

[0113] For an IV linear converter, memory cells (such as reference or peripheral memory cells) or transistors operating in the linear region, or resistors can be used to linearly convert input and output currents to input and output voltages.

[0114] Alternatively, the memory cells of the VMM arrays described herein can be configured to operate in the saturation region. Ids=1 / 2 * β * (Vgs-Vth) 2 ; β=u * Cox * Wt / L W α (Vgs-Vth) 2 , that is, the weight W is (Vgs-Vth) 2 is proportional to.

[0115] The word line, control gate, or erase gate can be used as the input of a memory cell operating in the saturation region, and the bit line or source line can be used as the output of an output neuron.

[0116] Alternatively, the memory cells of the VMM arrays described herein may be used in all regions or combinations thereof (sub-threshold, linear, or saturation).

[0117] Other embodiments for the VMM array 33 of Figure 9 are described in U.S. Patent Application Serial No. 15 / 826,345, which is incorporated herein by reference. As described in that application, the source lines or bit lines can be used as neuron outputs (current sum outputs).

[0118] FIG. 12 shows a neuron VMM array 1200 that is particularly suitable for the memory cells 210 shown in FIG. 2 and used as synapses between an input layer and the next layer. The VMM array 1200 includes a memory array 1203 of non-volatile memory cells, a reference array 1201 of a first non-volatile reference memory cell, and a reference array 1202 of a second non-volatile reference memory cell. The reference arrays 1201 and 1202 arranged in the columns of the array serve to convert the current inputs flowing into terminals BLR0, BLR1, BLR2, and BLR3 into voltage inputs WL0, WL1, WL2, and WL3. In practice, the first and second non-volatile reference memory cells are diode-connected through a multiplexer 1214 (only some of which are shown) with the current inputs flowing into them. The reference cells are adjusted (e.g., programmed) to a target reference level. The target reference level is provided by a reference mini-array matrix (not shown).

[0119] The memory array 1203 serves two purposes. First, it stores the weights used by the VMM array 1200 in each memory cell. Second, it effectively multiplies the inputs (i.e., the current inputs provided to terminals BLR0, BLR1, BLR2, and BLR3, which the reference arrays 1201 and 1202 convert to input voltages and provide to word lines WL0, WL1, WL2, and WL3) by the weights stored in the memory cell array 1203, and then adds all the results (memory cell currents) to generate outputs for each bit line (BL0-BLN), which are the inputs to the next layer or the last layer. With the memory array 1203 performing the multiplication and addition functions, the need for separate multiplication and addition logic is eliminated and is also power efficient. Here, voltage inputs are provided to word lines WL0, WL1, WL2, and WL3, and outputs appear on bit lines BL0-BLN, respectively, during read (inference) operations. The current placed on each bit line BL0-BLN performs a function of the sum of the currents from all the non-volatile memory cells connected to that particular bit line.

[0120] Table 5 shows the operating voltages for VMM array 1200. The columns in the table indicate the voltages applied to the word line of the selected cell, the word lines of the unselected cells, the bit lines of the selected cells, the bit lines of the unselected cells, the source lines of the selected cells, and the source lines of the unselected cells, and FLT indicates floating, i.e., no voltage is applied. The rows indicate read, erase, and program operations. Table 5: Operation of VMM Array 1200 in FIG. 12 [Table 5]

[0121] FIG. 13 illustrates a neuron VMM array 1300 that is particularly suited for the memory cells 210 shown in FIG. 2 and is utilized as part of synapses and neurons between an input layer and the next layer. The VMM array 1300 includes a memory array 1303 of non-volatile memory cells, a reference array 1301 of first non-volatile reference memory cells, and a reference array 1302 of second non-volatile reference memory cells. The reference arrays 1301 and 1302 run in the row direction of the VMM array 1300. The VMM array is similar to the VMM 1000, except that the word lines run vertically in the VMM array 1300. Here, inputs are provided to the word lines (WLA0, WLB0, WLA1, WLB2, WLA2, WLB2, WLA3, WLB3) and outputs appear on the source lines (SL0, SL1) during a read operation. The current applied to each source line performs a function of the sum of all the currents from the memory cells connected to that particular source line.

[0122] Table 6 shows the operating voltages for VMM array 1300. The columns in the table show the voltages applied to the word line of the selected cell, the word lines of the unselected cells, the bit lines of the selected cells, the bit lines of the unselected cells, the source lines of the selected cells, and the source lines of the unselected cells. The rows show the read, erase, and program operations. Table 6: Operation of VMM Array 1300 in FIG. 13 [Table 6]

[0123] 14 shows a neuron VMM array 1400 that is particularly suited for the memory cells 310 shown in FIG. 3 and is used as part of synapses and neurons between the input layer and the next layer. The VMM array 1400 includes a memory array 1403 of non-volatile memory cells, a reference array 1401 of a first non-volatile reference memory cell, and a reference array 1402 of a second non-volatile reference memory cell. The reference arrays 1401 and 1402 function to convert current inputs flowing into terminals BLR0, BLR1, BLR2, and BLR3 to voltage inputs CG0, CG1, CG2, and CG3. In practice, the first and second non-volatile reference memory cells are diode-connected through a multiplexer 1412 (only a portion of which is shown), with the current inputs flowing through BLR0, BLR1, BLR2, and BLR3. The multiplexers 1412 each include a respective multiplexer 1405 and cascoding transistor 1404 to ensure a constant voltage on the bit line (e.g., BLR0) of each of the first and second non-volatile reference memory cells during a read operation. The reference cells are adjusted to a target reference level.

[0124] Memory array 1403 serves two purposes. First, it stores the weights used by VMM array 1400. Second, it stores the weights used by VMM array 1400. Second, it stores the weights used by VMM array 1400. The inputs (current inputs provided to terminals BLR0, BLR1, BLR2, and BLR3) that reference arrays 1401 and 1402 convert to input voltages to be used by control gates (CG0, CG1, CG2, and CG3). ) ) by weights stored in the memory cell array and then all the results (cell currents) are added together to produce an output that appears on BL0-BLN and serves as the input to the next layer or the input to the last layer. Having the memory array perform the multiplication and addition functions removes the need for separate multiplication and addition logic circuits and is also power efficient. Here, the inputs are provided to the control gate lines (CG0, CG1, CG2, and CG3) and the outputs appear on the bit lines (BL0-BLN) during a read operation. The current applied to each bit line performs a function of the sum of all the currents from the memory cells connected to that particular bit line.

[0125] The VMM array 1400 performs one-way adjustment of the non-volatile memory cells in the memory array 1403. That is, each non-volatile memory cell is erased and then partially programmed until the desired charge on the floating gate is reached. This can be done, for example, using the precision programming techniques described below. If too much charge is added to the floating gate (such as an incorrect value being stored in the cell), the cell must be erased and the series of partial programming operations must be redone. As shown, two rows that share the same erase gate (such as EG0 or EG1) must be erased together (known as a page erase), and then each cell is partially programmed until the desired charge on the floating gate is reached.

[0126] Table 7 shows the operating voltages for the VMM array 1400. The columns in the table show the voltages applied to the word line of the selected cell, the word lines of the unselected cells, the bit line of the selected cell, the bit lines of the unselected cells, the control gate of the selected cell, the control gates of the unselected cells in the same sector as the selected cell, the control gates of the unselected cells in a different sector than the selected cell, the erase gate of the selected cell, the erase gates of the unselected cells, the source line of the selected cell, and the source lines of the unselected cells. The rows show the read, erase, and program operations. Table 7: Operation of VMM Array 1400 in FIG. 14 [Table 7]

[0127] 15 shows a neuron VMM array 1500 that is particularly suited for memory cells 310 shown in FIG. 3 and is used as part of synapses and neurons between an input layer and the next layer. The VMM array 1500 includes a memory array 1503 of non-volatile memory cells; The first non-volatile reference memory cell Reference Array 150 1 and, and a second reference array 1502 of non-volatile reference memory cells. EG lines EGR0, EG0, EG1, and EGR1 run vertically, while CG lines CG0, CG1, CG2, and CG3 and SL lines WL0, WL1, WL2, and WL3 run horizontally. VMM array 1500 is similar to VMM array 1400, except that VMM array 1500 implements bidirectional tuning, whereby each individual cell can be fully erased, partially programmed, and partially erased as needed to reach a desired amount of charge on the floating gate through the use of individual EG lines. As shown, reference arrays 1501 and 1502 convert input currents in terminals BLR0, BLR1, BLR2, and BLR3 into control gate voltages CG0, CG1, CG2, and CG3 (through the action of diode-connected reference cells via multiplexer 1514), which are applied to the memory cells in the row direction. The current outputs (neurons) are in bit lines BL0-BLN, each of which sums up all the currents from the non-volatile memory cells connected to that particular bit line.

[0128] Table 8 shows the operating voltages for the VMM array 1500. The columns in the table show the voltages applied to the word line of the selected cell, the word lines of the unselected cells, the bit line of the selected cell, the bit lines of the unselected cells, the control gate of the selected cell, the control gates of the unselected cells in the same sector as the selected cell, the control gates of the unselected cells in a different sector than the selected cell, the erase gate of the selected cell, the erase gates of the unselected cells, the source line of the selected cell, and the source lines of the unselected cells. The rows show the read, erase, and program operations. Table 8: Operation of VMM Array 1500 in FIG. 15 [Table 8]

[0129] 16 shows a neuron VMM array 1600 that is particularly suited for memory cells 210 shown in FIG. 2 and is used as part of synapses and neurons between an input layer and the next layer. In the VMM array 1600, inputs INPUT0, ..., INPUT Nare bit lines BL0, ..., BL N and outputs OUTPUT1, OUTPUT2, OUTPUT3, and OUTPUT4 are generated on source lines SL0, SL1, SL2, and SL3, respectively.

[0130] 17 illustrates a neuron VMM array 1700 that is particularly suited for memory cells 210 shown in FIG. 2 and that is utilized as part of the synapses and neurons between an input layer and the next layer. In this example, inputs INPUT0, INPUT1, INPUT2, and INPUT3 are received on source lines SL0, SL1, SL2, and SL3, respectively, and outputs OUTPUT0, ..., OUTPUT N are bit lines BL0, ..., BL N is generated.

[0131] 18 illustrates a neuron VMM array 1800 that is particularly suited for memory cells 210 shown in FIG. 2 and is used as part of the synapses and neurons between the input layer and the next layer. In this example, inputs INPUT0, ..., INPUT M are the word lines WL0, ..., WL M are received at OUTPUT0, ..., OUTPUT N are bit lines BL0, ..., BL N is generated.

[0132] 19 shows a neuron VMM array 1900 that is particularly suited for memory cells 310 shown in FIG. 3 and is used as part of synapses and neurons between an input layer and the next layer. In this example, inputs INPUT0, ..., INPUT M are the word lines WL0, ..., WL M are received at OUTPUT0, ..., OUTPUT N are bit lines BL0, ..., BL N is generated.

[0133] 20 shows a neuron VMM array 2000 that is particularly suited for the memory cells 410 shown in FIG. 4 and is used as part of the synapses and neurons between the input layer and the next layer. In this example, the input INPUT 0、 ..., INPUT n are the vertical control gate lines CG0, ..., CG N and outputs OUTPUT1 and OUTPUT2 are generated on source lines SL0 and SL1.

[0134] 21 shows a neuron VMM array 2100 that is particularly suited for memory cells 410 shown in FIG. 4 and is used as part of synapses and neurons between an input layer and the next layer. In this example, inputs INPUT0, ..., INPUT N are bit lines BL0, ..., BL N , 2901-(N-1) and 2901-N, which are coupled to the source lines SL0 and SL1, respectively. Exemplary outputs OUTPUT1 and OUTPUT2 are generated on source lines SL0 and SL1.

[0135] 22 shows a neuron VMM array 2200 that is particularly suitable for memory cells 310 shown in FIG. 3, memory cells 510 shown in FIG. 5, and memory cells 710 shown in FIG. 7, and is used as part of synapses and neurons between the input layer and the next layer. In this example, inputs INPUT0, ..., INPUT M are the word lines WL0, ..., WL M Received at OUTPUT0, ..., OUTPUT N are bit lines BL0, ..., BL N are generated respectively.

[0136] 23 shows a neuron VMM array 2300 that is particularly suitable for memory cells 310 shown in FIG. 3, memory cells 510 shown in FIG. 5, and memory cells 710 shown in FIG. 7, and is used as part of synapses and neurons between the input layer and the next layer. In this example, inputs INPUT0, ..., INPUT Mare the control gate lines CG0, ..., CG M Received at OUTPUT0, ..., OUTPUT N are the vertical source lines SL0, ..., SL N and each source line SL i is coupled to the source lines of all memory cells in column i.

[0137] 24 shows a neuron VMM array 2400 that is particularly suitable for memory cells 310 shown in FIG. 3, memory cells 510 shown in FIG. 5, and memory cells 710 shown in FIG. 7, and is used as part of synapses and neurons between the input layer and the next layer. In this example, inputs INPUT0, ..., INPUT M are the control gate lines CG0, ..., CG M Received at OUTPUT0, ..., OUTPUT N are the vertical bit lines BL0, ..., BL N and each bit line BL i is coupled to the bit lines of all memory cells in column i. [Long and short-term memory]

[0138] Prior art includes a concept known as long short-term memory (LSTM). LSTMs are often used in artificial neural networks. LSTMs allow artificial neural networks to remember information for any given period of time and use that information in subsequent operations. A traditional LSTM includes a cell, an input gate, an output gate, and a forget gate. The three gates regulate the flow of information into and out of the cell, and the period of time that information is stored within the LSTM. VMMs are particularly useful in LSTMs.

[0139] FIG. 25 illustrates an exemplary LSTM 2500. The LSTM 2500 in this example includes cells 2501, 2502, 2503, and 2504. Cell 2501 receives an input vector x0 and generates an output vector h0 and a cell state vector c0. Cell 2502 receives an input vector x1, an output vector (hidden state) h0 from cell 2501, and a cell state c0 from cell 2501, and generates an output vector h1 and a cell state vector c1. Cell 2503 receives an input vector x2, an output vector (hidden state) h1 from cell 2502, and a cell state c1 from cell 2502, and generates an output vector h2 and a cell state vector c2. Cell 2504 receives an input vector x3, an output vector (hidden state) h2 from cell 2503, and a cell state c2 from cell 2503, and generates an output vector h3. Additional cells can be used, the LSTM with four cells is just an example.

[0140] Figure 26 shows an example implementation of an LSTM cell 2600 that can be used for cells 2501, 2502, 2503, and 2504 in Figure 25. LSTM cell 2600 receives an input vector x(t), a cell state vector c(t-1) from a previous cell, and an output vector h(t-1) from a previous cell, and produces a cell state vector c(t) and an output vector h(t).

[0141] LSTM cell 2600 includes sigmoid function devices 2601, 2602, and 2603, each of which applies a number between 0 and 1 to control the degree to which each component of the input vector contributes to the output vector. LSTM cell 2600 also includes tanh devices 2604 and 2605 for applying a hyperbolic tangent function to the input vector, multiplier devices 2606, 2607, and 2608 for multiplying two vectors, and adder device 2609 for adding the two vectors. The output vector h(t) can be provided to the next LSTM cell in the system or can be accessed for other purposes.

[0142] 27 shows an example of an implementation of LSTM cell 2600, LSTM cell 2700. For the convenience of the reader, the same numbering scheme from LSTM cell 2600 is used in LSTM cell 2700. Sigmoid function devices 2601, 2602, and 2603, and tanh device 2604 each include multiple VMM arrays 2701 and activation circuit blocks 2702. It can therefore be seen that VMM arrays are particularly useful in LSTM cells used in certain neural network systems.

[0143] An alternative example of LSTM cell 2700 (and another example of one implementation of LSTM cell 2600) is shown in Figure 28. In Figure 28, sigmoid function devices 2601, 2602, and 2603, and tanh device 2604 may share the same physical hardware (VMM array 2801 and activation function block 2802) in a time-multiplexed manner. LSTM cell 2800 also includes a multiplier device 2803 for multiplying two vectors, an adder device 2808 for adding two vectors, a tanh device 2605 (including activation circuit block 2802), a register 2807 for storing the value i(t) output from sigmoid function block 2802, and a value f(t) output from multiplier device 2803 via multiplexer 2810. * A register 2804 stores c(t-1), and a value i(t) output from the multiplier device 2803 via a multiplexer 2810. * A register 2805 stores u(t) and the value o(t) output from the multiplier device 2803 via a multiplexer 2810. * It includes a register 2806 that stores {tilde over (c)}(t) and a multiplexer 2809.

[0144] Whereas LSTM cell 2700 includes multiple sets of VMM arrays 2701 and respective activation function blocks 2702, LSTM cell 2800 includes only one set of VMM arrays 2801 and activation function blocks 2802, which are used to represent multiple layers in embodiments of LSTM cell 2800. LSTM cell 2800 requires less space than LSTM cell 2700, as it requires ¼ the space for the VMMs and activation function blocks compared to LSTM cell 2700.

[0145] It can be further appreciated that an LSTM unit typically includes multiple VMM arrays, each of which requires functionality provided by specific circuit blocks outside the VMM array, such as an adder and activation circuit block and a high voltage generation block. Providing a separate circuit block for each VMM array would require a significant amount of space within a semiconductor device and would be somewhat inefficient. Thus, the embodiments described below attempt to minimize the circuitry required outside the VMM array itself. [Gated Recurrent Unit]

[0146] An analog VMM implementation can be used for gated recurrent units (GRUs), which are gating mechanisms in recurrent artificial neural networks. GRUs are similar to LSTMs, except that GRU cells generally contain fewer components than LSTM cells.

[0147] 29 illustrates an exemplary GRU 2900. The GRU 2900 in this example includes cells 2901, 2902, 2903, and 2904. Cell 2901 receives input vector x0 and generates output vector h0. Cell 2902 receives input vector x1 and output vector h0 from cell 2901 and generates output vector h1. Cell 2903 receives input vector x2 and output vector (hidden state) h1 from cell 2902 and generates output vector h2. Cell 2904 receives input vector x3 and output vector (hidden state) h2 from cell 2903 and generates output vector h3. Additional cells are possible, and a GRU with four cells is merely an example.

[0148] FIG. 30 illustrates an exemplary implementation of a GRU cell 3000 that can be used for cells 2901, 2902, 2903, and 2904 of FIG. 29. The GRU cell 3000 receives an input vector x(t) and an output vector h(t-1) from a preceding GRU cell and generates an output vector h(t). The GRU cell 3000 includes sigmoid function devices 3001 and 3002, each of which applies a number between 0 and 1 to the output vector h(t-1) and the components from the input vector x(t). The GRU cell 3000 also includes a tanh device 3003 for applying a hyperbolic tangent function to the input vector, a number of multiplier devices 3004, 3005, and 3006 for multiplying two vectors, an adder device 3007 for adding the two vectors, and a complementary device 3008 for subtracting the input from 1 to generate an output.

[0149] Figure 31 shows a GRU cell 3100, which is an example of one implementation of GRU cell 3000. For the convenience of the reader, the same numbering scheme from GRU cell 3000 is used in GRU cell 3100. As can be seen from Figure 31, sigmoid function devices 3001 and 3002 and tanh device 3003 each include multiple VMM arrays 3101 and activation function blocks 3102. It can therefore be seen that VMM arrays are particularly used in GRU cells used in certain neural network systems.

[0150] An alternative example of GRU cell 3100 (and another example of one implementation of GRU cell 3000) is shown in Figure 32. In Figure 32, GRU cell 3200 uses a VMM array 3201 and activation function block 3202, which, when configured as a sigmoid function, applies a number between 0 and 1 to control how much each component of the input vector contributes to the output vector. In Figure 32, sigmoid function devices 3001 and 3002 and tanh device 3003 share the same physical hardware (VMM array 3201 and activation function block 3202) in a time-multiplexed manner. The GRU cell 3200 also includes a multiplier device 3203 for multiplying two vectors together, an adder device 3205 for adding the two vectors together, a complementary device 3209 for subtracting an input from one to generate an output, a multiplexer 3204, and a value h(t-1) output from the multiplier device 3203 via the multiplexer 3204. * A register 3206 holds r(t) and a value h(t-1) output from the multiplier device 3203 via a multiplexer 3204. * A register 3207 holds z(t), and a value h^(t) output from the multiplier device 3203 via a multiplexer 3204. * and a register 3208 that holds (1-z((t)).

[0151] Whereas GRU cell 3100 includes multiple sets of VMM arrays 3101 and activation function blocks 3102, GRU cell 3200 includes only one set of VMM arrays 3201 and activation function blocks 3202 that are used to represent multiple tiers in embodiments of GRU cell 3200. GRU cell 3200 requires less space than GRU cell 3100 because it requires ⅓ the space for the VMM and activation function blocks compared to GRU cell 3100.

[0152] It can be further appreciated that a system utilizing a GRU typically includes multiple VMM arrays, each of which requires functionality provided by specific circuit blocks outside the VMM array, such as an adder and activation circuit block and a high voltage generation block. Providing separate circuit blocks for each VMM array would require a significant amount of space within a semiconductor device and would be somewhat inefficient. Thus, the embodiments described below attempt to minimize the circuitry required outside the VMM array itself.

[0153] The input to the VMM array can be analog levels, binary levels, timing pulses, or digital bits, and the output can be analog levels, binary levels, timing pulses, or digital bits (in this case an output ADC is required to convert the output analog level current or voltage to digital bits).

[0154] For each memory cell in the VMM array, each weight w can be implemented by a single memory cell, by a differential cell, or by two blended memory cells (an average of two or more cells). In the case of a differential cell, the weight w can be expressed as a differential weight (w=w + -w - ) for two blended memory cells, two memory cells are required to implement the weight w as the average of the two cells. [Embodiments for precision programming of cells in a VMM]

[0155] Embodiments are now described for precisely programming memory cells in a VMM by incrementing or decrementing programming voltages applied to different terminals of the memory cells.

[0156] FIG. 33A illustrates a programming method 3300. Initially, the method begins, typically in response to a program (conditioning) command received (step 3301). Next, a mass program operation programs all cells to a "0" state (step 3302). A soft erase operation then erases all cells to an intermediate weak erase level (step 3303), such that each cell draws, for example, approximately 3-5 μA of current during a read operation. This is in contrast to the highest deeply erased level, where each cell draws approximately 20-30 μA of current during a read operation. The soft erase is performed, for example, by applying incremental erase voltage pulses until an intermediate cell current is reached. The incremental erase voltage pulses are performed to limit the degradation to the memory cells experienced from a hard erase (i.e., maximum erase level). A hard program is then performed (step 3304) in all unselected cells to add electrons to the floating gates of the cells to a very deep programmed state, ensuring that those cells are truly "off", i.e., that those cells, e.g., unused memory cells, draw a negligible amount of current during a read operation. The hard program is performed, for example, with a higher incremental program voltage pulse and / or a longer program time.

[0157] A coarse programming method (bringing the cells fairly close to the target, e.g., 2x to 100x the target) is then performed on the selected cells (step 3305), followed by a fine programming method being performed on the selected cells (step 3306) to program the exact value desired into each selected cell.

[0158] FIG. 33B illustrates another programming method 3310 similar to programming method 3300. However, after the method begins (step 3301), instead of a program operation to program all cells to a "0" state as in step 3302 of FIG. 33A, a soft erase operation is used to erase all cells to a "1" state (step 3312). A soft program operation (step 3313) is then used to program all cells to an intermediate level so that each cell draws approximately 3-5 uA of current during a read operation. Thereafter, coarse programming method 3305 and fine programming method 3306 are performed as in FIG. 33A. A variation of the embodiment of FIG. 33B removes the soft program operation (step 3313) entirely. Multiple coarse programming methods can be used to accelerate programming, such as by targeting multiple progressively smaller coarse targets, before performing the fine programming step 3306. The fine programming method 3306 is performed, for example, with fine incremental program voltage pulses or constant program timing pulses.

[0159] Different terminals of the memory cell can be used for the coarse programming method 3305 and the fine programming method 3306. That is, during the coarse programming method 3305, the voltage applied to one of the terminals of the memory cell (which may be referred to as the coarse programming terminal) is changed until the desired voltage level is achieved in the floating gate 20, and during the fine programming method 3306, the voltage applied to one of the terminals of the memory cell (which may be referred to as the fine programming terminal) is changed until the desired level is achieved. Various combinations of terminals that can be used as coarse and fine programming terminals are shown in Table 9. Table 9: Memory Cell Terminals Used for Coarse and Fine Programming Methods [Table 9] Other combinations of terminals for the coarse and fine programming steps are possible.

[0160] 34 illustrates a first embodiment of the coarse programming method 3305, which is a lookup and execution method 3400. First, a lookup table lookup or function (IV curve) is performed to find the coarse target current value (I CT ) based on the value intended to be stored in the selected cell (step 3401). This table is created, for example, by silicon characterization or from wafer test calibration. The selected cell can be programmed to store one of N (e.g., 128, 64, 32, etc.) possible values. Each of the N values ​​represents a different desired current value (I D In one embodiment, the lookup table corresponds to a coarse target current value I for a cell selected during the search and execution method 3400. CT , where M is an integer less than N. For example, if N is 8, then M may be 4, meaning there are 8 possible values ​​that the selected cell can store, and one of the four coarse target current values ​​will be selected as the coarse target for the search and perform method 3400. That is, the search and perform method 3400 (which is again an embodiment of the coarse programming method 3305) programs the selected cell to store a desired value (I D ) somewhat close to the value (I CT ) and then the precision programming method 3306 programs the desired value (I D ) or achieve the desired value (I D The intention is to more precisely program the selected cells so that they are very close to the

[0161] Examples of cell values, desired current values, and coarse target current values ​​are shown in Tables 10 and 11 for a simple example of N=8 and M=4. Table 10: Example of N desired current values ​​for N=8 [Table 10] Table 11: Example of M target current values ​​for M=4 [Table 11] Offset value I CTOFFSETx is used to prevent overshooting the desired current value during coarse adjustment.

[0162] Rough target current value I CT Once selected, in step 3402, the selected cell is programmed by applying an initial voltage v0 to the selected cell's coarse programming terminal according to one of the sequences listed in Table 9 above. (The value of the initial voltage v0 and the appropriate coarse programming terminal are optionally programmed according to one of the sequences listed in Table 9 above, in accordance with the coarse target current value I CT This can be determined from a voltage lookup table that stores v0 in correspondence with Table 12. Initial voltage v0 applied to the coarse programming terminal [Table 12]

[0163] Next, in step 3403, the selected cell is charged to a voltage v i =v i-1 +v increment to the coarse programming terminal, where i starts at 1 and increments each time this step is repeated, and v increment is the coarse voltage increment that will cause programming to a degree commensurate with the granularity of the desired change. Thus, the first time step 3403 is performed at i=1, and v is v increment A read operation is then performed on the selected cell, and the current drawn through the selected cell (I cell ) is measured and a verification operation is performed (step 3404). cell I CT If so, the search and execute method 3400 is complete and the precision programming method 3306 can begin. cellI CT If not, step 3403 is repeated and i is incremented.

[0164] Thus, at the point where the coarse programming method 3305 ends and the fine programming method 3306 begins, the voltage v i is the final voltage applied to the coarse programming terminal to program the selected cell, and the selected cell is programmed to the coarse target current value I CT The values ​​associated with, in particular, I CT The goal of the fine programming method 3306 is to program a selected cell such that during a read operation, the selected cell receives a current I D The goal is to program the current to a point where the selected cell will pull in (plus or minus an acceptable margin of deviation, such as + / - 30%), where this current is the desired current value associated with the value intended to be stored in the selected cell.

[0165] FIG. 35 shows examples of different voltage progressions that may be applied to the coarse programming terminals of selected memory cells according to Table 9 during the coarse programming method 3305 and / or to the fine programming terminals of selected memory cells during the fine programming method 3306.

[0166] Under the first approach, increasing voltages are applied incrementally to the coarse programming terminal and / or the fine programming terminal to further program the selected memory cell. The starting point is v i which is the last voltage applied during the coarse programming method 3305. p1 V i Then, the voltage v i +v p1 is used to program the selected cell (indicated by the second pulse from the left in progression 3501). p1 v increment (the voltage increment used during the coarse programming method 3305). After each programming voltage is applied to the programming terminal, Icell PT1A verification step (similar to step 3404) is performed in which it is determined whether I is less than or equal to (the first precision target current value, here the second threshold value). PT1 =I D +I PT1OFFSET and I PT1OFFSET is an offset value that is added to prevent program overshoot. If the answer is no, another increment v p1 is added to the previously applied programming voltage and the process is repeated. This is cell I PT1 The following is repeated until a point at which this portion of the programming sequence stops. PT1 I D or with a sufficiently acceptable accuracy D If it is approximately equal to , the selected memory cell has been successfully programmed.

[0167] I PT1 I D If the voltage is not close enough to V, then further programming with smaller granularity can be done. Here, progression 3502 is used. The starting point of progression 3502 is the last voltage used for programming under progression 3501. The increment V p2 (v p1 A voltage smaller than I is applied to the memory cell and the combined voltage is applied to the precision programming terminal to program the selected memory cell. After each programming voltage is applied, I cell I PT2 A verification step (similar to step 3404) is performed in which it is determined whether I is less than or equal to (a second precision target current value, here a third threshold value). PT2 =ID+I PT2OFFSET and I PT2OFFSET is an offset value added to prevent program overshoot. If the answer is no, another increment V p2 is added to the previously applied programming voltage and the process is repeated. This is cell I PT2The following is repeated until a point at which this part of the programming sequence stops. Now, since the target value has been achieved with a sufficient acceptable accuracy, I PT2 I D or until programming can stop. D It is assumed that the 1601, 1602, and 1603 steps are sufficiently close to the 1601, 1602, and 1603 steps. Those skilled in the art will appreciate that additional progressions may be applied using smaller and smaller programming increments. For example, in Figure 36A, three progressions (3601, 3602, and 3603) are applied instead of just two.

[0168] A second approach is shown in sequence 3503. Here, instead of increasing the voltage applied during programming of a selected memory cell, the same voltage (V i , or V i +V p1 +V p1 , or V i +V p2 +V p2 ) are applied for durations of increasing duration. p1 and v in Progress 3502 p2 Instead of applying incremental voltages such as t, each applied pulse is t slower than the previous applied pulse. p1 Add an additional time increment t p1 is added to the programming pulse. After each programming pulse is applied to the precision programming terminal, the same verification step is performed as described above for step 3501. Optionally, additional steps can be applied, where the additional time increment added to the programming pulse is of shorter duration than the previous step used.

[0169] Optionally, additional program cycle progressions can be applied in which the programming pulses are of the same duration as the previous program cycle progression used. Although only one time progression is shown, one skilled in the art will understand that any number of different time progressions can be applied. That is, instead of changing the magnitude of the voltage used during programming, or instead of changing the period of the voltage pulses used during programming, the system can instead change the number of programming cycles used.

[0170] FIG. 36B shows a diagram of a complementary pulse program progression where the voltage applied to one precision programming terminal is increased and the voltage applied to another precision programming terminal is decreased. For example, an increasing voltage progression may be applied to the control gate of the selected cell and a decreasing voltage progression may be applied to the erase gate or source line of the selected cell. Or, alternatively, an increasing voltage progression may be applied to the erase gate or source line of the selected cell and a decreasing voltage progression may be applied to the control gate of the selected cell. These complementary progression program pulses result in greater precision in programming. For example, in a programming pulse cycle having a CG increment of 10 mV and an EG decrease of 20 mV, the resulting voltage in the FG after the precision programming pulse cycle will be 10 mV, assuming a 40% CG coupling ratio and a 10% EG coupling ratio to the FG. * 40%-20mV * 15%=approximately 1 mV. This complementary pulse programming method can be used in the fine program step 3306 after the coarse program step 3305 because the coarse program step 3305 typically uses only CG or EG increments during its programming operation.

[0171] Further details are now provided for the second and third embodiments of the coarse programming method 3305.

[0172] 37 shows a second embodiment of the coarse programming (adjustment) method 3305, which is an adaptive calibration method 3700. The method starts (step 3701). The cell is programmed by applying an initial voltage v0 to the coarse programming terminal according to one of the sequences shown in Table 9 (step 3702). Unlike the search and execute method 3400, here v0 is not obtained from a lookup table, but instead can be a relatively small initial value. The control gate (or erase gate) voltage of the cell (which may be referred to as CG1 or EG1) is measured at a first current value IR1 (e.g., 100 na) and the voltage on the same gate (which may be referred to as CG2 or EG2) is measured at a second current value IR2 (e.g., 10 na), i.e., in a non-limiting embodiment where IR2 is 10% of IR1, the subthreshold IV slope is determined and stored based on those measurements (e.g., 360 mV / dec or dV / d LOG(I) of current) (step 3703). The IV slope in the linear region is dV / dI.

[0173] New voltage v i The first time this step is performed, i=1, the program voltage v1 is determined based on the stored subthreshold slope value and the current target and offset values, for example, using a subthreshold equation as follows: vi=v i-1 +v increment , V increment is proportional to the slope of Vg Vg=n * Vt * log[Ids / wa * Io] Here, wa is the w of the memory cell, Ids is the cell current, and Io is the cell current when Vg=Vth, Vt is the thermal voltage, and using g from 1 to 2, V1 is determined by the current IR1, and V2 is determined by the current IR2. Slope=(V1-V2) / (LOG(IR1)-LOG(IR2)) In the formula, v increment = α * Tilt *(LOG(IR1)-LOG(I CT )) and I CT is the target current, and α is a predetermined constant (programming offset value)<1 to prevent overshoot, e.g., 0.9.

[0174] If the stored slope value is relatively steep, a relatively small current offset value can be used. If the stored slope value is relatively flat, a relatively high current offset value can be used. Thus, determining the slope information allows a current offset value to be selected that is customized to the particular cell in question. This generally makes the programming process shorter. As step 3704 is repeated, i is incremented and v is incremented. i =v i-1 +v increment Then the cell is i to the coarse adaptation programming terminal. increment Also, v corresponds to the target current value. increment It can also be determined from a look-up table that stores values ​​of

[0175] A read operation is then performed on the selected cell, and the current drawn through the selected cell (I cell ) is measured and a verification operation is performed (step 3705). cell I CT (here the coarse target threshold value), if I CT =I D +I CTOFFSET , I CTOFFSET is an offset value that is added to prevent program overshoot, the adaptive calibration method 3700 is completed, and the fine programming method 3306 can begin. cell I CT If not, then steps 3704-3705 (a new tilt measurement is taken using the new data point) or steps 3703-3705 (if the same tilt previously used is reused) are repeated and i is incremented.

[0176] 38 shows a high level block diagram of a circuit for implementation of method 3703. A current source 3801 is used to apply exemplary current values ​​IR1 and IR2 to a selected cell (here memory cell 3802), then the voltage at the coarse programming terminal of memory cell 3802 (V1 (VCGR1 or VEGR1) for IR1 and V2 (VCGR2 or VEGR2) for IR2) is measured and the coarse programming terminal is selected according to Table 9. The linear slope is (V1-V2) / decade of cell current on the LOGI-V curve, i.e., equal to (V1-V2) / (LOG(IR1)-LOG(IR2)).

[0177] 39 shows a third embodiment of a programming method 3305, which is an adaptive calibration method 3900. The method starts (step 3901). A cell is programmed with a default starting value v0 by applying v0 to the cell's coarse adaptive programming terminal (step 3902). v0 is obtained from a lookup table, such as that created from silicon characterization, and the table values ​​are offset so as not to overshoot the programmed target. An example of v0 is shown in Table 13. Table 13: Initial voltage v0 applied to the course adaptation programming terminal during the adaptation calibration method 3900 [Table 13]

[0178] Step 390 3 In the next step, an IV slope parameter is created for use in predicting the next programming voltage. A first voltage, V1, is applied to the control or erase gate of a selected cell, and the resulting cell current, IR1, is measured. Then, a second voltage, V2, is applied to the control or erase gate of a selected cell, and the resulting cell current, IR2, is measured. A slope is determined based on those measurements and stored, for example, according to the following equation in the subthreshold region (cells operating at subthreshold): Slope=(V1-V2) / (LOG(IR1)-LOG(IR2)) (Step 3903). Example values ​​for V1 and V2 are shown in Table 13 above.

[0179] Determining the IV slope information is customized to the particular cell in question. increment This allows the values ​​to be selected, which generally makes the programming process shorter.

[0180] Each time step 3904 is executed, i is incremented, initially at 0, and the desired programming voltage, v i is determined based on the stored slope values ​​and the current target and offset values ​​using the following equation: v i =v i-1 +v increment , In the formula, v increment = α * Tilt * (LOG(IR1)-LOG(I CT )) I CT is the target current, and α is a predetermined constant (programming offset value)<1 to prevent overshoot, e.g., 0.9.

[0181] The selected cell is then i (Step 3905)

[0182] A read operation is then performed on the selected cell, and the current drawn through the selected cell (I cell ) is measured and a verification operation is performed (step 3906). cell I CT (here the coarse target threshold value), if I CT =I D +I CTOFFSET , I CTOFFSETis an offset value that is added to prevent program overshoot, and the process proceeds to step 3907. If not, the process returns to step 3903 (new slope measurement) or 3904 (previous slope is reused) and i is incremented.

[0183] In step 3907, I cell I CT A smaller threshold value, I CT2 The purpose is to see if an overshoot has occurred. cell I CT It is to be less than cell I CT If it is much lower than I, then overshoot has occurred and the stored value may actually correspond to an incorrect value. cell I CT2 If not, then no overshoot has occurred and the adaptive calibration method 3900 is complete, at which point the process proceeds to the precision programming method 3306. cell I CT2 If so, then an overshoot has occurred. In that case, the selected cell is erased (step 3908) and the programming process resumes at step 3902. Optionally, if step 3908 is performed more than a predetermined number of times, the selected cell may be considered a bad cell that should not be used, and an error signal is output or a flag is set to identify the cell.

[0184] The precision program method 3306 may consist of multiple verify and program cycles, where the pulse width is fixed and the program voltage is incremented by a constant fine voltage to the next pulse, or the program voltage is fixed and the program pulse width is varied.

[0185] Optionally, during a read or verify operation, step 3906 of determining whether a current through a selected non-volatile memory cell is less than or equal to a first threshold current value includes applying a fixed bias to a terminal of the non-volatile memory cell, measuring and digitizing a current drawn by the selected non-volatile memory cell to generate a digital output bit, and digitizing the digital output bit relative to a first threshold current, I CT This can be done by comparing the digital bit representing

[0186] Optionally, during a read or verify operation, step 3907 of determining whether the current through the selected non-volatile memory cell is less than or equal to a second threshold current value includes applying a fixed bias to a terminal of the non-volatile memory cell, measuring and digitizing the current drawn by the selected non-volatile memory cell to generate a digital output bit, and digitizing the digital output bit relative to a second threshold current, I CT2 This can be done by comparing the digital bit representing

[0187] Optionally, steps 3906, 3907 of determining whether the current passing through the selected non-volatile memory cell during a read operation or a verify operation is below a first or second threshold current value, respectively, may be performed by applying an input to a terminal of the non-volatile memory cell, modulating the current drawn by the selected non-volatile memory cell with an output pulse to generate a modulated output, digitizing the modulated output to generate a digital output bit, and comparing the digital output bit to a digital bit representing the first or second threshold current, respectively.

[0188] Measuring the cell current for purposes of verifying or reading the current can be done by averaging multiple measurements, for example 8-32 measurements, to reduce the effects of noise.

[0189] 40 shows a fourth embodiment of the coarse programming method 3305, which is an absolute calibration method 4000. The method starts (step 4001). The relevant terminal of the cell is programmed with a default starting value v0 (step 4002). An example of v0 is shown in Table 14. Table 14: Initial voltages v0 applied to memory cell terminals during absolute calibration method 4000 [Table 14]

[0190] The voltage vTx on the coarse programming terminal is measured and stored (step 4003) at a current value Itarget driven through the cell as described above in connection with FIG. 38. A new coarse programming voltage, v1, is determined (step 4004) based on the stored voltage vTx and an offset value, vToffset (corresponding to Ioffset). For example, the new desired voltage v1 can be calculated as follows: v1=v0+(VTBIAS-vTx)-vToffset, where, for example, VTBIAS=about 1.5V, which is the default terminal voltage at the maximum target current (meaning the maximum current level the memory cell will tolerate). Essentially, the new target voltage is adjusted by an amount that is the difference between the current voltage vTx at the target current and the maximum voltage and offset.

[0191] The cell then i When i=1, the voltage v1 from step 4004 is used. When i>=2, the voltage v i =v i-1 +v increment is used. increment corresponds to the target current value v increment A read operation is then performed on the selected cell, and the current drawn through the selected cell (I cell ) is I CT is compared with (step 4006). cell I CTIf so, the absolute calibration method 4000 is complete and the fine programming method 3306 can begin. cell I CT If not, steps 4005-4006 are repeated and i is incremented.

[0192] Figure 41 shows a circuit 4100 for measuring vTx in step 4003 of the absolute calibration method 4000. vTx is measured at each memory cell 4103 (4103-0, 4103-1, 4103-2,... 4103-n). Here, n + 1 different current sources 4101 (4101-0, 4101-1, 4101-2,... 4101-n) generate different currents IO0, IO1, IO2,... IOn with increasing magnitudes. Each current source 4101 is connected to a respective inverter 4102 (4102-0, 4102-1, 4102-2,... 4102-n) and memory cell 4103 (4103-0, 4103-1, 4103-2,... 4103-n). The input to each inverter 4102 (4102-0, 4102-1, 4102-2,... 4102-n) is initially high, and the output of each inverter is initially low. Since IO0 < IO1 < IO2 <... < IOn, the output of inverter 4102-0 will switch from low to high first because the memory cell 4103-0 draws current from the current source 4101-0 and also draws current from the input node of inverter 4102-0, reducing the input voltage to inverter 4102-0 before the input voltages to the other inverters 4102. Next, the output of inverter 4102-1 switches from low to high, then the output of inverter 4102-2 switches in the same way, and so on until the output of inverter 4102-n switches from low to high. Each inverter 4102 controls a respective switch 4104 (4104-0, 4104-1, 4104-2,... 4104-n). As a result, when the output of the inverter 4102 is high, the switch 4104 is closed, and thereby, vTx is sampled by the capacitor 4105 (4105-0, 4105-1, 4105-2,... 4105-n). Therefore, the switches 4104 and the capacitors 4105 form a sample and hold circuit. In this way, vTx is measured using the sample and hold circuit.

[0193] Figure 42 shows an exemplary progression 4200 for programming a selected cell during the adaptive calibration method 3700 or the absolute calibration method 4000. A voltage VTP (programming voltage applied to the CG or EG terminal, corresponding to vi in ​​step 3704 of Figure 37 and step 4005 of Figure 40) is applied to a terminal of the selected memory cell using a bit line enable signal En_blx (x varies between 1 and n, where n is the number of bit lines).

[0194] Figure 43 shows another exemplary progression 4300 for programming a selected cell during the adaptive calibration method 3700 or the absolute calibration method 4000. A voltage VTP (programming voltage applied to the CG or EG terminal, corresponding to vi in ​​step 3704 of Figure 37 and step 4005 of Figure 40) is applied to a terminal of the selected memory cell using a bit line enable signal En_blx (x varies between 1 and n, where n is the number of bit lines).

[0195] In another embodiment, the voltage applied to the control gate terminal is incremented and the voltage applied to the erase gate terminal is also incremented.

[0196] In another embodiment, the voltage applied to the control gate terminal is incremented and the voltage applied to the erase gate terminal is decreased, as shown in Table 15. Table 15: Control gate terminal increment and erase gate terminal decrement [Table 15]

[0197] For comparison, examples of incrementing only the control gate terminal or incrementing only the erase gate terminal are included in Table 16. Table 16: Control gate terminal increment, erase gate terminal increment [Table 16]

[0198] FIG. 44 shows a system for implementing an input and output method for reading or verifying in a VMM array after precision programming. An input function circuit 4401 receives digital bit values ​​and converts them into analog signals to be used to apply voltages to the control gates of selected cells in the array 4404, which are determined via a control gate decoder 4402. At the same time, a word line decoder 4403 is also used to select the row in which the selected cell is located. An output neuron circuit block 4405 receives output currents from each column of cells in the array 4404. The output circuit block 4405 includes an integrating analog-to-digital converter (ADC), a successive approximation register (SAR) ADC, a sigma-delta ADC, or any other ADC scheme to provide a digital output.

[0199] In one embodiment, the digital value provided to input function circuit 4401 includes four bits (DIN3, DIN2, DIN1, and DIN0), or any number of bits, where the digital value represented by those bits corresponds to the number of input pulses applied to the control gate during a programming operation. More pulses result in a larger value being stored in the cell, which will cause a larger output current when the cell is read. Example input bit values ​​and pulse values ​​are shown in Table 17. Table 17: Digital bit input and number of generated pulses [Table 17]

[0200] In the above example, there are a maximum of 15 pulses for a 4-bit input digital. Each pulse is equal to one unit cell value (current), i.e., the precise programmed current. For example, if Icell unit = 1nA, then DIN[3~0] = 0001, Icell = 1 * 1nA=1nA, and for DIN[3~0]=1111, Icell=15 * 1nA=15nA.

[0201] In another embodiment, the digital bit inputs use digital bit position summation to read out the cell or neuron (e.g., the precisely programmed value of the bit line output) value, as shown in Table 18. Here, only four pulses or four fixed identical bias inputs (e.g., word line or control gate inputs) are needed to evaluate a four-bit digital value. For example, a first pulse or a first fixed bias is used to evaluate DIN0, a second pulse or a second fixed bias with the same value as the first value is used to evaluate DIN1, a third pulse or a third fixed bias with the same value as the first value is used to evaluate DIN2, and a fourth pulse or a fourth fixed bias with the same value as the first value is used to evaluate DIN3. Then, the results from the four pulses are summed according to the bit position, with each output result multiplied (scaled) by a multiplication coefficient that is 2^n (n is the digital bit position), as shown in Table 19. The digital bit summation formula that is realized is the following: Output=2^0 * DIN0+2^1 * DIN1+2^2 * DIN2+2^3 * DIN3) * In Icell units, where Icell represents the precision programmed current.

[0202] For example, if Icell unit = 1nA, DIN[3~0] = 0001, Icell total = 0+0+0+1 * 1nA=1nA, and for DIN[3~0]=1111, Icell total=8 * 1nA+4 * 1nA+2 * 1nA+1 * 1nA=15nA. Table 18: Digital bit input summation [Table 18] Table 19: Summation of digital input bits Dn and 2^n output multiplication coefficient [Table 19]

[0203] Another embodiment having a hybrid input with multiple digital input pulse ranges and sum of input digital ranges is shown in Table 20 for an exemplary 4-bit digital input. In this embodiment, DINn-0 can be divided into m different groups, each group is evaluated, and the output is scaled by a multiplication factor depending on the group binary position. For example, for a 4-bit DIN3-0, the groups can be DIN3-2 and DIN1-0, and the output of DIN1-0 is scaled by 1 (X1) and the output of DIN3-2 is scaled by 4 (X4). Table 20: Summation of Hybrid Inputs and Outputs with Multiple Input Ranges [Table 20]

[0204] Another embodiment combines a hybrid input range with a hybrid supercell. A hybrid supercell includes multiple physical x-bit cells to implement a logical n-bit cell with the x-cell output scaled by 2^n binary positions. For example, two 4-bit cells (cell 1, cell 0) are used to implement an 8-bit logical cell. The output of cell 0 is scaled by 1 (X1) and the output of cell 1 is scaled by 4 (X, 2^2). Other combinations of physical x-cells to implement an n-bit logical cell are possible, such as two 2-bit physical cells and one 4-bit physical cell to implement an 8-bit logical cell.

[0205] FIG. 45 illustrates a digital bit input using digital bit position summation to read out a cell or neuron (e.g., the value of the bit line output) current modulated by a modulator 4510 with an output pulse width designed according to the digital input bit position (e.g., converting the current into an output voltage (V=Current *45 shows another embodiment similar to the system of FIG. 44 except for a bias (applied to the input word line or control gate) for converting the DIN0 input to a pulse width / capacity (pulse width / capacity). For example, a first input bias (applied to the input word line or control gate) is used to evaluate DIN0 and the current (cell or neuron) output is modulated by modulator 4510 with a unit pulse width proportional to the DIN0 bit position, which is in units of 1 (x1), a second input bias is used to evaluate DIN1 and the current output is modulated by modulator 4510 with a pulse width proportional to the DIN1 bit position, which is in units of 2 (x2), a third input bias is used to evaluate DIN2 and the current output is modulated by modulator 4510 with a pulse width proportional to the DIN2 bit position, which is in units of 4 (x4), and a fourth input bias is used to evaluate DIN3 and the current output is modulated by modulator 4510 with a pulse width proportional to the DIN3 bit position, which is in units of 8 (x8). Each output is then converted to a digital bit for each digital input bit DIN0-DIN3 by ADC (analog-to-digital converter) 4511. The total output is then output by adder 4512 as the sum of the four digital outputs generated from the DIN0-3 inputs.

[0206] FIG. 46 shows an example of a charge adder 4600 that can be used to sum the output of the VMM, Icell, during a validation operation or during analog-to-digital conversion of an output neuron to obtain a single analog value representing the output of the VMM, which can then be optionally converted to a digital bit value. The charge adder 4600 can be used, for example, as the adder 4512. The charge adder 4600 includes a current source 4601 (here representing the current Icell output by the VMM) and a sample and hold circuit including a switch 4602 and a sample and hold (S / H) capacitor 4603. The example shown utilizes a 4-bit digital value for the output, but other numbers of bits can be used instead. There are four S / H circuits to hold the values ​​generated from the four evaluation pulses, and these values ​​are summed at the end of the process. The S / H capacitor 4603 is a 2^n of the S / H capacitor.* DINn bit position. For example, switch 4602 of C_DIN3 is closed when Icell>8×current threshold, switch 4602 of C_DIN2 is closed when Icell>4×current threshold, switch 4602 of C_DIN1 is closed when Icell>2×current threshold, and switch 4602 of C_DIN0 is closed when Icell>current threshold. Thus, the digital value stored by sample and hold capacitor 4603 reflects the value of Icell 4601.

[0207] FIG. 47 shows a current adder 4700 that can be used to sum the output of the VMM, Icell during a verification operation or during analog-to-digital conversion of an output neuron. The charge adder 4700 can be used, for example, as adder 4512. The current adder 4700 includes a current source 4701 (here representing the Icell output from the VMM), a switch 4702, switches 4703 and 4704, and a transistor 4705. The example shown utilizes a 4-bit digital value of the output, where the bit values ​​are represented by currents I_DIN0, I_DIN1, I_DIN2, and I_DIN3. The bit position of each transistor 4705 affects the value represented by that bit. Switch 4703 for I_DIN3 is closed when Icell>8×current threshold, switch 4703 for I_DIN2 is closed when Icell>4×current threshold, switch 4704 for I_DIN1 is closed when Icell>2×current threshold, and switch 4703 for I_DIN0 is closed when Icell>current threshold. Thus, the digital value output by transistor 4705 (where a “1” is represented by a positive current and a “0” is represented by no current, or vice versa) reflects the value of Icell 4601.

[0208] FIG. 48 shows a digital adder 4800 that receives multiple digital values ​​and sums them together to generate an output DOUT that represents the sum of the inputs. The digital adder 4600 can be used, for example, as adder 4512. The digital adder 4800 can be used during a validation operation or during analog-to-digital conversion of an output neuron. As shown in the example of a four-bit digital value, there are digital output bits to hold the values ​​from the four evaluation pulses, which are summed at the end of the process. The digital output is a sum of 2^n * It is digitally scaled based on the DINn bit position, e.g., DOUT3=x8 DOUT0, _DOUT2=x4 DOUT1, I_DOUT1=x2 DOUT0, I_DOUT0=DOUT0.

[0209] FIG. 49A shows a dual slope integrating ADC 4900 applied to an output neuron to convert the cell current to a digital output bit. An integrator consisting of an integrating op-amp 4901 and an integrating capacitor 4902 integrates the cell current ICELL against a reference current IREF. As shown in FIG. 49B, for a fixed time t1, switch S1 is closed and switch S2 is open, the cell current is up-integrated (Vout rises in waveform 4950), switch S1 is opened and switch S2 is closed, so that the reference current IREF is applied to be down-integrated at time t2 (Vout falls in waveform 4950). The value of the current Icell is =t2 / t1 * The target value VREF is applied to comparator 4904, and the output EC4905 of comparator 4904 can be used as a trigger to determine the number of cycles IREF is applied until VOUT falls below VREF. For example, for t1, with 10-bit digital bit resolution, 1024 cycles are used, and the number of cycles for t2 varies from 0 to 1024 cycles depending on the Icell value. When a target value is applied to comparator 4904 as VREF, the output EC4905 of comparator 4904 can be used as a trigger to determine the number of cycles IREF is applied until VOUT falls below VREF.

[0210] FIG. 49C shows a single slope integrating ADC 4960 applied to an output neuron 4966, ICELL, to convert the cell current into a digital output bit. teeth, An integrating op-amp 4961, an integrating capacitor 4962, an op-amp 4964, and switches S1 and S3 Includes The integrating opamp 4961 and integrating capacitor 4962 integrate the output neuron current, ICELL. As shown in FIG. 49D, during time t1, the cell current is up-integrated (Vout rises until it reaches Vref2), and during time t2, which starts simultaneously with time t1 but is greater than time t1, the cell current of the reference cell is up-integrated. The cell current ICELL is:=Cint * Vref2 / t. A pulse counter coupled to the output of the comparator 4965 is used to count the number of pulses (digital output bits) during each integration time t1, t2. For example, as shown, the digital output bits for t1 are less than the digital output bits for t2, which means that the cell current during t1 is greater than the cell current during t2. An initial calibration is performed to calibrate the integration capacitor value with a reference current and fixed time, Cint=Tref * Iref / Vref2.

[0211] FIG. 49E shows a dual slope integrating ADC 4980 applied to an output neuron 4984, ICELL, to convert the cell current to a digital output bit. The dual slope integrating ADC 4980 includes switches S1, S2, and S3, an operational amplifier 4981, a capacitor 4982, and a reference current source 4983. The dual slope integrating ADC 4980 does not utilize an integrating operational amplifier. The cell current or the reference current is directly integrated on the capacitor 4982. A pulse counter is used to count the pulses (digital output bits) during the integration time. The current Icell is given by:=t2 / t1 * It is an IREF.

[0212] FIG. 49F shows a single slope integrating ADC 4990 applied to output neuron 4994, ICELL, to convert the cell current to a digital output bit. The single slope integrating ADC 4990 includes switches S2 and S3, an opamp 4991, and a capacitor 4992. The single slope integrating ADC 4980 does not utilize an integrating opamp. The cell current is directly integrated on capacitor 4992. A pulse counter is used to count the pulses (digital output bits) during the integration time. Cell current Icell=Cint * Vref2 / t.

[0213] FIG. 50A shows a SAR (successive approximation register) ADC applied to the output neuron to convert the cell current to a digital output bit. The cell current can be dropped through a resistor to convert it to a voltage VCELL. Alternatively, the cell current can charge up a S / H capacitor to convert the cell current to a voltage VCELL. VCELL is provided to the inverting input of a comparator 5003, the output of which is fed to the selection input of the SAR 5001. A clock input CLK is also provided to the SAR 5001. A binary search is used to calculate the bits starting from the MSB bit (most significant bit). Based on the digital bit DN-D0 output from the SAR 5001 and received as an input to the DAC 5002, the output of the DAC 5002 is used to set the non-inverting input of the comparator 5003, i.e. the appropriate analog reference voltage for the comparator 5003. The output of the comparator 5003 is in turn fed back to the SAR 5001 to select the next analog level. As shown in FIG. 50B, in an example of four digital output bits, there are four evaluation periods, including, without limitation, a first pulse to evaluate DOUT3 by setting the analog level to the middle, then a second pulse to evaluate DOUT2 by setting the analog level to the middle of the upper half or the middle of the lower half.

[0214] Modified binary search such as cyclic (algorithmic) ADCs can be used for cell tuning (e.g., programming) verification or output neuron conversion. Modified binary search such as switched-cap (SC) charge redistribution ADCs can be used for cell tuning (e.g., programming) verification or output neuron conversion.

[0215] FIG. 51 shows a sigma-delta ADC 5100 applied to an output neuron to convert the cell currents to digital output bits. An integrator consisting of an opamp 5101 and a capacitor 5105 integrates the sum of a current ICELL from a selected cell current 5106 and a reference current IREF coming from a 1-bit current cDAC 5104. A comparator 5102 compares the integrated output voltage of the opamp 5101 against a reference voltage, VREF2. A clocked DFF 5103 provides a digital output stream according to the output of the comparator 5102 received at the D input of DFF 5103. The digital output stream typically goes to a digital filter before being output as a digital output bit.

[0216] FIG. 52A shows a ramp-type analog-to-digital converter 5200 including a current source 5201 (representing the received neuron current ICELL), a switch 5202, a variable configurable capacitor 5203, and a comparator 5204 that receives as a non-inverting input the voltage across the variable configurable capacitor 5203, labeled Vneu, and as an inverting input a configurable reference voltage Vreframp, and generates an output Cout. Vreframp is ramped up at discrete levels every comparison clock cycle. The comparator 5204 compares Vneu with Vreframp, resulting in an output Cout that is "1" when Vneu>Vreframp and "0" otherwise. Thus, the output Cout is a pulse whose width varies in response to Ineu. The larger Ineu is, the longer the period during which Cout is "1", resulting in a wider pulse width at the output Cout. A digital counter 5220 converts each pulse 522 at the output Cout into a count value 5221, which is a digital output bit, as shown in FIG. 52B, for two different ICELL currents, marked OT1A and OT2A, respectively.

[0217] Alternatively, the ramp voltage Vreframp is a continuous ramp voltage 5255 shown in graph 5250 of FIG. 52B.

[0218] Alternatively, a multiple ramp embodiment is shown in FIG. 52C to reduce conversion time by utilizing a coarse-fine ramp conversion algorithm. First, a coarse reference ramp reference voltage 5271 is ramped quickly to determine each ICELL subrange. Then, a fine reference ramp reference voltage 5272 for each subrange, Vreframp1 and Vreframp2, respectively, is used to convert the ICELL current in each subrange. As shown, there are two subranges of the fine reference ramp voltage. More than one coarse / fine step or two subranges are possible.

[0219] FIG. 53 shows an algorithmic analog-to-digital output converter 5300 including a switch 5301, a switch 5302, a sample and hold (S / H) circuit 5303, a 1-bit analog-to-digital converter (ADC) 5304, a 1-bit digital-to-analog converter (DAC) 5305, a summer 5306, and a gain of 2 residue operational amplifier (2x opamp) 5307. The algorithmic analog-to-digital output converter 5300 generates a converted digital output 5308 in response to an analog input Vin and control signals applied to the switches 5302 and 5303. An input received at the analog input Vin (e.g., Vneu in FIG. 52) is first sampled in the S / H circuit 5303 in response to the switch 5302, and then a conversion is performed in N clock cycles for N bits. For each conversion clock cycle, the 1-bit ADC 5304 compares the S / H voltage 5309 with a reference voltage VREF / 2 and outputs a digital bit (e.g., "0" if the input ≦ VREF / 2, and "1" if the input > VREF / 2). This digital output bit, which is the digital output signal 5308, is then converted to an analog voltage (e.g., either VREF / 2 or 0V) by the 1-bit DAC 5305 and fed to the summer 5306 to be subtracted from the S / H voltage 5309. The 2× residue opamp 5307 then amplifies the summer difference voltage output to a converted residue voltage 5310, which is fed to the S / H circuit 5303 via the switch 5301 for the next clock cycle. Instead of this 1-bit (i.e., 2-level) algorithmic ADC, a 1.5-bit (i.e., 3-level) algorithmic ADC can be used to reduce the effects of offsets from the ADC 5304 and the residue opamp 5307, etc. For use with a 1.5-bit algorithmic ADC, a 1.5-bit or 2-bit (ie, 4-level) DAC is preferred. In another embodiment, a hybrid ADC can be used, for example for a 9-bit ADC, the first 4 bits can be generated by a SAR ADC and the remaining 5 bits can be generated using a gradient or ramp ADC. [Programming and verifying multiple physical cells as a single logical multi-bit cell]

[0220] The programming and verify devices and methods described above are capable of operating simultaneously with multiple physical cells as a logical multi-bit cell.

[0221] FIG. 54 illustrates a logical multi-bit cell 5400 that includes i physical cells, labeled physical cells 5401-1, 5401-2, ..., 5401-i. In one embodiment, the physical cells 5401 have uniform diffusion widths (transistor widths). In another embodiment, the physical cells 5401 have non-uniform diffusion widths (different transistor widths, transistors with larger widths can store more levels and therefore a larger number of bits). In both embodiments, the physical cells 5401 are programmed, verified, and read as a unit, specifically as a single logical n-bit cell that can store more levels than each of the m-bit cells. For example, if m=2, then each physical cell 5401 can hold one of four levels (L0, L1, L2, L3). Two such cells can be treated as a single logical cell with n=3, so that a single logical cell can hold one of eight levels (L0, L1, L2, L3, L4, L5, L6, L7). As another example, for m=3, each physical cell 5401 can hold one of eight levels (L0, ..., L7). Four such cells can be treated as a single logical cell with n=5, so that a single logical cell can hold one of 32 levels (L0, ..., L31).

[0222] 55 shows a method 5500 of programming logic multi-bit cell 5400. First, j of the i physical cells 5401-1, ..., 5401-i, where j≦i, are programmed and verified (step 5501) using one of the coarse programming methods 3305 until the coarse current target for j physical cells is achieved. Next, k of the j physical cells, where k≦j, are programmed and verified (step 5502) using one of the fine programming methods 3306 until the fine current target for k physical cells is achieved.

[0223] The method 5500 may be performed on more than one subset of the i physical cells 5401-1, . . . 5401-i to achieve a desired overall level of the logical multi-bit cell 5400.

[0224] For example, if i=4, then there are four cells, 5401-1, 5401-2, 5401-3, and 5401-4. If we assume that each cell can hold one of eight different levels, then logical multi-bit cell 5400 can hold one of 32 different levels. If the desired programming value is L27, then that level (corresponding to the desired read current) can be achieved in any number of different ways.

[0225] For example, method 5500 can be performed on cells 5401-1, 5401-2, and 5401-3 until those cells collectively hold L23 (24th level), and then method 5500 is performed on cell 5401-4 to program that cell to the 4th level such that logical multi-bit cell 5400 achieves L27 (28th level).

[0226] As another example, method 5500 can be performed on cells 5401-1, 5401-2, 5401-3, and 5401-4 until those cells collectively hold L25 (26th level), and then method 5500 can be performed only on cell 5401-4 until the entire logical multi-bit cell 5400 stores a value that achieves L27 (28th level).

[0227] Other approaches are possible and method 5500 can be performed on different subsets of i physical cells until the desired level is achieved.

[0228] In another embodiment, in a situation where i physical cells have non-uniform diffusion widths, a coarse programming step 3305 can be performed with j1 physical cells having wider transistor widths until the j1 physical cells collectively achieve the coarse current target, and then a fine programming step 3306 can be performed with the j2 physical cells having the smallest transistor widths until the j1+j2 physical cells collectively achieve the fine current target.

[0229] It should be noted that, as used herein, both the terms "over" and "on" are inclusive of "directly on" (with no intermediate material, element, or gap disposed between them) and "indirectly on" (with an intermediate material, element, or gap disposed between them). Similarly, the term "adjacent" includes "directly adjacent" (with no intermediate material, element, or gap disposed between them) and "indirectly adjacent" (with an intermediate material, element, or gap disposed between them), "attached" includes "directly attached" (with no intermediate material, element, or gap disposed between them) and "indirectly attached to" (with an intermediate material, element, or gap disposed between them), and "electrically coupled" includes "directly electrically coupled" (without an intermediate material or element disposed between them that electrically connects the elements together) and "indirectly electrically coupled to" (with an intermediate material or element disposed between them that electrically connects the elements together). For example, forming an element "above a substrate" can include forming the element directly on the substrate, with no intermediate materials / elements between them, and forming the element indirectly on the substrate, with one or more intermediate materials / elements between them.

Claims

1. 1. A method of programming a selected non-volatile memory cell to store one of N possible values, where N is an integer greater than two, the selected non-volatile memory cell including a floating gate, a control gate terminal, an erase gate terminal, and a source line terminal, the method comprising: performing a first programming process including a plurality of program verify cycles, wherein programming durations of increasing duration are applied to terminals of the selected non-volatile memory cells in each program verify cycle after a first program verify cycle; Each program verification cycle: applying a first voltage to one of the erase gate terminal and the control gate terminal of the selected non-volatile memory cell; measuring a first current through the selected non-volatile memory cell as a result of the application; applying a second voltage to the one of the erase gate terminal and the control gate terminal of the selected non-volatile memory cell; measuring a second current through the selected non-volatile memory cell as a result of the application; determining a slope value based on the first voltage, the second voltage, the first current, and the second current; determining a next programming duration based on the slope value; programming the non-volatile memory cell with the next programming duration; The method includes repeating the steps of determining a next programming duration and programming the non-volatile memory cell with the next programming duration until a current passing through the selected non-volatile memory cell during a read or verify operation is below a first threshold current value.

2. 1. A method of programming a selected non-volatile memory cell to store one of N possible values, where N is an integer greater than two, the selected non-volatile memory cell including a floating gate, a control gate terminal, an erase gate terminal, and a source line terminal, the method comprising: performing a first programming process including a plurality of program verify cycles, wherein programming durations of increasing duration are applied to terminals of the selected non-volatile memory cells in each program verify cycle after a first program verify cycle; if a current through the selected non-volatile memory cell during a read or verify operation is less than or equal to a first threshold current value, performing a second programming process until the current through the selected non-volatile memory cell during the read or verify operation is less than or equal to a second threshold current value; the second programming process includes applying voltage pulses of increasing duration to the control gates of the selected non-volatile memory cells; The method, wherein the second programming process further comprises applying voltage pulses of increasing duration to the erase gates of the selected non-volatile memory cells.

3. 1. A method of programming a selected non-volatile memory cell to store one of N possible values, where N is an integer greater than two, the selected non-volatile memory cell including a floating gate, a control gate terminal, an erase gate terminal, and a source line terminal, the method comprising: performing a first programming process including a plurality of program verify cycles, wherein programming durations of increasing duration are applied to terminals of the selected non-volatile memory cells in each program verify cycle after a first program verify cycle; if a current through the selected non-volatile memory cell during a read or verify operation is less than or equal to a first threshold current value, performing a second programming process until the current through the selected non-volatile memory cell during the read or verify operation is less than or equal to a second threshold current value; the second programming process includes applying voltage pulses of increasing duration to the control gates of the selected non-volatile memory cells; The method, wherein the second programming process further comprises applying voltage pulses of decreasing duration to the erase gates of the selected non-volatile memory cells.

4. 1. A method of programming a selected non-volatile memory cell to store one of N possible values, where N is an integer greater than two, the selected non-volatile memory cell including a floating gate, a control gate terminal, an erase gate terminal, and a source line terminal, the method comprising:

11. A method comprising: performing a first programming process including a plurality of program verify cycles, during which a first programming voltage of increasing magnitude is applied to the control gate of the selected non-volatile memory cell and a second programming voltage of decreasing magnitude is applied to the erase gate of the selected non-volatile memory cell.

5. 5. The method of claim 4, wherein each program verify cycle includes verifying that a current through the selected non-volatile memory cell during a read or verify operation is less than or equal to a first threshold current value.

6. Each program verification cycle: applying a first voltage to one of the erase gate and the control gate of the selected non-volatile memory cell; measuring a first current through the selected non-volatile memory cell as a result of the application; applying a second voltage to the one of the erase gate and the control gate of the selected non-volatile memory cell; measuring a second current through the selected non-volatile memory cell as a result of the application; determining a slope value based on the first voltage, the second voltage, the first current, and the second current; determining a next programming voltage for the programming voltage increasing in magnitude and a next programming voltage for the programming voltage decreasing in magnitude, respectively, based on the slope value; programming the non-volatile memory cells with the next programming voltage; 5. The method of claim 4, further comprising repeating the steps of determining a next programming voltage and programming the non-volatile memory cell with the next programming voltage until a current through the selected non-volatile memory cell during a read or verify operation is below a first threshold current value.

7. 5. The method of claim 4, further comprising: performing a second programming process until a current through the selected non-volatile memory cell during a read or verify operation is below a second threshold current value.

8. The step of performing a first programming process includes:

5. The method of claim 4, further comprising the step of erasing the selected non-volatile memory cell and repeating the first programming process if the current through the selected non-volatile memory cell is less than or equal to a third threshold current value.

9. 8. The method of claim 7, further comprising: performing a third programming process until a current through the selected non-volatile memory cell during a read or verify operation is below a fourth threshold current value.

10. 1. A method of programming a selected non-volatile memory cell to store one of N possible values, where N is an integer greater than two, the selected non-volatile memory cell including a floating gate, a control gate terminal, an erase gate terminal, and a source line terminal, the method comprising:

11. A method comprising: performing a first programming process including a plurality of program verify cycles, in which first programming voltages of increasing duration are applied to the control gate terminals of the selected non-volatile memory cells and second programming voltages of decreasing duration are applied to the erase gate terminals of the selected non-volatile memory cells.

11. 11. The method of claim 10, wherein each program verify cycle includes verifying that a current through the selected non-volatile memory cell during a read or verify operation is less than or equal to a first threshold current value.

12. Each program verification cycle: applying a first voltage to one of the erase gate and the control gate of the selected non-volatile memory cell; measuring a first current through the selected non-volatile memory cell as a result of the application; applying a second voltage to the one of the erase gate and the control gate of the selected non-volatile memory cell; measuring a second current through the selected non-volatile memory cell as a result of the application; determining a slope value based on the first voltage, the second voltage, the first current, and the second current; determining next programming durations for the programming voltages of increasing duration and for the programming voltages of decreasing duration based on the slope values; programming the non-volatile memory cell with the next programming duration; 11. The method of claim 10, further comprising: repeating the steps of determining a next programming duration and programming the non-volatile memory cell with the next programming duration until a current through the selected non-volatile memory cell during a read or verify operation is below a first threshold current value.

13. 11. The method of claim 10, further comprising: performing a second programming process until a current through the selected non-volatile memory cell during a read or verify operation is below a second threshold current value.

14. The step of performing a first programming process includes:

11. The method of claim 10, further comprising the step of erasing the selected non-volatile memory cell and repeating the first programming process if the current through the selected non-volatile memory cell is less than or equal to a third threshold current value.

15. 14. The method of claim 13, further comprising: performing a third programming process until a current through the selected non-volatile memory cell during a read or verify operation is below a fourth threshold current value.

16. 1. A method of programming a selected non-volatile memory cell to store one of N possible values, where N is an integer greater than two, the selected non-volatile memory cell including a floating gate, a control gate, an erase gate, and a source line terminal, the method comprising:

11. A method comprising: performing a first programming process including a plurality of program verify cycles, in which a first programming voltage of increasing magnitude is applied to the erase gate of the selected non-volatile memory cell and a second programming voltage of decreasing magnitude is applied to the control gate of the selected non-volatile memory cell.

17. 17. The method of claim 16, wherein each program verify cycle includes verifying that a current through the selected non-volatile memory cell during a read or verify operation is less than or equal to a first threshold current value.

18. Each program verification cycle: applying a first voltage to one of the erase gate and the control gate of the selected non-volatile memory cell; measuring a first current through the selected non-volatile memory cell as a result of the application; applying a second voltage to the one of the erase gate and the control gate of the selected non-volatile memory cell; measuring a second current through the selected non-volatile memory cell as a result of the application; determining a slope value based on the first voltage, the second voltage, the first current, and the second current; determining a next programming voltage for the programming voltage increasing in magnitude and a next programming voltage for the programming voltage decreasing in magnitude, respectively, based on the slope value; programming the non-volatile memory cells with the next programming voltage; 17. The method of claim 16, further comprising repeating the steps of determining a next programming voltage and programming the non-volatile memory cell with the next programming voltage until a current through the selected non-volatile memory cell during a read or verify operation is below a first threshold current value.

19. 20. The method of claim 18, further comprising: performing a second programming process until a current through the selected non-volatile memory cell during a read or verify operation is below a second threshold current value.

20. The step of performing a first programming process includes:

20. The method of claim 18, further comprising the step of erasing the selected non-volatile memory cell and repeating the first programming process if the current through the selected non-volatile memory cell is less than or equal to a third threshold current value.

21. 20. The method of claim 19, further comprising: performing a third programming process until a current through the selected non-volatile memory cell during a read or verify operation is below a fourth threshold current value.

22. 1. A method of programming a selected non-volatile memory cell to store one of N possible values, where N is an integer greater than two, the selected non-volatile memory cell including a floating gate, a control gate, an erase gate, and a source line terminal, the method comprising:

11. A method comprising: performing a first programming operation including a plurality of program verify cycles, in which a first programming voltage of increasing duration is applied to the erase gate of the selected non-volatile memory cell and a second programming voltage of decreasing duration is applied to the control gate of the selected non-volatile memory cell.

23. 23. The method of claim 22, wherein each program verify cycle includes verifying that a current through the selected non-volatile memory cell during a read or verify operation is less than or equal to a first threshold current value.

24. Each program verification cycle: applying a first voltage to one of the erase gate and the control gate of the selected non-volatile memory cell; measuring a first current through the selected non-volatile memory cell as a result of the application; applying a second voltage to the one of the erase gate and the control gate of the selected non-volatile memory cell; measuring a second current through the selected non-volatile memory cell as a result of the application; determining a slope value based on the first voltage, the second voltage, the first current, and the second current; determining next programming durations for the programming voltages of increasing duration and for the programming voltages of decreasing duration based on the slope values; programming the non-volatile memory cells with the next programming duration; 24. The method of claim 23, further comprising: repeating the steps of determining a next programming duration and programming the non-volatile memory cell with the next programming duration until a current through the selected non-volatile memory cell during a read or verify operation is below a first threshold current value.

25. 23. The method of claim 22, further comprising: performing a second programming process until a current through the selected non-volatile memory cell during a read or verify operation is below a second threshold current value.

26. The step of performing a first programming process includes:

23. The method of claim 22, further comprising the step of erasing the selected non-volatile memory cell and repeating the first programming process if the current through the selected non-volatile memory cell is less than or equal to a third threshold current value.

27. 26. The method of claim 25, further comprising: performing a third programming process until a current through the selected non-volatile memory cell during a read or verify operation is below a fourth threshold current value.

28. 1. A method of programming a selected non-volatile memory cell to store one of N possible values, N being an integer greater than two, the selected non-volatile memory cell including a first gate, a first terminal, and a second terminal, the method comprising: performing a plurality of program verify cycles, during each of the program verify cycles a first programming voltage is applied to the first terminal of the selected non-volatile memory cell, the first programming voltage decreasing with each subsequent cycle; The method, wherein during each of the program verify cycles, a second programming voltage is applied to the second terminal of the selected non-volatile memory cell, the second programming voltage increasing with each subsequent cycle.

29. 30. The method of claim 28, wherein the selected non-volatile memory cells are split-gate memory cells.

30. 30. The method of claim 29, wherein the first gate is a floating gate.

31. 31. The method of claim 30, wherein the first terminal is a source line terminal.

32. 32. The method of claim 31 , wherein the second terminal is a control gate terminal.

33. 31. The method of claim 30, wherein the first terminal is an erase gate terminal.

34. 34. The method of claim 33, wherein the second terminal is a control gate terminal.

Citation Information

Patent Citations

  • Non-volatile semiconductor memory

    JP1995169284A

  • Nonvolatile semiconductor memory device

    JP2001067884A

  • Efficient verification for coarse / fine programming of non-volatile memory

    JP2007520845A

  • Deep Learning Neural Network Classifier Using Non-Volatile Memory Arrays

    JP2019517138A

  • High precision and highly efficient tuning mechanisms and algorithms for analog neuromorphic memory in artificial neural networks

    WO2019108334A1