Precise programming method and apparatus for analog neural memory in artificial neural network

A precision programming method and apparatus address the challenge of precise charge deposition on floating gates in non-volatile memory cells, improving the performance and efficiency of analog neuromorphic memories by enabling high-precision programming in neural networks.

JP2025124641AActive Publication Date: 2025-08-26SILICON STORAGE TECHNOLOGY INC

Patent Information

Application Number
JP2025075177
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-01-23
Filing Date
2025-04-30
Publication Date
2025-08-26
Estimated Expiration
2040-05-22

AI Technical Summary

Technical Problem

Existing artificial neural networks face challenges in programming non-volatile memory cells in vector matrix multiplication arrays with the required precision and granularity, particularly in analog neuromorphic memories, due to the need for precise and rapid charge deposition on floating gates.

Method used

A precision programming algorithm and apparatus are developed to accurately and swiftly deposit precise amounts of charge onto the floating gates of non-volatile memory cells in vector matrix multiplication arrays, enabling high-precision programming of selected cells to hold one of N different values.

Benefits of technology

This solution allows for precise and rapid programming of non-volatile memory cells, enhancing the performance and efficiency of analog neuromorphic memories by enabling fine-tuning of synaptic weights in neural networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025124641000001_ABST
    Figure 2025124641000001_ABST
Patent Text Reader

Abstract

To provide a method for programming and verifying multiple physical cells as a single logical multi-bit cell.SOLUTION: There is provided a method of programming a logical multi-bit cell 5400 in which physical cells 5401 are programmed, verified, and read as a single logical n-bit cell that can store more levels than each of m-bit cells. First, j of i physical cells 5401-1,..., 5401-i (where j≤i) are programmed and verified using a coarse programming method until a coarse current target for the j physical cells is achieved. Next, k out of the j physical cells (where k≤j), are programmed and verified using a precision programming method until a precision current target is achieved for k physical cells.SELECTED DRAWING: Figure 55
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] (Priority Claim) This application claims priority to U.S. Provisional Patent Application No. 62 / 933,809, filed November 11, 2019, entitled "PRECISE PROGRAMMING METHOD AND APPARATUS FOR ANALOG NEURAL MEMORY IN A DEEP LEARNING ARTIFICIAL NEURAL NETWORK," and to U.S. Patent Application No. 16 / 751,202, filed January 23, 2020, entitled "PRECISE PROGRAMMING METHOD AND APPARATUS FOR ANALOG NEURAL MEMORY IN A DEEP LEARNING ARTIFICIAL NEURAL NETWORK."

[0002] FIELD OF THE INVENTION Numerous embodiments are disclosed for a precision programming algorithm and apparatus for precisely and quickly depositing precise amounts of charge onto the floating gates of non-volatile memory cells in vector matrix multiplication (VMM) arrays within artificial neural networks. [Background technology]

[0003] Artificial neural networks mimic biological neural networks (the central nervous systems of animals, particularly the brain) and are used to estimate or approximate functions that may depend on multiple inputs and are generally unknown. Artificial neural networks generally contain layers of interconnected "neurons" that exchange messages.

[0004] Figure 1 shows an artificial neural network, where circles represent layers of inputs or neurons. Connections (called synapses) are represented by arrows and have numerical weights that can be adjusted based on experience. This allows the artificial neural network to adapt to the inputs and learn. Typically, an artificial neural network contains multiple layers of inputs. There are typically one or more hidden layers of neurons, and an output layer of neurons that provide the neural network's output. At each level, neurons make decisions, individually or collectively, based on the data they receive from the synapses.

[0005] One of the major challenges in developing artificial neural networks for high-performance information processing is the lack of suitable hardware technology. In practice, practical artificial neural networks rely on a very large number of synapses, which allows for high connectivity between neurons and therefore a very high degree of parallelization of computational processes. In principle, such complexity could be achieved using digital supercomputers or dedicated graphic processing unit clusters. However, in addition to high costs, these approaches also suffer from poor energy efficiency compared to biological networks, which primarily perform low-precision analog computations and therefore consume much less energy. While CMOS analog circuits have been used in artificial neural networks, most CMOS-implemented synapses are too bulky given the large number of neurons and synapses.

[0006] The applicant previously disclosed an artificial (analog) neural network utilizing one or more non-volatile memory arrays as synapses in U.S. Patent Application No. 15 / 594,439, published as U.S. Patent Publication No. 2017 / 0337466, which is incorporated by reference. The non-volatile memory array operates as an analog neuromorphic memory. As used herein, the term neuromorphic refers to a circuit that implements a model of a nervous system. The analog neuromorphic memory includes a first plurality of synapses configured to receive a first plurality of inputs and generate a first plurality of outputs therefrom, and a first plurality of neurons configured to receive the first plurality of outputs. The first plurality of synapses include a plurality of memory cells, each including spaced apart source and drain regions formed in a semiconductor substrate with a channel region extending therebetween, a floating gate disposed above and insulated from a first portion of the channel region, and a non-floating gate disposed above and insulated from a second portion of the channel region. Each of the plurality of memory cells is configured to store a weight value corresponding to a number of electrons on the floating gate. The plurality of memory cells are configured to multiply a first plurality of inputs by the stored weight values ​​to generate a first plurality of outputs. An array of memory cells arranged in this manner may be referred to as a vector matrix multiplication (VMM) array.

[0007] Each nonvolatile memory cell used in an analog neuromorphic memory array must store a very specific and precise amount of charge, i.e., the number of electrons, in its floating gate in response to erasure and programming. For example, each floating gate must store one of N different values, where N is the number of different weights that can be exhibited by each cell. Examples of N include 16, 32, 64, 128, and 256. One challenge in analog neuromorphic memory systems is the ability to program selected cells to the N different values ​​with the required precision and granularity.

[0008] What is needed is an improved programming system and method suitable for use with VMM arrays in analog neuromorphic memories. Summary of the Invention

[0009] Numerous embodiments are disclosed for precision programming algorithms and apparatus for precisely and rapidly depositing precise amounts of charge onto the floating gates of non-volatile memory cells in vector matrix multiplication (VMM) arrays in analog neuromorphic memories, such that selected cells can be programmed with extremely high precision to hold one of N different values.

[0010]

[0011]

[0012]

[0013]

[0014]

[0015]

[0016]

[0017]

[0018]

[0019]

[0020]

[0021]

[0022]

[0023]

[0024]

[0025]

[0026]

[0027]

[0028]

[0029]

[0030]

[0031]

[0032]

[0033]

[0034]

[0035]

[0036]

[0037]

[0038]

[0039]

[0040]

[0041]

[0042]

[0043]

[0044]

[0045]

[0046]

[0047]

[0048]

[0049]

[0050]

[0051]

[0052]

[0053]

[0054]

[0055]

[0056]

[0057]

[0058]

[0059]

[0060]

[0061]

[0062]

[0063]

[0064]

[0065]

[0066]

[0067]

[0068]

[0069]

[0070] [Brief explanation of the drawings]

[0071] [Figure 1] FIG. 1 illustrates a prior art artificial neural network. [Figure 2] 1 shows a prior art split-gate flash memory cell. [Figure 3] 1 illustrates another prior art split-gate flash memory cell. [Figure 4] 1 illustrates another prior art split-gate flash memory cell. [Figure 5]1 illustrates another prior art split-gate flash memory cell. [Figure 6] 1 illustrates another prior art split-gate flash memory cell. [Figure 7] 1 shows a prior art stacked gate flash memory cell. [Figure 8] FIG. 1 illustrates various levels of an exemplary artificial neural network that utilizes one or more non-volatile memory arrays. [Figure 9] FIG. 1 is a block diagram illustrating a vector matrix multiplication system. [Figure 10] FIG. 1 is a block diagram illustrating an example artificial neural network utilizing one or more vector-matrix multiplication systems. [Figure 11] 1 illustrates another embodiment of a vector matrix multiplication system. [Figure 12] 1 illustrates another embodiment of a vector matrix multiplication system. [Figure 13] 1 illustrates another embodiment of a vector matrix multiplication system. [Figure 14] 1 illustrates another embodiment of a vector matrix multiplication system. [Figure 15] 1 illustrates another embodiment of a vector matrix multiplication system. [Figure 16] 1 illustrates another embodiment of a vector matrix multiplication system. [Figure 17] 1 illustrates another embodiment of a vector matrix multiplication system. [Figure 18] 1 illustrates another embodiment of a vector matrix multiplication system. [Figure 19] 1 illustrates another embodiment of a vector matrix multiplication system. [Figure 20] 1 illustrates another embodiment of a vector matrix multiplication system. [Figure 21] 1 illustrates another embodiment of a vector matrix multiplication system. [Figure 22] 1 illustrates another embodiment of a vector matrix multiplication system. [Figure 23] 1 illustrates another embodiment of a vector matrix multiplication system. [Figure 24]1 illustrates another embodiment of a vector matrix multiplication system. [Figure 25] 1 shows a prior art long-term memory system. [Figure 26] An exemplary cell for use in a long-term memory system is shown. [Figure 27] 27 illustrates one embodiment of the exemplary cell of FIG. 26. [Figure 28] 27 illustrates another embodiment of the exemplary cell of FIG. 26. [Figure 29] 1 shows a prior art gated recurrent unit system. [Figure 30] 1 shows an exemplary cell for use in a gated recurrent unit system. [Figure 31] 31 illustrates one embodiment of the exemplary cell of FIG. 30. [Figure 32] 31 illustrates another embodiment of the exemplary cell of FIG. 30. [Figure 33A] 1 illustrates one embodiment of a method for programming a non-volatile memory cell. [Figure 33B] 1 illustrates another embodiment of a method for programming a non-volatile memory cell. [Figure 34] 1 illustrates one embodiment of a coarse programming method. [Figure 35] 1 illustrates exemplary pulses used in programming non-volatile memory cells. [Figure 36A] 1 illustrates exemplary pulses used in programming non-volatile memory cells. [Figure 36B] 1 illustrates exemplary complementary increment and decrement pulses used in programming non-volatile memory cells. [Figure 37] 1 illustrates a calibration algorithm for programming non-volatile memory cells that adjusts programming parameters based on the tilt characteristics of the cell. [Figure 38] 38 shows the circuit used in the calibration algorithm of FIG. 37. [Figure 39] 1 illustrates a calibration algorithm for programming non-volatile memory cells. [Figure 40]1 illustrates a calibration algorithm for programming non-volatile memory cells. [Figure 41] 42 shows the circuit used in the calibration algorithm of FIG. 41. [Figure 42] 3 illustrates an exemplary progression of voltages applied to the control gates of non-volatile memory cells during a programming operation. [Figure 43] 3 illustrates an exemplary progression of voltages applied to the control gates of non-volatile memory cells during a programming operation. [Figure 44] 1 illustrates a system for applying programming voltages during programming of non-volatile memory cells in a vector multiplication matrix system. [Figure 45] 1 shows a vector multiplication matrix system having a modulator, an analog-to-digital converter, and an output block including a summer. [Figure 46] 1 shows a charge adder circuit. [Figure 47] 1 shows a current adder circuit. [Figure 48] 1 shows a digital adder circuit. [Figure 49A] 1 illustrates one embodiment of an integrating analog-to-digital converter for neuron outputs. [Figure 49B] 49B shows a graph illustrating the voltage output over time of the integrating analog-to-digital converter of FIG. 49A. [Figure 49C] 10 illustrates another embodiment of an integrating analog-to-digital converter for neuron outputs. [Figure 49D] 49D shows a graph illustrating the voltage output over time of the integrating analog-to-digital converter of FIG. 49C. [Figure 49E] 10 illustrates another embodiment of an integrating analog-to-digital converter for neuron outputs. [Figure 49F] 10 illustrates another embodiment of an integrating analog-to-digital converter for neuron outputs. [Figure 50A] 1 shows a successive approximation analog-to-digital converter for neuron output. [Figure 50B] 1 shows a successive approximation analog-to-digital converter for neuron output. [Figure 51] 1 illustrates an embodiment of a sigma-delta analog-to-digital converter. [Figure 52A] 1 illustrates an embodiment of a ramp analog-to-digital converter. [Figure 52B] 1 illustrates an embodiment of a ramp analog-to-digital converter. [Figure 52C] 1 illustrates an embodiment of a ramp analog-to-digital converter. [Figure 53] 1 illustrates an embodiment of an algorithmic analog-to-digital converter. [Figure 54] 1 shows a logical multi-bit cell. [Figure 55] 55 illustrates a method for programming the logical multi-bit cell of FIG. DETAILED DESCRIPTION OF THE INVENTION

[0072] The artificial neural network of the present invention utilizes a combination of CMOS technology and non-volatile memory arrays. [Non-volatile memory cell]

[0073] Digital nonvolatile memories are well known. For example, U.S. Pat. No. 5,029,130 ​​(the "'130 patent"), incorporated herein by reference, discloses an array of split-gate nonvolatile memory cells, a type of flash memory cell. One such memory cell 210 is shown in FIG. 2. Each memory cell 210 includes a source region 14 and a drain region 16 formed in a semiconductor substrate 12, with a channel region 18 between the source region 14 and the drain region 16. A floating gate 20 is formed above and insulated from (and controls the conductivity of) a first portion of the channel region 18 and over a portion of the source region 14. A word line terminal 22 (typically coupled to a word line) has a first portion disposed above and insulated from (and controls the conductivity of) a second portion of the channel region 18, and a second portion extending upwardly and above the floating gate 20. The floating gate 20 and word line terminal 22 are insulated from the substrate 12 by a gate oxide. A bit line terminal 24 is coupled to the drain region 16.

[0074] The memory cell 210 is erased (electrons are removed from the floating gate) by applying a high positive voltage to the word line terminal 22, which causes electrons in the floating gate 20 to pass via Fowler-Nordheim tunneling from the floating gate 20 to the word line terminal 22 through the insulator between them.

[0075] The memory cell 210 is programmed (electrons are applied to the floating gate) by applying a positive voltage to the word line terminal 22 and a positive voltage to the source region 14. An electron current flows from the source region 14 (source line terminal) towards the drain region 16. The electrons accelerate and are heated when they reach the gap between the word line terminal 22 and the floating gate 20. Some of the heated electrons are injected into the floating gate 20 through the gate oxide due to electrostatic attraction from the floating gate 20.

[0076] The memory cell 210 is read by applying a positive read voltage to the drain region 16 and word line terminal 22 (turning on the portion of the channel region 18 below the word line terminal). When the floating gate 20 is positively charged (i.e., erased with electrons), the portion of the channel region 18 below the floating gate 20 is also turned on, and current flows through the channel region 18, which is sensed as an erased or "1" state. When the floating gate 20 is negatively charged (i.e., programmed with electrons), the portion of the channel region below the floating gate 20 is mostly or completely off, and no (or very little) current flows through the channel region 18, which is sensed as a programmed or "0" state.

[0077] Table 1 shows typical voltage ranges that may be applied to the terminals of memory cell 110 to perform read, erase, and program operations. Table 1: Operation of flash memory cell 210 of FIG. 2 [Table 1] "Read 1" is a read mode in which the cell current is output to the bit line. "Read 2" is a read mode in which the cell current is output from the source line terminal.

[0078] Figure 3 shows a memory cell 310 similar to memory cell 210 of Figure 2, with the addition of a control gate (CG) terminal 28. The control gate terminal 28 is biased at a high voltage (e.g., 10V) during programming, a low or negative voltage (e.g., 0V / -8V) during erasure, and a low or medium voltage (e.g., 0V / 2.5V) during reading. The other terminals are biased similarly to the terminals of Figure 2.

[0079] FIG. 4 shows a four-gate memory cell 410 including a source region 14, a drain region 16, a floating gate 20 above a first portion of a channel region 18, a select gate 22 (typically coupled to a word line, WL) above a second portion of the channel region 18, a control gate 28 above the floating gate 20, and an erase gate 30 above the source region 14. This configuration is described in U.S. Pat. No. 6,747,310, which is incorporated herein by reference for all purposes. Here, all gates, except for the floating gate 20, are non-floating gates, i.e., they are electrically connected or connectable to a voltage source. Programming is performed by heated electrons injecting themselves from the channel region 18 into the floating gate 20. Erasing is performed by electrons tunneling from the floating gate 20 to the erase gate 30.

[0080] Table 2 shows typical voltage ranges that may be applied to the terminals of memory cell 410 to perform read, erase, and program operations. Table 2: Operation of flash memory cell 410 of FIG. 4 [Table 2] "Read 1" is a read mode in which the cell current is output to the bit line. "Read 2" is a read mode in which the cell current is output from the source line terminal.

[0081] 5 shows a memory cell 510 similar to memory cell 410 of FIG. 4, except that memory cell 510 does not include an erase gate (EG) terminal. Erasing is accomplished by biasing substrate 18 to a high voltage and control gate CG terminal 28 to a low or negative voltage. Alternatively, erasure is accomplished by biasing word line terminal 22 to a positive voltage and control gate terminal 28 to a negative voltage. Programming and reading are similar to those of FIG. 4.

[0082] Figure 6 shows another type of flash memory cell, a three-gate memory cell 610. Memory cell 610 is identical to memory cell 410 of Figure 4, except that memory cell 610 does not have a separate control gate terminal. Erase and read operations (erasure occurs through use of the erase gate terminal) are similar to those of Figure 4, except that no control gate bias is applied. Programming operations are also performed without a control gate bias, and as a result, a higher voltage must be applied to the source line terminal during a program operation to compensate for the lack of control gate bias.

[0083] Table 3 shows typical voltage ranges that may be applied to the terminals of memory cell 610 to perform read, erase, and program operations. Table 3: Operation of Flash Memory Cell 610 of FIG. 6 [Table 3] "Read 1" is a read mode in which the cell current is output to the bit line. "Read 2" is a read mode in which the cell current is output from the source line terminal.

[0084] Figure 7 shows another type of flash memory cell, a stacked gate memory cell 710. Memory cell 710 is similar to memory cell 210 of Figure 2, except that the floating gate 20 extends over the entire channel region 18, and a control gate terminal 22 (coupled to a word line) extends above the floating gate 20, separated by an insulating layer (not shown). Erase, programming, and read operations operate in a similar manner to those described above for memory cell 210.

[0085] Table 4 shows typical voltage ranges that may be applied to the terminals of memory cell 710 and substrate 12 to perform read, erase, and program operations. Table 4: Operation of Flash Memory Cell 710 of FIG. 7 [Table 4]

[0086] "Read 1" is a read mode in which the cell current is output to the bit line. "Read 2" is a read mode in which the cell current is output at the source line terminal. Optionally, in an array including rows and columns of memory cells 210, 310, 410, 510, 610, or 710, a source line may be coupled to one row of memory cells or two adjacent rows of memory cells. That is, a source line terminal may be shared by adjacent rows of memory cells.

[0087] In order to utilize a memory array containing one of the non-volatile memory cell types in an artificial neural network as described above, two modifications are made. First, as explained further below, the lines are configured so that each memory cell can be individually programmed, erased, and read without adversely affecting the memory state of other memory cells in the array. Second, continuous (analog) programming of the memory cells is provided.

[0088] Specifically, the memory state (i.e., the charge on the floating gate) of each memory cell in the array can be continuously changed from a fully erased state to a fully programmed state, independently and with minimal disturbance to other memory cells. In another embodiment, the memory state (i.e., the charge on the floating gate) of each memory cell in the array can be continuously changed from a fully programmed state to a fully erased state, and vice versa, independently and with minimal disturbance to other memory cells. This means that the cell storage is analog, or at a minimum, capable of storing one of a large number of discrete values ​​(such as 16 or 64 different values), allowing every cell in a memory array to be very precisely and individually tunable, making memory arrays ideal for storage and fine-tuning the synaptic weights of a neural network.

[0089] The methods and means described herein can be applied to other non-volatile memory technologies, such as, but not limited to, SONOS (silicon-oxide-nitride-oxide-silicon, charge traps in nitride), MONOS (metal-oxide-nitride-oxide-silicon, metal charge traps in nitride), ReRAM (resistive random access memory), PCM (phase change memory), MRAM (magnetoresistive random access memory), FeRAM (ferroelectric random access memory), OTP (bi-level or multi-level one-time programmable), and CeRAM (strongly correlated electron memory). The methods and means described herein can be applied to, but not limited to, volatile memory technologies used in neural networks, such as SRAM, DRAM, and other volatile synapse cells. [Neural Networks Using Nonvolatile Memory Cell Arrays]

[0090] 8 conceptually illustrates a non-limiting example of a neural network utilizing the non-volatile memory array of the present embodiments. This example uses the non-volatile memory array neural network for a face recognition application, although other suitable applications can also be implemented using the non-volatile memory array-based neural network.

[0091] S0 is the input layer, which in this example is a 32x32 pixel RGB image with 5-bit precision (i.e., three 32x32 pixel arrays, one for each color R, G, and B, with each pixel having 5-bit precision). Synapse CB1 going from input layer S0 to layer C1 scans the input image with overlapping 3x3 pixel filters (kernels), applying different sets of weights to some instances and shared weights to other instances, and shifts the filters by one pixel (or two or more pixels, depending on the model). Specifically, the values ​​of nine pixels in the 3x3 portion of the image (i.e., called filters or kernels) are provided to synapse CB1, which multiplies these nine input values ​​by the appropriate weights, and after summing the outputs of the multiplications, a single output value is determined and provided by the first synapse of CB1 to generate one pixel of the layer of feature map C1. The 3x3 filter is then shifted one pixel to the right in input layer S0 (i.e., adding a column of three pixels to the right and dropping a column of three pixels on the left), so that the nine pixel values ​​of this newly positioned filter are provided to synapse CB1, where they are multiplied by the same weights as above to determine a second single output value by the associated synapse. This process continues until the 3x3 filter has scanned the entire 32x32 pixel image of input layer S0 for all three colors and all bits (precision values). The process is then repeated using different sets of weights to generate different feature maps for C1 until all of the feature maps for layer C1 have been calculated.

[0092] In this example, there are 16 feature maps in layer C1, each having 30x30 pixels. Each pixel is a new feature pixel extracted from the multiplication of the input and the kernel, and therefore each feature map is a two-dimensional array. Thus, in this example, layer C1 comprises 16 layers of two-dimensional arrays. (Note that the layers and arrays referred to herein are logical, not necessarily physical, relationships; i.e., the arrays are not necessarily oriented in a physical two-dimensional array.) Each of the 16 feature maps in layer C1 is generated by one of 16 different sets of synaptic weights applied to the filter scans. The C1 feature maps can all target different aspects of the same image feature, such as boundary identification. For example, a first map (generated using a first set of weights shared by all scans used to generate this first map) can identify circular edges, while a second map (generated using a second set of weights different from the first set of weights) can identify rectangular edges or the aspect ratio of a particular feature, etc.

[0093] From layer C1, an activation function P1 (pooling) is applied before layer S1, which pools values ​​from non-overlapping, contiguous 2x2 regions within each feature map. The purpose of the pooling function is to average nearby locations (or a max function could be used), e.g., to reduce dependency on edge locations, and to reduce data size before proceeding to the next stage. In layer S1, there are 16 15x15 feature maps (i.e., 16 different arrays of 15x15 pixels each). Synapse CB2 from layer S1 to layer C2 scans the maps in S1 with a 4x4 filter with a filter shift of 1 pixel. In layer C2, there are 22 12x12 feature maps. From layer C2, an activation function P2 (pooling) is applied before layer S2, which pools values ​​from non-overlapping, contiguous 2x2 regions within each feature map. In layer S2, there are 22 6x6 feature maps. At synapse CB3 going from layer S2 to layer C3, an activation function (pooling) is applied, where every neuron in layer C3 connects to every map in layer S2 through a respective synapse in CB3. There are 64 neurons in layer C3. Synapse CB4 going from layer C3 to output layer S3 fully connects C3 to S3, i.e., every neuron in layer C3 connects to every neuron in layer S3. The output at S3 includes 10 neurons, where the neuron with the highest output determines the class. This output can indicate, for example, the identification or classification of the content of the original image.

[0094] Each layer of synapses is implemented using an array or part of an array of non-volatile memory cells.

[0095] Figure 9 is a block diagram of a system that can be used for this purpose. A vector-matrix multiplication (VMM) system 32 contains nonvolatile memory cells and is utilized as a synapse between one layer and the next (such as CB1, CB2, CB3, and CB4 in Figure 6). Specifically, the VMM system 32 includes a VMM array 33 containing nonvolatile memory cells arranged in rows and columns, an erase gate and word line gate decoder 34, a control gate decoder 35, a bit line decoder 36, and a source line decoder 37, which decode the respective inputs to the nonvolatile memory cell array 33. Input to the VMM array 33 can come from the erase gate and word line gate decoder 34 or from the control gate decoder 35. The source line decoder 37 in this example also decodes the output of the VMM array 33. Alternatively, the bit line decoder 36 can decode the output of the VMM array 33.

[0096] The VMM array 33 serves two purposes. First, it stores the weights used by the VMM system 32. Second, the VMM array 33 effectively multiplies the inputs by the weights stored in the VMM array 33 and sums them for each output line (source line or bit line) to produce an output, which becomes the input to the next layer or the input to the last layer. By performing multiplication and addition functions, the VMM array 33 eliminates the need for separate multiplication and addition logic and is also power efficient due to in-place memory computation.

[0097] The outputs of the VMM array 33 are fed to a differential summer (such as a summing op-amp or summing current mirror) 38, which sums the outputs of the VMM array 33 to create a single value for the convolution. The differential summer 38 is arranged to perform a summation of both the positive and negative weight inputs to output a single value.

[0098] The summed output values ​​of the differential summer 38 are then provided to an activation function circuit 39, which rectifies the output. The activation function circuit 39 may provide a sigmoid function, a tanh function, a ReLU function, or any other nonlinear function. The rectified output values ​​of the activation function circuit 39 become elements of the feature map of the next layer (e.g., C1 in FIG. 8) and are then applied to the next synapse to generate the next feature map layer or the final layer. Thus, in this example, the VMM array 33 comprises multiple synapses (which receive input from a previous layer of neurons or from an input layer such as an image database), and the summers 38 and activation function circuit 39 comprise multiple neurons.

[0099] The inputs to the VMM system 32 of FIG. 9 (WLx, EGx, CGx, and optionally BLx and SLx) may be analog levels, binary levels, digital pulses (in which case a pulse-to-analog converter PAC may be required to convert the pulses to appropriate input analog levels), or digital bits (in which case a DAC is provided to convert the digital bits to appropriate input analog levels), and the outputs may be analog levels, binary levels, digital pulses, or digital bits (in which case an output ADC is provided to convert the output analog levels to digital bits).

[0100] FIG. 10 is a block diagram illustrating the use of multiple layers of VMM system 32, labeled in the figure as VMM systems 32a, 32b, 32c, 32d, and 32e. As shown in FIG. 10, input (denoted Inputx) is converted from digital to analog by digital-to-analog converter 31 and provided to input VMM system 32a. The converted analog input can be a voltage or current. The first layer of input D / A conversion can be performed by using a function or LUT (look-up table) that maps input Inputx to the appropriate analog level of the matrix multiplier of input VMM system 32a. Input conversion can also be performed by an analog-to-analog (A / A) converter to convert an external analog input to a mapped analog input to input VMM system 32a. Input conversion can also be performed by a digital-to-digital pulse (D / P) converter to convert an external digital input to a mapped digital pulse(s) to input VMM system 32a.

[0101] The output generated by input VMM system 32a is then provided as input to the next VMM system (hidden level 1) 32b, which in turn generates an output that is provided as input to input VMM system (hidden level 2) 32c, and so on. The various layers of VMM system 32 function as layers of synapses and neurons of a convolutional neural network (CNN). VMM systems 32a, 32b, 32c, 32d, and 32e can each be a standalone physical system that includes a corresponding non-volatile memory array, or multiple VMM systems can utilize different portions of the same physical non-volatile memory array, or multiple VMM systems can utilize overlapping portions of the same physical non-volatile memory array. Each VMM system 32a, 32b, 32c, 32d, and 32e can also be time-multiplexed to various portions of its array or neurons. The example shown in Figure 10 includes five layers (32a, 32b, 32c, 32d, 32e): one input layer (32a), two hidden layers (32b, 32c), and two fully connected layers (32d, 32e). Those skilled in the art will appreciate that this is merely an example and that the system may alternatively include more than two hidden layers and more than two fully connected layers. [VMM Array]

[0102] 11 shows a neuron VMM array 1100 that is particularly suited for memory cells 310 shown in FIG. 3 and is utilized as part of the synapses and neurons between the input layer and the next layer. VMM array 1100 includes a memory array 1101 of non-volatile memory cells and a reference array 1102 of non-volatile reference memory cells (located at the top of the array). Alternatively, a separate reference array can be located at the bottom.

[0103] In VMM array 1100, control gate lines, such as control gate line 1103, run vertically (thus row-oriented reference array 1102 is orthogonal to control gate line 1103), and erase gate lines, such as erase gate line 1104, run horizontally. Here, inputs to VMM array 1100 are provided to control gate lines (CG0, CG1, CG2, CG3), and outputs of VMM array 1100 appear on source lines (SL0, SL1). In one embodiment, only even rows are used, and in another embodiment, only odd rows are used. The current applied to each source line (SL0, SL1, respectively) performs a function of the sum of all currents from memory cells connected to that particular source line.

[0104] As described herein for neural networks, the non-volatile memory cells of VMM array 1100, ie, the flash memory of VMM array 1100, are preferably configured to operate in the sub-threshold region.

[0105] The nonvolatile reference memory cells and nonvolatile memory cells described herein are biased in weak inversion as follows: Ids=Io * e (Vg-Vth) / nVt =w * Io * e (Vg) / nVt In the formula, w=e (-Vth) / nVt and where Ids is the drain-source current, Vg is the gate voltage of the memory cell, Vth is the threshold voltage of the memory cell, and Vt is the thermal voltage = k * T / q, k is Boltzmann's constant, T is temperature in Kelvin, q is the electron charge, n is the slope coefficient = 1 + (Cdep / Cox), where Cdep = capacitance of the depletion layer and Cox is capacitance of the gate oxide layer, Io is the memory cell current at a gate voltage equal to the threshold voltage, and Io is (Wt / L) * u * Cox * (n-1) * Vt 2where u is the carrier mobility, and Wt and L are the width and length of the memory cell, respectively.

[0106] When using an IV log converter that converts the input current Ids to an input voltage Vg using a memory cell (such as a reference memory cell or a peripheral memory cell) or transistor: Vg=n * Vt * log[Ids / wp * Io] where wp is the w of the reference or peripheral memory cell.

[0107] When using an IV log converter that converts the input current Ids to an input voltage Vg using a memory cell (such as a reference memory cell or a peripheral memory cell) or transistor: Vg=n * Vt * log[Ids / wp * Io]

[0108] where wp is the w of the reference or peripheral memory cell.

[0109] For a memory array used as a vector matrix multiplier VMM array, the output current is: Iout=wa * Io * e (Vg) / nVt , i.e. Iout=(wa / wp) * Iin=W * Iin W=e (Vthp-Vtha) / nVt Iin=wp * Io * e (Vg) / nVt where wa=w for each memory cell in the memory array.

[0110] The word line or control gate can be used as the input of the memory cell for the input voltage.

[0111] Alternatively, the non-volatile memory cells of the VMM arrays described herein can be configured to operate in the linear region. Ids=β * (Vgs-Vth) * Vds; β=u * Cox * Wt / L W α (Vgs-Vth) That is, the weight W in the linear region is proportional to (Vgs-Vth).

[0112] The word line or control gate or bit line or source line can be used as the input of a memory cell operating in the linear region, and the bit line or source line can be used as the output of the memory cell.

[0113] For the IV linear converter, memory cells (such as reference or peripheral memory cells) or transistors operating in the linear region, or resistors can be used to linearly convert input and output currents to input and output voltages.

[0114] Alternatively, the memory cells of the VMM arrays described herein can be configured to operate in the saturation region. Ids=1 / 2 * β * (Vgs-Vth) 2 ; β=u * Cox * Wt / L W α (Vgs-Vth) 2 , that is, the weight W is (Vgs-Vth) 2 is proportional to.

[0115] The word line, control gate, or erase gate can be used as the input of a memory cell operating in the saturation region, and the bit line or source line can be used as the output of an output neuron.

[0116] Alternatively, the memory cells of the VMM arrays described herein can be used in all regions or combinations thereof (subthreshold, linear, or saturation).

[0117] Other embodiments for the VMM array 33 of Figure 9 are described in U.S. patent application Ser. No. 15 / 826,345, which is incorporated herein by reference. As described in that application, the source lines or bit lines can be used as neuron outputs (current sum outputs).

[0118] FIG. 12 shows a neuron VMM array 1200 particularly suited for the memory cells 210 shown in FIG. 2 and utilized as synapses between the input layer and the next layer. The VMM array 1200 includes a memory array 1203 of nonvolatile memory cells, a reference array 1201 of first nonvolatile reference memory cells, and a reference array 1202 of second nonvolatile reference memory cells. The reference arrays 1201 and 1202, arranged in columns of the array, function to convert current inputs flowing into terminals BLR0, BLR1, BLR2, and BLR3 into voltage inputs WL0, WL1, WL2, and WL3. In practice, the first and second nonvolatile reference memory cells are diode-connected through multiplexer 1214 (only a portion of which is shown) with current inputs flowing into them. The reference cells are adjusted (e.g., programmed) to target reference levels, which are provided by a reference mini-array matrix (not shown).

[0119] Memory array 1203 serves two purposes. First, it stores the weights used by VMM array 1200 in each memory cell. Second, memory array 1203 effectively multiplies the inputs (i.e., the current inputs provided to terminals BLR0, BLR1, BLR2, and BLR3, which reference arrays 1201 and 1202 convert to input voltages and provide to word lines WL0, WL1, WL2, and WL3) by the weights stored in memory array 1203, and then adds all the results (memory cell currents) to generate outputs for each bit line (BL0-BLN), which serve as inputs to the next layer or the last layer. Having memory array 1203 perform the multiplication and addition functions eliminates the need for separate multiplication and addition logic and is also power efficient. Here, voltage inputs are provided to word lines WL0, WL1, WL2, and WL3, and outputs appear on bit lines BL0-BLN, respectively, during read (inference) operations. The current placed on each bit line BL0-BLN performs a function of the sum of the currents from all non-volatile memory cells connected to that particular bit line.

[0120] Table 5 shows the operating voltages for VMM array 1200. The columns in the table indicate the voltages applied to the word line of the selected cell, the word lines of the unselected cells, the bit lines of the selected cell, the bit lines of the unselected cells, the source lines of the selected cell, and the source lines of the unselected cells, and FLT indicates floating, i.e., no voltage is applied. The rows indicate read, erase, and program operations. Table 5: Operation of VMM Array 1200 in Figure 12 [Table 5]

[0121] FIG. 13 illustrates a neuron VMM array 1300 that is particularly suited for the memory cells 210 shown in FIG. 2 and that is utilized as part of the synapses and neurons between the input layer and the next layer. The VMM array 1300 includes a memory array 1303 of nonvolatile memory cells, a reference array 1301 of first nonvolatile reference memory cells, and a reference array 1302 of second nonvolatile reference memory cells. The reference arrays 1301 and 1302 extend in the row direction of the VMM array 1300. The VMM array is similar to the VMM 1000, except that word lines extend vertically in the VMM array 1300. Here, inputs are provided to the word lines (WLA0, WLB0, WLA1, WLB2, WLA2, WLB2, WLA3, WLB3), and outputs appear on the source lines (SL0, SL1) during a read operation. The current applied to each source line performs a function of the sum of all currents from the memory cells connected to that particular source line.

[0122] Table 6 shows the operating voltages for VMM array 1300. The columns in the table indicate the voltages applied to the word line of the selected cell, the word lines of the unselected cells, the bit lines of the selected cell, the bit lines of the unselected cells, the source lines of the selected cell, and the source lines of the unselected cells. The rows indicate read, erase, and program operations. Table 6: Operation of VMM Array 1300 in Figure 13 [Table 6]

[0123] 14 illustrates a neuron VMM array 1400 that is particularly suited for memory cells 310 shown in FIG. 3 and that is utilized as part of synapses and neurons between an input layer and the next layer. VMM array 1400 includes a memory array 1403 of nonvolatile memory cells, a reference array 1401 of first nonvolatile reference memory cells, and a reference array 1402 of second nonvolatile reference memory cells. Reference arrays 1401 and 1402 function to convert current inputs flowing into terminals BLR0, BLR1, BLR2, and BLR3 into voltage inputs CG0, CG1, CG2, and CG3. In effect, the first and second nonvolatile reference memory cells are diode-connected through multiplexer 1412 (only a portion of which is shown), with the current inputs flowing through BLR0, BLR1, BLR2, and BLR3. The multiplexers 1412 each include a respective multiplexer 1405 and cascoding transistor 1404 to ensure a constant voltage on each bit line (e.g., BLR0) of the first and second non-volatile reference memory cells during a read operation, where the reference cells are adjusted to a target reference level.

[0124] Memory array 1403 serves two purposes: First, it stores the weights used by VMM array 1400. Second, memory array 1403 is a multiplication and addition logic circuitry architecture, where inputs are provided to the control gate lines (CG0, CG1, CG2, and CG3) and outputs are provided to the bit lines (BL0-BLN) during a read operation. The current applied to each bit line performs a function of the sum of all the currents from the memory cells connected to that particular bit line.

[0125] The VMM array 1400 performs one-way tuning of the non-volatile memory cells in the memory array 1403. That is, each non-volatile memory cell is erased and then partially programmed until the desired charge on the floating gate is reached. This can be done, for example, using the precision programming techniques described below. If too much charge is added to the floating gate (e.g., an incorrect value is stored in the cell), the cell must be erased and the series of partial programming operations must be redone. As shown, two rows that share the same erase gate (e.g., EG0 or EG1) must be erased together (known as a page erase), and then each cell is partially programmed until the desired charge on the floating gate is reached.

[0126] Table 7 shows the operating voltages for VMM array 1400. The columns in the table indicate the voltages applied to the word line of the selected cell, the word lines of the unselected cells, the bit line of the selected cell, the bit lines of the unselected cells, the control gate of the selected cell, the control gates of the unselected cells in the same sector as the selected cell, the control gates of the unselected cells in a different sector from the selected cell, the erase gate of the selected cell, the erase gates of the unselected cells, the source line of the selected cell, and the source lines of the unselected cells. The rows indicate read, erase, and program operations. Table 7: Operation of the VMM Array 1400 in Figure 14 [Table 7]

[0127] FIG. 15 illustrates a neuron VMM array 1500 that is particularly suited for memory cells 310 shown in FIG. 3 and that is utilized as part of synapses and neurons between an input layer and the next layer. VMM array 1500 includes a memory array 1503 of nonvolatile memory cells, a reference array 1501 or first nonvolatile reference memory cells, and a reference array 1502 of second nonvolatile reference memory cells. EG lines EGR0, EG0, EG1, and EGR1 extend vertically, while CG lines CG0, CG1, CG2, and CG3 and SL lines WL0, WL1, WL2, and WL3 extend horizontally. VMM array 1500 is similar to VMM array 1400, except that VMM array 1500 implements bidirectional tuning: each individual cell can be fully erased, partially programmed, and, if necessary, partially erased to reach a desired amount of charge on its floating gate through the use of separate EG lines. As shown, reference arrays 1501 and 1502 convert input currents in terminals BLR0, BLR1, BLR2, and BLR3 into control gate voltages CG0, CG1, CG2, and CG3 (through the action of diode-connected reference cells via multiplexer 1514), which are applied to the memory cells in the row direction. The current outputs (neurons) are in bit lines BL0 through BLN, each bit line summing all the currents from the non-volatile memory cells connected to that particular bit line.

[0128] Table 8 shows the operating voltages for VMM array 1500. The columns in the table indicate the voltages applied to the word line of the selected cell, the word lines of the unselected cells, the bit line of the selected cell, the bit lines of the unselected cells, the control gate of the selected cell, the control gates of the unselected cells in the same sector as the selected cell, the control gates of the unselected cells in a different sector from the selected cell, the erase gate of the selected cell, the erase gates of the unselected cells, the source line of the selected cell, and the source lines of the unselected cells. The rows indicate read, erase, and program operations. Table 8: Operation of VMM Array 1500 in Figure 15 [Table 8]

[0129] 16 shows a neuron VMM array 1600 that is particularly suited to the memory cells 210 shown in FIG. 2 and is used as part of the synapses and neurons between the input layer and the next layer. In the VMM array 1600, inputs INPUT0, ..., INPUT N are the bit lines BL0, ..., BL N and outputs OUTPUT1, OUTPUT2, OUTPUT3, and OUTPUT4 are generated on source lines SL0, SL1, SL2, and SL3, respectively.

[0130] 17 illustrates a neuron VMM array 1700 that is particularly suited for memory cells 210 shown in FIG. 2 and that is utilized as part of the synapses and neurons between an input layer and the next layer. In this example, inputs INPUT0, INPUT1, INPUT2, and INPUT3 are received on source lines SL0, SL1, SL2, and SL3, respectively, and outputs OUTPUT0, ..., OUTPUT N are the bit lines BL0, ..., BL N is generated.

[0131] 18 shows a neuron VMM array 1800 that is particularly suited for memory cells 210 shown in FIG. 2 and is used as part of the synapses and neurons between the input layer and the next layer. In this example, inputs INPUT0, ..., INPUT M are the word lines WL0, ..., WL M OUTPUT0, ..., OUTPUT N are the bit lines BL0, ..., BL N is generated.

[0132] 19 shows a neuron VMM array 1900 that is particularly suited for memory cells 310 shown in FIG. 3 and is used as part of the synapses and neurons between the input layer and the next layer. In this example, inputs INPUT0, ..., INPUT M are the word lines WL0, ..., WL M OUTPUT0, ..., OUTPUT N are the bit lines BL0, ..., BL Nis generated.

[0133] 20 shows a neuron VMM array 2000 that is particularly suited for the memory cells 410 shown in FIG. 4 and is used as part of the synapses and neurons between the input layer and the next layer. In this example, the input INPUT 0、 ..., INPUT n are the vertical control gate lines CG0, ..., CG N and outputs OUTPUT1 and OUTPUT2 are generated on source lines SL0 and SL1.

[0134] 21 shows a neuron VMM array 2100 that is particularly suited for memory cells 410 shown in FIG. 4 and is used as part of the synapses and neurons between the input layer and the next layer. In this example, inputs INPUT0,..., INPUT N are the bit lines BL0, ..., BL N , 2901-(N-1) and 2901-N, which are coupled to the source lines SL0 and SL1, respectively. Exemplary outputs OUTPUT1 and OUTPUT2 are generated on source lines SL0 and SL1.

[0135] 22 shows a neuron VMM array 2200 that is particularly suitable for memory cells 310 shown in FIG. 3, memory cells 510 shown in FIG. 5, and memory cells 710 shown in FIG. 7, and is used as part of synapses and neurons between the input layer and the next layer. In this example, inputs INPUT0, ..., INPUT M are the word lines WL0, ..., WL M Received and output OUTPUT0, ..., OUTPUT N are the bit lines BL0, ..., BL N are generated respectively.

[0136] 23 shows a neuron VMM array 2300 that is particularly suitable for memory cells 310 shown in FIG. 3, memory cells 510 shown in FIG. 5, and memory cells 710 shown in FIG. 7, and is used as part of the synapses and neurons between the input layer and the next layer. In this example, inputs INPUT0, ..., INPUT M are the control gate lines CG0, ..., CG M Received at OUTPUT0, ..., OUTPUT N are the vertical source lines SL0, ..., SL N are generated on each source line SL i is coupled to the source lines of all memory cells in column i.

[0137] 24 shows a neuron VMM array 2400 that is particularly suitable for memory cells 310 shown in FIG. 3, memory cells 510 shown in FIG. 5, and memory cells 710 shown in FIG. 7, and is used as part of the synapses and neurons between the input layer and the next layer. In this example, inputs INPUT0, ..., INPUT M are the control gate lines CG0, ..., CG M Received at OUTPUT0, ..., OUTPUT N are the vertical bit lines BL0, ..., BL N and each bit line BL i is coupled to the bit lines of all memory cells in column i. [Long-term and short-term memory]

[0138] Prior art includes a concept known as long short-term memory (LSTM). LSTMs are often used in artificial neural networks. LSTMs allow artificial neural networks to remember information for any given period of time and use that information in subsequent operations. A traditional LSTM includes a cell, an input gate, an output gate, and a forget gate. The three gates regulate the flow of information into and out of the cell and the duration for which information is stored within the LSTM. VMMs are particularly useful in LSTMs.

[0139] Figure 25 shows an example LSTM 2500. LSTM 2500 in this example includes cells 2501, 2502, 2503, and 2504. Cell 2501 receives input vector x0 and generates output vector h0 and cell state vector c0. Cell 2502 receives input vector x1, output vector (hidden state) h0 from cell 2501, and cell state c0 from cell 2501, and generates output vector h1 and cell state vector c1. Cell 2503 receives input vector x2, output vector (hidden state) h1 from cell 2502, and cell state c1 from cell 2502, and generates output vector h2 and cell state vector c2. Cell 2504 receives input vector x3, output vector (hidden state) h2 from cell 2503, and cell state c2 from cell 2503, and generates output vector h3. Additional cells can be used; an LSTM with four cells is just an example.

[0140] Figure 26 shows an example implementation of an LSTM cell 2600 that can be used for cells 2501, 2502, 2503, and 2504 in Figure 25. LSTM cell 2600 receives an input vector x(t), a cell state vector c(t-1) from a previous cell, and an output vector h(t-1) from a previous cell, and produces a cell state vector c(t) and an output vector h(t).

[0141] LSTM cell 2600 includes sigmoid function devices 2601, 2602, and 2603, each of which applies a number between 0 and 1 to control the degree to which each component of the input vector contributes to the output vector. LSTM cell 2600 also includes tanh devices 2604 and 2605 for applying a hyperbolic tangent function to the input vector, multiplier devices 2606, 2607, and 2608 for multiplying two vectors, and adder device 2609 for adding the two vectors. The output vector h(t) can be provided to the next LSTM cell in the system or can be accessed for other purposes.

[0142] 27 shows LSTM cell 2700, which is an example of one implementation of LSTM cell 2600. For the convenience of the reader, the same numbering scheme from LSTM cell 2600 is used in LSTM cell 2700. Sigmoid function devices 2601, 2602, and 2603 and tanh device 2604 each include multiple VMM arrays 2701 and activation circuit blocks 2702. It can therefore be seen that VMM arrays are particularly useful in LSTM cells used in certain neural network systems.

[0143] An alternative example of LSTM cell 2700 (and another example of one implementation of LSTM cell 2600) is shown in Figure 28. In Figure 28, sigmoid function devices 2601, 2602, and 2603 and tanh device 2604 may share the same physical hardware (VMM array 2801 and activation function block 2802) in a time-multiplexed manner. LSTM cell 2800 also includes a multiplier device 2803 for multiplying two vectors, an adder device 2808 for adding two vectors, a tanh device 2605 (including activation circuit block 2802), a register 2807 for storing the value i(t) output from sigmoid function block 2802, and a value f(t) output from multiplier device 2803 via multiplexer 2810. * a register 2804 for storing c(t-1) and a value i(t) output from the multiplier device 2803 via a multiplexer 2810; * A register 2805 stores u(t) and the value o(t) output from the multiplier device 2803 via a multiplexer 2810. * It includes a register 2806 that stores c(t), and a multiplexer 2809.

[0144] While LSTM cell 2700 includes multiple sets of VMM arrays 2701 and respective activation function blocks 2702, LSTM cell 2800 includes only one set of VMM arrays 2801 and activation function blocks 2802, which are used to represent multiple layers in embodiments of LSTM cell 2800. LSTM cell 2800 requires less space than LSTM cell 2700, as it requires one-quarter the space for the VMMs and activation function blocks compared to LSTM cell 2700.

[0145] It can be further appreciated that an LSTM unit typically includes multiple VMM arrays, each of which requires functionality provided by specific circuit blocks outside the VMM array, such as adder and activation circuit blocks and high voltage generation blocks. Providing a separate circuit block for each VMM array would require a significant amount of space within a semiconductor device and would be somewhat inefficient. Therefore, the embodiments described below attempt to minimize the circuitry required outside the VMM array itself. [Gated Recurrent Unit]

[0146] An analog VMM implementation can utilize a gated recurrent unit (GRU), which is a gating mechanism in recurrent artificial neural networks. GRUs are similar to LSTMs, except that GRU cells generally contain fewer components than LSTM cells.

[0147] Figure 29 shows an example GRU 2900. GRU 2900 in this example includes cells 2901, 2902, 2903, and 2904. Cell 2901 receives input vector x0 and generates output vector h0. Cell 2902 receives input vector x1 and output vector h0 from cell 2901 and generates output vector h1. Cell 2903 receives input vector x2 and output vector (hidden state) h1 from cell 2902 and generates output vector h2. Cell 2904 receives input vector x3 and output vector (hidden state) h2 from cell 2903 and generates output vector h3. Additional cells can be used; a GRU with four cells is merely an example.

[0148] FIG. 30 shows an example implementation of a GRU cell 3000 that can be used for cells 2901, 2902, 2903, and 2904 in FIG. 29. GRU cell 3000 receives an input vector x(t) and an output vector h(t-1) from a preceding GRU cell and generates an output vector h(t). GRU cell 3000 includes sigmoid function devices 3001 and 3002, each of which applies a number between 0 and 1 to components from the output vector h(t-1) and the input vector x(t). GRU cell 3000 also includes a tanh device 3003 for applying a hyperbolic tangent function to the input vector, multiple multiplier devices 3004, 3005, and 3006 for multiplying two vectors, an adder device 3007 for adding the two vectors, and a complementary device 3008 for subtracting the input from 1 to generate an output.

[0149] Figure 31 shows GRU cell 3100, which is an example of one implementation of GRU cell 3000. For the convenience of the reader, the same numbering scheme as GRU cell 3000 is used in GRU cell 3100. As can be seen from Figure 31, sigmoid function devices 3001 and 3002 and tanh device 3003 each include multiple VMM arrays 3101 and activation function blocks 3102. It can therefore be seen that VMM arrays are particularly used in GRU cells used in certain neural network systems.

[0150] An alternative example of GRU cell 3100 (and another example of one implementation of GRU cell 3000) is shown in Figure 32. In Figure 32, GRU cell 3200 uses VMM array 3201 and activation function block 3202, which, when configured as a sigmoid function, applies a number between 0 and 1 to control the degree to which each component of the input vector contributes to the output vector. In Figure 32, sigmoid function devices 3001 and 3002 and tanh device 3003 share the same physical hardware (VMM array 3201 and activation function block 3202) in a time-multiplexed manner. GRU cell 3200 also includes a multiplier device 3203 for multiplying two vectors, an adder device 3205 for adding the two vectors, a complementary device 3209 for subtracting the input from one to generate an output, a multiplexer 3204, and a value h(t-1) output from multiplier device 3203 via multiplexer 3204. * A register 3206 holds r(t) and a value h(t-1) output from the multiplier device 3203 via multiplexer 3204. * A register 3207 holds z(t) and a value h^(t) output from the multiplier device 3203 via multiplexer 3204. * and a register 3208 that holds (1-z((t)).

[0151] While GRU cell 3100 includes multiple sets of VMM arrays 3101 and activation function blocks 3102, GRU cell 3200 includes only one set of VMM arrays 3201 and activation function blocks 3202, which are used to represent multiple layers in embodiments of GRU cell 3200. GRU cell 3200 requires less space than GRU cell 3100, as it requires one-third the space for the VMM and activation function blocks compared to GRU cell 3100.

[0152] It will be further appreciated that a system utilizing a GRU typically includes multiple VMM arrays, each of which requires functionality provided by specific circuit blocks outside the VMM array, such as adder and activation circuit blocks and high voltage generation blocks. Providing separate circuit blocks for each VMM array would require a significant amount of space within a semiconductor device and would be somewhat inefficient. Therefore, the embodiments described below attempt to minimize the circuitry required outside the VMM array itself.

[0153] The input to the VMM array can be analog levels, binary levels, timing pulses, or digital bits, and the output can be analog levels, binary levels, timing pulses, or digital bits (in this case, an output ADC is required to convert the output analog level current or voltage to digital bits).

[0154] For each memory cell in the VMM array, each weight w can be implemented by a single memory cell, by a differential cell, or by two blended memory cells (an average of two or more cells). In the case of differential cells, the weight w is defined as a differential weight (w=w + -w - ), two memory cells are required to implement weight w as the average of the two cells. [Embodiments for Precision Programming of Cells in a VMM]

[0155] Embodiments will now be described for precisely programming memory cells in a VMM by incrementing or decrementing programming voltages applied to different terminals of the memory cells.

[0156] FIG. 33A illustrates a programming method 3300. Initially, the method typically begins in response to a received program (condition) command (step 3301). Next, a mass program operation programs all cells to a "0" state (step 3302). A soft erase operation then erases all cells to an intermediate weak erase level (step 3303), such that each cell draws, for example, approximately 3-5 μA of current during a read operation. This contrasts with the highest deeply erased level, where each cell draws approximately 20-30 μA of current during a read operation. Soft erase is performed, for example, by applying incremental erase voltage pulses until an intermediate cell current is reached. The incremental erase voltage pulses are used to limit the degradation experienced by memory cells from a hard erase (i.e., maximum erase level). A hard program is then performed (step 3304) on all unselected cells, adding electrons to the floating gates of the cells to a very deep programmed state, ensuring that those cells are truly "off," i.e., that those cells, e.g., unused memory cells, draw negligible amounts of current during a read operation. The hard program is performed, for example, with higher incremental program voltage pulses and / or a longer program time.

[0157] A coarse programming method (bringing the cells fairly close to the target, e.g., 2x-100x the target) is then performed on the selected cells (step 3305), followed by a fine programming method (step 3306) on the selected cells to program the exact value desired into each selected cell.

[0158] FIG. 33B illustrates another programming method 3310 similar to programming method 3300. However, after the method begins (step 3301), instead of a program operation that programs all cells to a “0” state as in step 3302 of FIG. 33A, a soft erase operation is used to erase all cells to a “1” state (step 3312). A soft program operation (step 3313) is then used to program all cells to an intermediate level so that each cell draws approximately 3-5 uA of current during a read operation. Thereafter, coarse programming method 3305 and fine programming method 3306 are performed as in FIG. 33A. A variation of the embodiment of FIG. 33B eliminates the soft program operation (step 3313) entirely. Multiple coarse programming methods can be used to accelerate programming, such as by targeting multiple progressively smaller coarse targets, before performing the fine programming step 3306. The fine programming method 3306 is performed, for example, with fine incremental program voltage pulses or constant program timing pulses.

[0159] Different terminals of the memory cell can be used for the coarse programming method 3305 and the fine programming method 3306. That is, during the coarse programming method 3305, the voltage applied to one of the terminals of the memory cell (which may be referred to as the coarse programming terminal) is varied until a desired voltage level is achieved in the floating gate 20, and during the fine programming method 3306, the voltage applied to one of the terminals of the memory cell (which may be referred to as the fine programming terminal) is varied until a desired level is achieved. Various combinations of terminals that can be used as coarse and fine programming terminals are shown in Table 9. Table 9: Memory Cell Terminals Used for Coarse and Fine Programming Methods [Table 9] Other combinations of terminals for the coarse and fine programming steps are possible.

[0160] 34 illustrates a first embodiment of the coarse programming method 3305, which is a lookup and execution method 3400. First, a lookup table lookup or function (IV curve) is performed to determine the coarse target current value (I CT ) based on the value intended to be stored in the selected cell (step 3401). This table is created, for example, by silicon characterization or from wafer test calibration. The selected cell can be programmed to store one of N possible values ​​(e.g., 128, 64, 32, etc.). Each of the N values ​​represents a different desired current value (I) to be drawn by the selected cell during a read operation. D In one embodiment, the lookup table corresponds to the coarse target current value I for the cell selected during the search and execution method 3400. CT , where M is an integer less than N. For example, if N is 8, then M may be 4, meaning there are 8 possible values ​​that the selected cell can store, and one of the four coarse target current values ​​will be selected as the coarse target for search and execute method 3400. That is, search and execute method 3400 (which is again an embodiment of coarse programming method 3305) programs the selected cell to store a desired value (I D ) somewhat close to the value (I CT ) and then the precision programming method 3306 programs the desired value (I D ) or achieve the desired value (I D The intention is to more precisely program the selected cells so that they are very close to the

[0161] Examples of cell values, desired current values, and coarse target current values ​​are shown in Tables 10 and 11 for a simple example with N=8 and M=4. Table 10: Example of N desired current values ​​for N=8 [Table 10] Table 11: Example of M target current values ​​for M=4 [Table 11] Offset value I CTOFFSETx is used to prevent overshooting the desired current value during coarse adjustment.

[0162] Rough target current value I CT Once selected, in step 3402 the selected cell is programmed by applying an initial voltage v to the selected cell's coarse programming terminal according to one of the sequences listed in Table 9 above. (The value of the initial voltage v and the appropriate coarse programming terminal are optionally associated with a coarse target current value I CT This can be determined from a voltage lookup table that stores v0 corresponding to Table 12. Initial voltage v0 applied to the coarse programming terminal [Table 12]

[0163] Next, in step 3403, the selected cell is charged to a voltage v i =v i-1 +v increment to the coarse programming terminal, where i starts at 1 and increments each time this step is repeated, and v increment is the coarse voltage increment that will cause programming commensurate with the granularity of the desired change. Thus, the first time step 3403 is performed at i=1, and v is the coarse voltage increment that will cause programming commensurate with the granularity of the desired change. increment A read operation is then performed on the selected cell, and the current drawn through the selected cell (I cell ) is measured and a verification operation is performed (step 3404). cell I CTIf so, the search and execute method 3400 is complete and the precision programming method 3306 can begin. cell I CT If not, step 3403 is repeated and i is incremented.

[0164] Thus, at the point where the coarse programming method 3305 ends and the fine programming method 3306 begins, the voltage v i is the final voltage applied to the coarse programming terminal to program the selected cell, which will then be programmed to the coarse target current value I CT The values ​​associated with CT The goal of the fine programming method 3306 is to program the selected cell so that during a read operation, the selected cell receives a current I D The current is programmed to the point where the selected cell will draw current (with an acceptable margin of deviation, such as + / - 30% or less), which is the desired current value associated with the value intended to be stored in the selected cell.

[0165] FIG. 35 shows examples of different voltage progressions that may be applied to the coarse programming terminals of selected memory cells according to Table 9 during the coarse programming method 3305 and / or to the fine programming terminals of selected memory cells during the fine programming method 3306.

[0166] Under the first approach, increasing voltages are applied to the coarse programming terminal and / or the fine programming terminal in a progressive manner to further program the selected memory cell. i , which is the last voltage applied during the coarse programming method 3305. p1 ga v i is applied to, and then the voltage v i +v p1 is used to program the selected cell (indicated by the second pulse from the left in progression 3501). p1 is v increment(the voltage increment used during the coarse programming method 3305). After each programming voltage is applied to the programming terminal, Icell PT1 A verification step (similar to step 3404) is performed in which it is determined whether I is less than or equal to (the first precision target current value, here the second threshold value). PT1 =I D +I PT1OFFSET and I PT1OFFSET is an offset value added to prevent program overshoot. If the answer is no, another increment v p1 is added to the previously applied programming voltage and the process is repeated. cell I PT1 The following is repeated up to a point where this portion of the programming sequence stops. PT1 I D or with a sufficient accuracy to D If the value is approximately equal to , the selected memory cell has been successfully programmed.

[0167] I PT1 I D If the voltages are not close enough to V, then further programming with smaller granularity can be done. Here, progression 3502 is used. The starting point of progression 3502 is the last voltage used for programming under progression 3501. The increment V p2 (v p1 (less than ) is applied to that voltage, and the combined voltage is applied to the precision programming terminal to program the selected memory cell. After each programming voltage is applied, I cell I PT2 A verification step (similar to step 3404) is performed in which it is determined whether I is less than or equal to (a second precision target current value, here a third threshold value). PT2 =ID+I PT2OFFSET and I PT2OFFSETis an offset value added to prevent program overshoot. If the decision is negative, another increment V p2 is added to the previously applied programming voltage and the process is repeated. cell I PT2 The following is repeated until a point at which this part of the programming sequence stops. Now, since the target value has been achieved with a sufficient acceptable accuracy, I PT2 I D or until programming can stop. D It is assumed that the sigma is sufficiently close to 3601. Those skilled in the art will appreciate that additional progressions may be applied using smaller and smaller programming increments. For example, in Figure 36A, three progressions (3601, 3602, and 3603) are applied instead of just two.

[0168] A second approach is shown in progression 3503. Here, instead of increasing the voltage applied during programming of the selected memory cell, the same voltage (V i , or V i +V p1 +V p1 , or V i +V p2 +V p2 ) are applied for durations of increasing duration. p1 and v in Progress 3502 p2 Instead of applying incremental voltages such as t, each applied pulse is greater than the previous pulse by t p1 Add an additional time increment t p1 is applied to the programming pulse. After each programming pulse is applied to the precision programming terminal, the same verification step is performed as described above for step 3501. Optionally, additional steps can be applied, where the additional time increment added to the programming pulse is of shorter duration than the previous step used.

[0169] Optionally, additional program cycle progressions can be applied in which the programming pulses are of the same duration as the progression of the previous program cycle used. Although only one time progression is shown, those skilled in the art will understand that any number of different time progressions can be applied. That is, instead of changing the magnitude of the voltage used during programming or instead of changing the period of the voltage pulses used during programming, the system can instead change the number of programming cycles used.

[0170] FIG. 36B shows a diagram of a complementary pulse program progression in which the voltage applied to one precision programming terminal is increased and the voltage applied to another precision programming terminal is decreased. For example, an increasing voltage progression can be applied to the control gate of a selected cell, and a decreasing voltage progression can be applied to the erase gate or source line of the selected cell. Or, alternatively, an increasing voltage progression can be applied to the erase gate or source line of a selected cell, and a decreasing voltage progression can be applied to the control gate of the selected cell. These complementary progression program pulses result in greater precision in programming. For example, in a programming pulse cycle having a CG increment of 10 mV and an EG decrement of 20 mV, the resulting voltage in FG after the precision programming pulse cycle will be 10 mV, assuming a 40% CG coupling ratio and a 10% EG coupling ratio to FG. * 40%-20mV * 15% = approximately 1 mV. This complementary pulse programming method can be used in the fine program step 3306 after the coarse program step 3305 because the coarse program step 3305 typically uses only CG or EG increments during its programming operation.

[0171] Further details will now be provided for the second and third embodiments of the coarse programming method 3305.

[0172] 37 shows a second embodiment of the coarse programming (adjustment) method 3305, which is an adaptive calibration method 3700. The method begins (step 3701). The cell is programmed by applying an initial voltage v to the coarse programming terminal (step 3702) according to one of the sequences shown in Table 9. Unlike the search and execute method 3400, here v is not obtained from a lookup table, but instead can be a relatively small initial value. The cell's control gate (or erase gate) voltage (which may be referred to as CG1 or EG1) is measured at a first current value IR1 (e.g., 100 nA), and the voltage on the same gate (which may be referred to as CG2 or EG2) is measured at a second current value IR2 (e.g., 10 nA), i.e., in a non-limiting embodiment where IR2 is 10% of IR1, the subthreshold IV slope is determined and stored based on those measurements (e.g., 360 mV / dec or dV / d LOG(I) of current) (step 3703). The IV slope in the linear region is dV / dI.

[0173] new voltage v i The first time this step is performed, i=1, the program voltage v1 is determined based on the stored subthreshold slope value and the current target and offset values, for example, using a subthreshold equation such as: vi=v i-1 +v increment , V increment is proportional to the slope of Vg Vg=n * Vt * log[Ids / wa * Io] Here, wa is the w of the memory cell, Ids is the cell current, Io is the cell current when Vg=Vth, Vt is the thermal voltage, and using g from 1 to 2, V1 is determined by the current IR1, and V2 is determined by the current IR2. Slope=(V1-V2) / (LOG(IR1)-LOG(IR2)) In the formula, v increment =α * Incline *(LOG(IR1)-LOG(I CT )) and I CT is the target current, and α is a predetermined constant (programming offset value)<1 to prevent overshoot, for example 0.9.

[0174] If the stored slope value is relatively steep, a relatively small current offset value can be used. If the stored slope value is relatively flat, a relatively high current offset value can be used. Thus, determining the slope information allows a current offset value to be selected that is customized for the particular cell in question. This generally makes the programming process shorter. As step 3704 is repeated, i is incremented and v is i =v i-1 +v increment Then the cell is i is programmed by applying v to the coarse adaptation programming terminal. increment is also related to the target current value, v increment It can also be determined from a look-up table that stores values ​​of

[0175] A read operation is then performed on the selected cell, drawing a current (I cell ) is measured and a verification operation is performed (step 3705). cell I CT (here the coarse target threshold value), if I CT =I D +I CTOFFSET , I CTOFFSET is an offset value added to prevent program overshoot, the adaptive calibration method 3700 is complete, and the fine programming method 3306 can begin. cell I CT If not, then steps 3704-3705 are repeated (a new tilt measurement is taken using the new data point) or steps 3703-3705 (if the same tilt previously used is reused) and i is incremented.

[0176] 38 shows a high-level block diagram of a circuit for implementing method 3703. A current source 3801 is used to apply exemplary current values ​​IR1 and IR2 to a selected cell (here, memory cell 3802), then the voltage at the coarse programming terminal of memory cell 3802 (V1 (VCGR1 or VEGR1) for IR1 and V2 (VCGR2 or VEGR2) for IR2) is measured and the coarse programming terminal is selected according to Table 9. The linear slope is (V1-V2) / decade of cell current on the LOGI-V curve, i.e., equal to (V1-V2) / (LOG(IR1)-LOG(IR2)).

[0177] 39 shows a third embodiment of the programming method 3305, which is an adaptive calibration method 3900. The method begins (step 3901). The cell is programmed with a default starting value, v0, by applying v0 to the cell's coarse adaptive programming terminal (step 3902). v0 is obtained from a lookup table, such as one created from silicon characterization, and the table values ​​are offset so as not to overshoot the programmed target. An example of v0 is shown in Table 13. Table 13: Initial voltage v0 applied to the course adaptation programming terminal during the adaptation calibration method 3900 [Table 13]

[0178] In step 390,3, an IV slope parameter is created for use in predicting the next programming voltage. A first voltage, V1, is applied to the control gate or erase gate of a selected cell, and the resulting cell current, IR1, is measured. Then, a second voltage, V2, is applied to the control gate or erase gate of the selected cell, and the resulting cell current, IR2, is measured. The slope is determined based on these measurements and stored, for example, according to the following equation in the subthreshold region (cells operating at subthreshold): Slope=(V1-V2) / (LOG(IR1)-LOG(IR2)) (Step 3903). Example values ​​for V1 and V2 are shown in Table 13 above.

[0179] Determining the IV slope information is customized to the particular cell in question. increment This allows values ​​to be selected, which generally makes the programming process shorter.

[0180] Each time step 3904 is executed, i is incremented, initially at 0, and the desired programming voltage, v i is determined based on the stored slope values ​​and current target and offset values ​​using the following equation: v i =v i-1 +v increment , In the formula, v increment =α * Incline * (LOG(IR1)-LOG(I CT ))、 I CT is the target current, and α is a predetermined constant (programming offset value)<1 to prevent overshoot, for example 0.9.

[0181] The selected cell is then i (Step 3905)

[0182] A read operation is then performed on the selected cell, drawing a current (I cell ) is measured and a verification operation is performed (step 3906). cell I CT (here the coarse target threshold value), if I CT =I D +I CTOFFSET , I CTOFFSETis an offset value that is added to prevent program overshoot, and the process proceeds to step 3907. Otherwise, the process returns to step 3903 (new slope measurement) or 3904 (previous slope is reused), and i is incremented.

[0183] In step 3907, I cell I CT A smaller threshold value, I CT2 The purpose is to check if an overshoot has occurred. cell I CT It is to be less than cell I CT If it is much lower than I, then overshoot has occurred and the stored value may actually correspond to an incorrect value. cell I CT2 If it is not, then no overshoot has occurred and the adaptive calibration method 3900 is complete, at which point the process proceeds to the fine programming method 3306. cell I CT2 If it is less than or equal to 1, then an overshoot has occurred. In that case, the selected cell is erased (step 3908) and the programming process resumes at step 3902. Optionally, if step 3908 is performed more than a predetermined number of times, the selected cell may be considered a bad cell that should not be used, and an error signal is output or a flag is set to identify the cell.

[0184] The fine program method 3306 may consist of multiple verify and program cycles, where the pulse width is fixed and the program voltage is incremented by a constant fine voltage to the next pulse, or the program voltage is fixed and the program pulse width is varied.

[0185] Optionally, during a read or verify operation, step 3906 of determining whether the current through the selected non-volatile memory cell is less than or equal to a first threshold current value includes applying a fixed bias to a terminal of the non-volatile memory cell, measuring and digitizing the current drawn by the selected non-volatile memory cell to generate a digital output bit, and digitizing the digital output bit relative to the first threshold current, I CT This can be done by comparing the digital bits representing

[0186] Optionally, during a read or verify operation, step 3907 of determining whether the current through the selected non-volatile memory cell is less than or equal to a second threshold current value includes applying a fixed bias to a terminal of the non-volatile memory cell, measuring and digitizing the current drawn by the selected non-volatile memory cell to generate a digital output bit, and digitizing the digital output bit relative to the second threshold current, I CT2 This can be done by comparing the digital bits representing

[0187] Optionally, steps 3906, 3907 of determining whether the current passing through the selected non-volatile memory cell during a read operation or a verify operation is less than or equal to a first or second threshold current value, respectively, may be performed by applying an input to a terminal of the non-volatile memory cell, modulating the current drawn by the selected non-volatile memory cell with an output pulse to generate a modulated output, digitizing the modulated output to generate a digital output bit, and comparing the digital output bit to a digital bit representing the first or second threshold current, respectively.

[0188] Measuring the cell current for purposes of verifying or reading the current can be done by averaging multiple measurements, for example 8-32 measurements, to reduce the effects of noise.

[0189] 40 shows a fourth embodiment of the coarse programming method 3305, which is an absolute calibration method 4000. The method starts (step 4001). The relevant terminal of the cell is programmed with a default starting value v0 (step 4002). An example of v0 is shown in Table 14. Table 14: Initial voltages v0 applied to memory cell terminals during absolute calibration method 4000 [Table 14]

[0190] The voltage vTx on the coarse programming terminal is measured and stored at a current value Itarget driven through the cell as described above in connection with FIG. 38 (step 4003). A new coarse programming voltage, v1, is determined based on the stored voltage vTx and an offset value, vToffset (corresponding to Ioffset) (step 4004). For example, the new desired voltage v1 can be calculated as follows: v1=v0+(VTBIAS-vTx)-vToffset, where, for example, VTBIAS=approximately 1.5V, which is the default terminal voltage at the maximum target current (meaning the maximum current level the memory cell will tolerate). Essentially, the new target voltage is adjusted by an amount that is the difference between the current voltage vTx at the target current and the maximum voltage and offset.

[0191] The cell then i When i=1, the voltage v1 from step 4004 is used. When i>=2, the voltage v i =v i-1 +v increment is used. increment corresponds to the target current value, v increment A read operation is then performed on the selected cell, and the current drawn through the selected cell (I cell ) is I CT (Step 4006). cell I CTIf so, the absolute calibration method 4000 is complete and the fine programming method 3306 can begin. cell I CT If not, steps 4005-4006 are repeated and i is incremented.

[0192] FIG. 41 shows a circuit 4100 for measuring vTx in step 4003 of the absolute calibration method 4000. vTx is measured at each memory cell 4103 (4103-0, 4103-1, 4103-2,... 4103-n). Here, n + 1 different current sources 4101 (4101-0, 4101-1, 4101-2,... 4101-n) generate different currents IO0, IO1, IO2,... IOn with increasing magnitudes. Each current source 4101 is connected to a respective inverter 4102 (4102-0, 4102-1, 4102-2,... 4102-n) and memory cell 4103 (4103-0, 4103-1, 4103-2,... 4103-n). The input to each inverter 4102 (4102-0, 4102-1, 4102-2,... 4102-n) is initially high, and the output of each inverter is initially low. Since IO0 < IO1 < IO2 <... < IOn, the output of inverter 4102-0 will first switch from low to high because the memory cell 4103-0 draws current from the current source 4101-0 and also draws current from the input node of inverter 4102-0, reducing the input voltage to inverter 4102-0 before the input voltages to the other inverters 4102. Next, the output of inverter 4102-1 switches from low to high, then the output of inverter 4102-2 switches similarly, and so on until the output of inverter 4102-n switches from low to high. Each inverter 4102 controls a respective switch 4104 (4104-0, 4104-1, 4104-2,... 4104-n), and as a result, when the output of inverter 4102 is high, the switch 4104 is closed, whereby vTx is sampled by the capacitor 4105 (4105-0, 4105-1, 4105-2,... 4105-n). Thus, the switches 4104 and the capacitors 4105 form a sample and hold circuit. In this way, vTx is measured using the sample and hold circuit.

[0193] Figure 42 shows an exemplary progression 4200 for programming a selected cell during adaptive calibration method 3700 or absolute calibration method 4000. A voltage VTP (the programming voltage applied to the CG or EG terminal, corresponding to vi in ​​step 3704 of Figure 37 and step 4005 of Figure 40) is applied to the terminal of the selected memory cell using a bit line enable signal En_blx (x varies between 1 and n, where n is the number of bit lines).

[0194] Figure 43 shows another exemplary progression 4300 for programming a selected cell during adaptive calibration method 3700 or absolute calibration method 4000. A voltage VTP (the programming voltage applied to the CG or EG terminal, corresponding to vi in ​​step 3704 of Figure 37 and step 4005 of Figure 40) is applied to the terminal of the selected memory cell using a bit line enable signal En_blx (x varies between 1 and n, where n is the number of bit lines).

[0195] In another embodiment, the voltage applied to the control gate terminal is incremented and the voltage applied to the erase gate terminal is also incremented.

[0196] In another embodiment, the voltage applied to the control gate terminal is incremented and the voltage applied to the erase gate terminal is decreased, as shown in Table 15. Table 15: Control gate terminal increment and erase gate terminal decrement [Table 15]

[0197] For comparison, examples of incrementing only the control gate terminal or incrementing only the erase gate terminal are included in Table 16. Table 16: Control gate terminal increment, erase gate terminal increment [Table 16]

[0198] Figure 44 shows a system for implementing input and output methods for reading or verifying within a VMM array after precision programming. Input function circuit 4401 receives digital bit values, converts them to analog signals, and uses them to apply voltages to the control gates of selected cells within array 4404, which are determined via control gate decoder 4402. At the same time, word line decoder 4403 is also used to select the row in which the selected cell is located. Output neuron circuit block 4405 receives output currents from each column of cells within array 4404. Output circuit block 4405 may include an integrating analog-to-digital converter (ADC), a successive approximation register (SAR) ADC, a sigma-delta ADC, or any other ADC scheme for providing a digital output.

[0199] In one embodiment, the digital value provided to input function circuit 4401 includes four bits (DIN3, DIN2, DIN1, and DIN0), or any number of bits, where the digital value represented by those bits corresponds to the number of input pulses applied to the control gate during a programming operation. More pulses result in a larger value being stored in the cell, which causes a larger output current when the cell is read. Example input bit values ​​and pulse values ​​are shown in Table 17. Table 17: Digital bit input and number of generated pulses [Table 17]

[0200] In the above example, there are a maximum of 15 pulses for a 4-bit input digital. Each pulse is equal to one unit cell value (current), i.e., the precision programmed current. For example, if Icell unit = 1nA, then DIN[3~0] = 0001, then Icell = 1 * 1nA=1nA, and for DIN[3~0]=1111, Icell=15 * 1nA=15nA.

[0201] In another embodiment, digital bit inputs use digital bit position summation to read out the value of a cell or neuron (e.g., the precise programmed value of the bit line output), as shown in Table 18. Here, only four pulses or four fixed, identical bias inputs (e.g., word line or control gate inputs) are required to evaluate a four-bit digital value. For example, a first pulse or a first fixed bias is used to evaluate DIN0, a second pulse or a second fixed bias having the same value as the first value is used to evaluate DIN1, a third pulse or a third fixed bias having the same value as the first value is used to evaluate DIN2, and a fourth pulse or a fourth fixed bias having the same value as the first value is used to evaluate DIN3. Then, the results from the four pulses are summed according to bit position, with each output result multiplied (scaled) by a multiplication coefficient of 2^n (n is the digital bit position), as shown in Table 19. The digital bit summation formula implemented is: Output = 2^0 * DIN0+2^1 * DIN1+2^2 * DIN2+2^3 * DIN3) * In Icell units, where Icell represents the precision programmed current.

[0202] For example, if Icell unit = 1nA, DIN[3~0] = 0001, Icell total = 0+0+0+1 * 1nA = 1nA, and DIN[3~0] = 1111, Icell total = 8 * 1nA+4 * 1nA+2 * 1nA+1 * 1nA=15nA. Table 18: Digital Bit Input Addition [Table 18] Table 19: Summation of Digital Input Bits Dn and 2^n Output Multiplication Factor [Table 19]

[0203] Another embodiment having a hybrid input with multiple digital input pulse ranges and a sum of input digital ranges is shown in Table 20 for an exemplary 4-bit digital input. In this embodiment, DINn-0 can be divided into m different groups, each group is evaluated, and the output is scaled by a multiplication factor depending on the group binary position. For example, for a 4-bit DIN3-0, the groups could be DIN3-2 and DIN1-0, with the output of DIN1-0 scaled by 1 (X1) and the output of DIN3-2 scaled by 4 (X4). Table 20: Hybrid Input / Output Summation with Multiple Input Ranges [Table 20]

[0204] Another embodiment combines a hybrid input range with a hybrid supercell. The hybrid supercell includes multiple physical x-bit cells to implement a logical n-bit cell with the x-cell output scaled by 2^n binary positions. For example, to implement an 8-bit logical cell, two 4-bit cells (cell 1, cell 0) are used. The output of cell 0 is scaled by 1 (X1) and the output of cell 1 is scaled by 4 (X, 2^2). Other combinations of physical x-cells to implement an n-bit logical cell are possible, such as two 2-bit physical cells and one 4-bit physical cell to implement an 8-bit logical cell.

[0205] FIG. 45 illustrates a cell or neuron where the digital bit input is modulated by a modulator 4510 with an output pulse width designed according to the digital input bit position using digital bit position summation (e.g., reading out the current of the cell or neuron (e.g., the value of the bit line output) (e.g., converting the current into an output voltage (V=Current *45 shows another embodiment similar to the system of FIG. 44 except for a bias (applied to the input word line or control gate) for converting the DIN0 signal into a pulse width / capacitance signal. For example, a first bias (applied to the input word line or control gate) is used to evaluate DIN0, and the current (cell or neuron) output is modulated by modulator 4510 with a unit pulse width proportional to the DIN0 bit position, which is in units of 1 (x1), a second input bias is used to evaluate DIN1, and the current output is modulated by modulator 4510 with a pulse width proportional to the DIN1 bit position, which is in units of 2 (x2), a third input bias is used to evaluate DIN2, and the current output is modulated by modulator 4510 with a pulse width proportional to the DIN2 bit position, which is in units of 4 (x4), and a fourth input bias is used to evaluate DIN3, and the current output is modulated by modulator 4510 with a pulse width proportional to the DIN3 bit position, which is in units of 8 (x8). Each output is then converted to a digital bit for each digital input bit DIN0-DIN3 by an ADC (analog-to-digital converter) 4511. The total output is then output by an adder 4512 as the sum of the four digital outputs generated from the DIN0-3 inputs.

[0206] FIG. 46 shows an example of a charge adder 4600 that can be used to sum the outputs of the VMM, Icell, during a verification operation or analog-to-digital conversion of an output neuron to obtain a single analog value representing the output of the VMM, which can then optionally be converted to a digital bit value. The charge adder 4600 can be used, for example, as adder 4512. The charge adder 4600 includes a current source 4601 (here representing the current Icell output by the VMM) and a sample-and-hold circuit including a switch 4602 and a sample-and-hold (S / H) capacitor 4603. The example shown utilizes a 4-bit digital value for the output, but other numbers of bits can be used instead. There are four S / H circuits to hold the values ​​generated from the four evaluation pulses, and these values ​​are summed at the end of the process. The S / H capacitor 4603 has a 2^n of its S / H capacitors.* The C_DINn bit positions are selected in a ratio associated with Icell 4601. For example, switch 4602 for C_DIN3 is closed when Icell > 8 x current threshold, switch 4602 for C_DIN2 is closed when Icell > 4 x current threshold, switch 4602 for C_DIN1 is closed when Icell > 2 x current threshold, and switch 4602 for C_DIN0 is closed when Icell > current threshold. Thus, the digital value stored by sample and hold capacitor 4603 reflects the value of Icell 4601.

[0207] 47 shows a current adder 4700 that can be used to sum the outputs of the VMM, Icell, during a verification operation or analog-to-digital conversion of an output neuron. The charge adder 4700 can be used, for example, as adder 4512. The current adder 4700 includes a current source 4701 (here representing Icell output from the VMM), a switch 4702, switches 4703 and 4704, and a transistor 4705. The example shown utilizes a 4-bit digital value at the output, with the bit values ​​represented by currents I_DIN0, I_DIN1, I_DIN2, and I_DIN3. The bit position of each transistor 4705 affects the value represented by that bit. Switch 4703 for I_DIN3 is closed when Icell > 8 x current threshold, switch 4703 for I_DIN2 is closed when Icell > 4 x current threshold, switch 4704 for I_DIN1 is closed when Icell > 2 x current threshold, and switch 4703 for I_DIN0 is closed when Icell > current threshold. Thus, the digital value output by transistor 4705 (where a "1" is represented by a positive current and a "0" is represented by no current, or vice versa) reflects the value of Icell 4601.

[0208] Figure 48 shows a digital adder 4800 that receives multiple digital values, sums them together, and generates an output DOUT that represents the sum of the inputs. Digital adder 4600 can be used, for example, as adder 4512. Digital adder 4800 can be used during validation operations or during analog-to-digital conversion of output neurons. As shown in the example of a 4-bit digital value, there are digital output bits to hold the values ​​from the four evaluation pulses, and these values ​​are summed at the end of the process. The digital output is a 2^n * It is digitally scaled based on the DINn bit position, for example, DOUT3=x8 DOUT0, _DOUT2=x4 DOUT1, I_DOUT1=x2 DOUT0, I_DOUT0=DOUT0.

[0209] Figure 49A shows a dual slope integrating ADC 4900 applied to an output neuron to convert the cell current into a digital output bit. An integrator consisting of an integrating op-amp 4901 and an integrating capacitor 4902 integrates the cell current ICELL with respect to a reference current IREF. As shown in Figure 49B, for a fixed time t1, switch S1 is closed and switch S2 is open, and the cell current is up-integrated (Vout rises as shown in waveform 4950), then switch S1 is opened and switch S2 is closed, and the reference current IREF is applied at time t2 to be down-integrated (Vout falls as shown in waveform 4950). The value of the current Icell is given by = t2 / t1 * For example, for t1, with a digital bit resolution of 10 bits, 1024 cycles are used, and the number of cycles for t2 varies from 0 to 1024 cycles depending on the Icell value. Once a target value is applied to comparator 4904 as VREF, the output EC4905 of comparator 4904 can be used as a trigger to determine the number of cycles to apply IREF until VOUT falls below VREF.

[0210] Figure 49C shows a single slope integrating ADC 4960 applied to output neuron 4966, ICELL, to convert the cell current to a digital output bit. The ADC 4960 includes an integrating operational amplifier 4961, an integrating capacitor 4962, an operational amplifier 4964, and switches S1 and S3. The integrating operational amplifier 4961 and integrating capacitor 4962 integrate the output neuron current, ICELL. As shown in Figure 49D, during time t1, the cell current is up-integrated (Vout rises until it reaches Vref2), and during time t2, which begins simultaneously with but is greater than time t1, the cell current of the reference cell is up-integrated. The cell current ICELL is given by: * Vref2 / t. A pulse counter coupled to the output of comparator 4965 is used to count the number of pulses (digital output bits) during each integration time t1, t2. For example, as shown, the digital output bits for t1 are less than the digital output bits for t2, which means that the cell current during t1 is greater than the cell current during t2. An initial calibration is performed to calibrate the integration capacitor value with a reference current and a fixed time, Cint=Tref * Iref / Vref2.

[0211] Figure 49E shows a dual-slope integrating ADC 4980 applied to output neuron 4984, ICELL, to convert the cell current into a digital output bit. Dual-slope integrating ADC 4980 includes switches S1, S2, and S3, an operational amplifier 4981, a capacitor 4982, and a reference current source 4983. Dual-slope integrating ADC 4980 does not utilize an integrating operational amplifier. The cell current or the reference current is directly integrated on capacitor 4982. A pulse counter is used to count pulses (digital output bits) during the integration time. The current Icell is given by: = t2 / t1 * It is an IREF.

[0212] Figure 49F shows a single slope integrating ADC 4990 applied to output neuron 4994, ICELL, to convert the cell current to a digital output bit. The single slope integrating ADC 4990 includes switches S2 and S3, an operational amplifier 4991, and a capacitor 4992. The single slope integrating ADC 4980 does not utilize an integrating operational amplifier. The cell current is integrated directly on capacitor 4992. A pulse counter is used to count pulses (digital output bits) during the integration time. Cell current Icell = Cint * Vref2 / t.

[0213] FIG. 50A shows a SAR (successive approximation register) ADC applied to an output neuron to convert the cell current to a digital output bit. The cell current can be dropped through a resistor to convert it to a voltage V. Alternatively, the cell current can charge up a S / H capacitor to convert the cell current to a voltage V. V is provided to the inverting input of comparator 5003, the output of which is fed to the select input of SAR 5001. A clock input CLK is also provided to SAR 5001. A binary search is used to calculate the bits starting from the MSB (most significant bit). Based on the digital bit DN-D0 output from SAR 5001 and received as an input to DAC 5002, the output of DAC 5002 is used to set the appropriate analog reference voltage to the non-inverting input of comparator 5003, i.e., comparator 5003. The output of comparator 5003 is in turn fed back to SAR 5001 to select the next analog level. As shown in FIG. 50B, in the example of four digital output bits, there are four evaluation periods, including, without limitation, a first pulse to evaluate DOUT3 by setting the analog level to the middle, and then a second pulse to evaluate DOUT2 by setting the analog level to the middle of the upper half or the middle of the lower half.

[0214] Modified binary searches, such as cyclic (algorithmic) ADCs, can be used for cell tuning (e.g., programming), verification, or output neuron conversion. Modified binary searches, such as switched-cap (SC) charge redistribution ADCs, can be used for cell tuning (e.g., programming), verification, or output neuron conversion.

[0215] Figure 51 shows a sigma-delta ADC 5100 applied to an output neuron to convert cell currents into digital output bits. An integrator consisting of an operational amplifier 5101 and a capacitor 5105 integrates the sum of a current ICELL from a selected cell current 5106 and a reference current IREF provided by a 1-bit current cDAC 5104. A comparator 5102 compares the integrated output voltage of the operational amplifier 5101 against a reference voltage, VREF2. A clocked DFF 5103 provides a digital output stream according to the output of the comparator 5102, which is received at the D input of DFF 5103. The digital output stream typically goes through a digital filter before being output as digital output bits.

[0216] FIG. 52A shows a ramp-type analog-to-digital converter 5200 including a current source 5201 (representing the received neuron current ICELL), a switch 5202, a variable configurable capacitor 5203, and a comparator 5204 that receives a voltage across the variable configurable capacitor 5203, labeled Vneu, as its non-inverting input, a configurable reference voltage Vreframp as its inverting input, and generates an output Cout. Vreframp is ramped up by discrete levels every comparison clock cycle. Comparator 5204 compares Vneu with Vreframp, resulting in an output Cout that is "1" when Vneu > Vreframp and "0" otherwise. Thus, the output Cout is a pulse whose width varies in response to Ineu. The larger Ineu is, the longer the period during which Cout is "1," resulting in a wider pulse width of the output Cout. A digital counter 5220 converts each pulse 522 at the output Cout into a count value 5221, which is a digital output bit, as shown in FIG. 52B, for two different ICELL currents, marked OT1A and OT2A, respectively.

[0217] Alternatively, the ramp voltage Vreframp is a continuous ramp voltage 5255 shown in graph 5250 of FIG. 52B.

[0218] Alternatively, a multiple ramp embodiment is shown in Figure 52C to reduce conversion time by utilizing a coarse-fine ramp conversion algorithm. First, a coarse reference ramp reference voltage 5271 is ramped quickly to determine each ICELL subrange. Then, a fine reference ramp reference voltage 5272 for each subrange, namely Vreframp1 and Vreframp2, respectively, is used to convert the ICELL current within each subrange. As shown, there are two subranges of the fine reference ramp voltage. More than two coarse / fine steps or two subranges are possible.

[0219] Figure 53 shows an algorithmic analog-to-digital output converter 5300 including a switch 5301, a switch 5302, a sample and hold (S / H) circuit 5303, a 1-bit analog-to-digital converter (ADC) 5304, a 1-bit digital-to-analog converter (DAC) 5305, a summer 5306, and a gain of 2 residue operational amplifier (2x opamp) 5307. The algorithmic analog-to-digital output converter 5300 generates a converted digital output 5308 in response to an analog input Vin and control signals applied to switches 5302 and 5302. An input received at the analog input Vin (e.g., Vneu in Figure 52) is first sampled by S / H circuit 5303 in response to switch 5302, and then a conversion is performed for N bits in N clock cycles. For each conversion clock cycle, 1-bit ADC 5304 compares S / H voltage 5309 with a reference voltage VREF / 2 and outputs a digital bit (e.g., "0" if the input ≦ VREF / 2, or "1" if the input > VREF / 2). This digital output bit, digital output signal 5308, is then converted to an analog voltage (e.g., either VREF / 2 or 0V) by 1-bit DAC 5305 and provided to summer 5306 to be subtracted from S / H voltage 5309. 2× residue op amp 5307 then amplifies the summer difference voltage output to become converted residue voltage 5310, which is provided to S / H circuit 5303 via switch 5301 for the next clock cycle. Instead of this 1-bit (i.e., 2-level) algorithmic ADC, a 1.5-bit (i.e., 3-level) algorithmic ADC can be used to reduce the effects of offsets from ADC 5304, residue op amp 5307, etc. For use with a 1.5-bit algorithmic ADC, a 1.5-bit or 2-bit (ie, 4-level) DAC is preferred. In another embodiment, a hybrid ADC can be used, for example, for a 9-bit ADC, the first 4 bits can be generated by a SAR ADC and the remaining 5 bits can be generated using a gradient or ramp ADC. [Programming and verifying multiple physical cells as a single logical multi-bit cell]

[0220] The programming and verifying devices and methods described above are capable of operating simultaneously with multiple physical cells as a logical multi-bit cell.

[0221] FIG. 54 illustrates a logical multi-bit cell 5400 including i physical cells, labeled physical cells 5401-1, 5401-2, ..., 5401-i. In one embodiment, physical cells 5401 have uniform diffusion widths (transistor widths). In another embodiment, physical cells 5401 have non-uniform diffusion widths (different transistor widths, with transistors having larger widths being capable of storing more levels and therefore a greater number of bits). In both embodiments, physical cells 5401 are programmed, verified, and read as a unit, specifically as a single logical n-bit cell capable of storing more levels than each of the m-bit cells. For example, if m=2, each physical cell 5401 can hold one of four levels (L0, L1, L2, L3). Two such cells can be treated as a single logical cell with n=3, so that a single logical cell can hold one of eight levels (L0, L1, L2, L3, L4, L5, L6, L7). As another example, with m=3, each physical cell 5401 can hold one of eight levels (L0, ..., L7). Four such cells can be treated as a single logical cell with n=5, so that a single logical cell can hold one of 32 levels (L0, ..., L31).

[0222] 55 shows a method 5500 of programming logic multi-bit cell 5400. First, j of i physical cells 5401-1, ..., 5401-i, where j≦i, are programmed and verified using one of the coarse programming methods 3305 until the coarse current target for j physical cells is achieved (step 5501). Next, k of j physical cells, where k≦j, are programmed and verified using one of the fine programming methods 3306 until the fine current target for k physical cells is achieved (step 5502).

[0223] The method 5500 may be performed on more than one subset of the i physical cells 5401-1, . . . 5401-i to achieve a desired overall level of the logical multi-bit cell 5400.

[0224] For example, if i=4, then there are four cells: 5401-1, 5401-2, 5401-3, and 5401-4. If we assume that each cell can hold one of eight different levels, then logical multi-bit cell 5400 can hold one of 32 different levels. If the desired programming value is L27, then that level (corresponding to the desired read current) can be achieved in any number of different ways.

[0225] For example, method 5500 can be performed on cells 5401-1, 5401-2, and 5401-3 until those cells collectively hold L23 (24th level), and then method 5500 is performed on cell 5401-4 to program that cell to the fourth level so that logical multi-bit cell 5400 achieves L27 (28th level).

[0226] As another example, method 5500 can be performed on cells 5401-1, 5401-2, 5401-3, and 5401-4 until those cells collectively hold L25 (26th level), and then method 5500 can be performed only on cell 5401-4 until the entire logical multi-bit cell 5400 stores a value that achieves L27 (28th level).

[0227] Other approaches are possible, and method 5500 can be performed on different subsets of i physical cells until the desired level is achieved.

[0228] In another embodiment, in a situation where i physical cells have non-uniform diffusion widths, coarse programming step 3305 can be performed on j1 ​​physical cells having wider transistor widths until the j1 physical cells collectively achieve the coarse current target, and then fine programming step 3306 can be performed on j2 physical cells having the smallest transistor widths until the j1+j2 physical cells collectively achieve the fine current target.

[0229] It should be noted that, as used herein, both the terms "over" and "on" inclusively include "directly on" (with no intermediate material, element, or gap disposed therebetween) and "indirectly on" (with an intermediate material, element, or gap disposed therebetween). Similarly, the term "adjacent" includes "directly adjacent" (with no intermediate material, element, or gap disposed therebetween) and "indirectly adjacent" (with an intermediate material, element, or gap disposed therebetween); "attached" includes "directly attached" (with no intermediate material, element, or gap disposed therebetween) and "indirectly attached to" (with an intermediate material, element, or gap disposed therebetween); and "electrically coupled" includes "directly electrically coupled" (with no intermediate material or element disposed therebetween that electrically connects the elements together) and "indirectly electrically coupled to" (with an intermediate material or element disposed therebetween that electrically connects the elements together). For example, forming an element "over a substrate" can include forming the element directly on the substrate with no intermediate materials / elements therebetween, and forming the element indirectly on the substrate with one or more intermediate materials / elements therebetween.

Claims

1. 1. A method of programming a selected non-volatile memory cell to store one of N possible values, where N is an integer greater than 2, and the selected non-volatile memory cell includes a floating gate, a control gate terminal, an erase gate terminal, and a source line terminal, the method comprising:

10. A method comprising: performing a first programming process including a plurality of program verify cycles, wherein programming voltages of increasing magnitude are applied to terminals of the selected non-volatile memory cells in each program verify cycle after a first program verify cycle.

2. 2. The method of claim 1, wherein each program verify cycle includes verifying that a current through the selected non-volatile memory cell during a read or verify operation is less than or equal to a first threshold current value.

3. Each program verification cycle: applying a first voltage to one of the erase gate terminal and the control gate terminal of the selected non-volatile memory cell; measuring a first current through the selected non-volatile memory cell as a result of said applying; applying a second voltage to the one of the erase gate terminal and the control gate terminal of the selected non-volatile memory cell; measuring a second current through the selected non-volatile memory cell as a result of said applying; determining a slope value based on the first voltage, the second voltage, the first current, and the second current; 2. The method of claim 1, further comprising: determining a next programming voltage of the increasing magnitude programming voltages for a next program verify cycle based on the determined slope value.

4. 4. The method of claim 3, further comprising programming the selected non-volatile memory cell using the next programming voltage.

5. 5. The method of claim 4, further comprising repeating the steps of determining a next programming voltage and programming the nonvolatile memory cell using the next programming voltage until a current through the selected nonvolatile memory cell during a read or verify operation is equal to or less than a first threshold current value.

6. 3. The method of claim 2, further comprising: if the current passing through the selected non-volatile memory cell during the read or verify operation is less than or equal to the first threshold current value, performing a second programming process until the current passing through the selected non-volatile memory cell during the read or verify operation is less than or equal to a second threshold current value.

7. The step of performing a first programming process comprises:

2. The method of claim 1, further comprising the step of erasing the selected non-volatile memory cell and repeating the first programming process if the current through the selected non-volatile memory cell is less than or equal to a third threshold current value.

8. 8. The method of claim 7, further comprising: performing a third programming process until a current through the selected non-volatile memory cell during a read or verify operation is below a fourth threshold current value.

9. 7. The method of claim 6, wherein the second programming process comprises applying voltage pulses of increasing magnitude to the control gates of the selected non-volatile memory cells.

10. 10. The method of claim 9, wherein the second programming process further comprises applying voltage pulses of increasing magnitude to the erase gates of the selected non-volatile memory cells.

11. 10. The method of claim 9, wherein the second programming process further comprises applying decreasing magnitude voltage pulses to the erase gates of the selected non-volatile memory cells.

12. 2. The method of claim 1, wherein the selected non-volatile memory cells are split-gate flash memory cells.

13. 2. The method of claim 1, wherein the selected non-volatile memory cell is in a vector matrix multiplication array in an analog neural network.

14. 1. A method of programming a selected non-volatile memory cell to store one of N possible values, where N is an integer greater than 2, and the selected non-volatile memory cell includes a floating gate, a control gate terminal, an erase gate terminal, and a source line terminal, the method comprising:

10. A method comprising: performing a first programming process including a plurality of program verify cycles, wherein programming voltage durations of increasing duration are applied to terminals of the selected non-volatile memory cells in each program verify cycle after a first program verify cycle.

15. 15. The method of claim 14, wherein each program verify cycle includes verifying that the current through the selected non-volatile memory cell during a read or verify operation is less than or equal to a first threshold current value.

16. Each program verification cycle: applying a first voltage to one of the erase gate terminal and the control gate terminal of the selected non-volatile memory cell; measuring a first current through the selected non-volatile memory cell as a result of said applying; applying a second voltage to the one of the erase gate terminal and the control gate terminal of the selected non-volatile memory cell; measuring a second current through the selected non-volatile memory cell as a result of said applying; determining a slope value based on the first voltage, the second voltage, the first current, and the second current; determining a programming duration next in the increasing duration programming durations based on the slope value; programming the non-volatile memory cells with the next programming duration; 15. The method of claim 14, further comprising repeating the steps of determining a next programming duration and programming the nonvolatile memory cell with the next programming duration until a current through the selected nonvolatile memory cell during a read or verify operation is equal to or less than a first threshold current value.

17. 15. The method of claim 14, further comprising: if the current passing through the selected non-volatile memory cell during the read or verify operation is less than or equal to the first threshold current value, performing a second programming process until the current passing through the selected non-volatile memory cell during the read or verify operation is less than or equal to a second threshold current value.

18. The step of performing a first programming process comprises:

15. The method of claim 14, further comprising erasing the selected non-volatile memory cell and repeating the first programming process if the current through the selected non-volatile memory cell is less than or equal to a third threshold current value.

19. 20. The method of claim 17, further comprising: performing a third programming process until a current through the selected non-volatile memory cell during a read or verify operation is less than or equal to a fourth threshold current value.

20. 18. The method of claim 17, wherein the second programming process comprises applying voltage pulses of increasing duration to the control gates of the selected non-volatile memory cells.

21. 21. The method of claim 20, wherein the second programming process further comprises applying voltage pulses of increasing duration to the erase gates of the selected non-volatile memory cells.

22. 21. The method of claim 20, wherein the second programming process further comprises applying voltage pulses of decreasing duration to the erase gates of the selected non-volatile memory cells.

23. 2. The method of claim 1, wherein the selected non-volatile memory cells are split-gate flash memory cells.

24. 2. The method of claim 1, wherein the selected non-volatile memory cell is in a vector matrix multiplication array in an analog neural network.

25. 1. A method of programming a selected non-volatile memory cell to store one of N possible values, where N is an integer greater than 2, and the selected non-volatile memory cell includes a floating gate, a control gate terminal, an erase gate terminal, and a source line terminal, the method comprising:

10. A method comprising: performing a first programming process including a plurality of program verify cycles, during which first programming voltages of increasing magnitude are applied to the control gates of the selected non-volatile memory cells and second programming voltages of decreasing magnitude are applied to the erase gates of the selected non-volatile memory cells.

26. 26. The method of claim 25, wherein each program verify cycle includes verifying that the current through the selected non-volatile memory cell during a read or verify operation is less than or equal to a first threshold current value.

27. Each program verification cycle: applying a first voltage to one of the erase gate and the control gate of the selected non-volatile memory cell; measuring a first current through the selected non-volatile memory cell as a result of said applying; applying a second voltage to the one of the erase gate and the control gate of the selected non-volatile memory cell; measuring a second current through the selected non-volatile memory cell as a result of said applying; determining a slope value based on the first voltage, the second voltage, the first current, and the second current; determining a next programming voltage for the programming voltages increasing in magnitude and a next programming voltage for the programming voltages decreasing in magnitude based on the slope value; programming the non-volatile memory cells with the next programming voltage; 26. The method of claim 25, further comprising repeating the steps of determining a next programming voltage and programming the nonvolatile memory cell with the next programming voltage until a current through the selected nonvolatile memory cell during a read or verify operation is equal to or less than a first threshold current value.

28. 26. The method of claim 25, further comprising: performing a second programming process until a current through the selected non-volatile memory cell during a read or verify operation is less than or equal to a second threshold current value.

29. The step of performing a first programming process comprises:

26. The method of claim 25, further comprising erasing the selected non-volatile memory cell and repeating the first programming process if the current through the selected non-volatile memory cell is less than or equal to a third threshold current value.

30. 30. The method of claim 28, further comprising: performing a third programming process until a current through the selected non-volatile memory cell during a read or verify operation is less than or equal to a fourth threshold current value.

31. 1. A method of programming a selected non-volatile memory cell to store one of N possible values, where N is an integer greater than 2, and the selected non-volatile memory cell includes a floating gate, a control gate terminal, an erase gate terminal, and a source line terminal, the method comprising:

10. A method comprising: performing a first programming process including a plurality of program verify cycles, wherein in the programming process, first programming voltages of increasing duration are applied to the control gate terminals of the selected non-volatile memory cells, and second programming voltages of decreasing duration are applied to the erase gate terminals of the selected non-volatile memory cells.

32. 32. The method of claim 31, wherein each program verify cycle includes verifying that the current through the selected non-volatile memory cell during a read or verify operation is less than or equal to a first threshold current value.

33. Each program verification cycle: applying a first voltage to one of the erase gate and the control gate of the selected non-volatile memory cell; measuring a first current through the selected non-volatile memory cell as a result of said applying; applying a second voltage to the one of the erase gate and the control gate of the selected non-volatile memory cell; measuring a second current through the selected non-volatile memory cell as a result of said applying; determining a slope value based on the first voltage, the second voltage, the first current, and the second current; determining next programming durations for the programming voltages of increasing duration and for the programming voltages of decreasing duration based on the slope values; programming the non-volatile memory cells with the next programming duration; 32. The method of claim 31, further comprising: repeating the steps of determining a next programming duration and programming the nonvolatile memory cell with the next programming duration until a current passing through the selected nonvolatile memory cell during a read or verify operation is equal to or less than a first threshold current value.

34. 32. The method of claim 31, further comprising: performing a second programming process until a current through the selected non-volatile memory cell during a read or verify operation is less than or equal to a second threshold current value.

35. The step of performing a first programming process comprises:

32. The method of claim 31, further comprising erasing the selected non-volatile memory cell and repeating the first programming process if the current through the selected non-volatile memory cell is less than or equal to a third threshold current value.

36. 35. The method of claim 34, further comprising: performing a third programming process until a current through the selected non-volatile memory cell during a read or verify operation is less than or equal to a fourth threshold current value.

37. 1. A method of programming a selected non-volatile memory cell to store one of N possible values, where N is an integer greater than 2, and the selected non-volatile memory cell includes a floating gate, a control gate, an erase gate, and a source line terminal, the method comprising:

10. A method comprising: performing a first programming process including a plurality of program verify cycles, wherein in the programming process, first programming voltages of increasing magnitude are applied to the erase gates of the selected non-volatile memory cells, and second programming voltages of decreasing magnitude are applied to the control gates of the selected non-volatile memory cells.

38. 38. The method of claim 37, wherein each program verify cycle includes verifying that the current through the selected non-volatile memory cell during a read or verify operation is less than or equal to a first threshold current value.

39. Each program verification cycle: applying a first voltage to one of the erase gate and the control gate of the selected non-volatile memory cell; measuring a first current through the selected non-volatile memory cell as a result of said applying; applying a second voltage to the one of the erase gate and the control gate of the selected non-volatile memory cell; measuring a second current through the selected non-volatile memory cell as a result of said applying; determining a slope value based on the first voltage, the second voltage, the first current, and the second current; determining a next programming voltage for the programming voltages increasing in magnitude and a next programming voltage for the programming voltages decreasing in magnitude based on the slope value; programming the non-volatile memory cells with the next programming voltage; 38. The method of claim 37, further comprising repeating the steps of determining a next programming voltage and programming the nonvolatile memory cell with the next programming voltage until a current passing through the selected nonvolatile memory cell during a read or verify operation is equal to or less than a first threshold current value.

40. 40. The method of claim 39, further comprising: performing a second programming process until a current through the selected non-volatile memory cell during a read or verify operation is less than or equal to a second threshold current value.

41. The step of performing a first programming process comprises:

40. The method of claim 39, further comprising erasing the selected non-volatile memory cell and repeating the first programming process if the current through the selected non-volatile memory cell is less than or equal to a third threshold current value.

42. 41. The method of claim 40, further comprising: performing a third programming process until a current through the selected non-volatile memory cell during a read or verify operation is less than or equal to a fourth threshold current value.

43. 1. A method of programming a selected non-volatile memory cell to store one of N possible values, where N is an integer greater than 2, and the selected non-volatile memory cell includes a floating gate, a control gate, an erase gate, and a source line terminal, the method comprising:

10. A method comprising: performing a first programming operation including a plurality of program verify cycles, wherein in the programming operation, first programming voltages of increasing duration are applied to the erase gates of the selected non-volatile memory cells, and second programming voltages of decreasing duration are applied to the control gates of the selected non-volatile memory cells.

44. 44. The method of claim 43, wherein each program verify cycle includes verifying that the current through the selected non-volatile memory cell during a read or verify operation is less than or equal to a first threshold current value.

45. Each program verification cycle: applying a first voltage to one of the erase gate and the control gate of the selected non-volatile memory cell; measuring a first current through the selected non-volatile memory cell as a result of said applying; applying a second voltage to the one of the erase gate and the control gate of the selected non-volatile memory cell; measuring a second current through the selected non-volatile memory cell as a result of said applying; determining a slope value based on the first voltage, the second voltage, the first current, and the second current; determining next programming durations for the programming voltages of increasing duration and for the programming voltages of decreasing duration based on the slope values; programming the non-volatile memory cells with the next programming duration; 45. The method of claim 44, further comprising repeating the steps of determining a next programming duration and programming the nonvolatile memory cell for the next programming duration until a current passing through the selected nonvolatile memory cell during a read or verify operation is equal to or less than a first threshold current value.

46. 44. The method of claim 43, further comprising: performing a second programming process until a current through the selected non-volatile memory cell during a read or verify operation is less than or equal to a second threshold current value.

47. The step of performing a first programming process comprises:

44. The method of claim 43, further comprising erasing the selected non-volatile memory cell and repeating the first programming process if the current through the selected non-volatile memory cell is less than or equal to a third threshold current value.

48. 47. The method of claim 46, further comprising: performing a third programming process until a current through the selected non-volatile memory cell during a read or verify operation is less than or equal to a fourth threshold current value.

49. 1. A method of programming a logical multi-bit cell including i physical non-volatile memory cells, each of the i physical non-volatile memory cells being capable of storing one of n possible values, where i is an integer greater than 1 and n is an integer greater than 1, the method comprising: performing a programming operation on j of the i physical non-volatile memory cells, where j is less than or equal to i, until a coarse current target for the logical multi-bit cell is achieved; performing programming operations on k of the j physical non-volatile memory cells, where k is less than or equal to j, until a precision current target for the logical multi-bit cell is achieved.

50. 50. The method of claim 49, wherein the i physical non-volatile memory cells have a uniform width.

51. 50. The method of claim 49, wherein the i physical non-volatile memory cells have non-uniform widths.

52. 1. A method of programming a logical multi-bit cell comprising i physical non-volatile memory cells of unequal widths, each of the i cells being capable of storing one of n possible values, where i is an integer greater than 1 and n is an integer greater than 1, the method comprising: performing programming operations on j of the i physical non-volatile memory cells containing the largest widths, where j is less than or equal to i, until a coarse current target for the logical multi-bit cell is achieved; performing programming operations on k of the j physical non-volatile memory cells having the smallest widths, where k is less than or equal to j, until a precision current target for the logical multi-bit cell is achieved.

53. 1. A method of programming a selected non-volatile memory cell to store one of N possible values, where N is an integer greater than two, and the selected non-volatile memory cell includes a first gate, a first terminal, and a second terminal, the method comprising:

10. A method comprising: performing a plurality of program verify cycles, during each of which a first programming voltage is applied to the first terminal of the selected non-volatile memory cell, the first programming voltage decreasing with each subsequent cycle.

54. 54. The method of claim 53, wherein during each of the program-verify cycles, a second programming voltage is applied to the second terminal of the selected non-volatile memory cell, the second programming voltage increasing with each subsequent cycle.

55. 55. The method of claim 54, wherein the selected non-volatile memory cells are split-gate memory cells.

56. 56. The method of claim 55, wherein the first gate is a floating gate.

57. 57. The method of claim 56, wherein the first terminal is a source line terminal.

58. 58. The method of claim 57, wherein the second terminal is a control gate terminal.

59. 57. The method of claim 56, wherein the first terminal is an erase gate terminal.

60. 60. The method of claim 59, wherein the second terminal is a control gate terminal.

Citation Information

Patent Citations

  • Memory cell

    JP2008103065A

  • Semiconductor memory

    JP2019145188A

  • System and method for storing multibit data in non-volatile memory

    WO2019089168A1

Cited By

  • Game machine

    JP2025158139A

  • Game machine

    JP2025158147A