Precise programming method and device for an analog neural memory in an artificial neural network

By using precise programming algorithms and devices in simulated neuromorphic memory, and programming nonvolatile memory cells using incremental or decreasing programming voltages, the problem of insufficient programming accuracy and granularity in the prior art is solved, and a high-precision programming effect is achieved.

CN114651307BActive Publication Date: 2025-06-27SILICON STORAGE TECHNOLOGY INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202080077970.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-01-23
Filing Date
2020-05-22
Publication Date
2025-06-27
Estimated Expiration
2040-05-22

AI Technical Summary

Technical Problem

The prior art is difficult to program the nonvolatile memory cells in neuromorphic memory with the required accuracy and granularity of different N values.

Method used

Accurate programming algorithms and devices are used to achieve accurate programming of nonvolatile memory cells in vector-matrix multiplication (VMM) arrays by incrementing or decreasing the programming voltages applied to different terminals of the memory cell.

Benefits of technology

High-precision programming of nonvolatile memory cells is realized, ensuring that one of N different values ​​is maintained in simulated neuromorphic memory, and improving programming accuracy and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114651307B_ABST
    Figure CN114651307B_ABST
Patent Text Reader

Abstract

The present invention discloses various embodiments of an accurate programming algorithm and apparatus for precisely and rapidly depositing a correct amount of charge on the floating gate of a non-volatile memory cell within a vector-matrix multiplication (VMM) array in an artificial neural network. Thus, selected cells can be programmed with extremely high precision to hold one of N different values.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Priority Claim

[0002] This application claims priority to U.S. Provisional Application No. 62 / 933,809, filed on November 11, 2019, entitled "PRECISE PROGRAMMING METHOD AND APPARATUS FOR ANALOG NEURAL MEMORY IN A DEEP LEARNING ARTIFICIAL NEURAL NETWORK" and U.S. Patent Application No. 16 / 751,202, filed on January 23, 2020, entitled "PRECISE PROGRAMMING METHOD AND APPARATUS FOR ANALOG NEURAL MEMORY IN A DEEP LEARNING ARTIFICIAL NEURAL NETWORK". Technical Field

[0003] The present invention discloses various embodiments of precise programming algorithms and apparatuses for precisely and rapidly depositing correct amounts of electric charge on the floating gates of non-volatile memory cells within a vector-matrix multiplication (VMM) array in an artificial neural network. Background Art

[0004] Artificial neural networks mimic biological neural networks (the central nervous system of animals, particularly the brain) and are used to estimate or approximate functions that can depend on a large number of inputs and are generally unknown. Artificial neural networks typically include layers of interconnected "neurons" that exchange messages with each other.

[0005] Figure 1 An artificial neural network is shown, where the circles represent the inputs or layers of neurons. The connections (called synapses) are represented by arrows and have numerical weights that can be adjusted according to experience. This enables the artificial neural network to adapt to the inputs and learn. Generally, an artificial neural network includes a layer of multiple inputs. There is usually one or more intermediate layers of neurons, and an output layer of neurons that provides the output of the neural network. The neurons at each level make decisions separately or jointly based on the data received from the synapses.

[0006] One of the main challenges in developing artificial neural networks for high-performance information processing is the lack of adequate hardware technology. In fact, physical artificial neural networks rely on a large number of synapses to achieve high connectivity between neurons, i.e., very high computational parallelism. In principle, such complexity can be achieved by digital supercomputers or clusters of dedicated graphics processing units. However, compared to biological networks, these methods are energy-inefficient in addition to being costly, as biological networks consume less energy mainly due to their performing low-precision analog computations. CMOS analog circuits have been used for artificial neural networks, but most CMOS-implemented synapses are too large given the large number of neurons and synapses.

[0007] The applicant previously disclosed in U.S. Patent Application No. 15 / 594,439 (published as U.S. Patent Publication 2017 / 0337466), which is incorporated herein by reference, an artificial (analog) neural network that utilizes one or more non-volatile memory arrays as synapses. The non-volatile memory arrays operate as analog neuromorphic memories. As used herein, the term "neuromorphic" refers to a circuit that implements a model of the nervous system. The analog neuromorphic memory includes a first plurality of synapses configured to receive a first plurality of inputs and generate a first plurality of outputs therefrom, and a first plurality of neurons configured to receive the first plurality of outputs. The first plurality of synapses includes a plurality of memory cells, where each memory cell of the memory cells includes: spaced-apart source and drain regions formed in a semiconductor substrate, where a channel region extends between the source and drain regions; a floating gate disposed over a first portion of the channel region and insulated from the first portion; and a non-floating gate disposed over a second portion of the channel region and insulated from the second portion. Each memory cell of the plurality of memory cells is configured to store a weight value corresponding to a plurality of electrons on the floating gate. The plurality of memory cells are configured to multiply the first plurality of inputs by the stored weight values to generate the first plurality of outputs. An array of memory cells arranged in this manner may be referred to as a vector matrix multiplication (VMM) array.

[0008] Each non-volatile memory cell used in the analog neuromorphic memory array must be erased and programmed to maintain a very specific and precise amount of electrical charge (i.e., number of electrons) in the floating gate. For example, each floating gate must maintain one of N different values, where N is the number of different weights that can be indicated by each cell. Examples of N include 16, 32, 64, 128, and 256. One challenge in analog neuromorphic memory systems is being able to program selected cells with the precision and granularity required for different N values.

[0009] Improved programming systems and methods are needed that are suitable for use with the VMM arrays in analog neuromorphic memories. SUMMARY OF THE INVENTION

[0010] The present invention discloses various embodiments of an accurate programming algorithm and apparatus for precisely and rapidly depositing a correct amount of charge on the floating gate of a non-volatile memory cell within a vector-matrix multiplication (VMM) array for an analog neuromorphic memory. Thus, selected cells can be programmed with extremely high precision to hold one of N different values. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 A schematic diagram showing an artificial neural network of the prior art.

[0012] Figure 2 A split-gate flash memory cell of the prior art is shown.

[0013] Figure 3 Another split-gate flash memory cell of the prior art is shown.

[0014] Figure 4 Another split-gate flash memory cell of the prior art is shown.

[0015] Figure 5 Another split-gate flash memory cell of the prior art is shown.

[0016] Figure 6 Another split-gate flash memory cell of the prior art is shown.

[0017] Figure 7 A stacked-gate flash memory cell of the prior art is shown.

[0018] Figure 8 A schematic diagram showing different levels of an exemplary artificial neural network using one or more non-volatile memory arrays.

[0019] Fig. 9 A block diagram showing a vector-matrix multiplication system.

[0020] Fig.10 A block diagram showing an exemplary artificial neural network using one or more vector-matrix multiplication systems.

[0021] Fig.11 Another embodiment of a vector-matrix multiplication system is shown.

[0022] Fig.12 Another embodiment of a vector-matrix multiplication system is shown.

[0023] Fig.13 Another embodiment of a vector-matrix multiplication system is shown.

[0024] Fig.14 Another embodiment of a vector-matrix multiplication system is shown.

[0025] Fig.15 Shows another embodiment of a vector-matrix multiplication system.

[0026] Fig.16 Shows another embodiment of a vector-matrix multiplication system.

[0027] Fig.17 Shows another embodiment of a vector-matrix multiplication system.

[0028] Fig.18 Shows another embodiment of a vector-matrix multiplication system.

[0029] Fig.19 Shows another embodiment of a vector-matrix multiplication system.

[0030] Fig. 20 Shows another embodiment of a vector-matrix multiplication system.

[0031] Fig.21 Shows another embodiment of a vector-matrix multiplication system.

[0032] Fig. 22 Shows another embodiment of a vector-matrix multiplication system.

[0033] Fig.23 Shows another embodiment of a vector-matrix multiplication system.

[0034] Fig.24 Shows another embodiment of a vector-matrix multiplication system.

[0035] Fig.25 Shows a prior art long short-term memory system.

[0036] Fig.26 Shows an exemplary cell used in a long short-term memory system.

[0037] Fig. 27 Shows Fig.26 an embodiment of an exemplary cell.

[0038] Fig.28 Shows Fig.26 another embodiment of an exemplary cell.

[0039] Fig.29 Shows a prior art gated recurrent unit system.

[0040] Fig.30 Shows an exemplary cell used in a gated recurrent unit system.

[0041] Fig.31 Shows Fig.30 an embodiment of an exemplary cell.

[0042] Fig.32 Shows Fig.30 Another embodiment of an exemplary cell.

[0043] Fig.33A Shows an embodiment of a method for programming a non - volatile memory cell.

[0044] Fig.33B Shows another embodiment of a method for programming a non - volatile memory cell.

[0045] Fig.34 Shows an embodiment of a rough programming method.

[0046] Fig.35 Shows exemplary pulses used in the programming of non - volatile memory cells.

[0047] Fig.36A Shows exemplary pulses used in the programming of non - volatile memory cells.

[0048] Fig.36B Shows exemplary complementary increment and decrement pulses used in the programming of non - volatile memory cells.

[0049] Fig.37 Shows a calibration algorithm for programming non - volatile memory cells, which adjusts programming parameters based on the slope characteristics of the cell.

[0050] Fig.38 Shows in Fig.37 The circuit used in the calibration algorithm.

[0051] Fig.39 Shows a calibration algorithm for programming non - volatile memory cells.

[0052] Fig.40 Shows a calibration algorithm for programming non - volatile memory cells.

[0053] Fig.41 Shows in Fig.41 The circuit used in the calibration algorithm.

[0054] Fig.42 Shows an exemplary progression of the voltage applied to the control gate of a non - volatile memory cell during a programming operation.

[0055] Fig.43 Shows an exemplary progression of the voltage applied to the control gate of a non - volatile memory cell during a programming operation.

[0056] Fig.44A system for applying a programming voltage during programming of a non-volatile memory cell within a vector-multiplication matrix system is shown.

[0057] Fig.45 A vector-matrix multiplication system is shown, which has an output module including a modulator, an analog-to-digital converter, and a summer.

[0058] Fig.46 A charge summing circuit is shown.

[0059] Fig.47 A current summing circuit is shown.

[0060] Fig.48 A digital summing circuit is shown.

[0061] Fig.49A An embodiment of an integrating analog-to-digital converter for neuron output is shown.

[0062] Fig.49B Shown shown Fig.49A A graph showing the variation of the voltage output of the integrating analog-to-digital converter with time is shown.

[0063] Fig.49C Another embodiment of an integrating analog-to-digital converter for neuron output is shown.

[0064] Fig.49D Shown shown Fig.49C A graph showing the variation of the voltage output of the integrating analog-to-digital converter with time is shown.

[0065] Fig.49E Another embodiment of an integrating analog-to-digital converter for neuron output is shown.

[0066] Fig.49F Another embodiment of an integrating analog-to-digital converter for neuron output is shown.

[0067] Fig.50A And Fig.50B A successive approximation analog-to-digital converter for neuron output is shown.

[0068] Fig.51 An embodiment of a Σ-Δ analog-to-digital converter is shown.

[0069] Fig.52A 、 Fig.52B And Fig.52C Embodiments of a ramp analog-to-digital converter are shown.

[0070] Fig.53 An embodiment of an algorithmic analog-to-digital converter is shown.

[0071] Fig.54Shows a logical multi-bit cell.

[0072] Fig.55 Shows programming Fig.54 method of the logical multi-bit cell. Detailed implementation

[0073] The artificial neural network of the present invention utilizes a combination of CMOS technology and a non-volatile memory array.

[0074] Non-volatile memory cell

[0075] Digital non-volatile memories are well known. For example, U.S. Patent 5,029,130 (“the '130 patent”), which is incorporated herein by reference, discloses an array of split-gate non-volatile memory cells, which is a type of flash memory cell. Such memory cells 210 are Figure 2 shown in. Each memory cell 210 includes a source region 14 and a drain region 16 formed in a semiconductor substrate 12, with a channel region 18 therebetween. A floating gate 20 is formed above and insulated from (and controls the conductivity of) a first portion of the channel region 18, and is formed above a portion of the source region 14. A word line terminal 22 (which is typically coupled to a word line) has a first portion disposed above a second portion of the channel region 18 and insulated from (and controls the conductivity of) the second portion of the channel region, and a second portion extending upward and located above the floating gate 20. The floating gate 20 and the word line terminal 22 are insulated from the substrate 12 by a gate oxide. A bit line terminal 24 is coupled to the drain region 16.

[0076] The memory cell 210 is erased by placing a high positive voltage on the word line terminal 22 (where electrons are removed from the floating gate), which causes electrons on the floating gate 20 to tunnel through the intermediate insulator from the floating gate 20 to the word line terminal 22 via Fowler-Nordheim tunneling.

[0077] The memory cell 210 is programmed by placing a positive voltage on the word line terminal 22 and a positive voltage on the source region 14 (where electrons are placed on the floating gate). An electron current will flow from the source region 14 (source line terminal) to the drain region 16. When the electrons reach the gap between the word line terminal 22 and the floating gate 20, the electrons will accelerate and heat up. Due to the electrostatic attraction from the floating gate 20, some of the heated electrons will be injected onto the floating gate 20 through the gate oxide.

[0078] The memory cell 210 is read by placing a positive read voltage on the drain region 16 and the word line terminal 22 that turns on the portion of the channel region 18 below the word line terminal. If the floating gate 20 is positively charged (i.e., electrons are erased), then the portion of the channel region 18 below the floating gate 20 is also turned on, and current will flow through the channel region 18, which is sensed as the erased state or the "1" state. If the floating gate 20 is negatively charged (i.e., programmed with electrons), then the portion of the channel region 18 below the floating gate 20 is mostly or completely turned off, and current will not (or very little current) flow through the channel region 18, which is sensed as the programmed state or the "0" state.

[0079] Table 1 shows the typical voltage ranges that can be applied to the terminals of the memory cell 110 for performing read, erase, and program operations:

[0080] Table 1: Figure 2 Operation of the flash memory cell 210

[0081] WL BL SL Read 1 0.5-3V 0.1-2V 0V Read 2 0.5-3V 0-2V 2-0.1V Erase About 11-13V 0V 0V programming 1-2V 1-3μA 9-10V

[0082] "Read 1" is a read mode in which the cell current is output on the bit line. "Read 2" is a read mode in which the cell current is output on the source line terminal.

[0083] Figure 3 The memory cell 310 is shown, which is similar to the Figure 2 memory cell 210, but with the addition of a control gate (CG) terminal 28. The control gate terminal 28 is biased at a high voltage (e.g., 10 V) during programming, at a low voltage or negative voltage (e.g., 0 V / -8 V) during erase, and at a low voltage or medium voltage (e.g., 0 V / 2.5 V) during read. The other terminals are biased similar to Figure 2 as such.

[0084] Figure 4 The four-gate memory cell 410 is shown, which includes a source region 14, a drain region 16, a floating gate 20 above a first portion of the channel region 18, a select gate 22 (commonly coupled to the word line WL) above a second portion of the channel region 18, a control gate 28 above the floating gate 20, and an erase gate 30 above the source region 14. This configuration is described in U.S. Patent 6,747,310, which is incorporated herein by reference for all purposes. Here, all gates except the floating gate 20 are non-floating gates, meaning they are electrically connected or can be electrically connected to a voltage source. Programming is performed by hot electrons from the channel region 18 that inject themselves into the floating gate 20. Erase is performed by electrons tunneling from the floating gate 20 to the erase gate 30.

[0085] Table 2 shows the typical voltage ranges that can be applied to the terminals of memory cell 410 for performing read operations, erase operations, and program operations:

[0086] Table 2: Figure 4 Operation of the flash memory cell 410

[0087] WL / SG BL CG EG SL Read 1 0.5-2V 0.1-2V 0-2.6V 0-2.6V 0V Read 2 0.5-2V 0-2V 0-2.6V 0-2.6V 2-0.1V Erase -0.5V / 0V 0V 0V / -8V 8-12V 0V programming 1V 1μA 8-11V 4.5-9V 4.5-5V

[0088] "Read 1" is a read mode in which the cell current is output on the bit line. "Read 2" is a read mode in which the cell current is output on the source line terminal.

[0089] Figure 5 Memory cell 510 is shown. Except for not including an erase gate EG terminal, memory cell 510 is similar to Figure 4 memory cell 410. Erasure is performed by biasing the substrate 18 to a high voltage and biasing the control gate CG terminal 28 to a low voltage or a negative voltage. Alternatively, erasure is performed by biasing the word line terminal 22 to a positive voltage and biasing the control gate terminal 28 to a negative voltage. Programming and reading are similar to Figure 4 that of

[0090] Figure 6 Trigate memory cell 610 is shown, which is another type of flash memory cell. Memory cell 610 is the same as Figure 4 memory cell 410, except that memory cell 610 does not have a separate control gate terminal. Except for not applying a control gate bias, the erase operation (erasing by using the erase gate terminal) and the read operation are similar to Figure 4 those of

[0091] Table 3 shows the typical voltage ranges that can be applied to the terminals of memory cell 610 for performing read operations, erase operations, and program operations:

[0092] Table 3: Figure 6 Operation of flash memory cell 610

[0093] WL / SG BL EG SL Read 1 0.5-2.2V 0.1-2V 0-2.6V 0V Read 2 0.5-2.2V 0-2V 0-2.6V 2-0.1V Erase -0.5V / 0V 0V 11.5V 0V programming 1V 2-3μA 4.5V 7-9V

[0094] "Read 1" is a read mode in which the cell current is output on the bit line. "Read 2" is a read mode in which the cell current is output on the source line terminal.

[0095] Figure 7 Stacked gate memory cell 710 is shown, which is another type of flash memory cell. Memory cell 710 is the same as Figure 2The memory cell 210 is similar, except that the floating gate 20 extends over the entire channel region 18, and the control gate terminal 22 (which will be coupled to the word line here) extends over the floating gate 20, separated by an insulating layer (not shown). The erase, program, and read operations operate in a manner similar to that described previously for the memory cell 210.

[0096] Table 4 shows the typical voltage ranges that can be applied to the terminals of the memory cell 710 and the substrate 12 for performing read, erase, and program operations:

[0097] Table 4: Figure 7 Operation of flash memory cell 710

[0098] CG BL SL Substrate Read 1 0-5V 0.1–2V 0-2V 0V Read 2 0.5-2V 0-2V 2-0.1V 0V Erase -8 to -10V / 0V FLT FLT 8-10V / 15-20V programming 8-12V 3-5V / 0V 0V / 3-5V 0V

[0099] "Read 1" is a read mode in which the cell current is output on the bit line. "Read 2" is a read mode in which the cell current is output on the source line terminal. Optionally, in an array including rows and columns of memory cells 210, 310, 410, 510, 610, or 710, the source line can be coupled to one row of memory cells or two adjacent rows of memory cells. That is, the source line terminal can be shared by memory cells in adjacent rows.

[0100] To utilize a memory array including one of the above types of non-volatile memory cells in an artificial neural network, two modifications are made. First, the circuitry is configured such that each memory cell can be individually programmed, erased, and read without adversely affecting the memory states of other memory cells in the array, as further explained below. Second, continuous (analog) programming of the memory cells is provided.

[0101] Specifically, the memory state of each memory cell in the array (i.e., the charge on the floating gate) can be continuously changed from a fully erased state to a fully programmed state independently and with minimal interference to other memory cells. In another embodiment, the memory state of each memory cell in the array (i.e., the charge on the floating gate) can be continuously changed from a fully programmed state to a fully erased state independently and with minimal interference to other memory cells, and vice versa. This means that the cell storage device is analog or can at least store one of a number of discrete values (such as 16 or 64 different values), which allows for very precise and individual tuning of all the cells in the memory array, and which makes the memory array ideal for storing and fine-tuning the synaptic weights of a neural network.

[0102] The methods and apparatuses described herein can be applied to other non-volatile memory technologies such as, but not limited to, SONOS (Silicon-Oxide-Nitride-Oxide-Silicon, charge trapping in nitride), MONOS (Metal-Oxide-Nitride-Oxide-Silicon, metal charge trapping in nitride), ReRAM (Resistive RAM), PCM (Phase Change Memory), MRAM (Magnetic RAM), FeRAM (Ferroelectric RAM), OTP (One-Time Programmable in bilayer or multilayer), and CeRAM (Correlated Electron RAM), etc. The methods and apparatuses described herein can be applied to volatile memory technologies for neural networks such as, but not limited to, SRAM, DRAM, and / or volatile synaptic units.

[0103] Neural Network Using Non-Volatile Memory Cell Array

[0104] Figure 8 Conceptually illustrates a non-limiting example of a neural network using a non-volatile memory array in this embodiment. This example uses a non-volatile memory array neural network for a face recognition application, but any other suitable application can also be implemented using a neural network based on a non-volatile memory array.

[0105] For this example, S0 is the input layer, which is a 32x32 pixel RGB image with 5-bit precision (i.e., three 32x32 pixel arrays, one for each color R, G, and B, and each pixel is 5-bit precision). The synapses CB1 from the input layer S0 to the layer C1 apply different weight sets in some cases and shared weights in other cases, and scan the input image with a 3x3 pixel overlapping filter (kernel), shifting the filter by 1 pixel (or more than 1 pixel as indicated by the model). Specifically, the values of 9 pixels in a 3x3 portion of the image (i.e., called the filter or kernel) are provided to the synapses CB1, where these 9 input values are multiplied by appropriate weights, and after summing the output of this multiplication, a single output value is determined and provided by the first synapse of CB1 for a pixel in one of the layers C1 of the generated feature map. Then the 3x3 filter is shifted one pixel to the right in the input layer S0 (i.e., adding the column of three pixels on the right and releasing the column of three pixels on the left), whereby the 9 pixel values in this newly positioned filter are provided to the synapses CB1, where they are multiplied by the same weights and a second single output value is determined by the associated synapse. This process continues until the 3x3 filter scans all three colors and all bits (precision values) of the entire 32x32 pixel image in the input layer S0. Then this process is repeated using different sets of weights to generate different feature maps of C1 until all the feature maps of layer C1 are calculated.

[0106] At layer C1, in this example, there are 16 feature maps, each feature map having 30x30 pixels. Each pixel is a new feature pixel extracted from the product of the input and the kernel, so each feature map is a two-dimensional array. Thus, in this example, layer C1 consists of 16 layers of two-dimensional arrays (remember that the layers and arrays referred to in this text are logical relationships and do not have to be physical relationships, i.e., the arrays do not have to be oriented as physical two-dimensional arrays). Each of the 16 feature maps in layer C1 is generated by one of a set of sixteen different sets of synaptic weights applied to the filter scan. The C1 feature maps can all relate to different aspects of the same image feature, such as boundary recognition. For example, the first map (generated using the first set of weights, shared for all scans used to generate that first map) can identify circular edges, the second map (generated using a second set of weights different from the first set) can identify rectangular edges, or the aspect ratio of certain features, and so on.

[0107] Before transitioning from layer C1 to layer S1, the activation function P1 (pooling) is applied, which pools the values from consecutive non-overlapping 2x2 regions in each feature map. The purpose of the pooling function is to take the mean of neighboring positions (or the max function can also be used) to, for example, reduce the dependence on edge positions and reduce the data size before entering the next stage. At layer S1, there are 16 15x15 feature maps (i.e., sixteen different arrays each with 15x15 pixels per feature map). The synapses CB2 from layer S1 to layer C2 scan the maps in S1 using a 4x4 filter, with the filter shifted by 1 pixel. At layer C2, there are 22 12x12 feature maps. Before transitioning from layer C2 to layer S2, the activation function P2 (pooling) is applied, which pools the values from consecutive non-overlapping 2x2 regions in each feature map. At layer S2, there are 22 6x6 feature maps. The activation function (pooling) is applied to the synapses CB3 from layer S2 to layer C3, where each neuron in layer C3 is connected via the corresponding synapses of CB3 to each map in layer S2. At layer C3, there are 64 neurons. The synapses CB4 from layer C3 to the output layer S3 fully connect C3 to S3, i.e., each neuron in layer C3 is connected to each neuron in layer S3. The output at S3 includes 10 neurons, and the neuron with the highest output determines the class. For example, this output can indicate the recognition or classification of the content of the original image.

[0108] Each layer of synapses is implemented using an array or a part of an array of non-volatile memory cells.

[0109] Fig. 9 Is a block diagram of a system that can be used for this purpose. The vector-matrix multiplication (VMM) system 32 includes non-volatile memory cells and serves as the synapses between one layer and the next layer (such as Figure 6CB1, CB2, CB3, and CB4 in). Specifically, the VMM system 32 includes a VMM array 33 (including non-volatile memory cells arranged in rows and columns), an erase gate and word line gate decoder 34, a control gate decoder 35, a bit line decoder 36, and a source line decoder 37, which decode corresponding inputs of the non-volatile memory cell array 33. The inputs to the VMM array 33 can come from the erase gate and word line gate decoder 34 or from the control gate decoder 35. In this example, the source line decoder 37 also decodes the output of the VMM array 33. Alternatively, the bit line decoder 36 can decode the output of the VMM array 33.

[0110] The VMM array 33 serves two purposes. First, it stores the weights to be used by the VMM system 32. Second, the VMM array 33 effectively multiplies the inputs with the weights stored in the VMM array 33 and each output line (source line or bit line) sums them to produce an output, which will be used as the input to the next layer or the input to the final layer. By performing the multiplication and addition functions, the VMM array 33 eliminates the need for separate multiplication and addition logic circuits and is also highly efficient due to its in-situ memory computing.

[0111] The output of the VMM array 33 is provided to a differential summing device (such as a summing operational amplifier or a summing current mirror) 38, which sums the output of the VMM array 33 to create a single value for this convolution. The differential summing device 38 is arranged to perform the summation of both positive and negative weight inputs to output a single value.

[0112] Then the output value of the differential summing device 38 is provided to an activation function circuit 39 after summation, and this activation function circuit corrects the output. The activation function circuit 39 can provide sigmoid, tanh, ReLU functions, or any other non-linear function. The corrected output value of the activation function circuit 39 becomes an element of the feature map for the next layer (e.g., Figure 8 layer C1 in), and then is applied to the next synapse to produce the next feature map layer or the final layer. Thus, in this example, the VMM array 33 constitutes multiple synapses (which receive their inputs from an existing neuron layer or from an input layer such as an image database), and the summing device 38 and the activation function circuit 39 constitute multiple neurons.

[0113] Fig. 9The inputs (WLx, EGx, CGx, and optionally BLx and SLx) to the VMM system 32 can be analog levels, binary levels, digital pulses (in which case a pulse - analog converter PAC may be required to convert the pulses to a suitable input analog level), or digital bits (in which case a DAC is provided to convert the digital bits to a suitable input analog level); the outputs can be analog levels, binary levels, digital pulses, or digital bits (in which case an output ADC is provided to convert the output analog level to digital bits).

[0114] Fig.10 A block diagram showing the use of a multi - layer VMM system 32 (here labeled VMM systems 32a, 32b, 32c, 32d, and 32e). As Fig.10 shown, the input (represented as Inputx) is converted from digital to analog by a digital - to - analog converter 31 and provided to the input VMM system 32a. The converted analog input can be voltage or current. The input D / A conversion of the first layer can be done by using a function or a LUT (look - up table) that maps the input Inputx to the appropriate analog level of a matrix multiplier for the input VMM system 32a. The input conversion can also be done by an analog - to - analog (A / A) converter to convert an external analog input to a mapped analog input to the input VMM system 32a. The input conversion can also be done by a digital - to - digital pulse (D / P) converter to convert an external digital input to one or more mapped digital pulses to the input VMM system 32a.

[0115] The output generated by the input VMM system 32a is provided as an input to the next VMM system (hidden level 1) 32b, which in turn generates an output provided as an input to the next VMM system (hidden level 2) 32c, and so on. Each layer of the VMM system 32 serves as a different layer of synapses and neurons of a convolutional neural network (CNN). Each of the VMM systems 32a, 32b, 32c, 32d, and 32e can be an independent physical system including a corresponding non - volatile memory array, or multiple VMM systems can utilize different parts of the same physical non - volatile memory array, or multiple VMM systems can utilize overlapping parts of the same physical non - volatile memory array. Each of the VMM systems 32a, 32b, 32c, 32d, and 32e can also be time - division multiplexed for different parts of its array or neurons. Fig.10 The example shown includes five layers (32a, 32b, 32c, 32d, 32e): an input layer (32a), two hidden layers (32b, 32c), and two fully - connected layers (32d, 32e). A person of ordinary skill in the art will know that this is merely exemplary, and conversely, the system can include more than two hidden layers and more than two fully - connected layers.

[0116] VMM Array

[0117] Fig.11 shows a neuron VMM array 1100, which is particularly suitable for Figure 3 the memory cell 310 shown, and serves as the synapse and component of the neuron between the input layer and the next layer. The VMM array 1100 includes a memory array 1101 of non-volatile memory cells and a reference array 1102 of non-volatile reference memory cells (at the top of the array). Alternatively, another reference array can be placed at the bottom.

[0118] In the VMM array 1100, control gate lines (such as control gate line 1103) extend in the vertical direction (so the reference array 1102 is orthogonal to the control gate line 1103 in the row direction), and erase gate lines (such as erase gate line 1104) extend in the horizontal direction. Here, the inputs of the VMM array 1100 are set on the control gate lines (CG0, CG1, CG2, CG3), and the outputs of the VMM array 1100 appear on the source lines (SL0, SL1). In one embodiment, only even rows are used, and in another embodiment, only odd rows are used. The current placed on each source line (SL0 and SL1 respectively) performs the summation function of all the currents from the memory cells connected to that particular source line.

[0119] As described herein for neural networks, the non-volatile memory cells of the VMM array 1100 (i.e., the flash memory of the VMM array 1100) are preferably configured to operate in the subthreshold region.

[0120] The non-volatile reference memory cells and non-volatile memory cells described herein are biased in weak inversion:

[0121] Ids = Io * e (Vg-Vth) / nVt = w * Io * e (Vg) / nVt ,

[0122] where w = e (-Vth) / nVt

[0123] where Ids is the drain-to-source current; Vg is the gate voltage on the memory cell; Vth is the threshold voltage of the memory cell; Vt is the thermal voltage = k * T / q, where k is the Boltzmann constant, T is the temperature in Kelvin, and q is the electron charge; n is the slope factor = 1 + (Cdep / Cox), where Cdep = the capacitance of the depletion layer, and Cox is the capacitance of the gate oxide layer; Io is the memory cell current at the gate voltage equal to the threshold voltage, and Io is related to (Wt / L) * u * Cox * (n - 1) * Vt 2Proportional, where u is the carrier mobility, and Wt and L are the width and length of the memory cell, respectively.

[0124] For an I-to-V logarithmic converter that uses a memory cell (such as a reference memory cell or a peripheral memory cell) or a transistor to convert an input current Ids into an input voltage Vg:

[0125] Vg = n * Vt * log[Ids / wp * Io]

[0126] Here, wp is the w of the reference memory cell or the peripheral memory cell.

[0127] For an I-to-V logarithmic converter that uses a memory cell (such as a reference memory cell or a peripheral memory cell) or a transistor to convert an input current Ids into an input voltage Vg:

[0128] Vg = n * Vt * log[Ids / wp * Io]

[0129] Here, wp is the w of the reference memory cell or the peripheral memory cell.

[0130] For a memory array used as a vector matrix multiplier (VMM) array, the output current is:

[0131] Iout = wa * Io * e (Vg) / nVt , that is

[0132] Iout = (wa / wp) * Iin = W * Iin

[0133] W = e (Vthp-Vtha) / nVt

[0134] Iin = wp * Io * e (Vg) / nVt

[0135] Here, wa = w of each memory cell in the memory array.

[0136] The word line or the control gate can be used as the input of the memory cell for the input voltage.

[0137] Alternatively, the non-volatile memory cells of the VMM array described herein can be configured to operate in the linear region:

[0138] Ids = β * (Vgs - Vth) * Vds; β = u * Cox * Wt / L,

[0139] W α (Vgs - Vth),

[0140] meaning that the weight W in the linear region is proportional to (Vgs - Vth)

[0141] A word line, a control gate, a bit line, or a source line can be used as an input to a memory cell operating in the linear region. A bit line or a source line can be used as an output of the memory cell.

[0142] For an I-to-V linear converter, a memory cell (such as a reference memory cell or a peripheral memory cell), a transistor, or a resistor operating in the linear region can be used to linearly convert an input / output current into an input / output voltage.

[0143] Alternatively, the memory cells of the VMM array described herein can be configured to operate in the saturation region:

[0144] Ids = 1 / 2 * β * (Vgs - Vth) 2 ; β = u * Cox * Wt / L

[0145] W α (Vgs - Vth) 2 , meaning that the weight W is proportional to (Vgs - Vth) 2 is proportional

[0146] A word line, a control gate, or an erase gate can be used as an input to a memory cell operating in the saturation region. A bit line or a source line can be used as an output of the output neuron.

[0147] Alternatively, the memory cells of the VMM array described herein can be used in all regions or combinations thereof (subthreshold, linear, or saturation regions).

[0148] Other embodiments of the VMM array 33 described in U.S. Patent Application No. 15 / 826,345 are described, which application is incorporated herein by reference. As described herein, a source line or a bit line can be used as a neuron output (current summing output). Fig. 9 is incorporated herein by reference. As described herein, a source line or a bit line can be used as a neuron output (current summing output).

[0149] Fig.12 shows a neuron VMM array 1200, which neuron VMM array is particularly suitable for Figure 2The memory cell 210 shown is used as a synapse between the input layer and the next layer. The VMM array 1200 includes a memory array 1203 of non-volatile memory cells, a reference array 1201 of first non-volatile reference memory cells, and a reference array 1202 of second non-volatile reference memory cells. The reference arrays 1201 and 1202 arranged along the column direction of the array are used to convert the current inputs flowing into the terminals BLR0, BLR1, BLR2, and BLR3 into voltage inputs WL0, WL1, WL2, and WL3. In fact, the first non-volatile reference memory cells and the second non-volatile reference memory cells are diode-connected through a multiplexer 1214 (only partially shown), into which the current inputs flow. The reference cells are tuned (e.g., programmed) to a target reference level. The target reference level is provided by a reference microarray matrix (not shown).

[0150] The memory array 1203 serves two purposes. First, it stores the weights to be used by the VMM array 1200 in its corresponding memory cells. Second, the memory array 1203 effectively multiplies the inputs (i.e., the current inputs provided at the terminals BLR0, BLR1, BLR2, and BLR3, which the reference arrays 1201 and 1202 convert into input voltages to be provided to the word lines WL0, WL1, WL2, and WL3) by the weights stored in the memory array 1203, and then sums all the results (memory cell currents) to produce an output on the corresponding bit lines (BL0 - BLN), which will be the input to the next layer or the input to the final layer. By performing the multiplication and addition functions, the memory array 1203 eliminates the need for separate multiplication logic circuits and addition logic circuits and is also highly efficient. Here, the voltage inputs are provided on the word lines (WL0, WL1, WL2, and WL3), and the outputs appear on the corresponding bit lines (BL0 - BLN) during a read (inference) operation. The current placed on each of the bit lines BL0 - BLN performs a summation function of the currents from all the non-volatile memory cells connected to that particular bit line.

[0151] Table 5 shows the operating voltages for the VMM array 1200. The columns in the table indicate the voltages placed on the word lines for the selected cells, the word lines for the unselected cells, the bit lines for the selected cells, the bit lines for the unselected cells, the source lines for the selected cells, and the source lines for the unselected cells, where FLT indicates floating, i.e., no voltage is applied. The rows indicate the read, erase, and program operations.

[0152] Table 5: Fig.12 Operation of the VMM array 1200

[0153] WL WL-Not selected BL BL-Not selected SL SL-Not selected Read 0.5-3.5V -0.5V / 0V 0.1-2V(Ineuron) 0.6V-2V / FLT 0V 0V Erase About 5-13V 0V 0V 0V 0V 0V programming 1-2V -0.5V / 0V 0.1-3uA Vinh about 2.5V 4-10V 0-1V / FLT

[0154] Fig.13A neuron VMM array 1300 is shown, which is particularly suitable for Figure 2 the memory cell 210 shown, and serves as the synapse and component of the neuron between the input layer and the next layer. The VMM array 1300 includes a memory array 1303 of non-volatile memory cells, a reference array 1301 of first non-volatile reference memory cells, and a reference array 1302 of second non-volatile reference memory cells. The reference arrays 1301 and 1302 extend in the row direction of the VMM array 1300. The VMM array is similar to the VMM 1000, except that in the VMM array 1300, the word lines extend in the vertical direction. Here, the inputs are set on the word lines (WLA0, WLB0, WLA1, WLB2, WLA2, WLB2, WLA3, WLB3), and the outputs appear on the source lines (SL0, SL1) during a read operation. The current placed on each source line performs a summation function of all the currents from the memory cells connected to that particular source line.

[0155] Table 6 shows the operating voltages for the VMM array 1300. The columns in the table indicate the voltages placed on the word lines for the selected cells, the word lines for the unselected cells, the bit lines for the selected cells, the bit lines for the unselected cells, the source lines for the selected cells, and the source lines for the unselected cells. The rows indicate the read, erase, and program operations.

[0156] Table 6: Fig.13 Operation of the VMM array 1300

[0157] WL WL-Not selected BL BL-Not selected SL SL-Not selected Read 0.5-3.5V -0.5V / 0V 0.1-2V 0.1V-2V / FLT About 0.3-1V (Ineuron) 0V Erase About 5-13V 0V 0V 0V 0V SL-Prohibit (about 4-8V) programming 1-2V -0.5V / 0V 0.1-3uA Vinh about 2.5V 4-10V 0-1V / FLT

[0158] Fig.14 A neuron VMM array 1400 is shown, which is particularly suitable for Figure 3The memory cell 310 shown and serves as the synapse and component of the neuron between the input layer and the next layer. The VMM array 1400 includes a memory array 1403 of non-volatile memory cells, a reference array 1401 of first non-volatile reference memory cells, and a reference array 1402 of second non-volatile reference memory cells. The reference arrays 1401 and 1402 are used to convert the current inputs flowing into the terminals BLR0, BLR1, BLR2, and BLR3 into voltage inputs CG0, CG1, CG2, and CG3. In fact, the first non-volatile reference memory cells and the second non-volatile reference memory cells are diode-connected through a multiplexer 1412 (only partially shown), and the current inputs flow into them through BLR0, BLR1, BLR2, and BLR3. Each multiplexer 1412 includes a corresponding multiplexer 1405 and a cascode transistor 1404 to ensure a constant voltage on the bit line (such as BLR0) of each of the first non-volatile reference memory cells and the second non-volatile reference memory cells during the read operation. The reference cells are tuned to a target reference level.

[0159] The memory array 1403 serves two purposes. First, it stores the weights to be used by the VMM array 1400. Second, the memory array 1403 effectively multiplies the inputs (the current inputs provided to the terminals BLR0, BLR1, BLR2, and BLR3, which the reference arrays 1401 and 1402 convert into input voltages to be provided to the control gates CG0, CG1, CG2, and CG3) by the weights stored in the memory array, and then sums all the results (cell currents) to produce an output that appears on BL0 - BLN and will be the input to the next layer or the input to the final layer. By performing the multiplication and addition functions, the memory array eliminates the need for separate multiplication and addition logic circuits and is also highly efficient. Here, the inputs are provided on the control gate lines (CG0, CG1, CG2, and CG3), and the outputs appear on the bit lines (BL0–BLN) during the read operation. The current placed on each bit line performs the summation function of all the currents from the memory cells connected to that particular bit line.

[0160] The VMM array 1400 implements unidirectional tuning for the non-volatile memory cells in the memory array 1403. That is, each non-volatile memory cell is erased and then partially programmed until the desired charge on the floating gate is reached. This can be performed, for example, using the precise programming techniques described below. If too much charge is placed on the floating gate (such that an incorrect value is stored in the cell), the cell must be erased and the sequence of partial programming operations must be restarted. As shown, two rows sharing the same erase gate (such as EG0 or EG1) need to be erased together (which is referred to as a page erase), and thereafter, each cell is partially programmed until the desired charge on the floating gate is reached.

[0161] Table 7 shows the operating voltages for the VMM array 1400. The columns in the table indicate the voltages placed on the word line for the selected cell, the word line for the unselected cell, the bit line for the selected cell, the bit line for the unselected cell, the control gate for the selected cell, the control gate for the unselected cell in the same sector as the selected cell, the control gate for the unselected cell in a different sector from the selected cell, the erase gate for the selected cell, the erase gate for the unselected cell, the source line for the selected cell, and the source line for the unselected cell. The rows indicate read, erase, and program operations.

[0162] Table 7: Fig.14 Operation of the VMM array 1400

[0163]

[0164] Fig.15 shows the neuron VMM array 1500, which is particularly suitable for Figure 3The memory cell 310 shown and serves as the synapse and component of the neuron between the input layer and the next layer. The VMM array 1500 includes a memory array 1503 of non-volatile memory cells, a reference array 1501 of first non-volatile reference memory cells, and a reference array 1502 of second non-volatile reference memory cells. The EG lines EGR0, EG0, EG1, and EGR1 extend vertically, while the CG lines CG0, CG1, CG2, and CG3 and the SL lines WL0, WL1, WL2, and WL3 extend horizontally. The VMM array 1500 is similar to the VMM array 1400, except that the VMM array 1500 implements bidirectional tuning, where each individual cell can be fully erased, partially programmed, and partially erased as needed to achieve the desired charge amount on the floating gate due to the use of individual EG lines. As shown, the reference arrays 1501 and 1502 convert the input current in the terminals BLR0, BLR1, BLR2, and BLR3 into the control gate voltages CG0, CG1, CG2, and CG3 to be applied to the memory cells in the row direction (by the action of the diode-connected reference cells via the multiplexer 1514). The current outputs (neurons) are in the bit lines BL0 - BLN, where each bit line sums all the currents from the non-volatile memory cells connected to that particular bit line.

[0165] Table 8 shows the operating voltages for the VMM array 1500. The columns in the table indicate the voltages placed on the word line for the selected cell, the word line for the unselected cell, the bit line for the selected cell, the bit line for the unselected cell, the control gate for the selected cell, the control gate for the unselected cell in the same sector as the selected cell, the control gate for the unselected cell in a different sector from the selected cell, the erase gate for the selected cell, the erase gate for the unselected cell, the source line for the selected cell, and the source line for the unselected cell. The rows indicate the read, erase, and program operations.

[0166] Table 8: Fig.15 Operation of the VMM array 1500

[0167]

[0168] Fig.16 Shows the neuron VMM array 1600, which is particularly suitable for Figure 2 the memory cell 210 shown and serves as the synapse and component of the neuron between the input layer and the next layer. In the VMM array 1600, the inputs INPUT0…, INPUT N are received on the bit lines BL0,... BL N respectively, and the outputs OUTPUT1, OUTPUT2, OUTPUT3, and OUTPUT4 are generated on the source lines SL0, SL1, SL2, and SL3 respectively.

[0169] Fig.17 Shows a neuron VMM array 1700, which is particularly suitable for Figure 2 the memory cell 210 shown, and serves as the synapse and component of the neuron between the input layer and the next layer. In this example, the inputs INPUT0, INPUT1, INPUT2, and INPUT3 are received on the source lines SL0, SL1, SL2, and SL3 respectively, and the outputs OUTPUT0,... OUTPUT N are generated on the bit lines BL0,…,BL N .

[0170] Fig.18 Shows a neuron VMM array 1800, which is particularly suitable for Figure 2 the memory cell 210 shown, and serves as the synapse and component of the neuron between the input layer and the next layer. In this example, the inputs INPUT0,…,INPUT M are received on the word lines WL0,…,WL M respectively, and the outputs OUTPUT0,..OUTPUT N are generated on the bit lines BL0,…,BL N .

[0171] Fig.19 Shows a neuron VMM array 1900, which is particularly suitable for Figure 3 the memory cell 310 shown, and serves as the synapse and component of the neuron between the input layer and the next layer. In this example, the inputs INPUT0,…,INPUT M are received on the word lines WL0,…,WL M respectively, and the outputs OUTPUT0,..OUTPUT N are generated on the bit lines BL0,…,BL N .

[0172] Fig. 20 Shows a neuron VMM array 2000, which is particularly suitable for Figure 4 the memory cell 410 shown, and serves as the synapse and component of the neuron between the input layer and the next layer. In this example, the inputs INPUT0,…,INPUT n are received on the vertical control gate lines CG0,…,CG N respectively, and the outputs OUTPUT1 and OUTPUT2 are generated on the source lines SL0 and SL1.

[0173] Fig.21Shows a neuron VMM array 2100, which is particularly suitable for Figure 4 the memory cell 410 shown, and serves as the synapse and component of the neuron between the input layer and the next layer. In this example, the inputs INPUT0 to INPUT N are received on the gates of bit line control gates 2901-1, 2901-2 to 2901-(N-1) and 2901-N respectively, and these gates are coupled to bit lines BL0 to BL N . Exemplary outputs OUTPUT1 and OUTPUT2 are generated on source lines SL0 and SL1.

[0174] Fig. 22 Shows a neuron VMM array 2200, which is particularly suitable for Figure 3 the memory cell 310 shown, Figure 5 the memory cell 510 shown, and Figure 7 the memory cell 710 shown, and serves as the synapse and component of the neuron between the input layer and the next layer. In this example, the inputs INPUT0,…,INPUT M are received on word lines WL0,…,WL M , and the outputs OUTPUT0,…,OUTPUT N are generated on bit lines BL0,…,BL N respectively.

[0175] Fig.23 Shows a neuron VMM array 2300, which is particularly suitable for Figure 3 the memory cell 310 shown, Figure 5 the memory cell 510 shown, and Figure 7 the memory cell 710 shown, and serves as the synapse and component of the neuron between the input layer and the next layer. In this example, the inputs INPUT0 to INPUT M are received on control gates CG0 to CG M . The outputs OUTPUT0,…,OUTPUT N are generated on vertical source lines SL0,…,SL N respectively, where each source line SL i is coupled to the source lines of all memory cells in column i.

[0176] Fig.24 Shows a neuron VMM array 2400, which is particularly suitable for Figure 3 the memory cell 310 shown, Figure 5 the memory cell 510 shown, and Figure 7The memory cell 710 shown and serves as the synapse and component of the neuron between the input layer and the next layer. In this example, the inputs INPUT0 to INPUT M are received on the control gate lines CG0 to CG M The outputs OUTPUT0, …, OUTPUT N are generated on the vertical bit lines BL0, …, BL N respectively, where each bit line BL i is coupled to the bit lines of all memory cells in column i.

[0177] Long Short-Term Memory

[0178] The prior art includes the concept known as long short-term memory (LSTM). LSTM is commonly used in artificial neural networks. LSTM allows artificial neural networks to remember information over an arbitrary predetermined time interval and use that information in subsequent operations. Conventional LSTM includes cells, an input gate, an output gate, and a forget gate. The three gates regulate the flow of information into and out of the cell and the time interval during which information is remembered in the LSTM. VMM is particularly useful in LSTM.

[0179] Fig.25 An exemplary LSTM 2500 is shown. The LSTM 2500 in this example includes cells 2501, 2502, 2503, and 2504. Cell 2501 receives the input vector x0 and generates the output vector h0 and the cell state vector c0. Cell 2502 receives the input vector x1, the output vector (hidden state) h0 from cell 2501, and the cell state c0 from cell 2501, and generates the output vector h1 and the cell state vector c1. Cell 2503 receives the input vector x2, the output vector (hidden state) h1 from cell 2502, and the cell state c1 from cell 2502, and generates the output vector h2 and the cell state vector c2. Cell 2504 receives the input vector x3, the output vector (hidden state) h2 from cell 2503, and the cell state c2 from cell 2503, and generates the output vector h3. Additional cells can be used, and an LSTM with four cells is merely an example.

[0180] Fig.26 An exemplary embodiment of the LSTM cell 2600 that can be used for Fig.25 the cells 2501, 2502, 2503, and 2504 is shown. The LSTM cell 2600 receives the input vector x(t), the cell state vector c(t−1) from the previous cell, and the output vector h(t−1) from the previous cell, and generates the cell state vector c(t) and the output vector h(t).

[0181] The LSTM cell 2600 includes sigmoid function devices 2601, 2602, and 2603, each of which applies a number between 0 and 1 to control the amount of each component in the input vector that is allowed to pass through to the output vector. The LSTM cell 2600 also includes tanh devices 2604 and 2605 for applying the hyperbolic tangent function to the input vector, multiplier devices 2606, 2607, and 2608 for multiplying two vectors together, and adder device 2609 for adding two vectors together. The output vector h(t) can be provided to the next LSTM cell in the system, or it can be accessed for other purposes.

[0182] Fig. 27 The LSTM cell 2700 is shown, which is an example of a specific implementation of the LSTM cell 2600. For the convenience of the reader, the same numbers are used in the LSTM cell 2700 as in the LSTM cell 2600. The sigmoid function devices 2601, 2602, and 2603 and the tanh device 2604 each include a plurality of VMM arrays 2701 and activation circuit blocks 2702. Thus, it can be seen that the VMM arrays are particularly useful in LSTM cells used in certain neural network systems.

[0183] An alternative form of the LSTM cell 2700 (and another example of a specific implementation of the LSTM cell 2600) is shown in Fig.28 In Fig.28 the sigmoid function devices 2601, 2602, and 2603 and the tanh device 2604 share the same physical hardware (VMM array 2801 and activation function block 2802) in a time-division multiplexing manner. The LSTM cell 2800 also includes multiplier device 2803 for multiplying two vectors together, adder device 2808 for adding two vectors together, tanh device 2605 (which includes activation circuit block 2802), register 2807 for storing the value i(t) when the value i(t) is output from the sigmoid function block 2802, register 2804 for storing the value when the value f(t)*c(t - 1) is output from the multiplier device 2803 through the multiplexer 2810, register 2805 for storing the value when the value i(t)*u(t) is output from the multiplier device 2803 through the multiplexer 2810, register 2806 for storing the value when the value o(t)*c~(t) is output from the multiplier device 2803 through the multiplexer 2810, and multiplexer 2809.

[0184] The LSTM cell 2700 includes multiple groups of VMM arrays 2701 and corresponding activation function blocks 2702, while the LSTM cell 2800 only includes one group of VMM arrays 2801 and activation function block 2802, which are used to represent multiple layers in the implementation of the LSTM cell 2800. The LSTM cell 2800 will require less space than the LSTM 2700 because, compared with the LSTM cell 2700, the LSTM cell 2800 only needs 1 / 4 of its space for the VMM and activation function blocks.

[0185] It can also be understood that an LSTM cell generally will include multiple VMM arrays, and each VMM array requires functions provided by some circuit blocks outside the VMM array, such as summing and activation circuit blocks and high-voltage generation blocks. Providing separate circuit blocks for each VMM array will require a large amount of space within the semiconductor device and will be somewhat inefficient. Therefore, the embodiments described below attempt to minimize the circuits required outside the VMM array itself.

[0186] Gate Controlled Recursive Unit

[0187] The analog VMM implementation can be used for GRUs (gated recurrent units). A GRU is a gating mechanism in a recurrent artificial neural network. A GRU is similar to an LSTM, except that a GRU cell generally includes fewer components than an LSTM cell.

[0188] Fig.29 An exemplary GRU 2900 is shown. The GRU 2900 in this example includes cells 2901, 2902, 2903, and 2904. The cell 2901 receives the input vector x0 and generates the output vector h0. The cell 2902 receives the input vector x1, the output vector h0 from the cell 2901, and generates the output vector h1. The cell 2903 receives the input vector x2 and the output vector (hidden state) h1 from the cell 2902 and generates the output vector h2. The cell 2904 receives the input vector x3 and the output vector (hidden state) h2 from the cell 2903 and generates the output vector h3. Additional cells can be used, and a GRU with four cells is just an example.

[0189] Fig.30 Shown can be used for Fig.29Exemplary specific implementations of GRU units 3000 of units 2901, 2902, 2903, and 2904. The GRU unit 3000 receives an input vector x(t) and an output vector h(t-1) from a previous GRU unit and generates an output vector h(t). The GRU unit 3000 includes sigmoid function devices 3001 and 3002, each of which applies a number between 0 and 1 to components from the output vector h(t-1) and the input vector x(t). The GRU unit 3000 also includes a tanh device 3003 for applying the hyperbolic tangent function to the input vector, a plurality of multiplier devices 3004, 3005, and 3006 for multiplying two vectors together, an adder device 3007 for adding two vectors together, and a complementary device 3008 for subtracting the input from 1 to generate the output.

[0190] Fig.31 FIG. shows a GRU unit 3100, which is an example of a specific implementation of the GRU unit 3000. For the convenience of the reader, the same numbers are used in the GRU unit 3100 as in the GRU unit 3000. As Fig.31 shown, the sigmoid function devices 3001 and 3002 and the tanh device 3003 each include a plurality of VMM arrays 3101 and activation function blocks 3102. Thus, it can be seen that VMM arrays are particularly useful in GRU units used in certain neural network systems.

[0191] An alternative form of the GRU unit 3100 (and another example of a specific implementation of the GRU unit 3000) is shown in Fig.32 In Fig.32 the GRU unit 3200 utilizes a VMM array 3201 and an activation function block 3202, which, when configured as a sigmoid function, applies a number between 0 and 1 to control how much of each component in the input vector is allowed to pass through to the output vector. In Fig.32In it, the sigmoid function devices 3001 and 3002, and the tanh device 3003 share the same physical hardware (the VMM array 3201 and the activation function block 3202) in a time-division multiplexing manner. The GRU unit 3200 also includes a multiplier device 3203 that multiplies two vectors together, an adder device 3205 that adds two vectors together, a complementary device 3209 that subtracts the input from 1 to generate an output, a multiplexer 3204, a register 3206 that holds the value when the value h(t - 1)*r(t) is output from the multiplier device 3203 through the multiplexer 3204, a register 3207 that holds the value when the value h(t - 1)*z(t) is output from the multiplier device 3203 through the multiplexer 3204, and a register 3208 that holds the value when the value h^(t)*(1 - z(t)) is output from the multiplier device 3203 through the multiplexer 3204.

[0192] The GRU unit 3100 includes multiple sets of VMM arrays 3101 and activation function blocks 3102, while the GRU unit 3200 only includes a single set of VMM arrays 3201 and activation function blocks 3202, which are used to represent multiple layers in the implementation of the GRU unit 3200. The GRU unit 3200 will require less space than the GRU unit 3100 because, compared with the GRU unit 3100, the GRU unit 3200 only needs 1 / 3 of its space for the VMM and activation function blocks.

[0193] It can also be understood that a system using GRUs will generally include multiple VMM arrays, and each VMM array requires functions provided by certain circuit blocks outside the VMM array (such as summing and activation circuit blocks and high-voltage generation blocks). Providing separate circuit blocks for each VMM array will require a large amount of space within the semiconductor device and will be somewhat inefficient. Therefore, the embodiments described below attempt to minimize the circuits required outside the VMM array itself.

[0194] The input of the VMM array can be an analog level, a binary level, a timing pulse, or a digital bit, and the output can be an analog level, a binary level, a timing pulse, or a digital bit (in this case, an output ADC is required to convert the output analog level current or voltage into digital bits).

[0195] For each memory cell in the VMM array, each weight w can be implemented by a single memory cell, or by a differential cell, or by two hybrid memory cells (the average of 2 or more cells). In the case of a differential cell, two memory cells are required to implement the weight w as a differential weight (w = w+ – w-). In two hybrid memory cells, two memory cells are required to implement the weight w as the average of the two cells.

[0196] Implementation for accurate programming of cells in a VMM

[0197] Embodiments for precisely programming memory cells within a VMM by incrementing or decrementing the programming voltage applied to different terminals of the memory cells will now be described.

[0198] Fig.33A Programming method 3300 is shown. First, the method starts (step 3301), which typically occurs in response to receiving a programming (tuning) command. Next, a bulk programming operation programs all cells to the "0" state (step 3302). Then, a soft erase operation erases all cells to an intermediate weak erase level such that each cell will consume, for example, approximately 3 μA - 5 μA of current during a read operation (step 3303). This is in contrast to the maximum deep erase level where each cell will consume approximately 20 μA - 30 μA of current during a read operation. For example, the soft erase is accomplished by applying incrementing erase voltage pulses until the intermediate cell current is reached. The incrementing erase voltage pulses are executed to limit the degradation caused to the memory cells by hard erase (i.e., the maximum erase level). Then, hard programming is performed on all unselected cells to a very deep programming state to add electrons to the floating gates of the cells (step 3304) to ensure that those cells are truly "off", which means those cells will consume a negligible amount of current during a read operation, such as unused memory cells. Hard programming is performed, for example, with higher incrementing programming voltage pulses and / or longer programming times.

[0199] Then, a coarse programming method is performed on the selected cells (to bring the cells closer to the target, e.g., 2 times - 100 times the target) (step 3305), after which a fine programming method is performed on the selected cells (step 3306) to program the exact value required for each selected cell.

[0200] Fig.33B Another programming method 3310 similar to programming method 3300 is shown. However, instead of the programming operation that programs all cells to the "0" state as in Fig.33A step 3302, after the method starts (step 3301), all cells are erased to the "1" state using a soft erase operation (step 3312). Then, all cells are programmed to an intermediate state using a soft programming operation (step 3313) such that each cell will consume approximately 3 μA - 5 μA of current during a read operation. Thereafter, the coarse programming method 3305 and the fine programming method 3306 are performed as Fig.33A shown. Fig.33BVariations of the implementation will completely remove the soft programming operation (step 3313). Multiple rough programming methods can be used to accelerate programming, such as aiming at multiple gradually smaller rough targets before performing the precise (fine) programming step 3306. The precise programming method 3306 is completed, for example, by programming voltage pulses or constant programming timing pulses in fine (precise) increments.

[0201] Different terminals of the memory cell can be used for the rough programming method 3305 and the precise programming method 3306. That is, during the rough programming method 3305, the voltage applied to one of the terminals of the memory cell (which can be referred to as the rough programming terminal) will change until the desired voltage level is achieved within the floating gate 20, and during the precise programming method 3306, the voltage applied to one of the terminals of the memory cell (which can be referred to as the precise programming terminal) will change until the desired level is achieved. Various combinations of terminals that can be used as rough programming terminals and precise programming terminals are shown in Table 9:

[0202] Table 9: Memory cell terminals for coarse programming method and fine programming method

[0203]

[0204] For the rough programming step and the precise programming step, other combinations of terminals are possible.

[0205] Fig.34 A first implementation of the rough programming method 3305 is shown, which is the search and execute method 3400. First, a lookup table search or function (I-V curve) is performed to determine the rough target current value (I CT )(step 3401) of the selected cell based on the value intended to be stored in the selected cell. For example, the table is created by silicon characteristics or wafer test calibration. The selected cell can be programmed to store one of N possible values (e.g., 128, 64, 32, etc.). Each of the N values corresponds to a different desired current value (I D ) consumed by the selected cell during a read operation. In one implementation, the lookup table can contain M possible current values to be used as the rough target current value I CT of the selected cell during the search and execute method 3400, where M is an integer less than N. For example, if N is 8, then M can be 4, which means there are 8 possible values that the selected cell can store, and one of the 4 rough target current values will be selected as the rough target for the search and execute method 3400. That is, the search and execute method 3400 (which is also an implementation of the rough programming method 3305) aims to quickly program the selected cell to a value (I D ) that is somewhat close to the desired value (I CT), and then the precise programming method 3306 is designed to more precisely program the selected cell to reach or be extremely close to the desired value (I D ).

[0206] For a simple example with N = 8 and M = 4, examples of cell values, desired current values, and rough target current values are shown in Tables 10 and 11:

[0207] Table 10: Example of N expected current values ​​when N = 8

[0208] <![CDATA The value to be stored in the selected cell > <![CDATA Desired current value (I D ) > 000 0.5nA 001 1nA 010 1.5nA 011 2nA 100 2.5nA 101 3nA 110 3.5nA 111 4nA

[0209] Table 11: Example of M target current values ​​when M = 4

[0210] <![CDATA Rough target current value (I CT ) > <![CDATA Associated cell value > <![CDATA[4nA+ I CTOFFSET1 > 000,001 <![CDATA[8nA+ I CTOFFSET2 > 010,011 <![CDATA[12nA+ I CTOFFSET3 > 100,101 <![CDATA[16nA+ I CTOFFSET4 > 110,111

[0211] Offset value I CTOFFSETx is used to prevent exceeding the desired current value during coarse adjustment.

[0212] In step 3402, once the rough target current value I CT is selected, the selected cell is programmed by applying the initial voltage v0 to the rough programming terminal of the selected cell according to one of the sequences listed in Table 9 above. (The value of the initial voltage v0 and the rough programming terminal can optionally be determined from a voltage lookup table storing v0 and the rough target current value I CT ):

[0213] Table 12: Initial voltage v0 applied to the coarse programming terminals

[0214] Program sequence name Coarse programming terminal <![CDATA[Initial voltage v applied to the coarse programming terminal o > CG-CG Control gate 28 (CG) 3-6V CG-EG Control gate 28 (CG) 3-6V EG-EG Erase gate 30 (EG) 0-3V CG-SL Control gate 28 (CG) 3-6V SL-CG Source line 14 (SL) 3-4V SL-EG Source line 14 (SL) 3-4V

[0215] Next, in step 3403, the selected cell is programmed by applying the voltage v i = v i-1 + v increment to the rough programming terminal, where i starts from 1 and increments each time this step is repeated, and where v increment is the increment of the rough voltage that will result in a programming degree suitable for the desired change granularity. Thus, for the first time step 3403, i = 1, and v1 will be v0 + v increment . Then the verification operation occurs (step 3404), where a read operation is performed on the selected cell, and the current (I cell ) consumed by the selected cell is measured. If I cell is less than or equal to I CT (here is the first threshold), then the search and execution of method 3400 are completed, and the precise programming method 3306 can be started. If I cell is not less than or equal to I CTThen, step 3403 is repeated and i is incremented.

[0216] Thus, at the moment when the rough programming method 3305 ends and the precise programming method 3306 begins, the voltage v i will be the final voltage applied to the rough programming terminal to program the selected cell, and the selected cell will store a value associated with the rough target current value I CT , specifically less than or equal to I CT . The purpose of the precise programming method 3306 is to program the selected cell to the extent that during a read operation its consumption current I D (plus or minus an acceptable deviation margin, such as + / - 30% or less), which is the desired current value associated with the value intended to be stored in the selected cell.

[0217] Figure 35 Shows an example of different voltage progressions that can be applied to the rough programming terminal of the selected memory cell during the rough programming method 3305 and / or to the precise programming terminal of the selected memory cell during the precise programming method 3306 according to Table 9.

[0218] In the first method, an increasing voltage is applied progressively to the rough programming terminal and / or the precise programming terminal to further program the selected memory cell. The starting point is v i , which is the final voltage applied during the rough programming method 3305. The increment v p1 is added to v i , and then the voltage v i + v p1 is used to program the selected cell (indicated by the second pulse from the left in progression 3501). v p1 is an increment that is less than v increment (the voltage increment used during the rough programming method 3305). After each programming voltage is applied to the programming terminal, a verification step (similar to step 3404) is performed, where it is determined whether I cell is less than or equal to I PT1 (which is the first precise target current value and here is the second threshold), where I PT1 = I D + I PT1OFFSET , where I PT1OFFSET is an offset value added to prevent programming overshoot. If not, another increment v p1 is added to the previously applied programming voltage, and the process is repeated. Repeat until the moment when I cell is less than or equal to I PT1 , at which point this part of the programming sequence stops. Optionally, if I PT1 is equal to ID or approximately equal to I with sufficient allowable precision D , then the selected memory cell has been successfully programmed.

[0219] If I PT1 is not close enough to I D , then further programming with a smaller granularity can be performed. Here, progressive 3502 is now used. The starting point of progressive 3502 is the final voltage used for programming under progressive 3501. The increment V p2 (which is less than v p1 ) is added to this voltage, and the combined voltage is applied to the precision programming terminal to program the selected memory cell. After each programming voltage is applied, a verification step (similar to step 3404) is performed, where it is determined whether I cell is less than or equal to I PT2 (which is the second precision target current value and here is the third threshold), where I PT2 = I D + I PT2OFFSET , where I PT2OFFSET is an offset value added to prevent programming overshoot. If not, another increment v p2 is added to the previously applied programming voltage, and the process is repeated. Repeat until the moment when I cell is less than or equal to I PT2 , at which point this part of the programming sequence stops. Here, it is assumed that I PT2 is equal to I D or close enough to I D so that programming can stop because the target value has been achieved with sufficient allowable precision. Those of ordinary skill in the art can understand that additional progressions can be applied by using smaller and smaller programming increments. For example, in Figure 36A , three progressions (3601, 3602, and 3603) are applied instead of just two progressions.

[0220] A second method is shown in progressive 3503. Here, instead of increasing the voltage applied during the programming of the selected memory cell, the same voltage (such as V i or V i + V p1 + V p1 or V i + V p2 + V p2 ) is applied for an increasingly long duration of the cycle. Instead of adding an incremental voltage such as v p1 in progressive 3501 and v p2 in progressive 3502, an additional time increment t p1 is added to the programming pulse so that each applied pulse is longer than the previously applied pulse by tp1 After each programming pulse is applied to the precision programming terminal, the same verification steps as previously described for progressive 3501 are performed. Optionally, if the additional time increment added to the programming pulse has a shorter duration than the progress previously used, then an additional progression can be applied.

[0221] Optionally, an additional programming cycle progression can be applied, where the programming pulse has the same duration as the previous programming cycle progression used. Although only one time progression is shown, those of ordinary skill in the art will understand that any number of different time progressions can be applied. That is, instead of changing the voltage amplitude used during programming or changing the period of the voltage pulse used during programming, the system can instead change the number of programming cycles used.

[0222] Figure 36B A diagram of complementary pulse programming progression is shown, where the voltage applied to one precision programming terminal increases while the voltage applied to the other precision programming terminal decreases. For example, an increasing voltage progression can be applied to the control gate of the selected cell, and a decreasing voltage progression can be applied to the erase gate or source line of the selected cell. Or in an alternative, an increasing voltage progression can be applied to the erase gate or source line of the selected cell, and a decreasing voltage progression can be applied to the control gate of the selected cell. These complementary progression programming pulses result in higher precision in programming. For example, in a programming pulse cycle with a 10 mV CG increment and a 20 mV EG decrement, the resulting voltage in FG after the precision programming pulse cycle is 10 mV * 40% - 20 mV * 15% = approximately 1 mV, assuming a 40% CG to FG coupling ratio and a 10% EG to FG coupling ratio. This complementary pulse programming method can be used in the precision programming step 3306 after the rough programming step 3305, because the rough programming step 3305 will typically only use a CG increment or an EG increment during its programming operation.

[0223] Additional details for the second and third embodiments of the rough programming method 3305 will now be provided.

[0224] Figure 37Shows a second embodiment of the coarse programming (tuning) method 3305, which is the adaptive calibration method 3700. The method starts (step 3701). The cell is programmed by applying an initial voltage v0 to the coarse programming terminal according to one of the sequences shown in Table 9 (step 3702). Different from the search and execution method 3400, here v0 is not derived from a look-up table, but can be a relatively small predetermined initial value. Measure the control gate (or erase gate) voltage of the cell (which can be called CG1 or EG1) at a first current value IR1 (e.g., 100 na), and measure the voltage on the same gate at a second current value IR2 (e.g., 10 na), i.e., in a non-limiting embodiment where IR2 is 10% of IR1 and the subthreshold I-V slope is determined and stored based on those measurements (e.g., a current of 360 mV / dec or dV / d LOG(I)). The I-V slope in the linear region will be dV / dI.

[0225] Determine the new programming voltage v i . When this step is executed for the first time, i = 1, and the programming voltage v1 is determined using the subthreshold formula based on the stored subthreshold slope value, as well as the current target value and offset value, such as the following:

[0226] vi = v i-1 + v increment ,

[0227] v increment Proportional to the slope of Vg

[0228] Vg = n * Vt * log[Ids / wa * Io]

[0229] Here, wa is the w of the memory cell, Ids is the cell current, Io is the cell current when Vg = Vth, Vt is the thermal voltage, and g ranges from 1 to 2, where V1 is determined at current IR1 and V2 is determined at current IR2.

[0230] Slope = (V1 – V2) / (LOG(IR1) – LOG(IR2))

[0231] Where v increment = α * slope * (LOG(IR1) – LOG(I CT ))

[0232] Where, I CT is the target current, and α is a predetermined constant <1 (programming offset value) to prevent overshoot, e.g., 0.9.

[0233] If the stored slope value is relatively steep, a relatively small current offset value can be used. If the stored slope value is relatively flat, a relatively high current offset value can be used. Thus, determining the slope information will allow selection of a current offset value customized for the particular cell under consideration. This will generally make the programming process shorter. When step 3704 is repeated, i is incremented, and v i = v i-1 + v increment . Then the cell is programmed by applying v i to the coarse adaptive programming terminals. v increment can also be determined from a look-up table storing v increment values versus target current values.

[0234] Next, a verification operation occurs, where a read operation is performed on the selected cell and the current (I cell ) consumed through the selected cell is measured (step 3705). If I cell is less than or equal to I CT (here it is the coarse target threshold), where I CT = I D + I CTOFFSET , where I CTOFFSET is an offset value added to prevent programming overshoot, then the adaptive calibration method 3700 is complete and the precise programming method 3306 can be started. If I cell is not less than or equal to I CT , then steps 3704 to 3705 (where new slope measurements are made with new data points) or steps 3703 to 3705 (where the previously used same slope is reused) are repeated, and i is incremented.

[0235] Figure 38 shows a high-level block diagram of a circuit for implementing method 3703. Exemplary current values IR1 and IR2 are applied to the selected cell (here the memory cell 3802) using a current source 3801, and then the voltage at the coarse programming terminals of the memory cell 3802 is measured (V1 (VCGR1 or VEGR1) for IR1 and V2 (VCGR2 or VEGR2) for IR2), where the coarse programming terminals are selected according to Table 9. The linear slope will be ten times the cell current on the I-V curve of (V1 - V2) / LOG, or equivalently equal to (V1 – V2) / (LOG(IR1) – LOG(IR2)).

[0236] Figure 39Shows a third implementation of programming method 3305, which is an adaptive calibration method 3900. The method starts (step 3901). The cell is programmed with a default starting value v0 by applying v0 to the coarse adaptive programming terminal of the cell (step 3902). v0 is derived from a look-up table created using silicon characteristics, and the values in the table are offset, such as to prevent overshoot of the programming target. An example of v0 is shown in Table 13:

[0237] Table 13: Initial voltage v0 applied to the coarse adaptive programming terminals during the adaptive calibration method 3900

[0238]

[0239] In step 3903, an I-V slope parameter is formed for predicting the next programming voltage. A first voltage V1 is applied to the control gate or erase gate of the selected cell, and the resulting cell current IR1 is measured. Then a second voltage V2 is applied to the control gate or erase gate of the selected cell, and the resulting cell current IR2 is measured. The slope is determined based on these measurements and stored, for example, according to an equation in the subthreshold region (for cells operating in the subthreshold):

[0240] Slope = (V1 – V2) / (LOG(IR1) – LOG(IR2))

[0241] (step 3903). Examples of the values of V1 and V2 are shown in Table 13 above.

[0242] Determining the I-V slope information allows the selection of a V value customized for the specific cell under discussion. This will generally make the programming process shorter. increment Value. This will generally make the programming process shorter.

[0243] Each time step 3904 is executed, i is incremented (starting value 0) to determine the desired programming voltage V using the following equation based on the stored slope value, the current target, and the offset value i :

[0244] v i = v i-1 + v increment ,

[0245] where v increment = α * slope * (LOG(IR1) – LOG(I CT ))), where I CT is the target current, and α is a predetermined constant < 1 (programming offset value) to prevent overshoot, for example, 0.9.

[0246] Then the selected cell is programmed using v i (step 3905).

[0247] Next, a verification operation occurs, where a read operation is performed on the selected cell and the current (I cell )(step 3906) consumed by the selected cell is measured. If I cell is less than or equal to I CT (where it is a rough target threshold here), where I CT = I D + I CTOFFSET , where I CTOFFSET is an offset value added to prevent programming overshoot, then the process proceeds to step 3907. If not, the process returns to step 3903 (new slope measurement) or 3904 (reusing the previous slope), and i is incremented.

[0248] At step 3907, I cell is compared with a threshold CT less than I CT2 . The purpose of this is to see if overshoot has occurred. That is, although the goal is to have I cell below ICT, if it is far below ICT, then overshoot has occurred and the stored value may actually correspond to the wrong value. If I cell is not less than or equal to I CT2 , then no overshoot has occurred and the adaptive calibration method 3900 is complete, and at this time the process proceeds to the precise programming method 3306. If I cell is less than or equal to I CT2 , then overshoot has occurred. Then the selected cell is erased (step 3908), and the programming process restarts at step 3902. Optionally, if step 3908 is executed more than a predetermined number of times, the selected cell can be considered a bad cell that should not be used, and an error signal is output, or a flag is set to identify the cell.

[0249] The precise programming method 3306 can consist of multiple verification loops and programming loops, where the programming voltage increases in constant fine voltages with a fixed pulse width, or where the programming voltage is fixed and the programming pulse width varies for each additional programming pulse.

[0250] Optionally, the step 3906 of determining whether the current through the selected non - volatile memory cell during a read or verification operation is less than or equal to a first threshold current value can be performed as follows: applying a fixed bias to the terminals of the non - volatile memory cell; measuring and digitizing the current consumed by the selected non - volatile memory cell to generate a digital output bit; and comparing the digital output bit with a digital bit representing the first threshold current I CT .

[0251] Optionally, step 3907 of determining whether the current through the selected non-volatile memory cell during a read or verify operation is less than or equal to a second threshold current value may be performed as follows: applying a fixed bias to the terminals of the non-volatile memory cell; measuring and digitizing the current consumed by the selected non-volatile memory cell to generate a digital output bit; and comparing the digital output bit with a digital bit representing the second threshold current I CT2 with that of the digital bit representing CT2 .

[0252] Optionally, each of steps 3906 and 3907 of determining whether the current through the selected non-volatile memory cell during a read or verify operation is correspondingly less than or equal to a first threshold current value or a second threshold current value may be performed as follows: applying an input to the terminals of the non-volatile memory cell; modulating the current consumed by the selected non-volatile memory cell with an output pulse to generate a modulated output; digitizing the modulated output to generate a digital output bit; and comparing the digital output bit correspondingly with a digital bit representing the first threshold current value or the second threshold current value.

[0253] Measuring the cell current for the purpose of verifying or reading the current may be done by taking the average of multiple times (e.g., 8 to 32 times) to reduce the impact of noise.

[0254] Figure 40 A fourth embodiment of the rough programming method 3305 is shown, which is the absolute calibration method 4000. The method starts (step 4001). The relevant terminals of the cell are programmed with a default starting value v0 (step 4002). Examples of v0 are shown in Table 14:

[0255] Table 14: Initial voltage v0 applied to the memory cell terminals during the absolute calibration method 4000

[0256] Program sequence name Coarse programming terminal <![CDATA[Initial voltage v applied to the coarse programming terminal o > CG-CG Control gate 28 (CG) Approximately 5 - 7V CG-EG Control gate 28 (CG) Approximately 5 - 7V EG-EG Erase gate 30 (EG) Approximately 0 - 2V CG-SL Control gate 28 (CG) Approximately 5 - 7V SL-CG Source line 14 (SL) Approximately 3 - 3.5V SL-EG Source line 14 (SL) Approximately 3 - 3.5V

[0257] The voltage vTx on the rough programming terminal is measured and stored with the current value Itarget driven through the cell as described above regarding Figure 38 (step 4003). A new rough programming voltage v1 is determined based on the stored voltage vTx and an offset value vToffset (which corresponds to Ioffset) (step 4004). For example, the new desired voltage v1 can be calculated as follows: v1 = v0 + (VTBIAS - vTx) - vToffset, where VTBIAS is, for example, about 1.5V, which is the default terminal voltage at the maximum target current (meaning the maximum current level allowed for the memory cell). Basically, the new target voltage is adjusted by the amount of the current voltage vTx at the target current from the difference between the maximum voltage and the offset.

[0258] Then v iProgram the cell (step 4005). When i = 1, use the voltage v1 obtained from step 4004. When i >= 2, use the voltage v i = v i-1 + v increment . v increment can be determined from a look-up table storing v increment values and target current values. Next, a verification operation occurs, in which a read operation is performed on the selected cell, and the current (I cell ) consumed by the selected cell is compared with I CT (step 4006). If I cell is less than or equal to I CT (where it is a threshold value), then the absolute calibration method 4000 is complete, and the precise programming method 3306 can be started. If I cell is not less than or equal to I CT , then repeat steps 4005 to 4006, and increment i.

[0259] Figure 41Illustrated is a circuit 4100 for measuring vTx in step 4003 of the absolute calibration method 4000. vTx is measured at each memory cell 4103 (4103-0, 4103-1, 4103-2, …, 4103-n). Here, n+1 different current sources 4101 (4101-0, 4101-1, 4101-2, ..., 4101-n) generate different currents IO0, IO1, IO2, …IOn with increasing amplitudes. Each current source 4101 is connected to a corresponding inverter 4102 (4102-0, 4102-1, 4102-2, ..., 4102-n) and memory cell 4103 (4103-0, 4103-1, 4103-2, ..., 4103-n). The input to each inverter 4102 (4102-0, 4102-1, 4102-..., 4102-n) is initially high, and the output of each inverter is initially low. Since IO0 < IO1 < IO2 < … < IOn, the output of inverter 4102-0 will first switch from low to high because memory cell 4103-0 will draw current from current source 4101-0 and also from the input node of inverter 4102-0, causing the input voltage to inverter 4102-0 to drop to low level before the input voltages to other inverters 4102. Next, the output of inverter 4102-1 will switch from low to high, then the output of inverter 4102-2, and so on until the output of inverter 4102-n switches from low to high. Each inverter 4102 controls a corresponding switch 4104 (4104-0, 4104-1, 4104-2, ..., 4104-n) such that when the output of inverter 4102 is high, switch 4104 closes, which will cause vTx to be sampled by capacitor 4105 (4105-0, 4105-1, 4105-2, ..., 4105-n). Thus, switches 4104 and capacitors 4105 can form a sample-and-hold circuit. In this way, a sample-and-hold circuit is used to measure vTx.

[0260] Figure 42 Illustrated is an exemplary progression 4200 for programming selected cells during the adaptive calibration method 3700 or the absolute calibration method 4000. A voltage VTP (which is the programming voltage applied to the CG or EG terminal and which corresponds to Figure 37 vi in step 3704 of Figure 40 and vi in step 4005 of

[0261] Figure 43Another exemplary progression 4300 for programming selected cells during the adaptive calibration method 3700 or the absolute calibration method 4000 is shown. A voltage VTP (which is the programming voltage applied to the CG or EG terminal and which corresponds to Figure 37 vi in step 3704 of Figure 40 and vi in step 4005 of

[0262]

[0263] In another embodiment, the voltage applied to the control gate terminal is incremented and the voltage applied to the erase gate terminal is also incremented.

[0264] Table 15: Control gate terminal increment and erase gate terminal decrement

[0265]

[0266]

[0267] Table 16: Control gate terminal increment; erase gate terminal increment

[0268]

[0269] Figure 44

[0270] A system for implementing input and output methods for reading or verifying within a VMM array after precise programming is shown. The input functional circuit 4401 receives digital bit values and converts these digital values into analog signals, which in turn are used to apply a voltage to the control gate of selected cells in the array 4404, where the selected cells are determined by the control gate decoder 4402. At the same time, the word line decoder 4403 is also used to select the row in which the selected cells are located. The output neuron circuit block 4405 receives the output current from each column of the cells in the array 4404. The output circuit block 4405 may include an integrating analog-to-digital converter (ADC), a successive approximation (SAR) ADC, a Σ-Δ ADC, or any other ADC scheme to provide a digital output.​​In one embodiment, the digital value provided to the input functional circuit 4401 includes four bits (DIN3, DIN2, DIN1, and DIN0) or any number of bits, and the digital value represented by these bits corresponds to the number of input pulses to be applied to the control gate during a programming operation. A greater number of pulses will result in a greater value being stored in the cell, which will result in a greater output current when the cell is read. Examples of input bit values and pulse values are shown in Table 17:

[0271] Table 17: Digital bit input and generated pulses

[0272] DIN3 DIN2 DIN1 DIN0 Generated input pulses 0 0 0 0 0 0 0 0 1 1 0 0 1 0 2 0 0 1 1 3 0 1 0 0 4 0 1 0 1 5 0 1 1 0 6 0 1 1 1 7 1 0 0 0 8 1 0 0 1 9 1 0 1 0 10 1 0 1 1 11 1 1 0 0 12 1 1 0 1 13 1 1 1 0 14 1 1 1 1 15

[0273] In the above example, there are up to 15 pulses for a 4-bit input digital. Each pulse is equal to one unit of cell value (current), i.e., the exact programming current. For example, if Icell unit = 1 nA, then for DIN[3-0] = 0001, Icell = 1 * 1 nA = 1 nA; and for DIN[3-0] = 1111, Icell = 15 * 1 nA = 15 nA.

[0274] In another embodiment, the digital bit input uses digital bit position summation to read out the cell value or neuron value (e.g., the exact programming value on the bit line output), as shown in Table 18. Here, only 4 pulses or 4 fixed identical bias inputs (e.g., inputs on the word line or control gate) are required to evaluate a 4-bit digital value. For example, the first pulse or the first fixed bias is used to evaluate DIN0, the second pulse or the second fixed bias having the same value as the first pulse or the second fixed bias is used to evaluate DIN1, the third pulse or the third fixed bias having the same value as the first pulse or the second fixed bias is used to evaluate DIN2, and the fourth pulse or the fourth fixed bias having the same value as the first pulse or the fourth fixed bias is used to evaluate DIN3. Then, the results of the four pulses are added according to the bit position, where each output result is multiplied (scaled down) by a multiplier, i.e., 2^n, where n is the digital bit position, as shown in Table 19. The implemented digital bit summation formula is as follows: Output = (2^0 * DIN0 + 2^1 * DIN1 + 2^2 * DIN2 + 2^3 * DIN3) * Icell unit, where Icell unit represents the exact programming current.

[0275] For example, if Icell unit = 1 nA, then for DIN[3-0] = 0001, the total Icell = 0 + 0 + 0 + 1 * 1 nA = 1 nA; and for DIN[3-0] = 1111, I cell total = 8 * 1 nA + 4 * 1 nA + 2 * 1 nA + 1 * 1 nA = 15 nA.

[0276] Table 18: Digital bit input summation

[0277] <![CDATA 2^3*DIN3 > <![CDATA 2^2*DIN2 > <![CDATA 2^1*DIN1 > <![CDATA 2^0*DIN0 > <![CDATA Total value > 0 0 0 0 0 0 0 0 1 1 0 0 2 0 2 0 0 2 1 3 0 4 0 0 4 0 4 0 1 5 0 4 2 0 6 0 4 2 1 7 8 0 0 0 8 8 0 0 1 9 8 0 2 0 10 8 0 2 1 11 8 4 0 0 12 8 4 0 1 13 8 4 2 0 14 8 4 2 1 15

[0278] Table 19: Summation of digital input bits Dn with 2^n output multiplication factor

[0279] DIN3 DIN2 DIN1 DIN0 Output X factor Y X1 Y X2 Y X4 Y X8

[0280] For an exemplary 4-digit input, Table 20 shows another embodiment with a mixed input having multiple digital input pulse ranges and a sum of input digital ranges. In this embodiment, DINn-0 can be divided into m different groups, where each group is evaluated and the output is scaled by a certain multiplication factor according to the group binary position. For example, for 4-bit DIN3-0, the groups can be DIN3-2 and DIN1-0, where the output for DIN1-0 is scaled by 1 (X1) and the output for DIN3-2 is scaled by 4 (X4).

[0281] Table 20: Hybrid input-output summation with multiple input ranges

[0282]

[0283] Another embodiment combines a mixed input range with a mixed supercell. The mixed supercell includes multiple physical x-bit units to implement a logical n-bit unit, where the x-bit output is scaled by 2^n binary positions. For example, to implement an 8-bit logic unit, two 4-bit units (Unit 1, Unit 0) are used. The output for Unit 0 is scaled by 1 (X1), and the output for Unit 1 is scaled by 4 (X, 2^2) to implement an n-bit logic unit. Other combinations of physical x units are possible, such as two 2-bit physical units and one 4-bit physical unit to implement an 8-bit logic unit.

[0284] Figure 45 is shown similar to Figure 44Another embodiment of the system differs in that the digital bit input uses digital bit position summation to read out the current of the cells or neurons (e.g., the value on the bit line output) modulated by the modulator 4510, where the output pulse width is designed according to the digital input bit position (e.g., converting the current to an output voltage V = current * pulse width / capacitance). For example, a first bias (applied to the input word line or control gate) is used to evaluate DIN0, and the current (cell or neuron) output is modulated by the modulator 4510 with a unit pulse width proportional to the DIN0 bit position, which is one (x1) unit, a second input bias is used to evaluate DIN1, and the current output is modulated by the modulator 4510 with a unit pulse width proportional to the DIN1 bit position, which is two (x2) units, a third input bias is used to evaluate DIN2, and the current output is modulated by the modulator 4510 with a pulse width proportional to the DIN2 bit position, which is four (x4) units, a fourth input bias is used to evaluate DIN3, and the current output is modulated by the modulator 4510 with a pulse width proportional to the DIN3 bit position, which is eight (x8) units. Then, each output is converted to a digital bit by an ADC (analog-to-digital converter) 4511 for each digital input bit DIN0 to DIN3. Then, the total output is output by the summer 4512 as the sum of the four digital outputs generated from the DIN0-3 inputs.

[0285] Figure 46 An example of a charge summer 4600 is shown, which can be used to sum the output Icell of the VMM during a verification operation or during the analog-to-digital conversion of the output neuron to obtain a single analog value representing the output of the VMM, and then optionally convert it to a digital bit value. For example, the charge summer 4600 can be used as the summer 4512. The charge summer 4600 includes a current source 4601 (here representing the current Icell output by the VMM) and a sample-and-hold circuit, which includes a switch 4602 and a sample-and-hold (S / H) capacitor 4603. The example shown uses a 4-bit digital value for the output, although other bit numbers can also be used instead. There are 4 S / H circuits to hold the values generated by 4 evaluation pulses, and these values are added at the end of the process. The S / H capacitor 4603 is selected to have a ratio associated with the 2^n*DINn bit position of this S / H capacitor. For example, the switch 4602 of C_DIN3 closes when Icell > 8x the current threshold, the switch 4602 of C_DIN2 closes when Icell > 4x the current threshold, the switch 4602 of C_DIN1 closes when Icell > 2x the current threshold, and the switch 4602 of C_DIN0 closes when Icell > the current threshold. Therefore, the digital value stored by the sample-and-hold capacitor 4603 will reflect the value of Icell 4601.

[0286] Figure 47 Shows a current summing device 4700, which can be used to sum the output Icell of the VMM during a verification operation or during the analog-to-digital conversion of the output neuron. For example, the charge summing device 4700 can be used as the summing device 4512. The current summing device 4700 includes a current source 4701 (here representing the Icell output from the VMM), switches 4702, 4703 and 4704, and a transistor 4705. The example shown uses a 4-bit digital value for the output, where the bit values are represented by currents I_DIN0, I_DIN1, I_DIN2, and I_DIN3. The bit position of each transistor 4705 affects the value represented by that bit. The switch 4703 for I_DIN3 closes when Icell > 8x the current threshold, the switch 4703 for I_DIN2 closes when Icell > 4x the current threshold, the switch 4704 for I_DIN1 closes when Icell > 2x the current threshold, and the switch 4703 for I_DIN0 closes when Icell > the current threshold. Thus, the digital value output by the transistor 4705 (where "1" is represented by a positive current and "0" is represented by no current, or vice versa) will reflect the value of Icell4601.

[0287] Figure 48 Shows a digital summing device 4800, which receives a plurality of digital values, sums them, and generates an output DOUT representing the sum of the inputs. For example, the digital summing device 4600 can be used as the summing device 4512. The digital summing device 4800 can be used during a verification operation or during the analog-to-digital conversion of the output neuron. As shown in the example for a 4-bit digital value, there are digital output bits to hold the values from 4 evaluation pulses, where these values are added at the end of the process. The digital output is digitally scaled based on the 2^n*DINn bit positions; for example, DOUT3 = x8 DOUT0, _DOUT2 = x4 DOUT1, I_DOUT1 = x2DOUT0, I_DOUT0 = DOUT0.

[0288] Figure 49A Shows an integrating dual-slope ADC 4900 applied to the output neuron to convert the cell current into digital output bits. An integrator composed of an integrating operational amplifier 4901 and an integrating capacitor 4902 integrates the cell current ICELL and the reference current IREF. As Figure 49BAs shown, during a fixed time t1, switch S1 is closed and switch S2 is open, and the cell current is integrally increased (Vout rises in waveform 4950), then switch S1 is open and switch S2 is closed, such that the reference current IREF is applied to integrally decrease it within time t2 (Vout falls in waveform 4950). The value of the current Icell is determined as = t2 / t1 * IREF. For example, for t1, for 10-bit digital bit resolution, 1024 cycles are used, and for t2, the number of cycles varies from 0 to 1024 cycles according to the Icell value. In the case where the target value is applied to comparator 4904 as VREF, the output EC 4905 of comparator 4904 can be used as a flip-flop to determine the number of cycles in which IREF is applied until VOUT drops below VREF.

[0289] Figure 49C Shown is an integrating single-slope ADC 4960 applied to output neuron 4966 ICELL to convert the cell current into digital output bits. ADC 4960 includes an integrating operational amplifier 4961, an integrating capacitor 4962, an operational amplifier 4964, and switches S1 and S3. The integrating operational amplifier 4961 and the integrating capacitor 4962 integrate the output neuron current ICELL. As Figure 49D shown, during time t1, the cell current is integrally increased (Vout rises until it reaches Vref2), and during time t2 (starting at the same time as time t1, but greater than time t1), the cell current of the reference cell is integrally increased. The cell current ICELL is determined as = Cint * Vref2 / t. A pulse counter coupled to the output of comparator 4965 is used to count the number of pulses (digital output bits) during the respective integration times t1, t2. For example, as shown, the digital output bits of t1 are less than those of t2, which means that the cell current during t1 is greater than the cell current during t2. Initial calibration is performed to calibrate the integrating capacitor value Cint = Tref * Iref / Vref2 using the reference current and the fixed time.

[0290] Figure 49E Shown is an integrating dual-slope ADC 4980 applied to output neuron 4984 ICELL to convert the cell current into digital output bits. The integrating dual-slope ADC 4980 includes switches S1, S2, and S3, an operational amplifier 4981, a capacitor 4982, and a reference current source 4983. The integrating dual-slope ADC 4980 does not use an integrating operational amplifier. The cell current or the reference current is directly integrated for capacitor 4982. A pulse counter is used to count the pulses (digital output bits) during the integration time. The current Icell = t2 / t1 * IREF.

[0291] Figure 49F An integrating single-slope ADC 4990 is shown that is applied to an output neuron 4994ICELL to convert a cell current into a digital output bit. The integrating single-slope ADC 4990 includes switches S2 and S3, an operational amplifier 4991, and a capacitor 4992. The integrating single-slope ADC 4980 does not use an integrating operational amplifier. The cell current is directly integrated against the capacitor 4992. A pulse counter is used to count pulses (digital output bits) during the integration time. The cell current Icell = Cint * Vref2 / t.

[0292] Figure 50A A SAR (successive approximation register) ADC is shown that is applied to an output neuron to convert a cell current into a digital output bit. The cell current can be dropped across a resistor to convert it into a voltage VCELL. Alternatively, the cell current can charge an S / H capacitor to convert the cell current into a voltage VCELL. VCELL is provided to the inverting input of a comparator 5003, and the output of the comparator is fed to the select input of a SAR 5001. A clock input CLK is further provided to the SAR 5001. A binary search is used to calculate the bits starting from the MSB (most significant bit). Based on the digital bits DN-D0 output from the SAR 5001 and received as the input to a DAC 5002, the output of the DAC 5002 is used to set the non-inverting input of the comparator 5003, i.e., the appropriate analog reference voltage for the comparator 5003. The output of the comparator 5003 is then fed back to the SAR 5001 to select the next analog level. As Figure 50B shown, for an example of 4-bit digital output bits, there are 4 evaluation cycles: the first pulse evaluates DOUT3 by setting the analog level in the middle, and then the second pulse evaluates DOUT2 by setting the analog level in the middle of the upper half or the middle of the lower half, without limitation.

[0293] A modified binary search such as a cyclic (algorithm) ADC can be used for cell tuning (e.g., programming) verification or output neuron conversion. An improved binary search such as a switched capacitor (SC) charge redistribution ADC can be used for cell tuning (e.g., programming) verification or output neuron conversion.

[0294] Figure 51Shown is a Σ-Δ type ADC 5100 applied to an output neuron to convert a cell current into a digital output bit. An integrator composed of an operational amplifier 5101 and a capacitor 5105 integrates the sum of the current ICELL from the selected cell current 5106 and the reference current IREF from a 1-bit current cDAC 5104. A comparator 5102 compares the integrated output voltage of the operational amplifier 5101 with a reference voltage VREF2. A clock-controlled DFF 5103 provides a digital output stream according to the output received at the D input of the DFF 5103 by the comparator 5102. The digital output stream typically enters a digital filter before being output as a digital output bit.

[0295] Figure 52A Shown is a ramp analog-to-digital converter 5200, which includes a current source 5201 (which represents the received neuron current ICELL), a switch 5202, a variable configurable capacitor 5203, and a comparator 5204 that receives the voltage (represented as Vneu) formed across the variable configurable capacitor 5203 as a non-inverting input and receives a configurable reference voltage Vreframp as an inverting input and generates an output Cout. Vreframp ramps up at discrete levels with each comparison clock cycle. The comparator 5204 compares Vneu with Vreframp, and thus when Vneu > Vreframp, the output Cout will be "1", and otherwise will be "0". Therefore, the output Cout will be a pulse whose width varies in response to Ineu. A larger Ineu will cause Cout to be "1" for a longer period of time, resulting in a wider pulse of the output Cout. A digital counter 5220 converts each pulse 522 of the output Cout into a count value 5221, which is a digital output bit, as Figure 52B shown respectively for two different ICELL currents (represented as OT1A and OT2A).

[0296] Alternatively, the ramp voltage Vreframp is a continuous ramp voltage 5255, as Figure 52B shown by the curve 5250 of

[0297] Alternatively, Figure 52CA multi-ramp implementation for reducing conversion time by utilizing a coarse-fine ramp conversion algorithm is shown. First, a coarse reference ramp reference voltage 5271 ramps in a fast manner to find the sub-range of each ICELL. Next, fine reference ramp reference voltages 5272 (i.e., Vreframp1 and Vreframp2) are used separately for each sub-range to convert the ICELL current within the corresponding sub-range. As shown, there are two sub-ranges of the fine reference ramp voltage. More than two coarse / fine steps or more than two sub-ranges are possible.

[0298] Figure 53 An algorithmic analog-to-digital output converter 5300 is shown, which includes switches 5301, 5302, a sample-and-hold (S / H) circuit 5303, a 1-bit analog-to-digital converter (ADC) 5304, a 1-bit digital-to-analog converter (DAC) 5305, a summer 5306, and the gain of two residue operational amplifiers (2 operational amplifiers) 5307. The algorithmic analog-to-digital output converter 5300 generates a converted digital output 5308 in response to an analog input Vin and control signals applied to switches 5302 and 5302. The input received at the analog input Vin (e.g., Vneu of FIG. 52) is first sampled by the S / H circuit 5303 in response to switch 5302 and then the conversion is performed over N clock cycles for N bits. For each conversion clock cycle, the 1-bit ADC 5304 compares the S / H voltage 5309 with the reference voltage VREF / 2 and outputs a digital bit (e.g., outputs "0" if the input <= VREF / 2 and outputs "1" if the input > VREF / 2). This digital output bit (which is the digital output signal 5308) is then converted by the 1-bit DAC 5305 into an analog voltage (e.g., converted into VREF / 2 or 0V) and fed to the summer 5306 to subtract from the S / H voltage 5309. The 2x residue opamp 5307 then amplifies the summer difference voltage output into a conversion residue voltage 5310, which is fed through switch 5301 to the S / H circuit 5303 for the next clock cycle. Instead of this 1-bit (i.e., 2-level) algorithmic ADC, a 1.5-bit (i.e., 3-level) algorithmic ADC can be used to reduce the effects of offsets such as those from the ADC 5304 and the residue operational amplifier 5307. For use with a 1.5-bit algorithmic ADC, a 1.5-bit or 2-bit (i.e., 4-level) DAC is preferred.

[0299] In another implementation, a hybrid ADC can be used. For example, for a 9-bit ADC, the first 4 bits can be generated by a SAR ADC and the remaining 5 bits can be generated using a slope ADC or a ramp ADC.

[0300] Programming and verifying multiple physical units as a single logical multi-bit unit

[0301] The above programming and verification devices and methods can operate simultaneously on multiple physical units that are logical multi-bit units.

[0302] Figure 54 A logical multi-bit unit 5400 is shown, which includes i physical units labeled as physical units 5401-1, 5401-2, …, 5401-i. In one embodiment, the physical units 5401 have a uniform diffusion width (transistor width). In another embodiment, the physical units 5401 have a non-uniform diffusion width (different transistor widths, where transistors with larger widths can store a larger number of levels and thus can store a larger number of bits). In both embodiments, the physical units 5401 are programmed, verified, and read as a single unit, specifically as a single logical n-bit unit that can store a larger number of levels than each m-bit unit. For example, if m = 2, each physical unit 5401 can hold one of four levels (L0, L1, L2, L3). Two such units can be considered as a single logical unit with n = 3, such that the single logical unit can hold one of eight levels (L0, L1, L2, L3, L4, L5, L6, L7). As another example, if m = 3, each physical unit 5401 can hold one of eight levels (L0, ..., L7). Four such units can be considered as a single logical unit with n = 5, such that the single logical unit can hold one of 32 levels (L0, ..., L31).

[0303] Figure 55 A method 5500 for programming the logical multi-bit unit 5400 is shown. First, j physical units (where j <= i) out of the i physical units 5401-1, …, 5401-i are programmed and verified using any one of the rough programming methods 3305 until a rough current target for the j physical units is achieved (step 5501). Next, k physical units (where k <= j) out of the j physical units are programmed and verified using any one of the precise programming methods 3306 until a precise current target for the k physical units is achieved (step 5502).

[0304] Method 5500 can be executed on more than one subset of the i physical units 5401-1, …, 5401-i to achieve the desired overall level of the logical multi-bit unit 5400.

[0305] For example, if i = 4, there will be four cells 5401-1, 5401-2, 5401-3, and 5401-4. If it is assumed that each cell can hold one of eight different levels, the logical multi-bit cell 5400 can hold one of 32 different levels. If the desired programmed value is L27, this level (which corresponds to the desired read current) can be achieved in any number of different ways.

[0306] For example, the method 5500 can be performed on cells 5401-1, 5401-2, and 5401-3 until these cells together hold L23 (the 24th level), and then the method 5500 can be performed on cell 5401-4 to program this cell to its fourth level so that the logical multi-bit cell 5400 reaches L27 (the 28th level).

[0307] As another example, the method 5500 can be performed on cells 5401-1, 5401-2, 5401-3, and 5401-4 until these cells together hold L25 (the 26th level), and then the method 5500 can be performed only on cell 5401-4 until it stores a value that causes the entire logical multi-bit cell 5400 to reach L27 (the 28th level).

[0308] Other methods are possible, and the method 5500 can be performed on different subsets of the i physical cells until the desired level is reached.

[0309] In another embodiment, in the case where the i physical cells have non-uniform diffusion widths, a coarse programming step 3305 can be performed on j1 physical cells having a wider transistor width until the j1 physical cells together achieve a coarse current target, and then a fine programming step 3306 can be performed on j2 physical cells having the smallest transistor width until the j1 + j2 physical cells together achieve a fine current target.

[0310] It should be noted that, as used herein, the terms "above" and "on" both inclusively encompass "directly on" (with no intervening material, element, or space therebetween) and "indirectly on" (with intervening material, element, or space therebetween). Similarly, the term "adjacent" includes "directly adjacent" (with no intervening material, element, or space therebetween) and "indirectly adjacent" (with intervening material, element, or space therebetween), "mounted to" includes "directly mounted to" (with no intervening material, element, or space therebetween) and "indirectly mounted to" (with intervening material, element, or space therebetween), and "electrically coupled to" includes "directly electrically coupled to" (with no intervening material or element electrically connecting the elements together) and "indirectly electrically coupled to" (with intervening material or element electrically connecting the elements together). For example, forming an element "above a substrate" may include directly forming the element on the substrate with no intervening material / element therebetween, and indirectly forming the element on the substrate with one or more intervening material / elements therebetween.

Claims

1. A method of programming a selected non - volatile memory cell to store one of N possible values, where N is an integer greater than 2, the selected non - volatile memory cell including a floating gate, a control gate terminal, an erase gate terminal, and a source line terminal, the method comprising: Performing a first programming process including a plurality of program verification cycles, wherein a programming voltage with an increasing amplitude is applied to the terminals of the selected non - volatile memory cell after the first program verification cycle and in each subsequent program verification cycle, and each program verification cycle includes: Applying a first voltage to one of the erase gate terminal and the control gate terminal of the selected non - volatile memory cell; Measuring a first current generated through the selected non - volatile memory cell; Applying a second voltage to the one of the erase gate terminal and the control gate terminal of the selected non - volatile memory cell; Measuring a second current generated through the selected non - volatile memory cell; Determining a slope value based on the first voltage, the second voltage, the first current, and the second current; and Determining the next programming voltage in the programming voltage with an increasing amplitude for the next program verification cycle based on the determined slope value.

2. The method according to claim 1, wherein each program verification cycle includes verifying that the current through the selected non - volatile memory cell during a read or verify operation is less than or equal to a first threshold current value.

3. The method according to claim 1, further comprising: Programming the selected non - volatile memory cell using the next programming voltage.

4. The method according to claim 3, further comprising: Repeating the steps of determining the next programming voltage and programming the non - volatile memory cell using the next programming voltage until the current through the selected non - volatile memory cell during a read or verify operation is less than or equal to a first threshold current value.

5. The method according to claim 2, further comprising: When the current through the selected non - volatile memory cell during the read or verify operation is less than or equal to the first threshold current value, performing a second programming process until the current through the selected non - volatile memory cell during a read or verify operation is less than or equal to a second threshold current value.

6. The method according to claim 1, further comprising: Performing a third programming process until the current through the selected non - volatile memory cell during a read or verify operation is less than or equal to a fourth threshold current value.

7. The method according to claim 5, wherein the second programming process includes applying voltage pulses with an increasing amplitude to the control gate of the selected non - volatile memory cell.

8. The method according to claim 7, wherein the second programming process further includes applying voltage pulses with an increasing amplitude to the erase gate of the selected non - volatile memory cell.

9. The method according to claim 1, wherein the selected non - volatile memory cell is a split - gate flash memory cell.

10. The method according to claim 1, wherein the selected non-volatile memory cell is in a vector-matrix multiplication array in an analog neural network.

11. A method of programming a selected non-volatile memory cell to store one of N possible values, where N is an integer greater than 2, the selected non-volatile memory cell including a floating gate, a control gate terminal, an erase gate terminal, and a source line terminal, the method comprising: Performing a first programming process including a plurality of program verification cycles, wherein during the first programming process, a first programming voltage with an increasing amplitude is applied to the control gate of the selected non-volatile memory cell and a second programming voltage with a decreasing amplitude is applied to the erase gate of the selected non-volatile memory cell, and wherein each program verification cycle includes: Applying a first voltage to one of the erase gate and the control gate of the selected non-volatile memory cell; Measuring a first current generated through the selected non-volatile memory cell; Applying a second voltage to the one of the erase gate and the control gate of the selected non-volatile memory cell; Measuring a second current generated through the selected non-volatile memory cell; Determining a slope value based on the first voltage, the second voltage, the first current, and the second current; Determining a next programming voltage in the first programming voltage with an increasing amplitude and the second programming voltage with a decreasing amplitude respectively based on the slope value; Programming the non-volatile memory cell using the next programming voltage; Repeating the steps of determining the next programming voltage and programming the non-volatile memory cell using the next programming voltage until the current through the selected non-volatile memory cell during a read or verify operation is less than or equal to a first threshold current value.

12. The method according to claim 11, wherein each program verification cycle includes verifying that the current through the selected non-volatile memory cell during a read or verify operation is less than or equal to a first threshold current value.

13. The method according to claim 11, further comprising: Performing a second programming process until the current through the selected non-volatile memory cell during a read or verify operation is less than or equal to a second threshold current value.

14. The method according to claim 11, wherein performing the first programming process further includes: When the current through the selected non-volatile memory cell is less than or equal to a third threshold current value, erasing the selected non-volatile memory cell and repeating the first programming process.

15. The method according to claim 13, further comprising: Performing a third programming process until the current through the selected non-volatile memory cell during a read or verify operation is less than or equal to a fourth threshold current value.

16. A method for programming a selected non-volatile memory cell to store one of N possible values, where N is an integer greater than 2, the selected non-volatile memory cell including a floating gate, a control gate, an erase gate, and a source line terminal, the method comprising: Performing a first programming process including a plurality of program verify cycles, wherein in the first programming process, a first programming voltage with an increasing amplitude is applied to the erase gate of the selected non-volatile memory cell and a second programming voltage with a decreasing amplitude is applied to the control gate of the selected non-volatile memory cell, and wherein each program verify cycle includes: Applying a first voltage to one of the erase gate and the control gate of the selected non-volatile memory cell; Measuring a first current generated through the selected non-volatile memory cell; Applying a second voltage to the one of the erase gate and the control gate of the selected non-volatile memory cell; Measuring a second current generated through the selected non-volatile memory cell; Determining a slope value based on the first voltage, the second voltage, the first current, and the second current; Based on the slope value, respectively determining the next programming voltage in the programming voltage with an increasing amplitude and the programming voltage with a decreasing amplitude; Programming the non-volatile memory cell using the next programming voltage; Repeating the steps of determining the next programming voltage and programming the non-volatile memory cell using the next programming voltage until the current through the selected non-volatile memory cell during a read or verify operation is less than or equal to a first threshold current value.

17. The method according to claim 16, wherein each program verify cycle includes verifying that the current through the selected non-volatile memory cell during a read or verify operation is less than or equal to a first threshold current value.

18. The method according to claim 16, further comprising: Performing a second programming process until the current through the selected non-volatile memory cell during a read or verify operation is less than or equal to a second threshold current value.

19. The method according to claim 16, wherein the step of performing the first programming process further includes: When the current through the selected non-volatile memory cell is less than or equal to a third threshold current value, erasing the selected non-volatile memory cell and repeating the first programming process.

20. The method according to claim 18, further comprising: Performing a third programming process until the current through the selected non-volatile memory cell during a read or verify operation is less than or equal to a fourth threshold current value.

Citation Information

Patent Citations

  • Deep learning neural network classifier using non-volatile memory array

    US11308383B2

  • Deep Learning Neural Network Classifier Using Non-volatile Memory Array

    US20170337466A1

  • Single transistor non-valatile electrically alterable semiconductor memory device

    US5029130A

  • Flash memory cells with separated self-aligned select and erase gates, and process of fabrication

    US6747310B2

  • High Precision And Highly Efficient Tuning Mechanisms And Algorithms For Analog Neuromorphic Memory In Artificial Neural Networks

    US20190164617A1