Precise Data Tuning Method and Device for Analog Neural Memory in Artificial Neural Networks

Through multi-step tuning algorithm and circuit design, the precise programming problem of non-volatile memory units in VMM arrays is solved, the calculation accuracy and energy efficiency of the neural network are improved, and it is suitable for high-performance information processing of simulated neuromorphic memory.

CN114930458BActive Publication Date: 2025-08-05SILICON STORAGE TECHNOLOGY INC
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202080091622.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-03-25
Filing Date
2020-07-02
Publication Date
2025-08-05
Estimated Expiration
2040-07-02

AI Technical Summary

Technical Problem

The prior art is difficult to program the high-precision and high-energy energy-efficient nonvolatile memory cells in artificial neural networks, especially in vector-matrix multiplication (VMM) arrays to accurately tune the charge amount on the floating gate.

Method used

A multi-step tuning algorithm is used, including initial current target setting, soft erasing, rough programming, fine programming, read operation and error calculation until the output error is less than a predetermined threshold. Combined with circuit designs such as adjustable current source and capacitors, precise programming of non-volatile memory cells is achieved.

Benefits of technology

It realizes accurate programming of nonvolatile memory units in VMM arrays, improves the computing accuracy and energy efficiency of neural networks, and is suitable for high-performance information processing that simulates neuromorphic memory.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114930458B_ABST
    Figure CN114930458B_ABST
Patent Text Reader

Abstract

The present invention discloses various embodiments of a precision programming algorithm and apparatus for accurately and rapidly depositing the correct amount of charge on the floating gates of nonvolatile memory cells within a vector-matrix multiplication (VMM) array in an artificial neural network. Thus, a selected cell can be programmed with extreme precision to hold one of N different values.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Priority Declaration

[0002] This application claims priority to U.S. Provisional Patent Application No. 62 / 957,013, filed on January 3, 2020, entitled “Precise Data Tuning Method And Apparatus For Analog Neuromorphic Memory In An Artificial Neural Network,” and U.S. Patent Application No. 16 / 829,757, filed on March 25, 2020, entitled “Precise Data Tuning Method And Apparatus For Analog Neural Memory In An Artificial Neural Network.” Technical Field

[0003] The present invention discloses various embodiments of precision-tuned methods and apparatus for accurately and rapidly depositing the correct amount of charge on the floating gates of nonvolatile memory cells within vector-matrix multiplication (VMM) arrays in artificial neural networks. Background Art

[0004] Artificial neural networks simulate biological neural networks (the central nervous system of animals, especially the brain) and are used to estimate or approximate functions that may depend on a large number of inputs and are generally unknown. Artificial neural networks typically consist of layers of interconnected "neurons" that exchange messages with each other.

[0005] Figure 1 An artificial neural network is shown, where circles represent inputs or layers of neurons. Connections (called synapses) are represented by arrows and have numerical weights that can be adjusted based on experience. This allows the artificial neural network to adapt to the input and learn. Typically, an artificial neural network includes multiple layers of inputs. There are typically one or more intermediate layers of neurons, and an output layer of neurons that provide the output of the neural network. Neurons at each level make decisions based on the data received from the synapses, either individually or collectively.

[0006] One of the main challenges in developing artificial neural networks for high-performance information processing is the lack of adequate hardware technology. In fact, practical artificial neural networks rely on a large number of synapses to achieve high connectivity between neurons, that is, very high computational parallelism. In principle, such complexity can be achieved using digital supercomputers or clusters of dedicated graphics processing units. However, in addition to being high-cost, these approaches are also mediocre in energy efficiency compared to biological networks, which consume less energy mainly due to the low-precision analog calculations they perform. CMOS analog circuits have been used in artificial neural networks, but given the large number of neurons and synapses, the synapses of most CMOS implementations are too large.

[0007] Applicant previously disclosed an artificial (simulated) neural network utilizing one or more non-volatile memory arrays as synapses in U.S. patent application Ser. No. 15 / 594,439 (published as U.S. Patent Publication No. 2017 / 0337466), which is incorporated herein by reference. The non-volatile memory array operates as a simulated neuromorphic memory. As used herein, the term "neuromorphic" refers to a circuit that implements a model of a neural system. The simulated neuromorphic memory includes a first plurality of synapses configured to receive a first plurality of inputs and generate a first plurality of outputs therefrom, and a first plurality of neurons configured to receive the first plurality of outputs. The first plurality of synapses includes a plurality of memory cells, wherein each of the memory cells includes: a source region and a drain region spaced apart formed in a semiconductor substrate, wherein a channel region extends between the source region and the drain region; a floating gate disposed over and insulated from a first portion of the channel region; and a non-floating gate disposed over and insulated from a second portion of the channel region. Each of the plurality of memory cells is configured to store a weight value corresponding to a plurality of electrons on the floating gate. The plurality of memory cells are configured to multiply the first plurality of inputs by the stored weight values to generate a first plurality of outputs.An array of memory cells arranged in this manner may be referred to as a vector-matrix multiplication (VMM) array.

[0008] Each nonvolatile memory cell used in a VMM must be erased and programmed to maintain a very specific and precise amount of charge (i.e., number of electrons) in the floating gate. For example, each floating gate must maintain one of N different values, where N is the number of different weights that can be represented by each cell. Examples of N include 16, 32, 64, 128, and 256. One challenge is being able to program the selected cell with the precision and granularity required for different values of N. For example, if the selected cell can contain one of 64 different values, extremely high precision is required in the programming operation.

[0009] There is a need for improved programming systems and methods suitable for use with VMM arrays in emulated neuromorphic memories. Summary of the Invention

[0010] The present invention discloses various embodiments of a precision-tuned algorithm and apparatus for accurately and rapidly depositing the correct amount of charge on the floating gates of nonvolatile memory cells within a VMM array in an analog neuromorphic memory system. Thus, a selected cell can be programmed with extreme precision to hold one of N different values.

[0011] In one embodiment, a method for tuning a selected nonvolatile memory cell in a vector-matrix multiplication array of nonvolatile memory cells is provided, the method comprising: (i) setting an initial current target for the selected nonvolatile memory cell; (ii) performing a soft erase on all nonvolatile memory cells in the vector-matrix multiplication array; (iii) performing a coarse programming operation on the selected memory cell; (iv) performing a fine programming operation on the selected memory cell; (v) performing a read operation on the selected memory cell and determining a current consumed by the selected memory cell during the read operation; (vi) calculating an output error based on a difference between the determined current and the initial current target; and repeating steps (i), (ii), (iii), (iv), (v), and (vi) until the output error is less than a predetermined threshold.

[0012] In another embodiment, a method for tuning a selected nonvolatile memory cell in a vector-matrix multiplication array of nonvolatile memory cells is provided, the method comprising: (i) setting an initial target for the selected nonvolatile memory cell; (ii) performing a programming operation on the selected memory cell; (iii) performing a read operation on the selected memory cell and determining a cell output consumed by the selected memory cell during the read operation; (iv) calculating an output error based on a difference between the determined output and the initial target; and (v) repeating steps (i), (ii), (iii), and (iv) until the output error is less than a predetermined threshold.

[0013] In another embodiment, a neuron output circuit for providing current to program weight values in selected memory cells in a vector-matrix multiplication array is provided, the neuron output circuit comprising: a first adjustable current source that generates a scaled current in response to the neuron current to achieve a positive weight; and a second adjustable current source that generates a scaled current in response to the neuron current to achieve a negative weight.

[0014] In another embodiment, a neuron output circuit for providing a current to program a weight value in a selected memory cell in a vector-matrix multiplication array is provided, the neuron output circuit comprising: an adjustable capacitor, the adjustable capacitor comprising a first terminal and a second terminal, the second terminal providing an output voltage for the neuron output circuit; a control transistor, the control transistor comprising a first terminal and a second terminal; a first switch, the first switch selectively coupled between the first terminal and the second terminal of the adjustable capacitor; a second switch, the second switch selectively coupled between the second terminal of the adjustable capacitor and the first terminal of the control transistor; and an adjustable current source, the adjustable current source coupled to the second terminal of the control transistor.

[0015] In another embodiment, a neuron output circuit for providing a current to program a weight value in a selected memory cell in a vector-matrix multiplication array is provided, the neuron output circuit comprising: an adjustable capacitor, the adjustable capacitor comprising a first terminal and a second terminal, the second terminal providing an output voltage for the neuron output circuit; a control transistor, the control transistor comprising a first terminal and a second terminal; a switch, the switch selectively coupled between the second terminal of the adjustable capacitor and the first terminal of the control transistor; and an adjustable current source, the adjustable current source coupled to the second terminal of the control transistor.

[0016] In another embodiment, a neuron output circuit for providing a current to program a weight value in a selected memory cell in a vector-matrix multiplication array is provided, the neuron output circuit comprising: an adjustable capacitor, the adjustable capacitor comprising a first terminal and a second terminal, the first terminal providing an output voltage for the neuron output circuit; a control transistor, the control transistor comprising a first terminal and a second terminal; a first switch, the first switch selectively coupled between the first terminal of the adjustable capacitor and the first terminal of the control transistor; and an adjustable current source, the adjustable current source coupled to the second terminal of the control transistor.

[0017] In another embodiment, a neuron output circuit for providing current to program a weight value in a selected memory cell in a vector-matrix multiplication array is provided, the neuron output circuit comprising: a first operational amplifier, the first operational amplifier comprising an inverting input, a non-inverting input, and an output; a second operational amplifier, the second operational amplifier comprising an inverting input, a non-inverting input, and an output; a first adjustable current source, the first adjustable current source coupled to the inverting input of the first operational amplifier; a second adjustable current source, the second adjustable current source coupled to the inverting input of the second operational amplifier; a first adjustable resistor, the first adjustable resistor coupled to the inverting input of the first operational amplifier; a second adjustable resistor, the second adjustable resistor coupled to the inverting input of the second operational amplifier; and a third adjustable resistor coupled between the output of the first operational amplifier and the inverting input of the second operational amplifier.

[0018] In another embodiment, a neuron output circuit for providing current to program a weight value in a selected memory cell in a vector-matrix multiplication array is provided, the neuron output circuit comprising: a first operational amplifier, the first operational amplifier comprising an inverting input, a non-inverting input, and an output; a second operational amplifier, the second operational amplifier comprising an inverting input, a non-inverting input, and an output; a first adjustable current source, the first adjustable current source coupled to the inverting input of the first operational amplifier; a second adjustable current source, the second adjustable current source coupled to the inverting input of the second operational amplifier; a first switch, the first switch coupled between the inverting input and the output of the first operational amplifier; a second switch, the second switch coupled between the inverting input and the output of the second operational amplifier; a first adjustable capacitor, the first adjustable capacitor coupled between the inverting input and the output of the first operational amplifier; a second adjustable capacitor, the second adjustable capacitor coupled between the inverting input and the output of the second operational amplifier; and a third adjustable capacitor coupled between the output of the first operational amplifier and the inverting input of the second operational amplifier. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 FIG. 1 is a schematic diagram showing an artificial neural network of the prior art.

[0020] Figure 2 A prior art split-gate flash memory cell is shown.

[0021] Figure 3 Another prior art split-gate flash memory cell is shown.

[0022] Figure 4 Another prior art split-gate flash memory cell is shown.

[0023] Figure 5 Another prior art split-gate flash memory cell is shown.

[0024] Figure 6 Another prior art split-gate flash memory cell is shown.

[0025] Figure 7 A prior art stacked gate flash memory cell is shown.

[0026] Figure 8 is a schematic diagram showing different levels of an exemplary artificial neural network using one or more VMM arrays.

[0027] Figure 9 FIG. 4 is a block diagram illustrating a VMM system including a VMM array and other circuits.

[0028] Figure 10 is a block diagram illustrating an exemplary artificial neural network using one or more VMM systems.

[0029] Figure 11 Another embodiment of a VMM array is shown.

[0030] Figure 12 Another embodiment of a VMM array is shown.

[0031] Figure 13 Another embodiment of a VMM array is shown.

[0032] Figure 14 Another embodiment of a VMM array is shown.

[0033] Figure 15 Another embodiment of a VMM array is shown.

[0034] Figure 16 Another embodiment of a VMM array is shown.

[0035] Figure 17 Another embodiment of a VMM array is shown.

[0036] Figure 18 Another embodiment of a VMM array is shown.

[0037] Figure 19 Another embodiment of a VMM array is shown.

[0038] Figure 20 Another embodiment of a VMM array is shown.

[0039] Figure 21 Another embodiment of a VMM array is shown.

[0040] Figure 22 Another embodiment of a VMM array is shown.

[0041] Figure 23 Another embodiment of a VMM array is shown.

[0042] Figure 24 Another embodiment of a VMM array is shown.

[0043] Figure 25 A prior art long short-term memory system is shown.

[0044] Figure 26 An exemplary cell used in a long short-term memory system is shown.

[0045] Figure 27 Show Figure 26 An embodiment of an exemplary unit of .

[0046] Figure 28 Show Figure 26 Another embodiment of an exemplary unit of .

[0047] Figure 29 A prior art gate-controlled recursive cell system is shown.

[0048] Figure 30 An exemplary cell for use in a gate-controlled recursive cell system is shown.

[0049] Figure 31 Show Figure 30 An embodiment of an exemplary unit of .

[0050] Figure 32 Show Figure 30 Another embodiment of an exemplary unit of .

[0051] Figure 33 A VMM system is shown.

[0052] Figure 34 A tuning correction method is shown.

[0053] Figure 35A A tuning correction method is shown.

[0054] Figure 35B The sector tuning correction method is shown.

[0055] Figure 36A Shows the effect of temperature on the value stored in a cell.

[0056] Figure 36BProblems caused by data drift during operation of a VMM system are shown.

[0057] Figure 36C Shown are blocks for compensating for data drift.

[0058] Figure 36D A data drift monitor is shown.

[0059] Figure 37 A bit line compensation circuit is shown.

[0060] Figure 38 Another line compensation circuit is shown.

[0061] Figure 39 Another line compensation circuit is shown.

[0062] Figure 40 Another line compensation circuit is shown.

[0063] Figure 41 Another line compensation circuit is shown.

[0064] Figure 42 Another line compensation circuit is shown.

[0065] Figure 43 A neuronal circuit is shown.

[0066] Figure 44 Another neuronal circuit is shown.

[0067] Figure 45 Another neuronal circuit is shown.

[0068] Figure 46 Another neuronal circuit is shown.

[0069] Figure 47 Another neuronal circuit is shown.

[0070] Figure 48 Another neuronal circuit is shown.

[0071] Figure 49A A block diagram showing the output circuit.

[0072] Figure 49B A block diagram of another output circuit is shown.

[0073] Figure 49C A block diagram of another output circuit is shown. DETAILED DESCRIPTION

[0074] The artificial neural network of the present invention utilizes a combination of CMOS technology and non-volatile memory arrays.

[0075] Non-volatile memory cells

[0076] Digital non-volatile memory is well known. For example, U.S. Patent No. 5,029,130 (“the '130 patent”), which is incorporated herein by reference, discloses an array of split-gate non-volatile memory cells, which is a type of flash memory cell. Such a memory cell 210 is Figure 2 . Each memory cell 210 includes a source region 14 and a drain region 16 formed in a semiconductor substrate 12, with a channel region 18 therebetween. A floating gate 20 is formed over and insulated from a first portion of the channel region 18 (and controls its electrical conductivity), and is formed over a portion of the source region 14. A wordline terminal 22 (which is typically coupled to a wordline) has a first portion disposed over and insulated from a second portion of the channel region 18 (and controls its electrical conductivity), and a second portion extending upward and over the floating gate 20. The floating gate 20 and wordline terminal 22 are insulated from the substrate 12 by a gate oxide. A bitline terminal 24 is coupled to the drain region 16.

[0077] Memory cell 210 is erased (where electrons are removed from the floating gate) by placing a high positive voltage on wordline terminal 22, which causes the electrons on floating gate 20 to tunnel through the intervening insulator via Fowler-Nordheim tunneling from floating gate 20 to wordline terminal 22.

[0078] Memory cell 210 is programmed by placing a positive voltage on wordline terminal 22 and a positive voltage on source region 14 (where electrons are placed on the floating gate). Electron current will flow from source region 14 (source line terminal) to drain region 16. When electrons reach the gap between wordline terminal 22 and floating gate 20, they will accelerate and become heated. Due to the electrostatic attraction from floating gate 20, some of the heated electrons will be injected through the gate oxide onto floating gate 20.

[0079] Memory cell 210 is read by placing a positive read voltage across drain region 16 and wordline terminal 22 (which turns on the portion of channel region 18 below the wordline terminal). If floating gate 20 is positively charged (i.e., electrons are erased), the portion of channel region 18 below floating gate 20 is also turned on, and current will flow through channel region 18, which is sensed as an erased state or a "1" state. If floating gate 20 is negatively charged (i.e., programmed by electrons), the portion of the channel region below floating gate 20 is mostly or completely turned off, and no current (or very little current) will flow through channel region 18, which is sensed as a programmed state or a "0" state.

[0080] Table 1 shows typical voltage ranges that may be applied to the terminals of the memory cell 110 for performing read operations, erase operations, and program operations:

[0081] Table 1: Figure 2 Operation of the flash memory unit 210

[0082] WL BL SL Read 1 0.5-3V 0.1-2V 0V Read 2 0.5-3V 0-2V 2-0.1V Erase About 11-13V 0V 0V programming 1V-2V 1-3μA 9-10V

[0083] "Read 1" is a read mode in which the cell current is output on the bit line. "Read 2" is a read mode in which the cell current is output on the source line terminal.

[0084] Figure 3 Memory cell 310 is shown, which is connected to Figure 2 The memory cell 210 is similar to the memory cell 210 of FIG, but with the addition of a control gate (CG) terminal 28. The control gate terminal 28 is biased at a high voltage (e.g., 10V) during programming, at a low voltage or negative voltage (e.g., 0V / -8V) during erasing, and at a low voltage or medium voltage (e.g., 0V / 2.5V) during reading. The other terminals are similar to Figure 2 That's biased.

[0085] Figure 4 A quad-gate memory cell 410 is shown, comprising a source region 14, a drain region 16, a floating gate 20 over a first portion of a channel region 18, a select gate 22 (typically coupled to a word line WL) over a second portion of the channel region 18, a control gate 28 over the floating gate 20, and an erase gate 30 over the source region 14. This configuration is described in U.S. Patent 6,747,310, which is incorporated herein by reference for all purposes. Here, except for the floating gate 20, all gates are non-floating, meaning they are electrically connected or capable of being electrically connected to a voltage source. Programming is performed by heated electrons from the channel region 18 that inject themselves into the floating gate 20. Erasing is performed by electrons tunneling from the floating gate 20 to the erase gate 30.

[0086] Table 2 shows typical voltage ranges that may be applied to the terminals of the memory cell 410 for performing read operations, erase operations, and program operations:

[0087] Table 2: Figure 4 Operation of the flash memory unit 410

[0088] WL / SG BL CG EG SL Read 1 0.5-2V 0.1-2V 0-2.6V 0-2.6V 0V Read 2 0.5-2V 0-2V 0-2.6V 0-2.6V 2-0.1V Erase -0.5V / 0V 0V 0V / -8V 8-12V 0V programming 1V 1μA 8-11V 4.5-9V 4.5-5V

[0089] "Read 1" is a read mode in which the cell current is output on the bit line. "Read 2" is a read mode in which the cell current is output on the source line terminal.

[0090] Figure 5 Memory cell 510 is shown, except that it does not have an erase gate EG terminal. Memory cell 510 is similar to Figure 4Erasing is performed by biasing the substrate 18 to a high voltage and biasing the control gate CG terminal 28 to a low or negative voltage. Alternatively, erasing is performed by biasing the word line terminal 22 to a positive voltage and biasing the control gate terminal 28 to a negative voltage. Programming and reading are similar to Figure 4 Like that.

[0091] Figure 6 A tri-gate memory cell 610 is shown, which is another type of flash memory cell. Figure 4 The memory cell 410 is the same as the memory cell 610 except that the memory cell 610 does not have a separate control gate terminal. Except that no control gate bias is applied, the erase operation (erasing by using the erase gate terminal) and the read operation are similar to Figure 4 The programming operation is also completed without a control gate bias, and as a result, a higher voltage must be applied on the source line terminal during the programming operation to compensate for the lack of control gate bias.

[0092] Table 3 shows typical voltage ranges that may be applied to the terminals of the memory cell 610 for performing read operations, erase operations, and program operations:

[0093] Table 3: Figure 6 Operation of the flash memory unit 610

[0094] WL / SG BL EG SL Read 1 0.5-2.2V 0.1-2V 0-2.6V 0V Read 2 0.5-2.2V 0-2V 0-2.6V 2-0.1V Erase -0.5V / 0V 0V 11.5V 0V programming 1V 2-3μA 4.5V 7-9V

[0095] "Read 1" is a read mode in which the cell current is output on the bit line. "Read 2" is a read mode in which the cell current is output on the source line terminal.

[0096] Figure 7 A stacked gate memory cell 710 is shown, which is another type of flash memory cell. Figure 2 Memory cell 210 is similar to that of FIG1 , except that floating gate 20 extends over the entire channel region 18, and control gate terminal 22 (which here will be coupled to a word line) extends over floating gate 20, separated by an insulating layer (not shown). Erase, program, and read operations operate in a similar manner as previously described for memory cell 210.

[0097] Table 4 shows typical voltage ranges that may be applied to the terminals of the memory cell 710 and substrate 12 for performing read, erase, and program operations:

[0098] Table 4: Figure 7 Operation of the flash memory unit 710

[0099] CG BL SL substrate Read 1 0-5V 0.1-2V 0-2V 0V Read 2 0.5-2V 0-2V 2-0.1V 0V Erase -8 to -10V / 0V FLT FLT 8-10V / 15-20V programming 8-12V 3-5V / 0V 0V / 3-5V 0V

[0100] "Read 1" is a read mode in which the cell current is output on the bit line. "Read 2" is a read mode in which the cell current is output on the source line terminal. Optionally, in an array including rows and columns of memory cells 210, 310, 410, 510, 610, or 710, the source line can be coupled to a row of memory cells or two adjacent rows of memory cells. That is, the source line terminal can be shared by memory cells in adjacent rows.

[0101] In order to utilize a memory array comprising one of the above-described types of nonvolatile memory cells in an artificial neural network, two modifications were made. First, the circuitry was configured so that each memory cell could be individually programmed, erased, and read without adversely affecting the memory states of other memory cells in the array, as explained further below. Second, continuous (analog) programming of the memory cells was provided.

[0102] Specifically, the memory state (i.e., charge on the floating gate) of each memory cell in the array can be changed continuously from a fully erased state to a fully programmed state independently and with minimal disturbance to other memory cells. In another embodiment, the memory state (i.e., charge on the floating gate) of each memory cell in the array can be changed continuously from a fully programmed state to a fully erased state, and vice versa, independently and with minimal disturbance to other memory cells. This means that the cell storage device is analog, or at least can store one of many discrete values (such as 16 or 64 different values), which allows very precise and individual tuning of all cells in the memory array, and makes the memory array ideal for storing and fine-tuning the synaptic weights of neural networks.

[0103] The methods and apparatus described herein can be applied to other non-volatile memory technologies, such as, but not limited to, SONOS (silicon-oxide-nitride-oxide-silicon, charge trapped in nitride), MONOS (metal-oxide-nitride-oxide-silicon, metal charge trapped in nitride), ReRAM (resistive RAM), PCM (phase change memory), MRAM (magnetic RAM), FeRAM (ferroelectric RAM), OTP (two-layer or multi-layer one-time programmable), and CeRAM (correlated electron RAM). The methods and apparatus described herein can be applied to volatile memory technologies used in neural networks, such as, but not limited to, SRAM, DRAM, and / or volatile synaptic cells.

[0104] Neural Networks Using Nonvolatile Memory Cell Arrays

[0105] Figure 8A non-limiting example of a neural network using a non-volatile memory array according to the present embodiment is conceptually illustrated. This example uses the non-volatile memory array neural network for a facial recognition application, but any other suitable application may also be implemented using a non-volatile memory array-based neural network.

[0106] For this example, S0 is the input layer, which is a 32x32 pixel RGB image with 5 bits of precision (i.e., three 32x32 pixel arrays, one for each color R, G, and B, with 5 bits of precision per pixel). Synapse CB1 from input layer S0 to layer C1 applies different sets of weights in some cases and shared weights in other cases, and scans the input image with a 3x3 pixel overlapping filter (kernel), shifting the filter by 1 pixel (or more than 1 pixel as dictated by the model). Specifically, the values of 9 pixels in a 3x3 portion of the image (i.e., called the filter or kernel) are provided to synapse CB1, where these 9 input values are multiplied by the appropriate weights, and after summing the outputs of these multiplications, a single output value is determined and provided by the first synapse of CB1 for use in generating a feature map for one of the pixels of layer C1. The 3x3 filter is then shifted one pixel to the right within the input layer S0 (i.e., a column of three pixels on the right is added and a column of three pixels on the left is released), whereby the nine pixel values in this newly positioned filter are provided to the synapse CB1, where they are multiplied by the same weights and a second single output value is determined by the associated synapse. This process continues until the 3x3 filter has scanned all three colors and all bits (precision values) across the entire 32x32 pixel image of the input layer S0. This process is then repeated using different sets of weights to generate different feature maps for C1 until all feature maps for layer C1 are calculated.

[0107] At layer C1, in this example, there are 16 feature maps, each with 30x30 pixels. Each pixel is a new feature pixel extracted from the product of the input and the kernel, so each feature map is a two-dimensional array, so in this example, layer C1 is composed of a two-dimensional array of 16 layers (remember that the layers and arrays referred to in this article are logical relationships, not necessarily physical relationships, that is, arrays do not have to be oriented to physical two-dimensional arrays). Each of the 16 feature maps in layer C1 is generated by one of sixteen different sets of synaptic weights applied to the filter scan. The C1 feature maps can all relate to different aspects of the same image features, such as edge recognition. For example, a first map (generated using a first set of weights, shared by all scans used to generate it) can identify circular edges, a second map (generated using a second set of weights different from the first) can identify rectangular edges, or the aspect ratio of certain features, and so on.

[0108] Before passing from layer C1 to layer S1, an activation function P1 (pooling) is applied, which pools the values from consecutive non-overlapping 2x2 regions in each feature map. The purpose of the pooling function is to average the values of adjacent locations (or a max function can be used), for example to reduce dependencies at edge locations, and to reduce the size of the data before entering the next stage. At layer S1, there are 16 15x15 feature maps (i.e., 16 different arrays of 15x15 pixels each). The synapse CB2 from layer S1 to layer C2 scans the map in S1 using a 4x4 filter, with the filter shifted by 1 pixel. At layer C2, there are 22 12x12 feature maps. Before passing from layer C2 to layer S2, an activation function P2 (pooling) is applied, which pools the values from consecutive non-overlapping 2x2 regions in each feature map. At layer S2, there are 22 6x6 feature maps. An activation function (pooling) is applied to the synapse CB3 from layer S2 to layer C3, where each neuron in layer C3 is connected to each map in layer S2 via a corresponding synapse on CB3. At layer C3, there are 64 neurons. Synapse CB4 from layer C3 to output layer S3 completely connects C3 to S3, that is, every neuron in layer C3 is connected to every neuron in layer S3. The output at S3 includes 10 neurons, where the highest output neuron determines the class. For example, this output can indicate the recognition or classification of the content of the original image.

[0109] The synapses at each layer are implemented using an array or a portion of an array of non-volatile memory cells.

[0110] Figure 9 A block diagram of a system that can be used for this purpose is shown in FIG. The VMM system 32 includes non-volatile memory units and serves as a synapse between one layer and the next (such as Figure 6 CB1, CB2, CB3, and CB4 in FIG. Specifically, the VMM system 32 includes a VMM array 33 (including nonvolatile memory cells arranged in rows and columns), an erase gate and word line gate decoder 34, a control gate decoder 35, a bit line decoder 36, and a source line decoder 37, which decode the corresponding inputs of the nonvolatile memory cell array 33. The inputs to the VMM array 33 can come from the erase gate and word line gate decoder 34 or from the control gate decoder 35. In this example, the source line decoder 37 also decodes the output of the VMM array 33. Alternatively, the bit line decoder 36 can decode the output of the VMM array 33.

[0111] The VMM array 33 serves two purposes. First, it stores weights to be used by the VMM system 32. Second, the VMM array 33 effectively multiplies the inputs by the weights stored in the VMM array 33 and each output line (source line or bit line) adds them together to produce an output, which will serve as the input to the next layer or the final layer. By performing multiplication and addition functions, the VMM array 33 eliminates the need for separate multiplication and addition logic circuits and is also highly power-efficient due to its in-situ memory calculations.

[0112] The output of the VMM array 33 is provided to a differential summer (such as a summing operational amplifier or a summing current mirror) 38, which sums the output of the VMM array 33 to create a single value for the convolution. The differential summer 38 is arranged to perform the summation of both the positive weight input and the negative weight input to output a single value.

[0113] The output values of the difference summer 38 are then summed and provided to the activation function circuit 39, which corrects the output. The activation function circuit 39 can provide a sigmoid, tanh, ReLU function or any other nonlinear function. The corrected output value of the activation function circuit 39 becomes the next layer (for example, Figure 8 The elements of the feature map of layer C1 in the image are then applied to the next synapse to produce the next feature map layer or the final layer. Thus, in this example, the VMM array 33 constitutes a plurality of synapses (which receive their inputs from existing neuron layers or from an input layer such as an image database), and the summer 38 and activation function circuit 39 constitute a plurality of neurons.

[0114] Figure 9 The inputs to the VMM system 32 (WLx, EGx, CGx, and optionally BLx and SLx) can be analog levels, binary levels, digital pulses (in which case a pulse-to-analog converter PAC may be required to convert the pulses to appropriate input analog levels), or digital bits (in which case a DAC is provided to convert the digital bits to appropriate input analog levels); the outputs can be analog levels, binary levels, digital pulses, or digital bits (in which case an output ADC is provided to convert the output analog levels into digital bits).

[0115] Figure 10 FIG. 1 is a block diagram illustrating the use of multiple layers of VMM systems 32 (labeled here as VMM systems 32a, 32b, 32c, 32d, and 32e). Figure 10As shown, the input (denoted as Inputx) is converted from digital to analog by a digital-to-analog converter 31 and provided to the input VMM system 32a. The converted analog input can be a voltage or a current. The first level of input D / A conversion can be accomplished by using a function or LUT (lookup table) that maps the input Inputx to the appropriate analog levels of the matrix multiplier of the input VMM system 32a. Input conversion can also be accomplished by an analog-to-analog (A / A) converter to convert the external analog input into a mapped analog input to the input VMM system 32a. Input conversion can also be accomplished by a digital-to-digital pulse (D / P) converter to convert the external digital input into one or more digital pulses that are mapped to the input VMM system 32a.

[0116] The output generated by input VMM system 32a is provided as input to the next VMM system (hidden level 1) 32b, which in turn generates an output that is provided as input to the next VMM system (hidden level 2) 32c, and so on. The layers of VMM system 32 serve as different layers of synapses and neurons of a convolutional neural network (CNN). Each VMM system 32a, 32b, 32c, 32d, and 32e can be a separate physical system including a corresponding non-volatile memory array, or multiple VMM systems can utilize different portions of the same physical non-volatile memory array, or multiple VMM systems can utilize overlapping portions of the same physical non-volatile memory array. Each VMM system 32a, 32b, 32c, 32d, and 32e can also be time-division multiplexed for different portions of its array or neurons. Figure 10 The example shown includes five layers (32a, 32b, 32c, 32d, 32e): one input layer (32a), two hidden layers (32b, 32c), and two fully connected layers (32d, 32e). Those skilled in the art will appreciate that this is merely exemplary and that, on the contrary, the system may include more than two hidden layers and more than two fully connected layers.

[0117] VMM array

[0118] Figure 11 A neuron VMM array 1100 is shown, which is particularly suitable for Figure 3 The memory cells 310 shown are used as synapses and components for neurons between the input layer and the next layer. The VMM array 1100 includes a memory array 1101 of nonvolatile memory cells and a reference array 1102 of nonvolatile reference memory cells (at the top of the array). Alternatively, another reference array can be placed at the bottom.

[0119] In VMM array 1100, control gate lines (such as control gate line 1103) extend in the vertical direction (so reference array 1102 is orthogonal to control gate line 1103 in the row direction), and erase gate lines (such as erase gate line 1104) extend in the horizontal direction. Here, the inputs of VMM array 1100 are provided on control gate lines (CG0, CG1, CG2, CG3), and the outputs of VMM array 1100 appear on source lines (SL0, SL1). In one embodiment, only even-numbered rows are used, and in another embodiment, only odd-numbered rows are used. The current placed on each source line (SL0, SL1, respectively) performs a summation function of all currents from the memory cells connected to that particular source line.

[0120] As described herein for neural networks, the non-volatile memory cells of VMM array 1100 (ie, the flash memory of VMM array 1100) are preferably configured to operate in the sub-threshold region.

[0121] Biasing the nonvolatile reference memory cell and the nonvolatile memory cell described herein in weak inversion:

[0122] Ids=Io*e (Vg-Vth) / nVt =w*Io*e (Vg) / nVt ,

[0123] where w = e (-Vth) / nVt

[0124] Where Ids is the drain-to-source current; Vg is the gate voltage on the memory cell; Vth is the threshold voltage of the memory cell; Vt is the thermal voltage = k*T / q, where k is the Boltzmann constant, T is the temperature in Kelvin, and q is the electron charge; n is the slope factor = 1+(Cdep / Cox), where Cdep = the capacitance of the depletion layer and Cox is the capacitance of the gate oxide layer; Io is the memory cell current at a gate voltage equal to the threshold voltage, and Io is the product of (Wt / L)*u*Cox*(n-1)*Vt 2 is proportional to where u is the carrier mobility, and Wt and L are the width and length of the memory cell, respectively.

[0125] For an I-to-V logarithmic converter that converts an input current Ids to an input voltage Vg using a memory cell (such as a reference memory cell or a peripheral memory cell) or a transistor:

[0126] Vg=n*Vt*log[Ids / wp*Io]

[0127] Here, wp is w of a reference memory cell or a peripheral memory cell.

[0128] For an I-to-V logarithmic converter that converts an input current Ids to an input voltage Vg using a memory cell (such as a reference memory cell or a peripheral memory cell) or a transistor:

[0129] Vg=n*Vt*log[Ids / wp*Io]

[0130] Here, wp is w of a reference memory cell or a peripheral memory cell.

[0131] For a memory array used as a vector matrix multiplier VMM array, the output current is:

[0132] Iout=wa*Io*e (Vg) / nVt ,Right now

[0133] Iout=(wa / wp)*Iin=W*Iin

[0134] W=e (Vthp-Vtha) / nVt

[0135] Iin=wp*Io*e (Vg) / nVt

[0136] Here, wa = w for each memory cell in the memory array.

[0137] The word line or control gate may be used as the input to the memory cell for the input voltage.

[0138] Alternatively, the non-volatile memory cells of the VMM array described herein may be configured to operate in the linear region:

[0139] Ids=β*(Vgs-Vth)*Vds; β=u*Cox*Wt / L,Wα(Vgs-Vth),

[0140] Meaning that the weight W in the linear region is proportional to (Vgs-Vth)

[0141] The word line or control gate or bit line or source line can serve as the input of the memory cell operating in the linear region. The bit line or source line can serve as the output of the memory cell.

[0142] For an I to V linear converter, a memory cell (eg, a reference memory cell or a peripheral memory cell) or a transistor or a resistor operating in a linear region may be used to linearly convert an input / output current into an input / output voltage.

[0143] Alternatively, the memory cells of the VMM array described herein may be configured to operate in the saturation region:

[0144] Ids=1 / 2*β*(Vgs-Vth) 2 ;β=u*Cox*Wt / L

[0145] Wα(Vgs-Vth) 2 , which means the weight W and (Vgs-Vth) 2 Proportional

[0146] The word line, control gate, or erase gate can be used as the input of a memory cell operating in the saturation region. The bit line or source line can be used as the output of an output neuron.

[0147] Alternatively, the memory cells of the VMM arrays described herein may be used in all regions or a combination thereof (subthreshold, linear, or saturation regions).

[0148] U.S. Patent Application No. 15 / 826,345 describes Figure 9 Other embodiments of the VMM array 33 of , which is incorporated herein by reference. As described herein, source lines or bit lines can be used as neuron outputs (current summing outputs).

[0149] Figure 12 A neuron VMM array 1200 is shown, which is particularly suitable for Figure 2 Memory cell 210 is shown and serves as a synapse between the input layer and the next layer. VMM array 1200 includes a memory array 1203 of nonvolatile memory cells, a reference array 1201 of first nonvolatile reference memory cells, and a reference array 1202 of second nonvolatile reference memory cells. Reference arrays 1201 and 1202, arranged along the columns of the array, are used to convert current inputs flowing into terminals BLR0, BLR1, BLR2, and BLR3 into voltage inputs WL0, WL1, WL2, and WL3. In practice, the first and second nonvolatile reference memory cells are diode-connected via a multiplexer 1214 (only partially shown), with the current input flowing therein. The reference cells are tuned (e.g., programmed) to a target reference level. The target reference level is provided by a reference microarray matrix (not shown).

[0150] The memory array 1203 serves two purposes. First, it stores the weights that the VMM array 1200 will use on its corresponding memory cells. Second, the memory array 1203 effectively multiplies the inputs (i.e., the current inputs provided at terminals BLR0, BLR1, BLR2, and BLR3, which the reference arrays 1201 and 1202 convert into input voltages to be provided to the word lines WL0, WL1, WL2, and WL3) by the weights stored in the memory array 1203, and then adds all the results (memory cell currents) to produce an output on the corresponding bit lines (BL0-BLN), which will be the input to the next layer or the final layer. By performing multiplication and addition functions, the memory array 1203 eliminates the need for separate multiplication logic circuits and addition logic circuits and is also highly power-efficient. Here, the voltage inputs are provided on the word lines (WL0, WL1, WL2, and WL3), and the outputs appear on the corresponding bit lines (BL0-BLN) during a read (inference) operation. The current placed on each of the bit lines BL0-BLN performs a summing function of the currents from all of the nonvolatile memory cells connected to that particular bit line.

[0151] Table 5 shows the operating voltages for VMM array 1200. The columns in the table indicate the voltages applied to the word line for a selected cell, the word line for an unselected cell, the bit line for a selected cell, the bit line for an unselected cell, the source line for a selected cell, and the source line for an unselected cell, where FLT indicates floating, i.e., no voltage applied. The rows indicate read, erase, and program operations.

[0152] Table 5: Figure 12 Operation of the VMM array 1200

[0153] WL WL-Not selected BL BL-Not selected SL SL-Not selected Read 0.5-3.5V -0.5V / 0V 0.1-2V(Ineuron) 0.6V-2V / FLT 0V 0V Erase About 5-13V 0V 0V 0V 0V 0V programming 1V-2V -0.5V / 0V 0.1-3uA Vinh about 2.5V 4-10V 0-1V / FLT

[0154] Figure 13 A neuron VMM array 1300 is shown, which is particularly suitable for Figure 2Memory cell 210 is shown and serves as a synapse and component for neurons between the input layer and the next layer. VMM array 1300 includes a memory array 1303 of nonvolatile memory cells, a reference array 1301 of first nonvolatile reference memory cells, and a reference array 1302 of second nonvolatile reference memory cells. Reference arrays 1301 and 1302 extend in the row direction of VMM array 1300. VMM array 1300 is similar to VMM 1000, except that in VMM array 1300, word lines extend in the vertical direction. Here, inputs are provided on word lines (WLA0, WLB0, WLA1, WLB2, WLA2, WLB2, WLA3, WLB3), and outputs appear on source lines (SL0, SL1) during a read operation. The current placed on each source line performs a summing function of all currents from the memory cells connected to that particular source line.

[0155] Table 6 shows the operating voltages for VMM array 1300. The columns in the table indicate the voltages placed on the word line for a selected cell, the word line for an unselected cell, the bit line for a selected cell, the bit line for an unselected cell, the source line for a selected cell, and the source line for an unselected cell. The rows indicate read, erase, and program operations.

[0156] Table 6: Figure 13 Operation of the VMM array 1300

[0157]

[0158] Figure 14 A neuron VMM array 1400 is shown, which is particularly suitable for Figure 3 Memory cell 310 is shown and serves as a synapse and component of neurons between the input layer and the next layer. VMM array 1400 includes a memory array 1403 of nonvolatile memory cells, a reference array 1401 of first nonvolatile reference memory cells, and a reference array 1402 of second nonvolatile reference memory cells. Reference arrays 1401 and 1402 are used to convert current inputs flowing into terminals BLR0, BLR1, BLR2, and BLR3 into voltage inputs CG0, CG1, CG2, and CG3. In practice, the first nonvolatile reference memory cell and the second nonvolatile reference memory cell are diode-connected via a multiplexer 1412 (only partially shown), with the current input flowing into them through BLR0, BLR1, BLR2, and BLR3. Multiplexers 1412 each include a respective multiplexer 1405 and a cascode transistor 1404 to ensure a constant voltage on a bit line (such as BLR0) of each of the first and second nonvolatile reference memory cells during a read operation. The reference cells are tuned to a target reference level.

[0159] The memory array 1403 serves two purposes. First, it stores the weights that will be used by the VMM array 1400. Second, the memory array 1403 effectively multiplies the inputs (current inputs provided to terminals BLR0, BLR1, BLR2, and BLR3, which the reference arrays 1401 and 1402 convert into input voltages to provide to the control gates CG0, CG1, CG2, and CG3) by the weights stored in the memory array and then adds all the results (cell currents) to produce an output, which appears in BL0-BLN and will be the input to the next layer or the final layer. By performing multiplication and addition functions, the memory array eliminates the need for separate multiplication and addition logic circuits and is also highly power-efficient. Here, the inputs are provided on the control gate lines (CG0, CG1, CG2, and CG3) and the outputs appear on the bit lines (BL0-BLN) during a read operation. The current placed on each bit line performs a summing function of all the currents from the memory cells connected to that particular bit line.

[0160] The VMM array 1400 implements unidirectional tuning for the nonvolatile memory cells in the memory array 1403. That is, each nonvolatile memory cell is erased and then partially programmed until the desired charge on the floating gate is reached. This can be performed, for example, using the precision programming techniques described below. If too much charge is placed on the floating gate (causing an incorrect value to be stored in the cell), the cell must be erased and the sequence of partial programming operations must be restarted. As shown, two rows sharing the same erase gate (such as EG0 or EG1) need to be erased together (which is called a page erase), and thereafter, each cell is partially programmed until the desired charge on the floating gate is reached.

[0161] Table 7 shows the operating voltages for VMM array 1400. The columns in the table indicate the voltages applied to the word line for a selected cell, the word line for an unselected cell, the bit line for a selected cell, the bit line for an unselected cell, the control gate for a selected cell, the control gate for an unselected cell in the same sector as the selected cell, the control gate for an unselected cell in a different sector from the selected cell, the erase gate for a selected cell, the erase gate for an unselected cell, the source line for a selected cell, and the source line for an unselected cell. The rows indicate read, erase, and program operations.

[0162] Table 7: Figure 14 Operation of the VMM array 1400

[0163]

[0164] Figure 15 A neuron VMM array 1500 is shown, which is particularly suitable for Figure 3Memory cells 310 are shown and serve as synapses and components for neurons between the input layer and the next layer. VMM array 1500 includes a memory array 1503 of nonvolatile memory cells, a reference array 1501 of first nonvolatile reference memory cells, and a reference array 1502 of second nonvolatile reference memory cells. EG lines EGR0, EG0, EG1, and EGR1 extend vertically, while CG lines CG0, CG1, CG2, and CG3 and SL lines WL0, WL1, WL2, and WL3 extend horizontally. VMM array 1500 is similar to VMM array 1400, except that VMM array 1500 implements bidirectional tuning, where each individual cell can be fully erased, partially programmed, and partially erased as needed to achieve a desired charge on the floating gate due to the use of separate EG lines. As shown, reference arrays 1501 and 1502 convert input currents in terminals BLR0, BLR1, BLR2, and BLR3 into control gate voltages CG0, CG1, CG2, and CG3 to be applied to the memory cells in the row direction (through the action of the reference cells connected via the diodes of multiplexer 1514). The current outputs (neurons) are in bit lines BL0-BLN, where each bit line sums all currents from the nonvolatile memory cells connected to that particular bit line.

[0165] Table 8 shows the operating voltages for VMM array 1500. The columns in the table indicate the voltages applied to the word line for a selected cell, the word line for an unselected cell, the bit line for a selected cell, the bit line for an unselected cell, the control gate for a selected cell, the control gate for an unselected cell in the same sector as the selected cell, the control gate for an unselected cell in a different sector from the selected cell, the erase gate for a selected cell, the erase gate for an unselected cell, the source line for a selected cell, and the source line for an unselected cell. The rows indicate read, erase, and program operations.

[0166] Table 8: Figure 15 Operation of the VMM array 1500

[0167]

[0168] Figure 16 A neuron VMM array 1600 is shown, which is particularly suitable for Figure 2 The memory cell 210 shown is used as a synapse and component of neurons between the input layer and the next layer. In the VMM array 1600, the input INPUT0..., INPUT N On bit lines BL0,...BL N The signals are received on the source lines SL0, SL1, SL2 and SL3, and outputs OUTPUT1, OUTPUT2, OUTPUT3 and OUTPUT4 are generated on the source lines SL0, SL1, SL2 and SL3, respectively.

[0169] Figure 17 A neuron VMM array 1700 is shown, which is particularly suitable for Figure 2 The memory unit 210 shown in FIG. 2 is used as a synapse and a component of a neuron between an input layer and a next layer. In this example, the inputs INPUT0, INPUT 1、 INPUT2 and INPUT3 are received on source lines SL0, SL1, SL2, and SL3, respectively, and output OUTPUT0, ... OUTPUT N On bit lines BL0,…,BL N Generate on.

[0170] Figure 18 A neuron VMM array 1800 is shown, which is particularly suitable for Figure 2 The memory unit 210 shown is used as a synapse and a component of the neurons between the input layer and the next layer. In this example, the inputs INPUT0, ..., INPUT M On word lines WL0,…,WL M is received and output OUTPUT0,…OUTPUT N On bit lines BL0,…,BL N Generate on.

[0171] Figure 19 A neuron VMM array 1900 is shown, which is particularly suitable for Figure 3 The memory unit 310 shown is used as a synapse and a component of the neurons between the input layer and the next layer. In this example, the inputs INPUT0, ..., INPUT M On word lines WL0,…,WL M is received and output OUTPUT0,…OUTPUT N On bit lines BL0,…,BL N Generate on.

[0172] Figure 20 FIG. 2 shows a neuron VMM array 2000, which is particularly suitable for Figure 4 The memory unit 410 shown in FIG. 4 is used as a synapse and a component of a neuron between the input layer and the next layer. In this example, the input INPUT 0, ...,INPUT n On the vertical control gate lines CG0,…,CG N The signals are received on the source lines SL0 and SL1, and outputs OUTPUT1 and OUTPUT2 are generated on the source lines SL0 and SL1.

[0173] Figure 21A neuron VMM array 2100 is shown, which is particularly suitable for Figure 4 The memory unit 410 shown in FIG. 4 is used as a synapse and a component of a neuron between an input layer and a next layer. In this example, inputs INPUT0 to INPUT N are received at the gates of bit line control gates 2901-1, 2901-2 to 2901-(N-1), and 2901-N, which are coupled to bit lines BL0 to BL1, respectively. N Exemplary outputs OUTPUT1 and OUTPUT2 are generated on source lines SL0 and SL1.

[0174] Figure 22 A neuron VMM array 2200 is shown, which is particularly suitable for Figure 3 The memory unit 310 shown, Figure 5 The memory cell 510 and Figure 7 The memory unit 710 shown is used as a synapse and a component of the neurons between the input layer and the next layer. In this example, the inputs INPUT0, ..., INPUT M On word lines WL0,…,WL M is received and output OUTPUT0,…,OUTPUT N On bit lines BL0, ..., BL N Generate on.

[0175] Figure 23 A neuron VMM array 2300 is shown, which is particularly suitable for Figure 3 The memory unit 310 shown, Figure 5 The memory cell 510 and Figure 7 The memory unit 710 shown in FIG. 1 is used as a synapse and a component of a neuron between an input layer and the next layer. In this example, inputs INPUT0 to INPUT M On the control gate lines CG0 to CG M The output is OUTPUT0, ..., OUTPUT N On the vertical source lines SL0, ..., SL N On the generation, each source line SL i The source line coupled to all memory cells in column i.

[0176] Figure 24 A neuron VMM array 2400 is shown, which is particularly suitable for Figure 3 The memory unit 310 shown, Figure 5 The memory cell 510 and Figure 7The memory unit 710 shown in FIG. 1 is used as a synapse and a component of a neuron between an input layer and the next layer. In this example, inputs INPUT0 to INPUT M On the control gate lines CG0 to CG M The output is OUTPUT0, ..., OUTPUT N On the vertical bit lines BL0, ..., BL N On the generated, where each bit line BL i The bit line coupled to all memory cells in column i.

[0177] Long Short-Term Memory

[0178] Existing technology includes a concept known as long short-term memory (LSTM). LSTM is commonly used in artificial neural networks. LSTM allows artificial neural networks to remember information for a predetermined, arbitrary time interval and use that information in subsequent operations. A typical LSTM consists of a cell, an input gate, an output gate, and a forget gate. These three gates regulate the flow of information into and out of the cell and the time interval over which information is remembered within the LSTM. VMMs are particularly useful within LSTMs.

[0179] Figure 25 An exemplary LSTM 2500 is shown. LSTM 2500 in this example includes units 2501, 2502, 2503, and 2504. Unit 2501 receives an input vector x0 and generates an output vector h0 and a unit state vector c0. Unit 2502 receives an input vector x1, an output vector (hidden state) h0 from unit 2501, and a unit state c0 from unit 2501, and generates an output vector h1 and a unit state vector c1. Unit 2503 receives an input vector x2, an output vector (hidden state) h1 from unit 2502, and a unit state c1 from unit 2502, and generates an output vector h2 and a unit state vector c2. Unit 2504 receives an input vector x3, an output vector (hidden state) h2 from unit 2503, and a unit state c2 from unit 2503, and generates an output vector h3. Additional units may be used, and an LSTM having four units is merely an example.

[0180] Figure 26 Shown available for Figure 25 FIG2 is an exemplary implementation of an LSTM unit 2600 for units 2501, 2502, 2503, and 2504 in FIG2. LSTM unit 2600 receives an input vector x(t), a cell state vector c(t-1) from the previous unit, and an output vector h(t-1) from the previous unit, and generates a cell state vector c(t) and an output vector h(t).

[0181] LSTM unit 2600 includes sigmoid function devices 2601, 2602, and 2603, each of which applies a number between 0 and 1 to control the amount of each component in the input vector that is allowed to pass through to the output vector. LSTM unit 2600 also includes tanh devices 2604 and 2605 for applying a hyperbolic tangent function to the input vector, multiplier devices 2606, 2607, and 2608 for multiplying two vectors together, and an addition device 2609 for adding two vectors together. The output vector h(t) can be provided to the next LSTM unit in the system, or it can be accessed for other purposes.

[0182] Figure 27 LSTM unit 2700 is shown, which is an example of a specific implementation of LSTM unit 2600. For the convenience of the reader, LSTM unit 2700 uses the same numbering as LSTM unit 2600. Sigmoid function devices 2601, 2602, and 2603, and tanh device 2604 each include multiple VMM arrays 2701 and activation circuit blocks 2702. Therefore, it can be seen that VMM arrays are particularly useful in LSTM units used in certain neural network systems.

[0183] An alternative form of LSTM cell 2700 (and another example of a specific implementation of LSTM cell 2600) is Figure 28 In Figure 28 In the example, sigmoid function devices 2601, 2602, and 2603 and tanh device 2604 share the same physical hardware (VMM array 2801 and activation function block 2802) in a time-division multiplexing manner. The LSTM unit 2800 also includes a multiplier device 2803 that multiplies two vectors together, an addition device 2808 that adds two vectors together, a tanh device 2605 (which includes an activation circuit block 2802), a register 2807 that stores the value i(t) when the value i(t) is output from the sigmoid function block 2802, a register 2804 that stores the value f(t)*c(t-1) when it is output from the multiplier device 2803 through the multiplexer 2810, a register 2805 that stores the value i(t)*u(t) when it is output from the multiplier device 2803 through the multiplexer 2810, a register 2806 that stores the value o(t)*c~(t) when it is output from the multiplier device 2803 through the multiplexer 2810, and a multiplexer 2809.

[0184] The LSTM unit 2700 includes multiple sets of VMM arrays 2701 and corresponding activation function blocks 2702, while the LSTM unit 2800 includes only one set of VMM arrays 2801 and activation function blocks 2802, which are used to represent multiple layers in the implementation of the LSTM unit 2800. The LSTM unit 2800 will require less space than the LSTM 2700 because the LSTM unit 2800 only requires 1 / 4 of its space for VMM and activation function blocks compared to the LSTM unit 2700.

[0185] It will also be appreciated that an LSTM cell will typically include multiple VMM arrays, each of which requires functionality provided by certain circuit blocks outside the VMM array (such as summer and activation circuit blocks, as well as high-voltage generation blocks). Providing a separate circuit block for each VMM array would require a significant amount of space within the semiconductor device and would be somewhat inefficient. Therefore, the embodiments described below attempt to minimize the circuitry required outside the VMM array itself.

[0186] Gate-controlled recursive unit

[0187] The specific implementation of the simulated VMM can be used for GRU (Gated Recurrent Unit). GRU is a gate-controlled mechanism in recurrent artificial neural networks. GRU is similar to LSTM, except that GRU cells generally contain fewer components than LSTM cells.

[0188] Figure 29 An exemplary GRU 2900 is shown. The GRU 2900 in this example includes units 2901, 2902, 2903, and 2904. Unit 2901 receives an input vector x0 and generates an output vector h0. Unit 2902 receives an input vector x1, an output vector h0 from unit 2901, and generates an output vector h1. Unit 2903 receives an input vector x2 and an output vector (hidden state) h1 from unit 2902 and generates an output vector h2. Unit 2904 receives an input vector x3 and an output vector (hidden state) h2 from unit 2903 and generates an output vector h3. Additional units may be used, and a GRU with four units is merely an example.

[0189] Figure 30 Shown available for Figure 292901, 2902, 2903, and 2904 of a GRU unit 3000. The GRU unit 3000 receives an input vector x(t) and an output vector h(t-1) from a previous GRU unit and generates an output vector h(t). The GRU unit 3000 includes sigmoid function devices 3001 and 3002, each of which applies a number between 0 and 1 to components from the output vector h(t-1) and the input vector x(t). The GRU unit 3000 also includes a tanh device 3003 for applying a hyperbolic tangent function to the input vector, a plurality of multiplier devices 3004, 3005, and 3006 for multiplying two vectors together, an addition device 3007 for adding two vectors together, and a complement device 3008 for subtracting the input from 1 to generate an output.

[0190] Figure 31 3100 is shown, which is an example of a specific implementation of the GRU unit 3000. For the convenience of the reader, the same numbering is used in the GRU unit 3100 as in the GRU unit 3000. Figure 31 As shown, the sigmoid function devices 3001 and 3002 and the tanh device 3003 each include a plurality of VMM arrays 3101 and an activation function block 3102. Therefore, it can be seen that the VMM array is particularly useful in the GRU unit used in some neural network systems.

[0191] An alternative form of GRU unit 3100 (and another example of a specific implementation of GRU unit 3000) is Figure 32 In Figure 32 In , the GRU unit 3200 utilizes a VMM array 3201 and an activation function block 3202 which, when configured as a sigmoid function, applies a number between 0 and 1 to control how much of each component in the input vector is allowed to pass through to the output vector. Figure 32In the example, sigmoid function devices 3001 and 3002 and tanh device 3003 share the same physical hardware (VMM array 3201 and activation function block 3202) in a time-division multiplexing manner. The GRU unit 3200 also includes a multiplier device 3203 that multiplies two vectors together, an addition device 3205 that adds two vectors together, a complement device 3209 that subtracts the input from 1 to generate an output, a multiplexer 3204, a register 3206 that holds the value h(t-1)*r(t) when it is output from the multiplier device 3203 through the multiplexer 3204, a register 3207 that holds the value h(t-1)*z(t) when it is output from the multiplier device 3203 through the multiplexer 3204, and a register 3208 that holds the value h^(t)*(1-z(t)) when it is output from the multiplier device 3203 through the multiplexer 3204.

[0192] The GRU unit 3100 includes multiple sets of VMM arrays 3101 and activation function blocks 3102, while the GRU unit 3200 includes only one set of VMM arrays 3201 and activation function blocks 3202, which are used to represent multiple layers in the implementation of the GRU unit 3200. The GRU unit 3200 will require less space than the GRU unit 3100 because the GRU unit 3200 only requires 1 / 3 of its space for VMMs and activation function blocks compared to the GRU unit 3100.

[0193] It will also be appreciated that a system utilizing a GRU will typically include multiple VMM arrays, each of which requires functionality provided by certain circuit blocks outside the VMM array (such as summer and activation circuit blocks, as well as high-voltage generation blocks). Providing a separate circuit block for each VMM array would require a significant amount of space within the semiconductor device and would be somewhat inefficient. Therefore, the embodiments described below attempt to minimize the circuitry required outside the VMM array itself.

[0194] The inputs to the VMM array can be analog levels, binary levels, timing pulses, or digital bits, and the outputs can be analog levels, binary levels, timing pulses, or digital bits (in this case, an output ADC is required to convert the output analog level current or voltage into digital bits).

[0195] For each memory cell in the VMM array, each weight w can be implemented by a single memory cell, a differential cell, or two hybrid memory cells (the average of two or more cells). In the case of a differential cell, two memory cells are required to implement the weight w as a differential weight (w = w + - w -). In the case of two hybrid memory cells, two memory cells are required to implement the weight w as the average of the two cells.

[0196] Implementation for fine tuning units in a VMM

[0197] Figure 33 A block diagram of a VMM system 3300 is shown. VMM system 3300 includes a VMM array 3301, a row decoder 3302, a high voltage decoder 3303, a column decoder 3304, a bit line driver 3305, an input circuit 3306, an output circuit 3307, a control logic unit 3308, and a bias generator 3309. VMM system 3300 also includes a high voltage generation block 3310, which includes a charge pump 3311, a charge pump regulator 3312, and a high voltage level generator 3313. VMM system 3300 also includes an algorithm controller 3314, an analog circuit 3315, a control logic unit 3316, and a test control logic unit 3317. The systems and methods described below can be implemented in VMM system 3300.

[0198] Input circuit 3306 may include circuits such as a DAC (digital-to-analog converter), a DPC (digital-to-pulse converter), an AAC (analog-to-analog converter, such as a current-to-voltage converter), a PAC (pulse-to-analog level converter), or any other type of converter. Input circuit 3306 may implement normalization, a scaling function, or an arithmetic function. Input circuit 3306 may implement a temperature compensation function for the input. Input circuit 3306 may implement an activation function, such as a ReLU or sigmoid function.

[0199] Output circuitry 3307 may include circuitry such as an ADC (analog-to-digital converter for converting neuron analog outputs into digital bits), an AAC (analog-to-analog converter, such as a current-to-voltage converter), an APC (analog-to-pulse converter), or any other type of converter. Output circuitry 3307 may implement an activation function, such as a ReLU or sigmoid function. Output circuitry 3307 may implement normalization, scaling, or arithmetic functions for neuron outputs. Output circuitry 3307 may implement a temperature compensation function for neuron outputs or array outputs (such as bitline outputs), as described below.

[0200] Figure 34A tuning correction method 3400 is shown, which can be executed by an algorithm controller 3314 in a VMM system 3300. The tuning correction method 3400 generates an adaptive target based on a final error generated from a cell output and a cell initial target. The method generally starts (step 3401) in response to receiving a tuning command. An initial current target (for a programming / verification algorithm) for a selected cell or a selected group of cells Itargetv(i) is determined using a predictive target model (such as by using a function or a look-up table), and a variable DeltaError is set to 0 (step 3402). The target function (if used) will be based on the I-V programming curve of the selected memory cell or group of cells. The target function also depends on various variations caused by array characteristics, such as the degree of programming interference exhibited by the cells (which depends on the cell address and cell hierarchy within a sector, where if a cell exhibits relatively large interference, the cell undergoes more programming time under a suppression condition, and cells with higher currents generally have more interference), coupling between cells, and various types of array noise. These variations for silicon can be characterized in terms of PVT (process, voltage, temperature). The look-up table (if used) can be characterized in the same way to simulate the I-V curve and various variations.

[0201] Then, a soft erase is performed on all cells in the VMM, which erases all cells to an intermediate weak erase level such that each cell will consume a current of, for example, about 3 μA - 5 μA during a read operation (step 3403). For example, the soft erase is performed by applying an incremental erase pulse voltage to the cells until an intermediate cell current is reached. Next, a deep programming operation is performed on all unused cells (step 3404) to reach a <pA current level. Then, a target adjustment (correction) based on an error result is performed. If DeltaError > 0, meaning the cell has experienced an overshoot during programming, then Itargetv(i+1) is set to Itarget + θ * DeltaError, where θ is, for example, a number 1 or close to 1 (step 3405A).

[0202] Itarget(i+1) can also be adjusted based on a proper error target adjustment / correction with respect to the previous Itarget(i). If DeltaError ≤ 0, meaning the cell has experienced an undershoot during programming, which means the cell current has not reached the target, then Itargetv(i+1) is set to the previous target Itargetv(i) (step 3405B).

[0203] Next, perform a coarse and / or fine programming and verification operation (step 3406). Multiple adaptive coarse programming methods can be used to accelerate programming, such as by targeting multiple progressively smaller coarse targets before performing the exact (fine) programming step. Adaptive exact programming is done, for example, with fine (exact) incremental programming voltage pulses or constant programming timing pulses. Embodiments of systems and methods for performing coarse programming and fine programming are described in U.S. Provisional Patent Application No. 62 / 933,809, filed Nov. 11, 2019, and titled "Precise Programming Method and Apparatus for Analog Neural Memory in a Deep Learning Artificial Neural Network," by the same assignee as this application, which is incorporated herein by reference.

[0204] Measure Icell in the selected cell (step 3407). For example, the cell current can be measured by a galvanometer circuit. For example, the cell current can be measured by an ADC (analog-to-digital converter) circuit, where in this case the output is represented by digital bits. For example, the cell current can be measured by an I-V (current-voltage converter) circuit, where in this case the output is represented by an analog voltage. Calculate DeltaError, which is Icell - Itarget, and which represents the difference between the actual current (Icell) and the target current (Itarget) in the measured cell. If |DeltaError| < DeltaMargin, the cell has reached the target current within a certain tolerance (DeltaMargin), and the method ends (step 3410). |DeltaError| = abs(DeltaError) = the absolute value of DeltaError. If not, the method returns to step 3403 and the steps are executed in sequence again (step 3410).

[0205] Figure 35A and Figure 35B illustrates a tuning correction method 3500, which can be executed by the algorithm controller 3314 in the VMM system 3300. Refer to Figure 35A, the start of the method (step 3501) is typically performed in response to receiving a tuning command. Erase the entire VMM array (step 3502) such as by a soft erase method. Perform a deep programming operation on all unused cells (step 3503) to achieve a cell current <pA level. Program all cells in the VMM array to an intermediate value, such as 0.5 μA - 1.0 μA, using coarse and / or fine programming loops (step 3504). Embodiments of systems and methods for performing coarse programming and fine programming are described in U.S. Provisional Patent Application No. 62 / 933,809, filed Nov. 11, 2019, and titled "Precise Programming Method and Apparatus for Analog Neural Memory in a Deep Learning Artificial Neural Network", by the same assignee as this application, which is incorporated herein by reference. Set the prediction target for the used cells using the function or look-up table as described above (step 3505). Then, perform the sector tuning method 3507 on each sector in the VMM (step 3506). A sector typically consists of two or more adjacent rows in the array.

[0206] Figure 35BShows the adaptive target sector tuning method 3507. All cells in the sector are programmed to the final expected value (e.g., 1 nA - 50 nA) using individual or combined programming / verification (P / V) methods such as: (1) coarse / fine / constant P / V cycling; (2) CG+ (only CG increment) or EG+ (only EG increment) or complementary CG+ / EG- (CG increment and EG decrement); and (3) first programming the deepest programmed cells (such as progressive grouping, meaning dividing the cells into different groups, and the group of cells with the lowest current is programmed first) (step 3508A). Next, determine whether Icell < Itarget. If so, the method proceeds to step 3509. If not, the method repeats step 3508A. In step 3509, measure DeltaError, which is equal to the measured Icell - Itarget(i + 1) (step 3509). Determine whether |DeltaError| < DeltaMargin (step 3510). If so, the method is complete (step 3511). If not, perform target adjustment. If DeltaError > 0, meaning the cell has experienced overshoot during programming, adjust the target by setting the new target to Itarget + θ * DeltaError, where θ is typically = 1 (step 3512A). Itarget(i + 1) can also be adjusted based on the appropriate error target adjustment / correction of the previous Itarget(i). If DeltaError < 0, meaning the cell has experienced undershoot during programming, which means the cell has not reached the target, adjust the target by maintaining the previous target, i.e., Itargetv(i + 1) = Itargetv(i) (step 3512B). Perform a soft erase on the sector (step 3513). Program all cells in the sector to an intermediate value (step 3514), and return to step 3509.

[0207] A typical neural network can have positive weights w+ and negative weights w- and a combined weight = w+ - w-. w+ and w- are implemented by memory cells (Iw+ and Iw- respectively), and the combined weight (Iw = Iw+ - Iw-, current subtraction) can be performed at the peripheral circuit level (such as at the array bit line output circuit). Thus, the weight tuning implementation for the combined weight can include, for example, tuning the w+ cells and w- cells simultaneously, tuning only the w+ cells, or tuning only the w- cells, as shown in Table 8. Using the previous reference Figure 34 / Figure 35A / Figure 35BThe described programming / verification and error target adjustment methods are used to perform tuning. Verification can be performed on only the combined weights (e.g., measuring / reading the combined weight currents instead of the individual positive w+ cell currents or w- cell currents), only on the w+ cell currents, or only on the w- cell currents.

[0208] For example, for a combination Iw of 3na, Iw+ can be 3na and Iw- can be 0na; or, Iw+ can be 13na and Iw- can be 10na, meaning that both the positive weight Iw+ and the negative weight Iw- are non-zero (e.g., where zero would represent a deeply programmed cell). Under certain operating conditions, this may be preferable because it will make both Iw+ and Iw- less susceptible to noise.

[0209] Table 8: Weight tuning methods

[0210]

[0211]

[0212] Figure 36A shows the behavior of the data (IV curve) as a function of temperature (e.g., in the subthreshold region), Figure 36B illustrates the problems caused by data drift during operation of a VMM system, and Figure 36C and Figure 36D Shows blocks for compensating for data drift and about Figure 36C , showing the block for compensating for temperature changes.

[0213] Figure 36A Illustrating a known characteristic of VMM systems as operating temperature increases, the sense current in any given selected nonvolatile memory cell in the VMM array increases in the subthreshold region, decreases in the saturation region, or generally decreases in the linear region.

[0214] Figure 36B The array current distribution (data drift) over time is shown, and it shows that the total output from the VMM array (which is the sum of the current from all bit lines in the VMM array) shifts to the right (or left, depending on the technology used) with operating time usage, which means that the total total output will drift over the lifetime of the VMM system. This phenomenon is called data drift because data drifts due to usage conditions and degrades due to environmental factors.

[0215] Figure 36C The bit line compensation circuit 3600 is shown. The bit line compensation circuit may include a compensation current i COMPThe output of the bit line output circuit 3610 is injected to compensate for data drift. The bit line compensation circuit 3600 may include a scaler circuit that amplifies or reduces the output based on a resistor or capacitor network. The bit line compensation circuit 3600 may include a shifter circuit that shifts or offsets the output based on its resistor or capacitor network.

[0216] Figure 36D A data drift monitor 3620 is shown that detects the amount of data drift. This information is then used as an input to the bit line compensation circuit 3600 so that an appropriate level of i can be selected. COMP .

[0217] Figure 37 36. The bit line compensation circuit 3700 is shown as an embodiment of the bit line compensation circuit 3600 in FIG. The bit line compensation circuit 3700 includes an adjustable current source 3701 and an adjustable current source 3702, which together generate i COMP , where i COMP Equal to the current generated by adjustable current source 3701 minus the current generated by adjustable current source 3702.

[0218] Figure 38 36. The bit line compensation circuit 3800 includes an operational amplifier 3801, an adjustable resistor 3802, and an adjustable resistor 3803. The operational amplifier 3801 receives a reference voltage VREF at its non-inverting terminal and receives V INPUT , where V INPUT It is from Figure 36C The bit line output circuit 3610 receives the voltage and generates the output V OUTPUT , where V OUTPUT It is V INPUT A scaled version of V can be created to compensate for data drift based on the ratio of resistors 3803 and 3802. By configuring the values of resistors 3803 and / or 3802, V OUTPUT .

[0219] Figure 39 A bit line compensation circuit 3900 is shown, which is an embodiment of the bit line compensation circuit 3600 in FIG. 36 . The bit line compensation circuit 3900 includes an operational amplifier 3901, a current source 3902, a switch 3904, and an adjustable integrating output capacitor 3903. Here, the current source 3902 is actually the output current on a single bit line or a collection of multiple bit lines (such as one for summing positive weights w+ and one for summing negative weights w-) in the VMM array. The operational amplifier 3901 receives a reference voltage VREF at its non-inverting terminal and V INPUT , where VINPUT It is from Figure 36C The bit line compensation circuit 3900 acts as an integrator that integrates the current Ineu through the capacitor 3903 within an adjustable integration time to generate an output voltage V OUTPUT , where V OUTPUT =Ineu*integral time / C 3903 , where C 3903 is the value of capacitor 3903. Therefore, the output voltage V OUTPUT is proportional to the (bit line) output current Ineu, proportional to the integration time, and inversely proportional to the capacitance of capacitor 3903. The bit line compensation circuit 3900 generates an output V OUTPUT , where V OUTPUT The value of is scaled based on the configured value of capacitor 3903 and / or the integration time to compensate for data drift.

[0220] Figure 40 A bit line compensation circuit 4000 is shown, which is one embodiment of the bit line compensation circuit 3600 in FIG 36. The bit line compensation circuit 4000 includes a current mirror 4010 having an M:N ratio, which means that I COMP =(M / N)*i input The current mirror 4010 receives the current i INPUT , and mirroring and optionally scaling this current to generate i COMP Therefore, by configuring the M parameter and / or N parameter, you can enlarge or reduce i COMP .

[0221] Figure 41 A bit line compensation circuit 4100 is shown, which is one implementation of the bit line compensation circuit 3600 in FIG 36. The bit line compensation circuit 4100 includes an operational amplifier 4101, an adjustable scaling resistor 4102, an adjustable shift resistor 4103, and an adjustable resistor 4104. The operational amplifier 4101 receives a reference voltage V at its non-inverting terminal. REF , and receives V at its inverting terminal IN . V IN In response to V INPUT and Vshft, where V INPUT It is from Figure 36C The bit line output circuit 3610 receives a voltage, and Vshft is intended to achieve V INPUT With V OUTPUT The voltage shifted between.

[0222] Therefore, V OUTPUT It is V INPUT A scaled and shifted version of to compensate for data drift.

[0223] Figure 42 36 . The bit line compensation circuit 4200 is shown as an embodiment of the bit line compensation circuit 3600 in FIG. The bit line compensation circuit 4200 includes an operational amplifier 4201, an input current source Ineu 4202, a current shifter 4203, switches 4205 and 4206, and an adjustable integrating output capacitor 4204. Here, the current source 4202 is actually the output current Ineu on a single bit line or multiple bit lines in the VMM array. The operational amplifier 4201 receives a reference voltage VREF at its non-inverting terminal and receives Ineu at its inverting terminal. IN , where I IN is the sum of Ineu and the current output of current shifter 4203, and generates the output V OUTPUT , where V OUTPUT Scaled (based on capacitor 4204) and shifted (based on Ishifter 4203) to compensate for data drift.

[0224] Figures 43 to 48 Various circuits are shown that may be used to provide the W value to be programmed or read into each selected cell during a program or read operation.

[0225] Figure 43 Neuron output circuit 4300 is shown, which includes adjustable current source 4301 and adjustable current source 4302, which together generate I OUT , where I OUT Equal to the current I generated by the adjustable current source 4301 W+ Subtract the current I generated by the adjustable current source 4302 W- Adjustable current Iw+ 4301 is a cell current or neuron current (such as a bit line current) used to scale current for positive weights. Adjustable current Iw- 4302 is a cell current or neuron current (such as a bit line current) used to scale current for negative weights. Current scaling is accomplished, for example, by an M:N ratio current mirror circuit, where Iout = (M / N)*Iin.

[0226] Figure 44 Neuron output circuit 4400 is shown, which includes an adjustable capacitor 4401, a control transistor 4405, a switch 4402, a switch 4403, and an adjustable current source 4404Iw+, which is a scaled output current of a cell current or a (bit line) neuron current such as an M:N current mirror circuit. Transistor 4405 is used, for example, to apply a fixed bias voltage to current 4404. Circuit 4404 generates V OUT , where V OUT Inversely proportional to capacitor 4401, proportional to the adjustable integration time (the time between switch 4403 closing and switch 4402 opening), and proportional to the current supplied by adjustable current source 4404IW+ The generated current is proportional to V OUT Equal to V+-((Iw+*integral time) / C 4401 ), where C 4401 is the value of capacitor 4401. The positive terminal V+ of capacitor 4401 is connected to the positive supply voltage, and the negative terminal V- of capacitor 4401 is connected to the output voltage V OUT .

[0227] Figure 45 Neuron circuit 4500 is shown, which includes capacitor 4401 and adjustable current source 4502, which is a scaled current of the cell current or (bit line) neuron current such as an M:N current mirror. Circuit 4500 generates V OUT , where V OUT Inversely proportional to capacitor 4401, proportional to the adjustable integration time (the time when switch 4501 is off), and proportional to the current supplied by adjustable current source 4502I Wi The capacitor 4401 is reused from the neuron output circuit 44 after completing its operation of integrating the current Iw+. Then, the positive terminal and the negative terminal (V+ and V-) are swapped in the neuron output circuit 45, where the positive terminal is connected to the output voltage V OUT , the output voltage is de-integrated by the current Iw-. The negative terminal is held at the previous voltage value by a clamping circuit (not shown). In practice, output circuit 44 is used for a positive weight implementation, and circuit 45 is used for a negative weight implementation, where the final charge on capacitor 4401 effectively represents the combined weight (Qw = Qw+ - Qw-).

[0228] Figure 46 A neuron circuit 4600 is shown, which includes an adjustable capacitor 4601, a switch 4602, a control transistor 4604, and an adjustable current source 4603. The circuit 4600 generates V OUT , where V OUT Inversely proportional to capacitor 4601, proportional to the adjustable integration time (the time when switch 4602 is off), and proportional to the current supplied by adjustable current source 4603I W- The generated current is proportional to the voltage. The negative terminal V- of capacitor 4601 is, for example, equal to ground. The positive terminal V+ of capacitor 4601 is, for example, initially precharged to a positive voltage before integrating the current Iw-. Neuron circuit 4600 can be used to replace neuron circuit 4500 and neuron circuit 4400 to implement the combined weight (Qw=Qw+-Qw-).

[0229] Figure 47Neuron circuit 4700 is shown, which includes operational amplifiers 4703 and 4706; adjustable current sources Iw+ 4701 and Iw- 4702; and adjustable resistors 4704, 4705, and 4707. Neuron circuit 4700 generates V OUT , this voltage is equal to R 4707 *(Iw+-Iw-). Adjustable resistor 4707 implements output scaling. Adjustable current sources Iw+ 4701 and Iw- 4702 also implement output scaling, for example, through an M:N ratio current mirror circuit (Iout = (M / N) * Iin).

[0230] Figure 48 Neuron circuit 4800 is shown, which includes operational amplifiers 4803 and 4806; switches 4808 and 4809; adjustable current sources Iw-4802 and Iw+4801; and adjustable capacitors 4804, 4805, and 4807. Neuron circuit 4800 generates V OUT , which is proportional to (Iw+ - Iw-), proportional to the integration time (the time switches 4808 and 4809 are off), and inversely proportional to the capacitance of capacitor 4807. Adjustable capacitor 4807 enables scaling of the output. Adjustable current sources Iw+ 4801 and Iw- 4802 also enable scaling of the output, for example, through an M:N ratio current mirror circuit (Iout = (M / N) * Iin). The integration time can also adjust the output scaling.

[0231] Figure 49A 、 Figure 49B and Figure 49C Showing output circuits such as Figure 33 Block diagram of the output circuit 3307 in FIG.

[0232] exist Figure 49A , output circuit 4901 includes ADC circuit 4911, which is used to directly digitize analog neuron output 4910 to provide digital output bits 4912.

[0233] exist Figure 49B , output circuit 4902 includes neuron output circuit 4921 and ADC 4911. Neuron output circuit 4921 receives neuron output 4920 and shapes it, which is then digitized by ADC circuit 4911 to generate output 4912. Neuron output circuit 4921 can be used for normalization, scaling, shifting, mapping, arithmetic operations, activation, and / or temperature compensation, such as previously described. The ADC circuit can be a serial (ramp-type or step-up or counting) ADC, a SAR ADC, a pipeline ADC, a sigma-delta ADC, or any other type of ADC.

[0234] exist Figure 49C, the output circuit includes a neuron output circuit 4921 and a converter circuit 4931 that receives the neuron output 4930 and is used to convert the output from the neuron output circuit 4921 into an output 4932. The converter 4931 may include an ADC, an AAC (similar to an analog-to-analog converter, such as a current-to-voltage converter), an APC (analog-to-pulse converter), or any other type of converter. The ADC 4911 or the converter 4931 may be used to implement an activation function by, for example, bit mapping (e.g., quantization) or clipping (e.g., clipped ReLU). The ADC 4911 and the converter 4931 may be configurable, such as for lower or higher precision (e.g., lower or higher number of bits), lower or higher performance (e.g., slower or faster speed), etc.

[0235] Another embodiment for scaling and shifting is by configuring an ADC (analog-to-digital) conversion circuit (such as a serial ADC, a SAR ADC, a pipeline ADC, a ramp ADC, etc.) to convert the array (bit line) output into digital bits with lower or higher bit precision, and then manipulating the digital output bits according to a certain function (e.g., linear or nonlinear, compression, nonlinear activation, etc.), such as by normalization (e.g., 12 bits to 8 bits), shifting, or remapping. An embodiment of the ADC conversion circuit is described in U.S. Provisional Patent Application No. 62 / 933,809, filed on November 11, 2019 by the same assignee as the present application and entitled “Precise Programming Method and Apparatus for Analog Neural Memory in a Deep Learning Artificial Neural Network,” which is incorporated herein by reference.

[0236] Table 9 shows an alternative method of performing read, erase, and program operations:

[0237] Table 9: Flash memory cell operations

[0238]

[0239]

[0240] Read and erase operations are similar to the previous table. However, the two methods used for programming are implemented by the Fowler-Nordheim (FN) tunneling mechanism.

[0241] An implementation for scaling the inputs may be accomplished, for example, by enabling a certain number of rows of the VMM at a time and then fully combining the results.

[0242] Another embodiment is to scale the input voltage and rescale the output appropriately to achieve normalization.

[0243] Another embodiment for scaling pulse width modulated inputs is by modulating the timing of the pulse width. An embodiment of this technique is described in U.S. patent application Ser. No. 16 / 449,201, filed on Jun. 21, 2019, by the same assignee as the present application, and entitled “Configurable Input Blocks and Output Blocks and Physical Layout for Analog Neural Memory in Deep Learning Artificial Neural Network,” which is incorporated herein by reference.

[0244] Another embodiment for scaling the input is by enabling one input bit at a time, for example, for an 8-bit input IN7:0, evaluating IN0, IN1, ..., IN7 in sequence, and then combining the output results with the appropriate bit weighting. An embodiment of this technique is described in U.S. patent application Ser. No. 16 / 449,201, filed on June 21, 2019, by the same assignee as the present application, and entitled “Configurable Input Blocks and Output Blocks and Physical Layout for Analog Neural Memory in Deep Learning Artificial Neural Network,” which is incorporated herein by reference.

[0245] Optionally, in the above embodiments, the cell current measured for the purpose of verifying or reading the current can be averaged or measured multiple times, for example 8 to 32 measurements, to reduce the effects of noise (such as RTN or any random noise) and / or detect any outlier bits that are defective and need to be replaced by redundant bits.

[0246] It should be noted that, as used herein, the terms "above" and "on" both inclusively include "directly on" (no intervening material, element, or space disposed therebetween) and "indirectly on" (intervening material, element, or space disposed therebetween). Similarly, the term "adjacent" includes "directly adjacent" (no intervening material, element, or space disposed therebetween) and "indirectly adjacent" (intervening material, element, or space disposed therebetween), "mounted to" includes "directly mounted to" (no intervening material, element, or space disposed therebetween) and "indirectly mounted to" (intervening material, element, or space disposed therebetween), and "electrically coupled to" includes "directly electrically coupled to" (no intervening material or element electrically connecting the elements together) and "indirectly electrically coupled to" (intervening material or element electrically connecting the elements together). For example, forming an element "above a substrate" may include forming the element directly on the substrate without an intervening material / element therebetween, as well as forming the element indirectly on the substrate with one or more intervening materials / elements therebetween.

Claims

1. A method of tuning a selected non-volatile memory cell in a vector-matrix multiplication array of non-volatile memory cells, the method comprising: (i) setting an initial target current of the selected nonvolatile memory cell and setting an output error to zero; (ii) performing a soft erase on all non-volatile memory cells in the vector-matrix multiplication array; (iii) performing a deep programming operation on all non-volatile memory cells except the selected non-volatile memory cell; (iv) if the output error is positive, meaning that the selected nonvolatile memory cell has experienced an overshoot during programming, revising the initial target current to a new target current based on the output error; If the output error is equal to zero or negative, setting the new target current to the previously set target current; (v) performing one or more of a coarse programming operation and a fine programming operation on the selected nonvolatile memory cells; (vi) performing a read operation on the selected nonvolatile memory cell and determining a cell output current consumed by the selected nonvolatile memory cell during the read operation; (vii) calculating the output error as a difference between the determined cell output current and the new target current; as well as (viii) Repeat steps (ii), (iii), (iv), (v), (vi), and (vii) until the output error is less than a predetermined threshold.

2. The method according to claim 1, further comprising: Perform a coarse program-verify cycle.

3. The method according to claim 1, further comprising: Perform a refined program-verify cycle. The method of claim 1 , wherein the selected non-volatile memory cells store positive weights. The method of claim 1 , wherein the selected nonvolatile memory cell stores a negative weight.

6. The method of claim 1, wherein the selected nonvolatile memory cell stores a positive weight and another selected nonvolatile memory cell stores a negative weight. The method of claim 6 , wherein the positive weights are non-zero and the negative weights are non-zero.

8. The method of claim 1, wherein one or more cells in the vector-matrix multiplication array are programmed using Fowler-Nordheim tunneling.

9. The method of claim 1, wherein one or more cells in the vector-matrix multiplication array are programmed using source side injection.

Citation Information

Patent Citations

  • Deep learning neural network classifier using non-volatile memory array

    US11308383B2

  • Deep Learning Neural Network Classifier Using Non-volatile Memory Array

    US20170337466A1

  • Configurable input blocks and output blocks and physical layout for analog neural memory in deep learning artificial neural network

    US20200349421A1

  • Single transistor non-valatile electrically alterable semiconductor memory device

    US5029130A

  • Flash memory cells with separated self-aligned select and erase gates, and process of fabrication

    US6747310B2