Precise data tuning method and apparatus for simulated neural memory in artificial neural network

By adopting tuning algorithms and neuron output circuits in artificial neural networks, the problem of accurate programming of non-volatile memory cells is solved, the computing efficiency and energy efficiency are improved, and it is suitable for high-performance information processing.

CN120808841APending Publication Date: 2025-10-17SILICON STORAGE TECHNOLOGY INC
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511008308.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2020-03-25
Filing Date
2020-07-02
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing technologies have difficulty in achieving precise programming of non-volatile memory cells in artificial neural networks, especially depositing a specific and precise amount of charge on the floating gate, resulting in low computing efficiency and high energy consumption.

Method used

A tuning algorithm is used, including initial current target setting, soft erasing, coarse programming, fine programming, reading operation and error calculation. This process is repeated until the output error is less than a threshold. The neuron output circuit is combined to provide current to achieve accurate programming of the weight value.

Benefits of technology

It achieves precise programming of non-volatile memory cells, improves computing efficiency and energy efficiency, and is suitable for high-performance information processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120808841A_ABST
    Figure CN120808841A_ABST
Patent Text Reader

Abstract

Various embodiments of a precise programming algorithm and apparatus for precisely and rapidly depositing the correct amount of charge on floating gates of non-volatile memory cells within a vector-matrix multiplication (VMM) array in an artificial neural network are disclosed. Accordingly, the selected cell can be programmed very precisely to maintain one of the N different values.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This patent application is a continuation-in-part of International Application No. PCT / US2020 / 040755, International Filing Date, July 2, 2020, entered into the National Stage in the United States Patent and Trademark Office as Application No. 202080091622.3, entitled “Precise Data Tuning Method And Apparatus For Analog Neuromorphic Memory In An Artificial Neural Network,” filed on July 2, 2020.

[0002] CLAIM OF PRIORITY

[0003] This application claims priority to U.S. Provisional Patent Application No. 62 / 957,013, filed January 3, 2020, entitled “Precise Data Tuning Method And Apparatus For Analog Neuromorphic Memory In An Artificial Neural Network,” and U.S. Patent Application No. 16 / 829,757, filed March 25, 2020, entitled “Precise Data Tuning Method And Apparatus For Analog Neural Memory In An Artificial Neural Network.” TECHNICAL FIELD

[0004] Disclosed are multiple embodiments of a precise tuning method and apparatus for depositing the correct amount of charge precisely and quickly on floating gates of non-volatile memory cells within a vector-matrix multiplication (VMM) array in an artificial neural network. BACKGROUND

[0005] Artificial neural networks mimic biological neural networks (the central nervous system of animals, particularly the brain), and are used to estimate or approximate functions that can depend on a large number of inputs and are typically unknown. Artificial neural networks typically include layers of interconnected “neurons” that exchange messages with each other.

[0006] Figure 1 An artificial neural network is shown, where the circles represent the inputs or layers of neurons. The connections, called synapses, are shown with arrows, and have numeric weights that can be tuned based on experience. This makes the artificial neural network adaptable to inputs and enables it to learn. Typically, an artificial neural network includes a layer of multiple inputs. There are usually one or more intermediate layers of neurons, and an output layer of neurons that provide the output of the neural network. The neurons at each level make decisions individually or collectively based on the data received from the synapses.

[0007] One of the major challenges in developing artificial neural networks for high performance information processing is the lack of adequate hardware technology. Indeed, practical artificial neural networks rely on a large number of synapses, thereby enabling high connectivity between neurons, i.e., very high computational parallelism. In principle, such complexity can be achieved by digital supercomputers or dedicated clusters of graphics processing units. However, compared to biological networks, these approaches are not only high cost, but also generally energy inefficient, biological networks consuming much less energy primarily due to their performing low precision analog computations. CMOS analog circuits have been used for artificial neural networks, but given the large number of neurons and synapses, most CMOS implemented synapses are too bulky.

[0008] Applicant previously disclosed in U.S. Patent Application No. 15 / 594,439 (published as U.S. Patent Publication 2017 / 0337466), which is incorporated by reference herein, an artificial (analog) neural network that utilizes one or more non-volatile memory arrays as synapses. The non-volatile memory arrays operate as analog neuromorphic memory. The term “neuromorphic” as used herein refers to a circuit that implements a model of a nervous system. The analog neuromorphic memory includes a first plurality of synapses configured to receive a first plurality of inputs and generate a first plurality of outputs therefrom, and a first plurality of neurons configured to receive the first plurality of outputs. The first plurality of synapses includes a plurality of memory cells, where each of the memory cells includes: spaced apart source and drain regions formed in a semiconductor substrate, with a channel region extending between the source and drain regions; a floating gate disposed over and insulated from a first portion of the channel region; and a non-floating gate disposed over and insulated from a second portion of the channel region. Each of the plurality of memory cells is configured to store a weight value corresponding to a number of electrons on the floating gate. The plurality of memory cells is configured to multiply the first plurality of inputs by the stored weight values to generate the first plurality of outputs. An array of memory cells arranged in this manner can be referred to as a vector matrix multiplication (VMM) array.

[0009] Each non-volatile memory cell used in the VMM must be erased and programmed to hold a very specific and precise amount of charge (i.e., number of electrons) in the floating gate. For example, each floating gate must hold one of N different values, where N is the number of different weights that can be indicated by each cell. Examples of N include 16, 32, 64, 128, and 256. One challenge is to be able to program the selected cells with the precision and granularity required for the different N values. For example, if the selected cells can include one of 64 different values, then a very high precision is required in the programming operation.

[0010] Improved programming systems and methods are needed that are suitable for use with a VMM array in an analog neuromorphic memory. Summary of the Invention

[0011] The present invention discloses various embodiments of a precision-tuned algorithm and apparatus for accurately and rapidly depositing the correct amount of charge on the floating gates of nonvolatile memory cells within a VMM array in an analog neuromorphic memory system. Thus, a selected cell can be programmed with extreme precision to hold one of N different values.

[0012] In one embodiment, a method for tuning a selected nonvolatile memory cell in a vector-matrix multiplication array of nonvolatile memory cells is provided, the method comprising: (i) setting an initial current target for the selected nonvolatile memory cell; (ii) performing a soft erase on all nonvolatile memory cells in the vector-matrix multiplication array; (iii) performing a coarse programming operation on the selected memory cell; (iv) performing a fine programming operation on the selected memory cell; (v) performing a read operation on the selected memory cell and determining a current consumed by the selected memory cell during the read operation; (vi) calculating an output error based on a difference between the determined current and the initial current target; and repeating steps (i), (ii), (iii), (iv), (v), and (vi) until the output error is less than a predetermined threshold.

[0013] In another embodiment, a method for tuning a selected nonvolatile memory cell in a vector-matrix multiplication array of nonvolatile memory cells is provided, the method comprising: (i) setting an initial target for the selected nonvolatile memory cell; (ii) performing a programming operation on the selected memory cell; (iii) performing a read operation on the selected memory cell and determining a cell output consumed by the selected memory cell during the read operation; (iv) calculating an output error based on a difference between the determined output and the initial target; and (v) repeating steps (i), (ii), (iii), and (iv) until the output error is less than a predetermined threshold.

[0014] In another embodiment, a neuron output circuit for providing current to program weight values ​​in selected memory cells in a vector-matrix multiplication array is provided, the neuron output circuit comprising: a first adjustable current source that generates a scaled current in response to the neuron current to achieve a positive weight; and a second adjustable current source that generates a scaled current in response to the neuron current to achieve a negative weight.

[0015] In another embodiment, a neuron output circuit for providing a current to program a weight value in a selected memory cell in a vector-matrix multiplication array is provided, the neuron output circuit comprising: a tunable capacitor comprising a first terminal and a second terminal, the second terminal providing an output voltage for the neuron output circuit; a control transistor comprising a first terminal and a second terminal; a first switch selectively coupled between the first terminal and the second terminal of the tunable capacitor; a second switch selectively coupled between the second terminal of the tunable capacitor and the first terminal of the control transistor; and a tunable current source coupled to the second terminal of the control transistor.

[0016] In another embodiment, a neuron output circuit for providing a current to program a weight value in a selected memory cell in a vector-matrix multiplication array is provided, the neuron output circuit comprising: a tunable capacitor comprising a first terminal and a second terminal, the second terminal providing an output voltage for the neuron output circuit; a control transistor comprising a first terminal and a second terminal; a switch selectively coupled between the second terminal of the tunable capacitor and the first terminal of the control transistor; and a tunable current source coupled to the second terminal of the control transistor.

[0017] In another embodiment, a neuron output circuit for providing a current to program a weight value in a selected memory cell in a vector-matrix multiplication array is provided, the neuron output circuit comprising: a tunable capacitor comprising a first terminal and a second terminal, the first terminal providing an output voltage for the neuron output circuit; a control transistor comprising a first terminal and a second terminal; a first switch selectively coupled between the first terminal and the first terminal of the control transistor; and a tunable current source coupled to the second terminal of the control transistor.

[0018] In another embodiment, a neuron output circuit for providing current to program a weight value in a selected memory cell in a vector-matrix multiplication array is provided, the neuron output circuit comprising: a first operational amplifier comprising an inverting input, a non-inverting input, and an output; a second operational amplifier comprising an inverting input, a non-inverting input, and an output; a first adjustable current source coupled to the inverting input of the first operational amplifier; a second adjustable current source coupled to the inverting input of the second operational amplifier; a first adjustable resistor coupled to the inverting input of the first operational amplifier; a second adjustable resistor coupled to the inverting input of the second operational amplifier; and a third adjustable resistor coupled between the output of the first operational amplifier and the inverting input of the second operational amplifier.

[0019] In another embodiment, a neuron output circuit for providing current to program a weight value in a selected memory cell in a vector-matrix multiplication array is provided, the neuron output circuit comprising: a first operational amplifier comprising an inverting input, a non-inverting input, and an output; a second operational amplifier comprising an inverting input, a non-inverting input, and an output; a first adjustable current source coupled to the inverting input of the first operational amplifier; a second adjustable current source coupled to the inverting input of the second operational amplifier; a first switch coupled between the inverting input and the output of the first operational amplifier; a second switch coupled between the inverting input and the output of the second operational amplifier; a first adjustable capacitor coupled between the inverting input and the output of the first operational amplifier; a second adjustable capacitor coupled between the inverting input and the output of the second operational amplifier; and a third adjustable capacitor coupled between the output of the first operational amplifier and the inverting input of the second operational amplifier. BRIEF DESCRIPTION OF DRAWINGS

[0020] Figure 1 To illustrate a schematic diagram of a prior art artificial neural network.

[0021] Figure 2 To illustrate a prior art split gate flash memory cell.

[0022] Figure 3 To illustrate another prior art split gate flash memory cell.

[0023] Figure 4 Another prior art split gate flash memory cell is shown.

[0024] Figure 5 Another prior art split gate flash memory cell is shown.

[0025] Figure 6 Another prior art split gate flash memory cell is shown.

[0026] Figure 7 A prior art stack gate flash memory cell is shown.

[0027] Figure 8 A schematic diagram to illustrate different levels of an exemplary artificial neural network using one or more VMM arrays.

[0028] Figure 9 A block diagram to illustrate a VMM system including a VMM array and other circuitry.

[0029] Figure 10 A block diagram to illustrate an exemplary artificial neural network using one or more VMM systems.

[0030] Figure 11 Another embodiment of a VMM array is shown.

[0031] Figure 12 Another embodiment of a VMM array is shown.

[0032] Figure 13 Another embodiment of a VMM array is shown.

[0033] Figure 14 Another embodiment of a VMM array is shown.

[0034] Figure 15 Another embodiment of a VMM array is shown.

[0035] Figure 16 Another embodiment of a VMM array is shown.

[0036] Figure 17 Another embodiment of a VMM array is shown.

[0037] Figure 18 Another embodiment of a VMM array is shown.

[0038] Figure 19 Another embodiment of a VMM array is shown.

[0039] Figure 20 Another embodiment of a VMM array is shown.

[0040] Figure 21 Another embodiment of a VMM array is shown.

[0041] Figure 22 Another embodiment of a VMM array is shown.

[0042] Figure 23 Another embodiment of a VMM array is shown.

[0043] Figure 24 Another embodiment of a VMM array is shown.

[0044] Figure 25 A prior art long short term memory system is shown.

[0045] Figure 26 An exemplary cell used in a long short term memory system is shown.

[0046] Figure 27 An embodiment of an exemplary cell of Figure 26 is shown.

[0047] Figure 28 Another embodiment of an exemplary cell of Figure 26 is shown.

[0048] Figure 29 A prior art gated recurrent cell system is shown.

[0049] Figure 30 An exemplary cell used in a gated recurrent cell system is shown.

[0050] Figure 31 An embodiment of an exemplary cell of Figure 30 is shown.

[0051] Figure 32 Another embodiment of an exemplary cell of Figure 30 is shown.

[0052] Figure 33 A VMM system is shown.

[0053] Figure 34 A tuning correction method is shown.

[0054] Figure 35A A tuning correction method is shown.

[0055] Figure 35B A sector tuning correction method is shown.

[0056] Figure 36A The effect of temperature on values stored in a cell is shown.

[0057] Figure 36BProblems caused by data drift during operation of a VMM system are shown.

[0058] Figure 36C A block for compensating for data drift is shown.

[0059] Figure 36D A data drift monitor is shown.

[0060] Figure 37 A bit line compensation circuit is shown.

[0061] Figure 38 Another bit line compensation circuit is shown.

[0062] Figure 39 Another bit line compensation circuit is shown.

[0063] Figure 40 Another bit line compensation circuit is shown.

[0064] Figure 41 Another bit line compensation circuit is shown.

[0065] Figure 42 Another bit line compensation circuit is shown.

[0066] Figure 43 A neuron circuit is shown.

[0067] Figure 44 Another neuron circuit is shown.

[0068] Figure 45 Another neuron circuit is shown.

[0069] Figure 46 Another neuron circuit is shown.

[0070] Figure 47 Another neuron circuit is shown.

[0071] Figure 48 Another neuron circuit is shown.

[0072] Figure 49A A block diagram of an output circuit is shown.

[0073] Figure 49B A block diagram of another output circuit is shown.

[0074] Figure 49C A block diagram of another output circuit is shown. DETAILED DESCRIPTION

[0075] The artificial neural network of the present invention utilizes a combination of CMOS technology and non-volatile memory arrays.

[0076] Non-volatile memory cell

[0077] Digital non-volatile memory is well known. For example, U.S. Patent 5,029,130 ("the '130 patent"), which is incorporated herein by reference, discloses an array of split-gate non-volatile memory cells, which is a type of flash memory cell. Such a memory cell 210 is shown in FIG. 1. Each memory cell 210 includes a source region 14 and a drain region 16 formed in a semiconductor substrate 12 with a channel region 18 therebetween. A floating gate 20 is formed over and insulated from (and controls the conductivity of) a first portion of the channel region 18, and over a portion of the source region 14. A word line terminal 22 (which is typically coupled to a word line) has a first portion disposed over and insulated from (and controls the conductivity of) a second portion of the channel region 18, and a second portion that extends upward and is located over the floating gate 20. The floating gate 20 and the word line terminal 22 are insulated from the substrate 12 by a gate oxide. A bit line terminal 24 is coupled to the drain region 16. Figure 2

[0078] Erase of the memory cell 210 (where electrons are removed from the floating gate) is performed by placing a high positive voltage on the word line terminal 22, which causes the electrons on the floating gate 20 to tunnel through the intervening insulator from the floating gate 20 to the word line terminal 22 via Fowler-Nordheim tunneling.

[0079] Programming of the memory cell 210 (where electrons are placed on the floating gate) is performed by placing a positive voltage on the word line terminal 22 and a positive voltage on the source region 14. Electron current will flow from the source region 14 (the source line terminal) to the drain region 16. As the electrons reach the gap between the word line terminal 22 and the floating gate 20, the electrons will accelerate and become hot. Some of the hot electrons will be injected through the gate oxide onto the floating gate 20 due to the electrostatic attraction from the floating gate 20.

[0080] Reading of the memory cell 210 is performed by placing a positive read voltage on the drain region 16 and the word line terminal 22, which turns on the portion of the channel region 18 under the word line terminal. If the floating gate 20 is positively charged (i.e., erased of electrons), then the portion of the channel region 18 under the floating gate 20 is also turned on, and current will flow through the channel region 18, which is sensed as an erased state or "1" state. If the floating gate 20 is negatively charged (i.e., programmed with electrons), then the portion of the channel region under the floating gate 20 is mostly or completely turned off, and current will not (or very little current) flow through the channel region 18, which is sensed as a programmed state or "0" state.

[0081] Table 1 shows typical voltage ranges that can be applied to the terminals of the memory cell 110 for performing read, erase, and program operations:

[0082] Table 1: Figure 2 Operation of the flash memory cell 210 of Figure 2 ​

[0083] WL BL SL Read 1 0.5-3V 0.1-2V 0V Read 2 0.5-3V 0-2V 2-0.1V Erase Approx. 11-13V 0V 0V Program 1V-2V 1-3pA 9-10V

[0084] “Read 1” is a read mode in which the cell current is output on the bit line. “Read 2” is a read mode in which the cell current is output on the source line terminal.

[0085] Figure 3 A memory cell 310 is shown, which is similar to Figure 2 memory cell 210, but with the addition of a control gate (CG) terminal 28. The control gate terminal 28 is biased at a high voltage during programming (e.g., 10V), at a low or negative voltage during erase (e.g., 0v / -8V), and at a low or intermediate voltage during read (e.g., 0v / 2.5V). The other terminals are biased similarly to Figure 2 .

[0086] Figure 4 A four-gate memory cell 410 is shown, which includes a source region 14, a drain region 16, a floating gate 20 over a first portion of a channel region 18, a select gate 22 (typically coupled to a word line WL) over a second portion of the channel region 18, a control gate 28 over the floating gate 20, and an erase gate 30 over the source region 14. This configuration is described in U.S. Patent 6,747,310, which is incorporated herein by reference for all purposes. Here, all of the gates are non-floating, meaning that they are electrically connected to or capable of being electrically connected to a voltage source, except for the floating gate 20. Programming is performed by heated electrons from the channel region 18 that inject themselves into the floating gate 20. Erase is performed by electrons tunneling from the floating gate 20 to the erase gate 30.

[0087] Table 2 shows typical voltage ranges that can be applied to the terminals of the memory cell 410 for performing read, erase, and program operations:

[0088] Table 2: Figure 4 Operation of the flash memory cell 410 of FIG. 4

[0089] WL / SG BL CG EG SL Read 1 0.5-2V 0.1-2V 0-2.6V 0-2.6V 0V Read 2 0.5-2V 0-2V 0-2.6V 0-2.6V 2-0.1V Erase -0.5V / 0V 0V 0V / -8V 8-12V 0V Program 1V 1pA 8-11V 4.5-9V 4.5-5V

[0090] “Read 1” is a read mode in which the cell current is output on the bit line. “Read 2” is a read mode in which the cell current is output on the source line terminal.

[0091] Figure 5 A memory cell 510 is shown, which is similar to Figure 4memory cell 410. Erase is performed by biasing the substrate 18 to a high voltage and biasing the control gate CG terminal 28 to a low voltage or negative voltage. Alternatively, erase is performed by biasing the word line terminal 22 to a positive voltage and biasing the control gate terminal 28 to a negative voltage. Programming and reading are similar to Figure 4

[0092] Figure 6 A three gate memory cell 610 is shown, which is another type of flash memory cell. The memory cell 610 is the same as the memory cell 410 of Figure 4 , except that the memory cell 610 does not have a separate control gate terminal. Except for the lack of a control gate bias, the erase operation (erase is performed using the erase gate terminal) and the read operation are similar to Figure 4 the operations of. In the absence of a control gate bias, the program operation is also accomplished, and as a result, a higher voltage must be applied on the source line terminal during the program operation to compensate for the lack of control gate bias.

[0093] Table 3 shows typical voltage ranges that can be applied to the terminals of the memory cell 610 for performing read, erase, and program operations:

[0094] Table 3: Figure 6 Operation of the flash memory cell 610 of FIG. 1

[0095]

[0096]

[0097] "Read 1" is a read mode in which the cell current is output on the bit line. "Read 2" is a read mode in which the cell current is output on the source line terminal.

[0098] Figure 7 A stacked gate memory cell 710 is shown, which is another type of flash memory cell. The memory cell 710 is similar to the memory cell 210 of Figure 2 , except that the floating gate 20 extends over the entire channel region 18, and the control gate terminal 22 (which here will be coupled to a word line) extends over the floating gate 20, separated by an insulating layer (not shown). The erase, program, and read operations operate in a similar manner as previously described for the memory cell 210.

[0099] Table 4 shows typical voltage ranges that can be applied to the terminals of the memory cell 710 and the substrate 12 for performing read, erase, and program operations:

[0100] Table 4: Figure 7 Operation of the flash memory cell 710

[0101] CG BL SL Substrate Read 1 0-5V 0.1-2V 0-2V 0V Read 2 0.5-2V 0-2V 2-0.1V 0V Erase -8 to -10V / 0V FLT FLT 8-10V / 15-20V Program 8-12V 3-5V / 0V 0V / 3-5V 0V

[0102] “Read 1” is a read mode in which the cell current is output on the bit line. “Read 2” is a read mode in which the cell current is output on the source line terminal. Optionally, in an array including rows and columns of memory cells 210, 310, 410, 510, 610, or 710, the source line can be coupled to a row of memory cells or to two adjacent rows of memory cells. That is, the source line terminal can be shared by memory cells of adjacent rows.

[0103] To utilize a memory array including one of the above types of non-volatile memory cells in an artificial neural network, two modifications are made. First, the circuitry is configured so that each memory cell can be individually programmed, erased, and read without adversely affecting the memory state of other memory cells in the array, as explained further below. Second, continuous (analog) programming of the memory cells is provided.

[0104] In particular, the memory state of each memory cell in the array (i.e., the charge on the floating gate) can be continuously changed from a fully erased state to a fully programmed state independently and with minimal interference to other memory cells. In another embodiment, the memory state of each memory cell in the array (i.e., the charge on the floating gate) can be continuously changed from a fully programmed state to a fully erased state independently and with minimal interference to other memory cells, or vice versa. This means that the cell storage is analog, or at least can store one of many discrete values (such as 16 or 64 different values), which allows very precise and individual tuning of all cells in the memory array, and which makes the memory array ideal for storing and fine-tuning the synaptic weights of a neural network.

[0105] The methods and apparatus described herein can be applied to other non-volatile memory technologies, such as but not limited to SONOS (silicon-oxide-nitride-oxide-silicon, charge trapping in nitride), MONOS (metal-oxide-nitride-oxide-silicon, metal charge trapping in nitride), ReRAM (resistive ram), PCM (phase change memory), MRAM (magnetic ram), FeRAM (ferroelectric ram), OTP (two or more layers of one-time programmable), and CeRAM (correlated electron ram), among others. The methods and apparatus described herein can be applied to volatile memory technologies for neural networks, such as but not limited to SRAM, DRAM, and / or volatile synapse cells.

[0106] Neural network employing an array of non-volatile memory cells

[0107] Figure 8A non-limiting example of a neural network using a non-volatile memory array is conceptually illustrated for this embodiment. This example uses a non-volatile memory array neural network for a facial recognition application, but any other suitable application can be implemented using a non-volatile memory array based neural network.

[0108] For this example, S0 is an input layer that is a 32x32 pixel RGB image with 5-bit precision (i.e., three 32x32 pixel arrays, one for each color R, G, and B, each pixel being 5-bit precision). The synapses CB1 from the input layer S0 to layer C1 apply different sets of weights in some cases, shared weights in other cases, and scan the input image with a 3x3 pixel overlap filter (kernel), shifting the filter by 1 pixel (or more than 1 pixel as indicated by the model). Specifically, the values of the 9 pixels in a 3x3 portion of the image (i.e., referred to as a filter or kernel) are provided to the synapses CB1, where these 9 input values are multiplied by appropriate weights, and after summing the outputs of this multiplication, a single output value is determined and provided by the first synapse of CB1 for generating a pixel of one of the layers C1 of feature maps. The 3x3 filter is then shifted one pixel to the right within the input layer S0 (i.e., adding a column of three pixels to the right and freeing a column of three pixels to the left), whereby the 9 pixel values in this newly positioned filter are provided to the synapses CB1, where they are multiplied by the same weights and a second single output value is determined by the associated synapses. This process continues until the 3x3 filter scans all three colors and all bits (precision values) across the entire 32x32 pixel image of the input layer S0. The process is then repeated using a different set of weights to generate a different feature map of C1 until all of the feature maps of layer C1 are computed.

[0109] At layer C1, in this example, there are 16 feature maps, each having 30x30 pixels. Each pixel is a new feature pixel extracted from the multiplication of the input and kernel, so each feature map is a two-dimensional array, and thus layer C1 is composed of a two-dimensional array of 16 layers in this example (remember that the layers and arrays referenced herein are logical relationships and not necessarily physical relationships, i.e., the arrays do not necessarily orient to a physical two-dimensional array). Each of the 16 feature maps in layer C1 is generated from one of sixteen different sets of synapse weights applied to the filter scan. The C1 feature maps can all relate to different aspects of the same image feature, such as boundary identification. For example, a first map (generated using a first set of weights, shared for all scans used to generate this first map) can identify circular edges, a second map (generated using a second set of weights different from the first set of weights) can identify rectangular edges, or aspect ratios of certain features, and so on.

[0110] Before transitioning from layer CI to layer SI, an activation function PI (pooling) is applied that pools values from consecutive non-overlapping 2x2 regions in each feature map. The purpose of the pooling function is to average (or max function can also be used) over neighboring locations, for example, to reduce the dependence on edge locations and to reduce the data size before entering the next stage. At layer SI, there are 16 feature maps of 15x15 (i.e., sixteen different arrays of 15x15 pixels per feature map). Synapses CB2 from layer SI to layer C2 scan the maps in SI with 4x4 filters, where the filters are shifted by 1 pixel. At layer C2, there are 22 feature maps of 12x12. Before transitioning from layer C2 to layer S2, an activation function P2 (pooling) is applied that pools values from consecutive non-overlapping 2x2 regions in each feature map. At layer S2, there are 22 feature maps of 6x6. An activation function (pooling) is applied to synapses CB3 from layer S2 to layer C3, where each neuron in layer C3 is connected to each map in layer S2 via a respective synapse of CB3. At layer C3, there are 64 neurons. Synapses CB4 from layer C3 to output layer S3 connect C3 to S3 completely, i.e., each neuron in layer C3 is connected to each neuron in layer S3. The output at S3 includes 10 neurons, where the highest output neuron determines the class. For example, the output can indicate a recognition or classification of the content of the original image.

[0111] An array or portion of an array of non-volatile memory cells is used to implement the synapses of each layer.

[0112] Figure 9 A block diagram of a system that can be used for this purpose. The VMM system 32 includes non-volatile memory cells and is used as a synapse between a layer and the next layer (such as CB1, CB2, CB3, and CB4 in Figure 6 Specifically, the VMM system 32 includes a VMM array 33 (including non-volatile memory cells arranged in rows and columns), erase gate and word line gate decoders 34, a control gate decoder 35, a bit line decoder 36, and a source line decoder 37, which decode respective inputs to the VMM array 33. The inputs to the VMM array 33 can come from the erase gate and word line gate decoders 34 or from the control gate decoder 35. In this example, the source line decoder 37 also decodes outputs from the VMM array 33. Alternatively, the bit line decoder 36 can decode outputs from the VMM array 33.

[0113] The VMM array 33 serves two purposes. First, it stores the weights to be used by the VMM system 32. Second, the VMM array 33 effectively multiplies the inputs by the weights stored in the VMM array 33 and each output line (source line or bit line) adds them to produce an output that will be the input to the next layer or the final layer. By performing the multiplication and addition functions, the VMM array 33 eliminates the need for separate multiplication and addition logic circuits and is also highly power efficient due to its in-place memory computation.

[0114] The output of the VMM array 33 is provided to a difference summer (such as a summing operational amplifier or summing current mirror) 38 that sums the outputs of the VMM array 33 to create a single value for the convolution. The difference summer 38 is arranged to perform the summing of both positive and negative weight inputs to output a single value.

[0115] The output value of the difference summer 38 is then summed and provided to an activation function circuit 39 that modifies the output. The activation function circuit 39 can provide a sigmoid, tanh, ReLU function or any other non-linear function. The modified output value of the activation function circuit 39 becomes an element of the feature map that is the next layer (e.g., layer Cl) in Figure 8 The VMM array 33 constitutes a plurality of synapses (which receive their inputs from an existing neuron layer or from an input layer such as an image database) and the summer 38 and activation function circuit 39 constitute a plurality of neurons in this example. Thus, the VMM system 32 is a multi-layer neural network.

[0116] Figure 9 The inputs (WLx, EGx, CGx and optionally BLx and SLx) to the VMM system 32 in FIG. 1 can be analog levels, binary levels, digital pulses (in which case a pulse-to-analog converter PAC can be needed to convert the pulses to the appropriate input analog levels) or digital bits (in which case a DAC is provided to convert the digital bits to the appropriate input analog levels); the outputs can be analog levels, binary levels, digital pulses or digital bits (in which case an output ADC is provided to convert the output analog levels to digital bits).

[0117] Figure 10 A block diagram to show the use of a multi-layer VMM system 32 (here labeled VMM systems 32a, 32b, 32c, 32d and 32e). As Figure 10As shown, input (denoted as Inputx) is converted from digital to analog by digital-to-analog converter 31 and provided to input VMM system 32a. The converted analog input can be voltage or current. The input D / A conversion of the first layer can be done by using a function or LUT (look-up table) that maps the input Inputx to the appropriate analog level of the matrix multiplier to the input VMM system 32a. The input conversion can also be done by an analog-to-analog (A / A) converter to convert an external analog input to a mapped analog input to the input VMM system 32a. The input conversion can also be done by a digital-to-digital pulse (D / P) converter to convert an external digital input to one or more mapped digital pulses to the input VMM system 32a.

[0118] The output generated by the input VMM system 32a is provided as input to the next VMM system (hidden layer 1) 32b, which in turn generates an output that is provided as input to the next VMM system (hidden layer 2) 32c, and so on. The layers of VMM systems 32 serve as different layers of synapses and neurons of a convolutional neural network (CNN). Each VMM system 32a, 32b, 32c, 32d, and 32e can be a separate physical system that includes a respective non-volatile memory array, or multiple VMM systems can utilize different portions of the same physical non-volatile memory array, or multiple VMM systems can utilize overlapping portions of the same physical non-volatile memory array. Each VMM system 32a, 32b, 32c, 32d, and 32e can also be time-division multiplexed for different portions of its array or neurons. Figure 10 The example shown includes five layers (32a, 32b, 32c, 32d, 32e): one input layer (32a), two hidden layers (32b, 32c), and two fully connected layers (32d, 32e). Those of ordinary skill in the art will appreciate that this is merely exemplary, and that, instead, a system can include more than two hidden layers and more than two fully connected layers.

[0119] VMM array

[0120] Figure 11 A neuron VMM array 1100 is shown, which is particularly suitable for use in a neural network system 1000 as shown in FIG. 10. Figure 3 The memory cells 310 shown, and serve as synapses and components of neurons between an input layer and a next layer. The VMM array 1100 includes a memory array 1101 of non-volatile memory cells and a reference array 1102 of non-volatile reference memory cells (at the top of the array). Alternatively, another reference array can be placed at the bottom.

[0121] In the VMM array 1100, the control gate lines, such as control gate line 1103, extend in the vertical direction (hence the reference array 1102 is orthogonal to the control gate lines 1103 in the row direction), and the erase gate lines, such as erase gate line 1104, extend in the horizontal direction. Here, the inputs to the VMM array 1100 are set on the control gate lines (CG0, CG1, CG2, CG3), and the outputs of the VMM array 1100 appear on the source lines (SL0, SL1). In one embodiment, only the even rows are used, and in another embodiment, only the odd rows are used. The current placed on each source line (SL0, SL1, respectively) performs a summation function of all the currents from the memory cells connected to that particular source line.

[0122] As described herein for neural networks, the non-volatile memory cells of the VMM array 1100 (i.e., the flash memory of the VMM array 1100) are preferably configured to operate in the sub-threshold region.

[0123] Biasing the non-volatile reference memory cells and non-volatile memory cells described herein in weak inversion:

[0124] Ids = Io * e (Vg-Vth) / nVt = w * Io * e (Vg) / nVt ,

[0125] where w = e (-Vth) / nVt

[0126] where Ids is the drain to source current; Vg is the gate voltage on the memory cell; Vth is the threshold voltage of the memory cell; Vt is the thermal voltage = k*T / q, where k is the Boltzmann constant, T is the temperature in Kelvin, and q is the electron charge; n is the slope factor = 1 + (Cdep / Cox), where Cdep = the capacitance of the depletion layer, and Cox is the capacitance of the gate oxide layer; Io is the memory cell current at the gate voltage equal to the threshold voltage, Io is proportional to (Wt / L)*u*Cox*(n-1)*Vt 2 where u is the carrier mobility, and Wt and L are the width and length of the memory cell, respectively.

[0127] For an I to V logarithmic converter that uses a memory cell (such as a reference memory cell or a peripheral memory cell) or a transistor to convert an input current Ids to an input voltage Vg:

[0128] Vg = n * Vt * log[Ids / wp*Io]

[0129] Here, wp is the w of the reference memory cell or peripheral memory cell.

[0130] For an I-to-V logarithmic converter using memory cells (such as reference memory cells or peripheral memory cells) or transistors to convert input current Ids to input voltage Vg:

[0131] Vg = n * Vt * log [Ids / wp * Io]

[0132] Here, wp is the w of the reference memory cells or peripheral memory cells.

[0133] For a memory array used as a vector matrix multiplier VMM array, the output current is:

[0134] Iout = wa * Io * e (Vg) / nVt i.e.

[0135] Iout = (wa / wp) * Iin = W * Iin

[0136] W = e (Vthp-Vtha) / nVt

[0137] Iin = wp * Io * e (Vg) / nVt

[0138] Here, wa = the w of each memory cell in the memory array.

[0139] A word line or control gate can be used as an input to the memory cells of the input voltage.

[0140] Alternatively, the non-volatile memory cells of the VMM array described herein can be configured to operate in a linear region:

[0141] Ids = β * (Vgs - Vth) * Vds; β = u * Cox * Wt / L,

[0142] Wα (Vgs - Vth),

[0143] meaning that the weight W in the linear region is proportional to (Vgs - Vth)

[0144] A word line or control gate or bit line or source line can be used as an input to the memory cells operating in the linear region. A bit line or source line can be used as an output from the memory cells.

[0145] For an I-to-V linear converter, memory cells (such as reference memory cells or peripheral memory cells) or transistors or resistors operating in a linear region can be used to linearly convert input / output currents to input / output voltages.

[0146] Alternatively, the memory cells of the VMM array described herein can be configured to operate in a saturation region:

[0147] Ids = 1 / 2 * β * (Vgs - Vth)2 ;β=u*Cox*Wt / L

[0148] Wα(Vgs-Vth) 2 , which means the weight W and (Vgs-Vth) 2 Proportional

[0149] The word line, control gate, or erase gate can be used as the input of a memory cell operating in the saturation region. The bit line or source line can be used as the output of an output neuron.

[0150] Alternatively, the memory cells of the VMM arrays described herein may be used in all regions or a combination thereof (subthreshold, linear, or saturation regions).

[0151] U.S. Patent Application No. 15 / 826,345 describes Figure 9 Other embodiments of the VMM array 33 of EMBODIMENTS 1 are incorporated herein by reference. As described herein, source lines or bit lines can be used as neuron outputs (current summing outputs).

[0152] Figure 12 A neuron VMM array 1200 is shown, which is particularly suitable for Figure 2 Memory cell 210 is shown and serves as a synapse between the input layer and the next layer. VMM array 1200 includes a memory array 1203 of nonvolatile memory cells, a reference array 1201 of first nonvolatile reference memory cells, and a reference array 1202 of second nonvolatile reference memory cells. Reference arrays 1201 and 1202, arranged along the columns of the array, are used to convert current inputs flowing into terminals BLR0, BLR1, BLR2, and BLR3 into voltage inputs WL0, WL1, WL2, and WL3. In practice, the first and second nonvolatile reference memory cells are diode-connected via a multiplexer 1214 (only partially shown), with the current input flowing therein. The reference cells are tuned (e.g., programmed) to a target reference level. The target reference level is provided by a reference microarray matrix (not shown).

[0153] Memory array 1203 serves two purposes. First, it stores the weights that VMM array 1200 will use on their respective memory cells. Second, memory array 1203 effectively multiplies the inputs (i.e., current inputs provided in terminals BLR0, BLR1, BLR2, and BLR3, which array 1201 and 1202 convert into input voltages to provide to word lines WL0, WL1, WL2, and WL3) by the weights stored in memory array 1203, and then adds all the results (memory cell currents) to produce an output on the respective bit lines (BL0-BLN), which will be the input to the next layer or the input to the final layer. By performing the multiplication and addition functions, memory array 1203 eliminates the need for separate multiplication logic circuits and addition logic circuits, and is also highly power efficient. Here, voltage inputs are provided on the word lines (WL0, WL1, WL2, and WL3), and outputs appear on the respective bit lines (BL0-BLN) during a read (inference) operation. The current placed on each bit line BL0-BLN performs a summation function of the currents from all non-volatile memory cells connected to that particular bit line.

[0154] Table 5 shows the operating voltages for VMM array 1200. The columns in the table indicate the voltages placed on the word line for selected cells, the word line for unselected cells, the bit line for selected cells, the bit line for unselected cells, the source line for selected cells, and the source line for unselected cells, where FLT indicates floating, i.e., no voltage is applied. The rows indicate read, erase, and program operations.

[0155] Table 5: Figure 12 Operation of the VMM array 1200

[0156] WL WL - unselected BL BL - unselected SL SL - unselected Read 0.5-3.5V -0.5V / 0V 0.1-2V (Ineuron) 0.6V-2V / FLT 0V 0V Erase Approx. 5-13V 0V 0V 0V 0V 0V Program 1V-2V -0.5V / 0V 0.1-3uA Vinh approx. 2.5V 4-10V 0-1V / FLT

[0157] Figure 13 A neuron VMM array 1300 is shown, which is particularly suitable for use in a neural network, such as a convolutional neural network. Figure 2The illustrated memory cells 210 and serve as synapses and components of neurons between the input layer and the next layer. The VMM array 1300 includes a memory array 1303 of non-volatile memory cells, a reference array 1301 of first non-volatile reference memory cells, and a reference array 1302 of second non-volatile reference memory cells. The reference arrays 1301 and 1302 extend in the row direction of the VMM array 1300. The VMM array is similar to the VMM 1000, except that in the VMM array 1300, the word lines extend in the vertical direction. Here, the inputs are placed on the word lines (WLA0, WLB0, WLA1, WLB2, WLA2, WLB2, WLA3, WLB3), and the outputs appear on the source lines (SL0, SL1) during a read operation. The current placed on each source line performs a summing function of all the currents from the memory cells connected to that particular source line.

[0158] Table 6 shows the operating voltages for the VMM array 1300. The columns in the table indicate the voltages placed on the word lines for selected cells, the word lines for unselected cells, the bit lines for selected cells, the bit lines for unselected cells, the source lines for selected cells, and the source lines for unselected cells. The rows indicate read, erase, and program operations.

[0159] Table 6: Figure 13 Operation of the VMM array 1300

[0160]

[0161] Figure 14 A neuron VMM array 1400 is shown that is particularly suitable for use as a synapse between the input layer and the next layer. Figure 3 The illustrated memory cells 310 and serve as synapses and components of neurons between the input layer and the next layer. The VMM array 1400 includes a memory array 1403 of non-volatile memory cells, a reference array 1401 of first non-volatile reference memory cells, and a reference array 1402 of second non-volatile reference memory cells. The reference arrays 1401 and 1402 are used to convert current inputs into voltage inputs CG0, CG1, CG2, and CG3 in the terminals BLR0, BLR1, BLR2, and BLR3 that flow into them. In effect, the first and second non-volatile reference memory cells are diode-connected through multiplexers 1412 (shown in part only), in which the current inputs flow into them through BLR0, BLR1, BLR2, and BLR3. The multiplexers 1412 each include a respective multiplexer 1405 and a common-source common-gate transistor 1404 to ensure a constant voltage on the bit line (such as BLR0) of each of the first and second non-volatile reference memory cells during a read operation. The reference cells are tuned to a target reference level.

[0162] Memory array 1403 serves two purposes. First, it stores the weights to be used by VMM array 1400. Second, memory array 1403 effectively multiplies the inputs (current inputs provided to terminals BLR0, BLR1, BLR2, and BLR3, which are converted to input voltages by reference arrays 1401 and 1402 to be provided to control gates CG0, CG1, CG2, and CG3) by the weights stored in the memory array, and then adds all the results (cell currents) to produce an output, which appears on BL0-BLN and will be the input to the next layer or the final layer. By performing the multiplication and addition functions, the memory array eliminates the need for separate multiplication and addition logic circuits and is also highly power efficient. Here, the inputs are provided on the control gate lines (CG0, CG1, CG2, and CG3), and the outputs appear on the bit lines (BL0-BLN) during a read operation. The current placed on each bit line performs a summing function of all the currents from the memory cells connected to that particular bit line.

[0163] VMM array 1400 implements one-way tuning for the non-volatile memory cells in memory array 1403. That is, each non-volatile memory cell is erased and then programmed partially until the desired charge on the floating gate is reached. This can be performed, for example, using the precise programming technique described below. If too much charge is placed on the floating gate (so that an incorrect value is stored in the cell), the cell must be erased and the sequence of partial programming operations must start over. As shown, two rows that share the same erase gate (such as EG0 or EG1) need to be erased together (which is referred to as a page erase), and thereafter each cell is programmed partially until the desired charge on the floating gate is reached.

[0164] Table 7 shows the operating voltages for VMM array 1400. The columns in the table indicate the voltages on the word line for the selected cell, the word line for the unselected cell, the bit line for the selected cell, the bit line for the unselected cell, the control gate for the selected cell, the control gate for the unselected cell in the same sector as the selected cell, the control gate for the unselected cell in a different sector than the selected cell, the erase gate for the selected cell, the erase gate for the unselected cell, the source line for the selected cell, and the source line for the unselected cell. The rows indicate the read, erase, and program operations.

[0165] Table 7: Figure 14 Operation of the VMM array 1400

[0166]

[0167] Figure 15 A neuron VMM array 1500 is shown that is particularly suitable for use in a neural network Figure 3The illustrated memory cells 310, and serve as synapses and components of neurons between the input layer and the next layer. The VMM array 1500 includes a memory array 1503 of non-volatile memory cells, a reference array 1501 of first non-volatile reference memory cells, and a reference array 1502 of second non-volatile reference memory cells. The EG lines EGR0, EGO, EG1, and EGR1 extend vertically, while the CG lines CG0, CG1, CG2, and CG3, and the SL lines WL0, WL1, WL2, and WL3 extend horizontally. The VMM array 1500 is similar to the VMM array 1400, except that the VMM array 1500 implements bidirectional tuning, in which each individual cell can be fully erased, partially programmed, and partially erased as needed to achieve a desired amount of charge on the floating gate due to the use of separate EG lines. As shown, the reference arrays 1501 and 1502 convert input current in the terminals BLR0, BLR1, BLR2, and BLR3 into control gate voltages CG0, CG1, CG2, and CG3 to be applied to the memory cells in the row direction (by action of the reference cells connected via diodes of multiplexer 1514). The current outputs (neurons) are in the bit lines BL0-BLN, where each bit line sums all the current from the non-volatile memory cells connected to that particular bit line.

[0168] Table 8 shows the operating voltages for the VMM array 1500. The columns in the table indicate the voltages on the word line for the selected cell, the word line for the unselected cell, the bit line for the selected cell, the bit line for the unselected cell, the control gate for the selected cell, the control gate for the unselected cell in the same sector as the selected cell, the control gate for the unselected cell in a different sector than the selected cell, the erase gate for the selected cell, the erase gate for the unselected cell, the source line for the selected cell, and the source line for the unselected cell. The rows indicate read, erase, and program operations.

[0169] Table 8: Figure 15 Operation of the VMM array 1500

[0170]

[0171] Figure 16 A neuron VMM array 1600 is shown that is particularly suitable for use as a synapse between the input layer and the next layer. Figure 2 The illustrated memory cells 210, and serve as synapses and components of neurons between the input layer and the next layer. In the VMM array 1600, the input INPUT0..., INPUT N are received on the bit lines BL0,... BL N and the outputs OUTPUT1, OUTPUT2, OUTPUT3, and OUTPUT4 are generated on the source lines SL0, SL1, SL2, and SL3, respectively.

[0172] Figure 17 A neuron VMM array 1700 is shown, which is particularly suitable for use in a neural network as described above. Figure 2 The memory cells 210 shown, and serve as synapses and components of neurons between the input layer and the next layer. In this example, the inputs INPUT0, INPUT1, INPUT2, and INPUT3 are received on the source lines SL0, SL1, SL2, and SL3, respectively, and the outputs OUTPUT0, …, OUTPUT N on the bit lines BL0, …, BL N are generated.

[0173] Figure 18 A neuron VMM array 1800 is shown, which is particularly suitable for use in a neural network as described above. Figure 2 The memory cells 210 shown, and serve as synapses and components of neurons between the input layer and the next layer. In this example, the inputs INPUT0, …, INPUT M are received on the word lines WL0, …, WL M respectively, and the outputs OUTPUT0, …, OUTPUT N are generated on the bit lines BL0, …, BL N .

[0174] Figure 19 A neuron VMM array 1900 is shown, which is particularly suitable for use in a neural network as described above. Figure 3 The memory cells 310 shown, and serve as synapses and components of neurons between the input layer and the next layer. In this example, the inputs INPUT0, …, INPUT M are received on the word lines WL0, …, WL M respectively, and the outputs OUTPUT0, …, OUTPUT N are generated on the bit lines BL0, …, BL N .

[0175] Figure 20 A neuron VMM array 2000 is shown, which is particularly suitable for use in a neural network as described above. Figure 4 The memory cells 410 shown, and serve as synapses and components of neurons between the input layer and the next layer. In this example, the inputs INPUT 0, , …, INPUT n are received on the vertical control gate lines CG0, …, CG N respectively, and the outputs OUTPUT1 and OUTPUT2 are generated on the source lines SL0 and SL1.

[0176] Figure 21A neuron VMM array 2100 is shown, which is particularly suitable for use with Figure 4 The memory cells 410 shown, and serve as synapses and components of neurons between an input layer and a next layer. In this example, inputs INPUT0,..., INPUT N are received on gates of bit line control gates 2901-1, 2901-2,..., 2901-(N-1), and 2901-N, respectively, which are coupled to bit lines BL0,..., BL N Exemplary outputs OUTPUT1 and OUTPUT2 are generated on source lines SL0 and SL1.

[0177] Figure 22 A neuron VMM array 2200 is shown, which is particularly suitable for use with Figure 3 The memory cells 310 shown, Figure 5 The memory cells 510 shown, and Figure 7 The memory cells 710 shown, and serve as synapses and components of neurons between an input layer and a next layer. In this example, inputs INPUT0,..., INPUT M are received on word lines WL0,..., WL M and outputs OUTPUT0,..., OUTPUT N are generated on bit lines BL0,..., BL N respectively.

[0178] Figure 23 A neuron VMM array 2300 is shown, which is particularly suitable for use with Figure 3 The memory cells 310 shown, Figure 5 The memory cells 510 shown, and Figure 7 The memory cells 710 shown, and serve as synapses and components of neurons between an input layer and a next layer. In this example, inputs INPUT0,..., INPUT M are received on control gate lines CG0,..., CG M Outputs OUTPUT0,..., OUTPUT N are generated on vertical source lines SL0,..., SL N respectively, where each source line SL i is coupled to source lines of all memory cells in column i.

[0179] Figure 24 A neuron VMM array 2400 is shown, which is particularly suitable for use with Figure 3 The memory cells 310 shown, Figure 5 The memory cells 510 shown, and Figure 7The memory unit 710 shown in FIG. 1 is used as a synapse and a component of a neuron between an input layer and the next layer. In this example, inputs INPUT0 to INPUT M On the control gate lines CG0 to CG M The output is OUTPUT0, ..., OUTPUT N On the vertical bit lines BL0, ..., BL N On the generated, where each bit line BL i The bit line coupled to all memory cells in column i.

[0180] Long short-term memory

[0181] Existing technology includes a concept known as long short-term memory (LSTM). LSTM is commonly used in artificial neural networks. LSTM allows artificial neural networks to remember information for a predetermined, arbitrary time interval and use that information in subsequent operations. A typical LSTM consists of a cell, an input gate, an output gate, and a forget gate. These three gates regulate the flow of information into and out of the cell and the time interval over which information is remembered within the LSTM. VMMs are particularly useful within LSTMs.

[0182] Figure 25 An exemplary LSTM 2500 is shown. LSTM 2500 in this example includes units 2501, 2502, 2503, and 2504. Unit 2501 receives an input vector x0 and generates an output vector h0 and a unit state vector c0. Unit 2502 receives an input vector x1, an output vector (hidden state) h0 from unit 2501, and a unit state c0 from unit 2501, and generates an output vector h1 and a unit state vector c1. Unit 2503 receives an input vector x2, an output vector (hidden state) h1 from unit 2502, and a unit state c1 from unit 2502, and generates an output vector h2 and a unit state vector c2. Unit 2504 receives an input vector x3, an output vector (hidden state) h2 from unit 2503, and a unit state c2 from unit 2503, and generates an output vector h3. Additional units may be used, and an LSTM having four units is merely an example.

[0183] Figure 26 Shown available for Figure 25 FIG2 is an exemplary implementation of an LSTM unit 2600 for units 2501, 2502, 2503, and 2504 in FIG2. LSTM unit 2600 receives an input vector x(t), a cell state vector c(t-1) from the previous unit, and an output vector h(t-1) from the previous unit, and generates a cell state vector c(t) and an output vector h(t).

[0184] LSTM unit 2600 includes sigmoid function devices 2601, 2602, and 2603, each of which applies a number between 0 and 1 to control the amount of each component in the input vector that is allowed to pass through to the output vector. LSTM unit 2600 also includes tanh devices 2604 and 2605 for applying a hyperbolic tangent function to the input vector, multiplier devices 2606, 2607, and 2608 for multiplying two vectors together, and an addition device 2609 for adding two vectors together. The output vector h(t) can be provided to the next LSTM unit in the system, or it can be accessed for other purposes.

[0185] Figure 27 LSTM unit 2700 is shown, which is an example of a specific implementation of LSTM unit 2600. For the convenience of the reader, LSTM unit 2700 uses the same numbering as LSTM unit 2600. Sigmoid function devices 2601, 2602, and 2603, and tanh device 2604 each include multiple VMM arrays 2701 and activation circuit blocks 2702. Therefore, it can be seen that VMM arrays are particularly useful in LSTM units used in certain neural network systems.

[0186] An alternative form of LSTM cell 2700 (and another example of a specific implementation of LSTM cell 2600) is Figure 28 As shown in Figure 28 In the example, sigmoid function devices 2601, 2602, and 2603 and tanh device 2604 share the same physical hardware (VMM array 2801 and activation function block 2802) in a time-division multiplexing manner. The LSTM unit 2800 also includes a multiplier device 2803 that multiplies two vectors together, an addition device 2808 that adds two vectors together, a tanh device 2605 (which includes an activation circuit block 2802), a register 2807 that stores the value i(t) when the value i(t) is output from the sigmoid function block 2802, a register 2804 that stores the value f(t)*c(t-1) when it is output from the multiplier device 2803 through the multiplexer 2810, a register 2805 that stores the value i(t)*u(t) when it is output from the multiplier device 2803 through the multiplexer 2810, a register 2806 that stores the value o(t)*c~(t) when it is output from the multiplier device 2803 through the multiplexer 2810, and a multiplexer 2809.

[0187] The LSTM unit 2700 contains multiple sets of VMM arrays 2701 and corresponding activation function blocks 2702, while the LSTM unit 2800 contains only one set of VMM arrays 2801 and an activation function block 2802, which are used to represent multiple layers in an implementation of the LSTM unit 2800. The LSTM unit 2800 will require less space than the LSTM 2700 because the LSTM unit 2800 requires only ¼ of the space for VMMs and activation function blocks as compared to the LSTM unit 2700.

[0188] It can also be appreciated that an LSTM unit will typically include multiple VMM arrays, each of which requires functionality provided by certain circuit blocks outside of the VMM arrays, such as summation and activation circuit blocks and high voltage generation blocks. Providing separate circuit blocks for each VMM array will require a large amount of space within the semiconductor device and will be somewhat inefficient. Thus, the implementations described below seek to minimize the circuitry required outside of the VMM arrays themselves.

[0189] Gate-controlled recurrent unit

[0190] Analog VMM implementations can be used for GRUs (Gated Recurrent Units). GRUs are gated mechanisms in recurrent artificial neural networks. GRUs are similar to LSTMs, except that a GRU unit generally contains fewer components than an LSTM unit.

[0191] Figure 29 An example GRU 2900 is shown. The GRU 2900 in this example includes units 2901, 2902, 2903, and 2904. The unit 2901 receives an input vector xo and generates an output vector ho. The unit 2902 receives an input vector xi, the output vector ho from the unit 2901, and generates an output vector hi. The unit 2903 receives an input vector x2 and the output vector (hidden state) hi from the unit 2902 and generates an output vector h2. The unit 2904 receives an input vector x3 and the output vector (hidden state) h2 from the unit 2903 and generates an output vector h3. Additional units can be used, and a GRU with four units is merely an example.

[0192] Figure 30 Analog VMM implementations can be used for GRUs (Gated Recurrent Units). GRUs are gated mechanisms in recurrent artificial neural networks. GRUs are similar to LSTMs, except that a GRU unit generally contains fewer components than an LSTM unit. Figure 29An exemplary implementation of a GRU cell 3000 of the cells 2901, 2902, 2903, and 2904. The GRU cell 3000 receives an input vector x(t) and an output vector h(t-1) from a previous GRU cell and generates an output vector h(t). The GRU cell 3000 includes sigmoid function devices 3001 and 3002, each of which applies a number between 0 and 1 to a component from the output vector h(t-1) and the input vector x(t). The GRU cell 3000 also includes a tanh device 3003 for applying a hyperbolic tangent function to the input vector, a plurality of multiplier devices 3004, 3005, and 3006 for multiplying two vectors together, an addition device 3007 for adding two vectors together, and a complement device 3008 for subtracting an input from 1 to generate an output.

[0193] Figure 31 A GRU cell 3100 is shown, which is an example of an implementation of the GRU cell 3000. For the convenience of the reader, the same numbering is used in the GRU cell 3100 as in the GRU cell 3000. As shown, the sigmoid function devices 3001 and 3002 and the tanh device 3003 each include a plurality of VMM arrays 3101 and an activation function block 3102. Thus, it can be seen that VMM arrays are particularly useful in GRU cells used in some neural network systems. Figure 31

[0194] An alternative form of the GRU cell 3100 (and another example of an implementation of the GRU cell 3000) is shown in FIG. 32. In this case, the GRU cell 3200 utilizes VMM arrays 3201 and activation function blocks 3202 that, when configured as sigmoid functions, apply a number between 0 and 1 to control how much of each component in the input vector is allowed to pass through to the output vector. In this case, the tanh device 3003 is replaced by a sigmoid function device 3203. Figure 32 Figure 32 Figure 32 ​​​In this case, the sigmoid function devices 3001 and 3002 and the tanh device 3003 share the same physical hardware (VMM array 3201 and activation function block 3202) in a time- multiplexed fashion. The GRU unit 3200 also includes a multiplier device 3203 that multiplies two vectors together, an adder device 3205 that adds two vectors together, a complement device 3209 that subtracts an input from 1 to generate an output, a multiplexer 3204, a register 3206 that holds the value h(t-1)*r(t) when output by the multiplexer 3204 from the multiplier device 3203, a register 3207 that holds the value h(t-1)*z(t) when output by the multiplexer 3204 from the multiplier device 3203, and a register 3208 that holds the value h^(t)*(1-z(t)) when output by the multiplexer 3204 from the multiplier device 3203.

[0195] The GRU unit 3100 contains multiple sets of VMM arrays 3101 and activation function blocks 3102, while the GRU unit 3200 contains only one set of VMM array 3201 and activation function block 3202, which are used to represent multiple layers in an implementation of the GRU unit 3200. The GRU unit 3200 will require less space than the GRU unit 3100 because the GRU unit 3200 only requires 1 / 3 of the space of the GRU unit 3100 for VMMs and activation function blocks.

[0196] It can also be appreciated that a system utilizing a GRU will typically include multiple VMM arrays, each of which will require functionality provided by certain circuit blocks outside of the VMM arrays, such as summers and activation circuit blocks and high voltage generation blocks. Providing separate circuit blocks for each VMM array will require a large amount of space within the semiconductor device and will be somewhat inefficient. Thus, the implementations described below seek to minimize the circuitry required outside of the VMM arrays themselves.

[0197] The inputs to the VMM arrays can be analog levels, binary levels, timing pulses, or digital bits, and the outputs can be analog levels, binary levels, timing pulses, or digital bits (in which case an output ADC is required to convert the output analog level current or voltage to a digital bit).

[0198] For each memory cell in the VMM array, each weight w can be implemented by a single memory cell or by a differential cell or by two hybrid memory cells (an average of 2 or more cells). In the case of a differential cell, two memory cells are required to implement the weight w as a differential weight (w = w+ - w-). In the two hybrid memory cells, two memory cells are required to implement the weight w as an average of the two cells.

[0199] Embodiments for fine tuning of cells in a VMM

[0200] Figure 33 A block diagram of a VMM system 3300 is shown. The VMM system 3300 includes a VMM array 3301, a row decoder 3302, a high voltage decoder 3303, a column decoder 3304, a bit line driver 3305, an input circuit 3306, an output circuit 3307, a control logic 3308, and a bias generator 3309. The VMM system 3300 also includes a high voltage generation block 3310 including a charge pump 3311, a charge pump regulator 3312, and a high voltage level generator 3313. The VMM system 3300 also includes an algorithm controller 3314, an analog circuit 3315, a control logic 3316, and a test control logic 3317. The systems and methods described below can be implemented in the VMM system 3300.

[0201] The input circuit 3306 can include circuitry such as a DAC (digital to analog converter), a DPC (digital to pulse converter), an AAC (analog to analog converter such as a current to voltage converter), a PAC (pulse to analog level converter), or any other type of converter. The input circuit 3306 can implement a normalization, a scaling function, or an arithmetic function. The input circuit 3306 can implement a temperature compensation function on the input. The input circuit 3306 can implement an activation function such as a ReLU or sigmoid function.

[0202] The output circuit 3307 can include circuitry such as an ADC (analog to digital converter to convert neuron analog output to digital bits), an AAC (analog to analog converter such as a current to voltage converter), an APC (analog to pulse converter), or any other type of converter. The output circuit 3307 can implement an activation function such as a ReLU or sigmoid function. The output circuit 3307 can implement a normalization, a scaling function, or an arithmetic function for neuron output. The output circuit 3307 can implement a temperature compensation function for neuron output or array output such as a bit line output, as described below.

[0203] Figure 34A tuning correction method 3400 is shown, which can be performed by the algorithm controller 3314 in the VMM system 3300. The tuning correction method 3400 generates adaptive targets based on the final error generated by the cell output and the cell initial target. The method generally starts in response to receiving a tuning command (step 3401). The initial current target for the selected cell or selected cell group Itargetv(i) (for the program / verify algorithm) is determined using a predictive target model (such as by using a function or a lookup table), and the variable DeltaError is set to 0 (step 3402). The target function (if used) will be based on the I-V programming curve for the selected memory cell or cell group. The target function also depends on various variations caused by array characteristics such as the degree of program disturbance exhibited by the cell (which depends on the cell address within the sector and the cell level, where if the cell exhibits relatively greater disturbance, the cell is subjected to more program time under suppression conditions, where cells with higher current generally have more disturbance), coupling between cells, and various types of array noise. These variations can be characterized for silicon in terms of PVT (process, voltage, temperature). The lookup table (if used) can be characterized in the same way to model the I-V curve and the various variations.

[0204] Then, a soft erase is performed on all cells in the VMM, which erases all cells to an intermediate weak erase level such that each cell will consume, for example, about 3-5 mA of current during a read operation (step 3403). The soft erase is performed, for example, by applying an incremental erase pulse voltage to the cells until the intermediate cell current is reached. Next, a deep program operation is performed on all unused cells (step 3404) in order to reach the <pA current level. Then, target adjustment (correction) based on the error result is performed. If DeltaError > 0, meaning that the cell has experienced an overshoot in programming, Itargetv(i+1) is set to Itarget + 0*DeltaError, where 0 is, for example, 1 or a number close to 1 (step 3405A).

[0205] Itarget(i+1) can also be adjusted with the appropriate error target adjustment / correction based on the previous Itarget(i). If DeltaError < 0, meaning that the cell has experienced an undershoot in programming, which means that the cell current has not reached the target, Itargetv(i+1) is set to the previous target Itargetv(i) (step 3405B).

[0206] Next, a coarse and / or fine programming and verification operation is performed (step 3406). Multiple adaptive coarse programming methods can be used to speed up programming, such as by targeting multiple progressively smaller coarse targets prior to performing the precise (fine) programming step. Adaptive precise programming is accomplished, for example, with fine (precise) increment programming voltage pulses or constant programming timing pulses. Embodiments of systems and methods for performing coarse programming and fine programming are described in U.S. Provisional Patent Application No. 62 / 933,809, filed November 11, 2019, and entitled “Precise Programming Method and Apparatus for Analog Neural Memory in a Deep Learning Artificial Neural Network,” by the same assignee as the present application, which is incorporated by reference herein.

[0207] Icell in the selected cell is measured (step 3407). For example, the cell current can be measured by a galvanometer circuit. For example, the cell current can be measured by an ADC (analog-to-digital converter) circuit, in which case the output is represented by digital bits. For example, the cell current can be measured by an I-V (current-to-voltage converter) circuit, in which case the output is represented by an analog voltage. DeltaError is calculated, which is Icell - Itarget, which represents the difference between the actual current in the measured cell (Icell) and the target current (Itarget). If |DeltaError| < DeltaMargin, then the cell has reached the target current within some tolerance (DeltaMargin), and the method ends (step 3410). |DeltaError| = abs(DeltaError) = the absolute value of DeltaError. If not, then the method returns to step 3403 and the steps are performed again in order (step 3410).

[0208] Figure 35A and Figure 35B A tuning correction method 3500 is shown, which can be performed by the algorithm controller 3314 in the VMM system 3300. Reference is made to Figure 35AThe beginning of the method (step 3501) is typically performed in response to receiving a tune command. The entire VMM array is erased, such as by a soft erase method (step 3502). A deep program operation is performed on all unused cells (step 3503) in order to reach a cell current < pA level. All cells in the VMM array are programmed to an intermediate value, such as 0.5 pA - 1.0 pA, using coarse and / or fine program loops (step 3504). Embodiments of systems and methods for performing coarse and fine programming are described in U.S. Provisional Patent Application No. 62 / 933,809, filed November 11, 2019, and entitled “Precise Programming Method and Apparatus for Analog Neural Memory in a Deep Learning Artificial Neural Network,” by the same assignee as the present application, which is incorporated by reference herein. A prediction target is set for used cells using a function or lookup table as described above (step 3505). Then, a sector tune method 3507 is performed on each sector in the VMM (step 3506). A sector is typically composed of two or more adjacent rows in the array.

[0209] Figure 35BAn adaptive target sector tuning method 3507 is shown. All cells in a sector are programmed to a final desired value (e.g., 1-50 nA) using separate or combined program / verify (P / V) methods such as: (1) coarse / fine / constant P / V cycles; (2) CG+ (CG increment only) or EG+ (EG increment only) or complementary CG+ / EG- (CG increment and EG decrement); and (3) first do the deepest programmed cells (such as progressive grouping, meaning dividing the cells into different groups, with the group of cells having the lowest current being programmed first) (step 3508A). Next, determine if Icell < Itarget. If yes, the method proceeds to step 3509. If no, the method repeats step 3508A. In step 3509, measure DeltaError, which equals the measured Icell - Itarget(i+1) (step 3509). Determine if |DeltaError| < DeltaMargin (step 3510). If yes, the method is complete (step 3511). If no, perform target adjustment. If DeltaError > 0, meaning the cell has experienced overshoot in programming, adjust the target by setting the new target to Itarget + Θ * DeltaError, where Θ is typically = 1 (step 3512A). Itarget(i+1) can also be adjusted based on the previous Itarget(i) with appropriate error target adjustment / correction. If DeltaError < 0, meaning the cell has experienced undershoot in programming, meaning the cell has not reached the target, adjust the target by keeping the previous target, i.e., Itargetv(i+1) = Itargetv(i) (step 3512B). Soft erase the sector (step 3513). Program all cells in the sector to an intermediate value (step 3514) and return to step 3509.

[0210] A typical neural network can have positive weights w+ and negative weights w- and a combined weight = w+ - w-. w+ and w- are implemented by memory cells (Iw+ and Iw- respectively) and the combined weight (Iw = Iw+ - Iw-, current subtraction) can be performed at the periphery level (such as at the array bit line output circuit). Thus, weight tuning implementations for the combined weight can include, for example, tuning both w+ cells and w- cells simultaneously, tuning only w+ cells, or tuning only w- cells, as shown in Table 8. Using the previous reference Figure 34 / Figure 35A / Figure 35BThe described program / verify and error target adjustment methods to perform tuning. The verify can be performed for only the combined weight (e.g., measuring / reading the combined weight current instead of the individual positive w+ cell current or w- cell current), only for the w+ cell current, or only for the w- cell current.

[0211] For example, for a combined Iw of 3na, Iw+ can be 3na and Iw- can be 0na; or, Iw+ can be 13na and Iw- can be 10na, meaning that neither the positive weight Iw+ nor the negative weight Iw- is zero (e.g., where zero would represent a deep program cell). This can be preferable under certain operating conditions because it would make it less likely for both Iw+ and Iw- to be affected by noise.

[0212] Table 8: Weight tuning methods

[0213] Iw Iw+ Iw- Explanation Initial target 3na 3na 0na Tune Iw+ and Iw- Initial target -2na 0na 2na Tune Iw+ and Iw- Initial target 3na 13na 10na Tune Iw+ and Iw- New target 2na 12na 10na Tune Iw+ only New target 2na 11na 9na Tune Iw+ and Iw- New target 4na 13na 9na Tune Iw- only New target 4na 12na 8na Tune Iw+ and Iw- New target -2na 8na 10na Tune Iw+ and Iw- New target -2na 7na 9na Tune Iw+ and Iw-

[0214] Figure 36A showing data behavior (I-V curves) as a function of temperature (e.g., in the subthreshold region), Figure 36B showing problems resulting from data drift during operation of the VMM system, and Figure 36C and Figure 36D showing blocks for compensating for data drift and regarding Figure 36C showing blocks for compensating for temperature variations.

[0215] Figure 36A showing known characteristics of the VMM system as a function of operating temperature, the sense current in any given selected non-volatile memory cell in the VMM array increases in the subthreshold region, decreases in the saturation region, or generally decreases in the linear region.

[0216] Figure 36B showing array current distribution as a function of time usage (data drift), and it shows that the aggregate output from the VMM array (which is the sum of the current from all bit lines in the VMM array) shifts to the right (or left, depending on the technology used) as a function of operating time usage, meaning that the total aggregate output will drift as a function of the life usage of the VMM system. This phenomenon is referred to as data drift because the data drifts due to usage conditions and degrades due to environmental factors.

[0217] Figure 36C showing bit line compensation circuit 3600, which can include a compensation current i COMPThe output of the bit line output circuit 3610 is injected to compensate for data drift. The bit line compensation circuit 3600 may include a scaler circuit that amplifies or reduces the output based on a resistor or capacitor network. The bit line compensation circuit 3600 may include a shifter circuit that shifts or offsets the output based on its resistor or capacitor network.

[0218] Figure 36D A data drift monitor 3620 is shown that detects the amount of data drift. This information is then used as an input to the bit line compensation circuit 3600 so that an appropriate level of i can be selected. COMP .

[0219] Figure 37 36. The bit line compensation circuit 3700 is shown as an embodiment of the bit line compensation circuit 3600 in FIG. The bit line compensation circuit 3700 includes an adjustable current source 3701 and an adjustable current source 3702, which together generate i COMP , where i COMP Equal to the current generated by adjustable current source 3701 minus the current generated by adjustable current source 3702.

[0220] Figure 38 36. The bit line compensation circuit 3800 includes an operational amplifier 3801, an adjustable resistor 3802, and an adjustable resistor 3803. The operational amplifier 3801 receives a reference voltage VREF at its non-inverting terminal and receives V INPUT , where V INPUT It is from Figure 36C The bit line output circuit 3610 receives the voltage and generates the output V OUTPUT , where V OUTPUT It is V INPUT A scaled version of V can be created to compensate for data drift based on the ratio of resistors 3803 and 3802. By configuring the values ​​of resistors 3803 and / or 3802, V OUTPUT .

[0221] Figure 39 A bit line compensation circuit 3900 is shown, which is an embodiment of the bit line compensation circuit 3600 in FIG. 36 . The bit line compensation circuit 3900 includes an operational amplifier 3901, a current source 3902, a switch 3904, and an adjustable integrating output capacitor 3903. Here, the current source 3902 is actually the output current on a single bit line or a collection of multiple bit lines (such as one for summing positive weights w+ and one for summing negative weights w-) in the VMM array. The operational amplifier 3901 receives a reference voltage VREF at its non-inverting terminal and V INPUT , where VINPUT is the voltage received from the bit line output circuit 3610 in Figure 36C The bit line compensation circuit 3900 acts as an integrator that integrates the current Ineu through the capacitor 3903 over an adjustable integration time to generate an output voltage V OUTPUT where V OUTPUT = Ineu * integration time / C 3903 where C 3903 is the value of the capacitor 3903. Thus, the output voltage V OUTPUT is proportional to the (bit line) output current Ineu, proportional to the integration time, and inversely proportional to the capacitance of the capacitor 3903. The bit line compensation circuit 3900 generates an output V OUTPUT where V OUTPUT is scaled in value based on the configured value of the capacitor 3903 and / or the integration time to compensate for data drift.

[0222] Figure 40 A bit line compensation circuit 4000 is shown that is one implementation of the bit line compensation circuit 3600 in FIG. 36. The bit line compensation circuit 4000 includes a current mirror 4010 with an M:N ratio, which means that I COMP = (M / N) * i input . The current mirror 4010 receives a current i INPUT and mirrors and optionally scales that current to generate i COMP . Thus, by configuring the M parameter and / or the N parameter, i COMP can be amplified or scaled down.

[0223] Figure 41 A bit line compensation circuit 4100 is shown that is one implementation of the bit line compensation circuit 3600 in FIG. 36. The bit line compensation circuit 4100 includes an operational amplifier 4101, an adjustable scaling resistor 4102, an adjustable shift resistor 4103, and an adjustable resistor 4104. The operational amplifier 4101 receives a reference voltage V REF on its non-inverting terminal and V IN on its inverting terminal. V IN is generated in response to V INPUT and Vshft, where V INPUT is the voltage received from the bit line output circuit 3610 in Figure 36C and Vshft is a voltage intended to achieve a shift between V INPUT and V OUTPUT .

[0224] Thus, V OUTPUT is a scaled and shifted version of V INPUT to compensate for data drift.

[0225] Figure 42 A bit line compensation circuit 4200 is shown, which is one implementation of the bit line compensation circuit 3600 in FIG. 36. The bit line compensation circuit 4200 includes an operational amplifier 4201, an input current source Ineu 4202, a current shifter 4203, switches 4205 and 4206, and an adjustable integration output capacitor 4204. Here, the current source 4202 is actually the output current Ineu on a single bit line or multiple bit lines in the VMM array. The operational amplifier 4201 receives a reference voltage VREF on its non-inverting terminal and receives I IN where I IN is the sum of Ineu and the current output by the current shifter 4203, and generates an output V OUTPUT where V OUTPUT is scaled (based on capacitor 4204) and shifted (based on Ishifter 4203) to compensate for data drift.

[0226] Figures 43-48 Various circuits that can be used to provide the W value to be programmed or read into each selected cell during a program or read operation are shown.

[0227] Figure 43 A neuron output circuit 4300 is shown, which includes an adjustable current source 4301 and an adjustable current source 4302, which together generate I OUT where I OUT is equal to the current I W+ generated by the adjustable current source 4301 minus the current I W- generated by the adjustable current source 4302. The adjustable current Iw+ 4301 is a scaled current of a cell current or neuron current (such as a bit line current) to implement a positive weight. The adjustable current Iw- 4302 is a scaled current of a cell current or neuron current (such as a bit line current) to implement a negative weight. Current scaling is accomplished such as by a M:N ratio current mirror circuit, where Iout = (M / N)*Iin.

[0228] Figure 44 A neuron output circuit 4400 is shown, which includes an adjustable capacitor 4401, a control transistor 4405, a switch 4402, a switch 4403, and an adjustable current source 4404 Iw+, which is a scaled output current of a cell current or (bit line) neuron current such as a M:N current mirror circuit. The transistor 4405 is used to apply a fixed bias voltage to the current 4404, for example. The circuit 4404 generates V OUT where V OUT is inversely proportional to the capacitor 4401, is proportional to the adjustable integration time (the time when the switch 4403 is closed and the switch 4402 is open), and is proportional to the adjustable current source 4404 IW+ The generated current is proportional to V OUT is equal to V+ - ((Iw+ * integration time) / C 4401 ), where C 4401 is the value of capacitor 4401. The positive terminal V+ of capacitor 4401 is connected to a positive supply voltage, and the negative terminal V- of capacitor 4401 is connected to the output voltage V OUT .

[0229] Figure 45 A neuron circuit 4500 is shown, which includes a capacitor 4401 and a tunable current source 4502, which current is a scaled current of a unit current or a (bitline) neuron current such as an M:N current mirror. The circuit 4500 generates V OUT , where V OUT is inversely proportional to capacitor 4401, is proportional to a tunable integration time (time for which switch 4501 is open), and is proportional to a current I Wi generated by tunable current source 4502. The capacitor 4401 is reused from neuron output circuit 44 after it has completed its operation of integrating the current Iw+. The positive and negative terminals (V+ and V-) are then swapped in neuron output circuit 45, where the positive terminal is connected to the output voltage V OUT , which is de-integrated by the current Iw-. The negative terminal is held at the previous voltage value by a clamping circuit (not shown). In practice, output circuit 44 is used for positive weight implementations, and circuit 45 is used for negative weight implementations, where the final charge on capacitor 4401 effectively represents the combined weight (Qw = Qw+ - Qw-).

[0230] Figure 46 A neuron circuit 4600 is shown, which includes a tunable capacitor 4601, a switch 4602, a control transistor 4604, and a tunable current source 4603. The circuit 4600 generates V OUT , where V OUT is inversely proportional to capacitor 4601, is proportional to a tunable integration time (time for which switch 4602 is open), and is proportional to a current I W- generated by tunable current source 4603. The negative terminal V- of capacitor 4601 is, for example, equal to ground. The positive terminal V+ of capacitor 4601 is, for example, initially pre-charged to a positive voltage before integrating the current Iw-. The neuron circuit 4600 can be used in place of neuron circuit 4500 and neuron circuit 4400 to implement a combined weight (Qw = Qw+ - Qw-).

[0231] Figure 47A neuron circuit 4700 is shown that includes operational amplifiers 4703 and 4706; adjustable current sources Iw+ 4701 and Iw- 4702; and adjustable resistors 4704, 4705, and 4707. The neuron circuit 4700 generates V OUT , which is equal to R 4707 *(Iw+-Iw-). The adjustable resistor 4707 implements scaling of the output. The adjustable current sources Iw+ 4701 and Iw- 4702 also implement scaling of the output, for example through an M:N ratio current mirror circuit (Iout = (M / N)*Iin).

[0232] Figure 48 A neuron circuit 4800 is shown that includes operational amplifiers 4803 and 4806; switches 4808 and 4809; adjustable current sources Iw- 4802 and Iw+ 4801; adjustable capacitors 4804, 4805, and 4807. The neuron circuit 4800 generates V OUT , which is proportional to (Iw+-Iw-), proportional to the integration time (time when switches 4808 and 4809 are off), and inversely proportional to the capacitance of capacitor 4807. The adjustable capacitor 4807 implements scaling of the output. The adjustable current sources Iw+ 4801 and Iw- 4802 also implement scaling of the output, for example through an M:N ratio current mirror circuit (Iout = (M / N)*Iin). The integration time can also adjust the output scaling.

[0233] Figure 49A 、 Figure 49B and Figure 49C A block diagram of an output circuit such as the output circuit 3307 in Figure 33 is shown.

[0234] In Figure 49A , the output circuit 4901 includes an ADC circuit 4911 that digitizes the analog neuron output 4910 directly to provide a digital output bit 4912.

[0235] In Figure 49B , the output circuit 4902 includes a neuron output circuit 4921 and an ADC 4911. The neuron output circuit 4921 receives the neuron output 4920 and shapes it, which is then digitized by the ADC circuit 4911 to generate the output 4912. The neuron output circuit 4921 can be used for normalization, scaling, shifting, mapping, arithmetic operations, activation, and / or temperature compensation, such as previously described. The ADC circuit can be a serial (ramp or up or counting) ADC, a SAR ADC, a pipeline ADC, a sigma-delta ADC, or any type of ADC.

[0236] In Figure 49CIn particular embodiments, the output circuit includes a neuron output circuit 4921 that receives the neuron output 4930, and a converter circuit 4931 for converting the output from the neuron output circuit 4921 to the output 4932. The converter 4931 can include an ADC, an AAC (analog-to-analog converter such as a current-to-voltage converter), an APC (analog-to-pulse converter), or any other type of converter. The ADC 4911 or the converter 4931 can be used to implement an activation function by, for example, bit-mapping (e.g., quantization) or clipping (e.g., clipped ReLU). The ADC 4911 and the converter 4931 can be configurable, such as for lower or higher precision (e.g., lower or higher number of bits), lower or higher performance (e.g., slower or faster speed), etc.

[0237] Another implementation for scaling and shifting is by configuring an ADC (analog-to-digital) conversion circuit (such as a serial ADC, a SAR ADC, a pipelined ADC, a ramp ADC, etc.) for converting the array (bit line) output to digital bits, such as with lower or higher bit precision, and then manipulating the digital output bits according to some function (e.g., linear or non-linear, compression, non-linear activation, etc.), such as by normalizing (e.g., 12 bits to 8 bits), shifting, or re-mapping. Embodiments of ADC conversion circuits are described in U.S. Provisional Patent Application No. 62 / 933,809, filed November 11, 2019, and titled “Precise Programming Method and Apparatus for Analog Neural Memory in a Deep Learning Artificial Neural Network,” by the same assignee as the present application, which is incorporated by reference herein.

[0238] Table 9 shows alternative methods of performing read, erase, and program operations:

[0239] Table 9: Operation of Flash Memory Cell

[0240] SL BL WL CG EG P-Sub Read 0 0.5 1 0 0 0 Erase 0 0 0 0 / -8V 10-12V / +8V 0 Program1 0-5V 0 0 8V -10 to -12V 0 - Program2 0 0 0 8V 0-5V -10V

[0241] Read and erase operations are similar to the previous tables. However, the two methods for programming are implemented by the Fowler-Nordheim (FN) tunneling mechanism.

[0242] An implementation for scaling the input can be done such as by enabling a certain number of rows of the VMM at a time and then fully combining the results together.

[0243] Another implementation is to scale the input voltage and appropriately rescale the output to achieve normalization.

[0244] Another implementation for scaling a pulse width modulated input is by modulating the timing of the pulse width. An embodiment of this technique is described in U.S. Patent Application No. 16 / 449,201, filed June 21, 2019, and entitled “Configurable Input Blocks and Output Blocks and Physical Layout for Analog Neural Memory in Deep Learning Artificial Neural Network,” by the same assignee as the present application, which is incorporated by reference herein.

[0245] Another implementation for scaling an input is by enabling one input bit at a time, e.g., for an 8-bit input IN7:0, evaluating IN0, IN1, …, IN7, respectively, in sequence, and then combining the output results together with the appropriate binary weight. An embodiment of this technique is described in U.S. Patent Application No. 16 / 449,201, filed June 21, 2019, and entitled “Configurable Input Blocks and Output Blocks and Physical Layout for Analog Neural Memory in Deep Learning Artificial Neural Network,” by the same assignee as the present application, which is incorporated by reference herein.

[0246] Optionally, in the above implementations, the measurement of the cell current for the purpose of verifying or reading the current can be averaged or measured multiple times, e.g., 8 to 32 times, to reduce the impact of noise, such as RTN or any random noise, and / or to detect any outlier bits that are defective and need to be replaced by a redundant bit.

[0247] It should be noted that as used herein, the terms "over," "on," "onto," "above," "down," "up," "on top," "adjacent," "attached to," "coupled to," "electrically coupled to," and the like can encompass both a direct connection between elements and an indirect connection between elements in which one or more intermediate elements are present. For example, an element formed "over" another element can include being formed directly over the other element without any intermediate elements intervening therebetween, as well as being indirectly formed over the other element with one or more intermediate elements intervening therebetween.

Claims

1. A neuron output circuit for providing a current to program a weight value in a selected memory cell in a vector-matrix multiplication array, the neuron output circuit comprising: a first adjustable current source that generates a scaled current in response to neuronal current to achieve a positive weight; as well as A second adjustable current source generates a scaled current in response to the neuron current to achieve the negative weight.

2. A neuron output circuit for providing a current to program a weight value in a selected memory cell in a vector-matrix multiplication array, the neuron output circuit comprising: an adjustable capacitor comprising a first terminal and a second terminal, the second terminal providing an output voltage for the neuron output circuit; a control transistor comprising a first terminal and a second terminal; a first switch selectively coupled between the first terminal and the second terminal of the adjustable capacitor; a second switch selectively coupled between the second terminal of the adjustable capacitor and the first terminal of the control transistor; as well as An adjustable current source is coupled to the second terminal of the control transistor.

3. A neuron output circuit for providing a current to program a weight value in a selected memory cell in a vector-matrix multiplication array, the neuron output circuit comprising: an adjustable capacitor comprising a first terminal and a second terminal, the second terminal providing an output voltage for the neuron output circuit; a control transistor comprising a first terminal and a second terminal; A switch selectively coupled between the second terminal of the adjustable capacitor and the between the first terminals of the control transistor; as well as An adjustable current source is coupled to the second terminal of the control transistor.

4. A neuron output circuit for providing a current to program a weight value in a selected memory cell in a vector-matrix multiplication array, the neuron output circuit comprising: an adjustable capacitor comprising a first terminal and a second terminal, the first terminal providing an output voltage for the neuron output circuit; a control transistor comprising a first terminal and a second terminal; a first switch selectively coupled between the first terminal of the adjustable capacitor and the first terminal of the control transistor; as well as An adjustable current source is coupled to the second terminal of the control transistor.

5. A neuron output circuit for providing a current to program a weight value in a selected memory cell in a vector-matrix multiplication array, the neuron output circuit comprising: a first operational amplifier comprising an inverting input, a non-inverting input, and an output; a second operational amplifier, the second operational amplifier comprising an inverting input, a non-inverting input, and an output; a first adjustable current source coupled to the inverting input of the first operational amplifier; a second adjustable current source coupled to the inverting input of the second operational amplifier; a first adjustable resistor coupled to the inverting input of the first operational amplifier; a second adjustable resistor coupled to the inverting input of the second operational amplifier; as well as A third adjustable resistor is coupled between the output of the first operational amplifier and the inverting input of the second operational amplifier.

6. A neuron output circuit for providing a current to program a weight value in a selected memory cell in a vector-matrix multiplication array, the neuron output circuit comprising: a first operational amplifier comprising an inverting input, a non-inverting input, and an output; a second operational amplifier, the second operational amplifier comprising an inverting input, a non-inverting input, and an output; a first adjustable current source coupled to the inverting input of the first operational amplifier; a second adjustable current source coupled to the inverting input of the second operational amplifier; a first switch coupled between the inverting input and the output of the first operational amplifier; a second switch coupled between the inverting input and the output of the second operational amplifier; a first adjustable capacitor coupled between the inverting input and the output of the first operational amplifier; a second adjustable capacitor coupled between the inverting input and the output of the second operational amplifier; as well as A third adjustable capacitor is coupled between the output of the first operational amplifier and the inverting input of the second operational amplifier.

Citation Information

Patent Citations

  • Deep learning neural network classifier using non-volatile memory array

    US11308383B2

  • Deep Learning Neural Network Classifier Using Non-volatile Memory Array

    US20170337466A1

  • High Precision And Highly Efficient Tuning Mechanisms And Algorithms For Analog Neuromorphic Memory In Artificial Neural Networks

    US20190164617A1

  • Configurable input blocks and output blocks and physical layout for analog neural memory in deep learning artificial neural network

    US20200349421A1

  • Single transistor non-valatile electrically alterable semiconductor memory device

    US5029130A