Circuit for compensating data drift in an analog neural memory in an artificial neural network
By introducing data drift monitoring circuit and bit line compensation circuit in the simulated neuromorphic memory system, the drift error problem during reading operation in the vector matrix multiplication array is solved, achieving higher accuracy and stability.
Patent Information
- Application Number
- CN202080091624.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-03-26
- Filing Date
- 2020-09-03
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2040-09-03
AI Technical Summary
There is data drift error in the nonvolatile memory cells in the vector matrix multiplication array in the simulated neuromorphic memory system, resulting in significant errors in the read operation.
A circuit is provided, including a data drift monitoring circuit and a bit line compensation circuit, for monitoring and compensating drift errors during read operations in a vector matrix multiplication array. The bit line compensation circuit generates a compensation current and injects it into the bit lines of the array to offset the drift error.
It effectively compensates for the drift error during reading operations in the vector matrix multiplication array, and improves the accuracy and stability of the memory cells.
Smart Images

Figure CN114902339B_ABST
Abstract
Description
[0001] Priority Claim
[0002] This application claims priority to U.S. Provisional Patent Application No. 62 / 957,013, filed on January 3, 2020, titled "Precise Data Tuning Method and Apparatus for Analog Neuromorphic Memory in an Artificial Neural Network", and U.S. Patent Application No. 16 / 830,733, filed on March 26, 2020, titled "Circuitry to Compensate for Data Drift in Analog Neural Memory in an Artificial Neural Network". Technical Field
[0003] A number of implementations are provided for compensating for drift errors in non-volatile memory cells within a VMM array in an analog neuromorphic memory system. Background Art
[0004] Artificial neural networks mimic biological neural networks (the central nervous system of animals, particularly the brain), and are used to estimate or approximate functions that may depend on a large number of inputs and are generally unknown. Artificial neural networks typically include layers of interconnected "neurons" that exchange messages with each other.
[0005] Figure 1 An artificial neural network is shown, where the circles represent the inputs or layers of neurons. The connections (called synapses) are represented by arrows and have numerical weights that can be adjusted according to experience. This enables the artificial neural network to adapt to the inputs and learn. Generally, an artificial neural network includes a layer of multiple inputs. There is typically one or more intermediate layers of neurons, and an output layer of neurons that provides the output of the neural network. The neurons at each level make decisions separately or jointly based on the data received from the synapses.
[0006] One of the main challenges in developing artificial neural networks for high-performance information processing is the lack of sufficient hardware technology. In fact, practical artificial neural networks rely on a large number of synapses to achieve high connectivity between neurons, that is, very high computational parallelism. In principle, such complexity can be achieved by digital supercomputers or clusters of dedicated graphics processing units. However, compared with biological networks, these methods are not only costly but also have mediocre energy efficiency. Biological networks consume less energy mainly because they perform low-precision analog computations. CMOS analog circuits have been used in artificial neural networks, but due to the large number of neurons and synapses given, most CMOS-implemented synapses are too large.
[0007] The applicant previously disclosed in U.S. Patent Application No. 15 / 594,439 (published as U.S. Patent Publication 2017 / 0337466) an artificial (analog) neural network that uses one or more non-volatile memory arrays as synapses, and this patent application is incorporated herein by reference. The non-volatile memory array operates as an analog neuromorphic memory. As used herein, the term "neuromorphic" refers to a circuit that implements a model of the nervous system. The analog neuromorphic memory includes a first plurality of synapses configured to receive a first plurality of inputs and generate a first plurality of outputs therefrom, and a first plurality of neurons configured to receive the first plurality of outputs. The first plurality of synapses includes a plurality of memory cells, and each memory cell of the memory cells includes: spaced-apart source regions and drain regions formed in a semiconductor substrate, wherein a channel region extends between the source region and the drain region; a floating gate disposed above a first portion of the channel region and insulated from the first portion; and a non-floating gate disposed above a second portion of the channel region and insulated from the second portion. Each memory cell of the plurality of memory cells is configured to store a weight value corresponding to a plurality of electrons on the floating gate. The plurality of memory cells are configured to multiply the first plurality of inputs by the stored weight values to generate the first plurality of outputs. An array of memory cells arranged in this way can be referred to as a vector matrix multiplication (VMM) array.
[0008] Each non-volatile memory cell used in the VMM must be erased and programmed to maintain a very specific and precise amount of electric charge (i.e., the number of electrons) in the floating gate. For example, each floating gate must maintain one of N different values, where N is the number of different weights that can be indicated by each cell. Examples of N include 16, 32, 64, 128, and 256. One challenge lies in being able to program the selected cells with the precision and granularity required for different N values. For example, if the selected cell can include one value out of 64 different values, extremely high precision is required in the programming operation.
[0009] Since these systems require such high precision, any errors caused by phenomena such as data drift can be significant.
[0010] Improved compensation circuits and methods are needed for compensating data drift in a VMM array in an analog neuromorphic memory. SUMMARY OF THE INVENTION
[0011] Numerous embodiments are provided for compensating drift errors in non-volatile memory cells within a VMM array in an analog neuromorphic memory system.
[0012] In one embodiment, a circuit for compensating drift errors during a read operation in a vector matrix multiplication array is provided. The circuit includes a data drift monitoring circuit coupled to the array to generate an output indicative of data drift; and a bit line compensation circuit for generating a compensation current in response to the output from the data drift monitoring circuit and injecting the compensation current into one or more bit lines of the array.
[0013] In another embodiment, a circuit for compensating drift errors during a read operation in a vector matrix multiplication array is provided. The circuit includes a bit line compensation circuit for generating a compensation current and injecting the compensation current into one or more bit lines of the array to compensate for drift errors.
[0014] In another embodiment, a circuit for compensating drift errors during a read operation in a vector matrix multiplication array is provided. The circuit includes a bit line compensation circuit for scaling the output of the array to compensate for drift errors.
[0015] In another embodiment, a circuit for compensating drift errors during a read operation in a vector matrix multiplication array is provided. The circuit includes a bit line compensation circuit for shifting the output of the array to compensate for drift errors.
[0016] In another embodiment, a method for compensating drift errors during a read operation in a vector matrix multiplication array is provided. The method includes monitoring data drift in the vector matrix multiplication array. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 A schematic diagram showing an artificial neural network of the prior art.
[0018] Figure 2 Showing a split-gate flash memory cell of the prior art.
[0019] Figure 3 Showing another split-gate flash memory cell of the prior art.
[0020] Figure 4Shows a split-gate flash memory cell of another prior art.
[0021] Figure 5 Shows a split-gate flash memory cell of another prior art.
[0022] Figure 6 Shows a split-gate flash memory cell of another prior art.
[0023] Figure 7 Shows a stacked-gate flash memory cell of the prior art.
[0024] Figure 8 Is a schematic diagram showing different levels of an exemplary artificial neural network using one or more VMM arrays.
[0025] Fig. 9 Is a block diagram showing a VMM system including a VMM array and other circuits.
[0026] Fig.10 Is a block diagram showing an exemplary artificial neural network using one or more VMM systems.
[0027] Fig.11 Shows another embodiment of a VMM array.
[0028] Fig.12 Shows another embodiment of a VMM array.
[0029] Fig.13 Shows another embodiment of a VMM array.
[0030] Fig.14 Shows another embodiment of a VMM array.
[0031] Fig.15 Shows another embodiment of a VMM array.
[0032] Fig.16 Shows another embodiment of a VMM array.
[0033] Fig.17 Shows another embodiment of a VMM array.
[0034] Fig.18 Shows another embodiment of a VMM array.
[0035] Fig.19 Shows another embodiment of a VMM array.
[0036] Fig. 20 Shows another embodiment of a VMM array.
[0037] Fig.21 Shows another embodiment of a VMM array.
[0038] Fig. 22 Shows another embodiment of the VMM array.
[0039] Fig.23 Shows another embodiment of the VMM array.
[0040] Fig.24 Shows another embodiment of the VMM array.
[0041] Fig.25 Shows a prior art long short-term memory system.
[0042] Fig.26 Shows an exemplary cell used in a long short-term memory system.
[0043] Fig. 27 Shows Fig.26 an embodiment of the exemplary cell.
[0044] Fig.28 Shows Fig.26 another embodiment of the exemplary cell.
[0045] Fig.29 Shows a prior art gated recurrent unit system.
[0046] Fig.30 Shows an exemplary cell used in a gated recurrent unit system.
[0047] Fig.31 Shows Fig.30 an embodiment of the exemplary cell.
[0048] Fig.32 Shows Fig.30 another embodiment of the exemplary cell.
[0049] Fig.33 Shows a VMM system.
[0050] Fig.34 Shows a tuning correction method.
[0051] Fig.35A Shows a tuning correction method.
[0052] Fig.35B Shows a sector tuning correction method.
[0053] Fig.36A Shows the effect of temperature on the value stored in the cell.
[0054] Fig.36B Shows the problems caused by data drift during the operation of the VMM system.
[0055] Fig.36C Blocks for compensating data drift are shown.
[0056] Fig.36D A data drift monitor is shown.
[0057] Fig.37 A bit line compensation circuit is shown.
[0058] Fig.38 Another bit line compensation circuit is shown.
[0059] Fig.39 Another bit line compensation circuit is shown.
[0060] Fig.40 Another bit line compensation circuit is shown.
[0061] Fig.41 Another bit line compensation circuit is shown.
[0062] Fig.42 Another bit line compensation circuit is shown.
[0063] Fig.43 A neuron circuit is shown.
[0064] Fig.44 Another neuron circuit is shown.
[0065] Fig.45 Another neuron circuit is shown.
[0066] Fig.46 Another neuron circuit is shown.
[0067] Fig.47 Another neuron circuit is shown.
[0068] Fig.48 Another neuron circuit is shown.
[0069] Fig.49A A block diagram of an output circuit is shown.
[0070] Fig.49B A block diagram of another output circuit is shown.
[0071] Fig.49C A block diagram of another output circuit is shown. Detailed Description
[0072] The artificial neural network of the present invention utilizes a combination of CMOS technology and a non - volatile memory array.
[0073] Non-volatile memory cell
[0074] Digital non-volatile memories are well known. For example, U.S. Patent 5,029,130 ("the '130 patent"), which is incorporated herein by reference, discloses an array of split-gate non-volatile memory cells, which are a type of flash memory cell. Such memory cells 210 are shown in Figure 2 Each memory cell 210 includes a source region 14 and a drain region 16 formed in a semiconductor substrate 12, with a channel region 18 therebetween. A floating gate 20 is formed over and insulated from (and controls the conductivity of) a first portion of the channel region 18, and is formed over a portion of the source region 14. A word line terminal 22 (which is typically coupled to a word line) has a first portion disposed over and insulated from (and controls the conductivity of) a second portion of the channel region 18, and a second portion that extends upward and is located over the floating gate 20. The floating gate 20 and the word line terminal 22 are insulated from the substrate 12 by a gate oxide. A bit line terminal 24 is coupled to the drain region 16.
[0075] The memory cell 210 is erased by placing a high positive voltage on the word line terminal 22 (where electrons are removed from the floating gate), which causes electrons on the floating gate 20 to tunnel through an intervening insulator from the floating gate 20 to the word line terminal 22 via Fowler-Nordheim tunneling.
[0076] The memory cell 210 is programmed by placing a positive voltage on the word line terminal 22 and a positive voltage on the source region 14 (where electrons are placed on the floating gate). An electron current will flow from the source region 14 (source line terminal) to the drain region 16. When the electrons reach the gap between the word line terminal 22 and the floating gate 20, the electrons will accelerate and heat up. Due to the electrostatic attraction from the floating gate 20, some of the heated electrons will be injected onto the floating gate 20 through the gate oxide.
[0077] The memory cell 210 is read by placing a positive read voltage on the drain region 16 and the word line terminal 22 (which turns on the portion of the channel region 18 under the word line terminal). If the floating gate 20 is positively charged (i.e., electrons have been erased), then the portion of the channel region 18 under the floating gate 20 is also turned on, and current will flow through the channel region 18, which is sensed as an erased state or a "1" state. If the floating gate 20 is negatively charged (i.e., programmed with electrons), then the portion of the channel region 18 under the floating gate 20 is mostly or completely turned off, and current will not (or very little current) flow through the channel region 18, which is sensed as a programmed state or a "0" state.
[0078] Table 1 shows the typical voltage ranges that can be applied to the terminals of the memory cell 110 for performing read operations, erase operations, and program operations:
[0079] Table 1: Figure 2 Operation of the flash memory cell 210
[0080] WL BL SL Read 1 0.5-3V 0.1-2V 0V Read 2 0.5-3V 0-2V 2-0.1V Erase About 11-13V 0V 0V programming 1V-2V 1-3μA 9-10V
[0081] "Read 1" is a read mode in which the cell current is output on the bit line. "Read 2" is a read mode in which the cell current is output on the source line terminal.
[0082] Figure 3 A memory cell 310 is shown, which is similar to Figure 2 the memory cell 210, but with the addition of a control gate (CG) terminal 28. The control gate terminal 28 is biased at a high voltage (e.g., 10V) during programming, at a low voltage or a negative voltage (e.g., 0V / -8V) during erasure, and at a low voltage or a medium voltage (e.g., 0V / 2.5V) during read. The other terminals are biased similar to Figure 2 that.
[0083] Figure 4 A four-gate memory cell 410 is shown, which includes a source region 14, a drain region 16, a floating gate 20 above a first portion of the channel region 18, a select gate 22 (usually coupled to the word line WL) above a second portion of the channel region 18, a control gate 28 above the floating gate 20, and an erase gate 30 above the source region 14. This configuration is described in U.S. Patent 6,747,310, which is incorporated herein by reference for all purposes. Here, all gates except the floating gate 20 are non-floating gates, which means they are electrically connected or can be electrically connected to a voltage source. Programming is performed by hot electrons from the channel region 18 that inject themselves into the floating gate 20. Erasure is performed by electrons tunneling from the floating gate 20 to the erase gate 30.
[0084] Table 2 shows the typical voltage ranges that can be applied to the terminals of the memory cell 410 for performing read operations, erase operations, and programming operations:
[0085] Table 2: Figure 4 Operation of the flash memory cell 410
[0086] WL / SG BL CG EG SL Read 1 0.5-2V 0.1-2V 0-2.6V 0-2.6V 0V Read 2 0.5-2V 0-2V 0-2.6V 0-2.6V 2-0.1V Erase -0.5V / 0V 0V 0V / -8V 8-12V 0V programming 1V 1μA 8-11V 4.5-9V 4.5-5V
[0087] "Read 1" is a read mode in which the cell current is output on the bit line. "Read 2" is a read mode in which the cell current is output on the source line terminal.
[0088] Figure 5 A memory cell 510 is shown, which is similar to Figure 4The memory cell 410 is similar. Erasure is performed by biasing the substrate 18 to a high voltage and biasing the control gate CG terminal 28 to a low voltage or a negative voltage. Alternatively, erasure is performed by biasing the word line terminal 22 to a positive voltage and biasing the control gate terminal 28 to a negative voltage. Programming and reading are similar to Figure 4 as described.
[0089] Figure 6 FIG. shows a triple-gate memory cell 610, which is another type of flash memory cell. The memory cell 610 is the same as Figure 4 the memory cell 410, except that the memory cell 610 does not have a separate control gate terminal. Except for not applying a control gate bias, the erasure operation (erasure using the erase gate terminal) and the read operation are similar to Figure 4 the operations described. Without a control gate bias, the programming operation is also completed, and as a result, a higher voltage must be applied to the source line terminal during the programming operation to compensate for the lack of a control gate bias.
[0090] Table 3 shows the typical voltage ranges that can be applied to the terminals of the memory cell 610 to perform read operations, erase operations, and programming operations:
[0091] Table 3: Figure 6 Operation of the flash memory cell 610
[0092] WL / SG BL EG SL Read 1 0.5-2.2V 0.1-2V 0-2.6V 0V Read 2 0.5-2.2V 0-2V 0-2.6V 2-0.1V Erase -0.5V / 0V 0V 11.5V 0V programming 1V 2-3μA 4.5V 7-9V
[0093] "Read 1" is a read mode in which the cell current is output on the bit line. "Read 2" is a read mode in which the cell current is output on the source line terminal.
[0094] Figure 7 FIG. shows a stacked-gate memory cell 710, which is another type of flash memory cell. The memory cell 710 is similar to Figure 2 the memory cell 210, except that the floating gate 20 extends over the entire channel region 18, and the control gate terminal 22 (which will be coupled to the word line here) extends over the floating gate 20, separated by an insulating layer (not shown). The erase, program, and read operations operate in a manner similar to that previously described for the memory cell 210.
[0095] Table 4 shows the typical voltage ranges that can be applied to the terminals of the memory cell 710 and the substrate 12 to perform read, erase, and program operations:
[0096] Table 4: Figure 7 Operation of the flash memory cell 710
[0097] CG BL SL Substrate Read 1 0-5V 0.1–2V 0-2V 0V Read 2 0.5-2V 0-2V 2-0.1V 0V Erase -8 to -10V / 0V FLT FLT 8-10V / 15-20V programming 8-12V 3-5V / 0V 0V / 3-5V 0V
[0098] "Read 1" is a read mode in which the cell current is output on the bit line. "Read 2" is a read mode in which the cell current is output on the source line terminal. Optionally, in an array of rows and columns including memory cells 210, 310, 410, 510, 610, or 710, the source line may be coupled to one row of memory cells or two adjacent rows of memory cells. That is, the source line terminal may be shared by memory cells of adjacent rows.
[0099] To utilize a memory array including one of the above types of non-volatile memory cells in an artificial neural network, two modifications were made. First, the circuitry was configured such that each memory cell could be individually programmed, erased, and read without adversely affecting the memory state of other memory cells in the array, as further explained below. Second, continuous (analog) programming of the memory cells was provided.
[0100] Specifically, the memory state of each memory cell in the array (i.e., the charge on the floating gate) can be continuously changed from a fully erased state to a fully programmed state independently and with minimal interference to other memory cells. In another embodiment, the memory state of each memory cell in the array (i.e., the charge on the floating gate) can be continuously changed from a fully programmed state to a fully erased state independently and with minimal interference to other memory cells, and vice versa. This means that the cell storage device is analog, or can at least store one of a number of discrete values (such as 16 or 64 different values), which allows for very precise and individual tuning of all the cells in the memory array, and which makes the memory array ideal for storing and fine-tuning the synaptic weights of a neural network.
[0101] The methods and devices described herein can be applied to other non-volatile memory technologies, such as but not limited to SONOS (silicon-oxide-nitride-oxide-silicon, charge trapping in nitride), MONOS (metal-oxide-nitride-oxide-silicon, metal charge trapping in nitride), ReRAM (resistive ram), PCM (phase change memory), MRAM (magnetic ram), FeRAM (ferroelectric ram), OTP (one-time programmable bilayer or multilayer), and CeRAM (correlated electron ram), etc. The methods and devices described herein can be applied to volatile memory technologies for neural networks, such as but not limited to SRAM, DRAM, and / or volatile synaptic units.
[0102] Neural Network Using Non-Volatile Memory Cell Array
[0103] Figure 8A non - limiting example of a neural network using a non - volatile memory array according to this embodiment is conceptually shown. This example uses the non - volatile memory array neural network for a face recognition application, but any other suitable application can also be implemented using a neural network based on a non - volatile memory array.
[0104] For this example, S0 is the input layer, which is a 32×32 pixel RGB image with 5 - bit precision (i.e., three 32×32 pixel arrays, one for each color R, G, and B, and each pixel is 5 - bit precision). The synapses CB1 from the input layer S0 to the layer C1 apply different sets of weights in some cases and shared weights in other cases, and scan the input image with a 3×3 pixel overlapping filter (kernel), shifting the filter by 1 pixel (or more than 1 pixel as indicated by the model). Specifically, the values of 9 pixels in a 3×3 portion of the image (i.e., what is called the filter or kernel) are provided to the synapses CB1, where these 9 input values are multiplied by appropriate weights, and after summing the output of this multiplication, a single output value is determined and provided by the first synapse of CB1 for a pixel in one of the layers C1 of the feature map. Then the 3×3 filter is shifted one pixel to the right in the input layer S0 (i.e., the column of three pixels on the right is added and the column of three pixels on the left is released), whereby the 9 pixel values in this newly positioned filter are provided to the synapses CB1, where they are multiplied by the same weights and a second single output value is determined by the associated synapse. This process continues until the 3×3 filter has scanned all three colors and all bits (precision values) of the entire 32×32 pixel image in the input layer S0. Then this process is repeated using different sets of weights to generate different feature maps of C1 until all the feature maps of layer C1 are calculated.
[0105] At layer C1, in this example, there are 16 feature maps, each with 30×30 pixels. Each pixel is a new feature pixel extracted from the product of the input and the kernel, so each feature map is a two - dimensional array, and thus in this example, layer C1 is composed of 16 layers of two - dimensional arrays (remember that the layers and arrays referred to herein are logical relationships and not necessarily physical relationships, i.e., the arrays do not have to be oriented as a physical two - dimensional array). Each of the 16 feature maps in layer C1 is generated by one of the sixteen different sets of synaptic weights applied to the filter scan. The C1 feature maps can all relate to different aspects of the same image feature, such as boundary recognition. For example, the first map (generated using the first set of weights, shared for all scans used to generate this first map) can identify circular edges, the second map (generated using a second set of weights different from the first set) can identify rectangular edges, or the aspect ratio of certain features, and so on.
[0106] Before transitioning from layer C1 to layer S1, an activation function P1 (pooling) is applied, which pools the values from consecutive non-overlapping 2×2 regions in each feature map. The purpose of the pooling function is to take the mean (or alternatively the max function) of neighboring locations to, for example, reduce the dependence on edge locations and reduce the data size before entering the next stage. At layer S1, there are 16 feature maps of 15×15 (i.e., sixteen different arrays of 15×15 pixels per feature map). The synapse CB2 from layer S1 to layer C2 scans the maps in S1 using a 4×4 filter, where the filter is shifted by 1 pixel. At layer C2, there are 22 feature maps of 12×12. Before transitioning from layer C2 to layer S2, an activation function P2 (pooling) is applied, which pools the values from consecutive non-overlapping 2×2 regions in each feature map. At layer S2, there are 22 feature maps of 6×6. An activation function (pooling) is applied to the synapse CB3 from layer S2 to layer C3, where each neuron in layer C3 is connected via the corresponding synapse of CB3 to each map in layer S2. At layer C3, there are 64 neurons. The synapse CB4 from layer C3 to the output layer S3 fully connects C3 to S3, i.e., each neuron in layer C3 is connected to each neuron in layer S3. The output at S3 includes 10 neurons, where the neuron with the highest output determines the class. For example, this output can indicate the recognition or classification of the content of the original image.
[0107] Each layer's synapse is implemented using an array of non-volatile memory cells or a portion of the array.
[0108] Fig. 9 is a block diagram of a system that can be used for this purpose. The VMM system 32 includes non-volatile memory cells and serves as the synapse between one layer and the next (such as Figure 6 CB1, CB2, CB3, and CB4 in ). Specifically, the VMM system 32 includes a VMM array 33 (including non-volatile memory cells arranged in rows and columns), an erase gate and word line decoder 34, a control gate decoder 35, a bit line decoder 36, and a source line decoder 37, which decode the corresponding inputs to the non-volatile memory cell array 33. The input to the VMM array 33 can come from the erase gate and word line decoder 34 or from the control gate decoder 35. In this example, the source line decoder 37 also decodes the output of the VMM array 33. Alternatively, the bit line decoder 36 can decode the output of the VMM array 33.
[0109] The VMM array 33 serves two purposes. First, it stores the weights to be used by the VMM system 32. Second, the VMM array 33 effectively multiplies the inputs with the weights stored in the VMM array 33 and each output line (source line or bit line) sums them up to produce an output, which will be used as the input for the next layer or the final layer. By performing the multiplication and addition functions, the VMM array 33 eliminates the need for separate multiplication and addition logic circuits and is also highly efficient due to its in-situ memory calculation.
[0110] The output of the VMM array 33 is provided to a differential summing device (such as a summing operational amplifier or a summing current mirror) 38, which sums up the output of the VMM array 33 to create a single value for this convolution. The differential summing device 38 is arranged to perform the summation of both positive and negative weight inputs to output a single value.
[0111] Then the output value of the differential summing device 38 is provided to an activation function circuit 39 after summation, and this activation function circuit corrects the output. The activation function circuit 39 can provide sigmoid, tanh, ReLU functions or any other non-linear function. The corrected output value of the activation function circuit 39 becomes an element of the feature map for the next layer (e.g., Figure 8 layer C1 in ) and is then applied to the next synapse to generate the next feature map layer or the final layer. Thus, in this example, the VMM array 33 constitutes multiple synapses (which receive their inputs from an existing neuron layer or from an input layer such as an image database), and the summing device 38 and the activation function circuit 39 constitute multiple neurons.
[0112] Fig. 9 The inputs (WLx, EGx, CGx and optionally BLx and SLx) to the VMM system 32 in can be analog levels, binary levels, digital pulses (in which case a pulse - analog converter PAC may be required to convert the pulses to a suitable input analog level) or digital bits (in which case a DAC is provided to convert the digital bits to a suitable input analog level); the output can be an analog level, binary level, digital pulse or digital bit (in which case an output ADC is provided to convert the output analog level to digital bits).
[0113] Fig.10 A block diagram showing the use of a multi-layer VMM system 32 (here labeled as VMM systems 32a, 32b, 32c, 32d and 32e). As Fig.10As shown, the input (denoted as Inputx) is converted from digital to analog by the digital-to-analog converter 31 and provided to the input VMM system 32a. The converted analog input can be either a voltage or a current. The input D / A conversion of the first layer can be accomplished by using a function or a LUT (look-up table) that maps the input Inputx to an appropriate analog level of a matrix multiplier for the input VMM system 32a. The input conversion can also be done by an analog-to-analog (A / A) converter to convert an external analog input into a mapped analog input to the input VMM system 32a. The input conversion can also be done by a digital-to-digital pulse (D / P) converter to convert an external digital input into one or more mapped digital pulses to the input VMM system 32a.
[0114] The output generated by the input VMM system 32a is provided as an input to the next VMM system (hidden level 1) 32b, which in turn generates an output provided as an input to the next VMM system (hidden level 2) 32c, and so on. Each layer of the VMM system 32 serves as different layers of synapses and neurons of a convolutional neural network (CNN). Each of the VMM systems 32a, 32b, 32c, 32d, and 32e can be an independent physical system including a corresponding non-volatile memory array, or multiple VMM systems can utilize different parts of the same physical non-volatile memory array, or multiple VMM systems can utilize overlapping parts of the same physical non-volatile memory array. Each of the VMM systems 32a, 32b, 32c, 32d, and 32e can also be time-division multiplexed for different parts of its array or neurons. Fig.10 The example shown includes five layers (32a, 32b, 32c, 32d, 32e): an input layer (32a), two hidden layers (32b, 32c), and two fully connected layers (32d, 32e). Those of ordinary skill in the art will know that this is merely exemplary, and conversely, the system can include more than two hidden layers and more than two fully connected layers.
[0115] VMM Array
[0116] Fig.11 A neuron VMM array 1100 is shown, which is particularly applicable to Figure 3 the memory cell 310 shown and serves as the synapses and components of neurons between the input layer and the next layer. The VMM array 1100 includes a memory array 1101 of non-volatile memory cells and a reference array 1102 of non-volatile reference memory cells (at the top of the array). Alternatively, another reference array can be placed at the bottom.
[0117] In the VMM array 1100, control gate lines (such as control gate line 1103) extend in the vertical direction (thus the reference array 1102 is orthogonal to the control gate line 1103 in the row direction), and erase gate lines (such as erase gate line 1104) extend in the horizontal direction. Here, the inputs of the VMM array 1100 are set on the control gate lines (CG0, CG1, CG2, CG3), and the outputs of the VMM array 1100 appear on the source lines (SL0, SL1). In one embodiment, only even rows are used, and in another embodiment, only odd rows are used. The current placed on each source line (SL0 and SL1 respectively) performs a summation function of all the currents from the memory cells connected to that particular source line.
[0118] As described herein for neural networks, the non-volatile memory cells of the VMM array 1100 (i.e., the flash memory of the VMM array 1100) are preferably configured to operate in the subthreshold region.
[0119] Biasing the non-volatile reference memory cells and non-volatile memory cells described herein in weak inversion:
[0120] Ids = Io*e (Vg-Vth) / nVt = w*Io*e (Vg) / nVt ,
[0121] where w = e (-Vth) / nVt
[0122] where Ids is the drain-to-source current; Vg is the gate voltage on the memory cell; Vth is the threshold voltage of the memory cell; Vt is the thermal voltage = k*T / q, where k is the Boltzmann constant, T is the temperature in Kelvin, and q is the electron charge; n is the slope factor = 1+(Cdep / Cox), where Cdep = the capacitance of the depletion layer, and Cox is the capacitance of the gate oxide layer; Io is the memory cell current at the gate voltage equal to the threshold voltage, and Io is proportional to (Wt / L)*u*Cox*(n - 1)*Vt 2 is proportional to, where u is the carrier mobility, and Wt and L are the width and length of the memory cell respectively.
[0123] For an I-to-V logarithmic converter that uses a memory cell (such as a reference memory cell or a peripheral memory cell) or a transistor to convert an input current Ids to an input voltage Vg:
[0124] Vg = n*Vt*log[Ids / wp*Io]
[0125] Here, wp is the w of the reference memory cell or the peripheral memory cell.
[0126] For an I-to-V logarithmic converter that uses a memory cell (such as a reference memory cell or a peripheral memory cell) or a transistor to convert an input current Ids to an input voltage Vg:
[0127] Vg = n * Vt * log[Ids / wp * Io]
[0128] Here, wp is the w of the reference memory cell or the peripheral memory cell.
[0129] For a memory array used as a vector matrix multiplier (VMM) array, the output current is:
[0130] Iout = wa * Io * e (Vg) / nVt , that is
[0131] Iout = (wa / wp) * Iin = W * Iin
[0132] W = e (Vthp-Vtha) / nVt
[0133] Iin = wp * Io * e (Vg) / nVt
[0134] Here, wa = the w of each memory cell in the memory array.
[0135] The word line or the control gate can be used as the input of the memory cell for the input voltage.
[0136] Alternatively, the non-volatile memory cells of the VMM array described herein can be configured to operate in the linear region:
[0137] Ids = β * (Vgs - Vth) * Vds; β = u * Cox * Wt / L,
[0138] W α (Vgs - Vth),
[0139] which means that the weight W in the linear region is proportional to (Vgs - Vth)
[0140] The word line or the control gate or the bit line or the source line can be used as the input of the memory cell operating in the linear region. The bit line or the source line can be used as the output of the memory cell.
[0141] For an I-to-V linear converter, a memory cell (such as a reference memory cell or a peripheral memory cell) or a transistor or a resistor operating in the linear region can be used to linearly convert the input / output current to the input / output voltage.
[0142] Alternatively, the memory cells of the VMM array described herein can be configured to operate in the saturation region:
[0143] Ids = 1 / 2 * β * (Vgs - Vth)2 ; β = u * Cox * Wt / L
[0144] Wα(Vgs - Vth) 2 , meaning the weight W is proportional to (Vgs - Vth) 2 is proportional
[0145] The word line, control gate or erase gate can be used as the input of a memory cell operating in the saturation region. The bit line or source line can be used as the output of the output neuron.
[0146] Alternatively, the memory cells of the VMM array described herein can be used in all regions or combinations thereof (subthreshold, linear or saturation regions).
[0147] Other embodiments of the VMM array 33 described in U.S. Patent Application No. 15 / 826,345 are incorporated herein by reference. As described herein, the source line or bit line can be used as the neuron output (current summing output). Fig. 9 is incorporated herein by reference. As described herein, the source line or bit line can be used as the neuron output (current summing output).
[0148] Fig.12 shows a neuron VMM array 1200, which is particularly suitable for Figure 2 the memory cell 210 shown, and serves as a synapse between the input layer and the next layer. The VMM array 1200 includes a memory array 1203 of non - volatile memory cells, a reference array 1201 of first non - volatile reference memory cells, and a reference array 1202 of second non - volatile reference memory cells. The reference arrays 1201 and 1202 arranged along the column direction of the array are used to convert the current inputs flowing into the terminals BLR0, BLR1, BLR2, and BLR3 into voltage inputs WL0, WL1, WL2, and WL3. In fact, the first non - volatile reference memory cells and the second non - volatile reference memory cells are diode - connected through a multiplexer 1214 (only partially shown), into which the current inputs flow. The reference cells are tuned (e.g., programmed) to a target reference level. The target reference level is provided by a reference micro - array matrix (not shown).
[0149] The memory array 1203 serves two purposes. First, it stores the weights to be used by the VMM array 1200 on its respective memory cells. Second, the memory array 1203 effectively multiplies the inputs (i.e., the current inputs provided at terminals BLR0, BLR1, BLR2, and BLR3, which are converted by reference arrays 1201 and 1202 into input voltages to be provided to word lines WL0, WL1, WL2, and WL3) by the weights stored in the memory array 1203 and then sums all the results (memory cell currents) to produce an output on the respective bit lines (BL0 - BLN), which will be the input to the next layer or the input to the final layer. By performing the multiplication and addition functions, the memory array 1203 eliminates the need for separate multiplication logic circuitry and addition logic circuitry and is also highly efficient. Here, the voltage inputs are provided on the word lines (WL0, WL1, WL2, and WL3), and the outputs appear on the respective bit lines (BL0 - BLN) during a read (inference) operation. The current placed on each of the bit lines in BL0 - BLN performs a summing function of the currents from all the non-volatile memory cells connected to that particular bit line.
[0150] Table 5 shows the operating voltages for the VMM array 1200. The columns in the table indicate the voltages placed on the word lines for selected cells, the word lines for unselected cells, the bit lines for selected cells, the bit lines for unselected cells, the source lines for selected cells, and the source lines for unselected cells, where FLT indicates floating, i.e., no voltage is applied. The rows indicate the read, erase, and program operations.
[0151] Table 5: Fig.12 Operation of the VMM array 1200
[0152] WL WL-Not selected BL BL-Not selected SL SL-Not selected Read 0.5-3.5V -0.5V / 0V 0.1-2V(Ineuron) 0.6V-2V / FLT 0V 0V Erase About 5-13V 0V 0V 0V 0V 0V programming 1V-2V -0.5V / 0V 0.1-3uA Vinh about 2.5V 4-10V 0-1V / FLT
[0153] Fig.13 A neuron VMM array 1300 is shown, which is particularly applicable to Figure 2The memory cell 210 shown is used as a synapse and component of a neuron between the input layer and the next layer. The VMM array 1300 includes a memory array 1303 of non-volatile memory cells, a reference array 1301 of first non-volatile reference memory cells, and a reference array 1302 of second non-volatile reference memory cells. The reference arrays 1301 and 1302 extend in the row direction of the VMM array 1300. The VMM array is similar to the VMM 1000, except that in the VMM array 1300, the word lines extend in the vertical direction. Here, the inputs are set on the word lines (WLA0, WLB0, WLA1, WLB2, WLA2, WLB2, WLA3, WLB3), and the outputs appear on the source lines (SL0, SL1) during a read operation. The current placed on each source line performs a summation function of all the currents from the memory cells connected to that particular source line.
[0154] Table 6 shows the operating voltages for the VMM array 1300. The columns in the table indicate the voltages placed on the word lines for selected cells, the word lines for unselected cells, the bit lines for selected cells, the bit lines for unselected cells, the source lines for selected cells, and the source lines for unselected cells. The rows indicate read, erase, and program operations.
[0155] Table 6: Fig.13 Operation of the VMM array 1300
[0156] WL WL-Not selected BL BL-Not selected SL SL-Not selected Read 0.5-3.5V -0.5V / 0V 0.1-2V 0.1V-2V / FLT About 0.3-1V (Ineuron) 0V Erase About 5-13V 0V 0V 0V 0V SL-Prohibit (about 4-8V) programming 1V-2V -0.5V / 0V 0.1-3uA Vinh about 2.5V 4-10V 0-1V / FLT
[0157] Fig.14 The neuron VMM array 1400 is shown, which is particularly suitable for Figure 3 the memory cell 310 shown and is used as a synapse and component of a neuron between the input layer and the next layer. The VMM array 1400 includes a memory array 1403 of non-volatile memory cells, a reference array 1401 of first non-volatile reference memory cells, and a reference array 1402 of second non-volatile reference memory cells. The reference arrays 1401 and 1402 are used to convert the current inputs flowing into the terminals BLR0, BLR1, BLR2, and BLR3 into voltage inputs CG0, CG1, CG2, and CG3. In fact, the first non-volatile reference memory cells and the second non-volatile reference memory cells are diode-connected through multiplexers 1412 (only partially shown), where the current inputs flow in through BLR0, BLR1, BLR2, and BLR3. Each multiplexer 1412 includes a corresponding multiplexer 1405 and a cascode transistor 1404 to ensure a constant voltage on the bit line (such as BLR0) of each of the first non-volatile reference memory cells and the second non-volatile reference memory cells during a read operation. The reference cells are tuned to a target reference level.
[0158] The memory array 1403 serves two purposes. First, it stores the weights to be used by the VMM array 1400. Second, the memory array 1403 effectively multiplies the inputs (the current inputs provided to terminals BLR0, BLR1, BLR2, and BLR3, which the reference arrays 1401 and 1402 convert into input voltages to be provided to control gates CG0, CG1, CG2, and CG3) by the weights stored in the memory array, and then sums all the results (cell currents) to produce an output that appears at BL0 - BLN and will be the input to the next layer or the input to the final layer. By performing the multiplication and addition functions, the memory array eliminates the need for separate multiplication and addition logic circuits and is also highly efficient. Here, the inputs are provided on the control gate lines (CG0, CG1, CG2, and CG3), and the outputs appear on the bit lines (BL0–BLN) during a read operation. The current placed on each bit line performs the summation function of all the currents from the memory cells connected to that particular bit line.
[0159] The VMM array 1400 implements unidirectional tuning for the non-volatile memory cells in the memory array 1403. That is, each non-volatile memory cell is erased and then partially programmed until the desired charge on the floating gate is reached. This can be performed, for example, using the precise programming techniques described below. If too much charge is placed on the floating gate (such that an incorrect value is stored in the cell), the cell must be erased and the sequence of partial programming operations must be restarted. As shown, two rows sharing the same erase gate (such as EG0 or EG1) need to be erased together (which is called a page erase), and thereafter, each cell is partially programmed until the desired charge on the floating gate is reached.
[0160] Table 7 shows the operating voltages for the VMM array 1400. The columns in the table indicate the voltages placed on the word line for the selected cell, the word line for the unselected cell, the bit line for the selected cell, the bit line for the unselected cell, the control gate for the selected cell, the control gate for the unselected cell in the same sector as the selected cell, the control gate for the unselected cell in a different sector from the selected cell, the erase gate for the selected cell, the erase gate for the unselected cell, the source line for the selected cell, and the source line for the unselected cell. The rows indicate the read, erase, and program operations.
[0161] Table 7: Fig.14 Operation of the VMM array 1400
[0162]
[0163] Fig.15 The neuron VMM array 1500 is shown, which is particularly suitable for Figure 3The memory cell 310 shown and serves as a synapse and component of a neuron between the input layer and the next layer. The VMM array 1500 includes a memory array 1503 of non-volatile memory cells, a reference array 1501 of first non-volatile reference memory cells, and a reference array 1502 of second non-volatile reference memory cells. The EG lines EGR0, EG0, EG1, and EGR1 extend vertically, while the CG lines CG0, CG1, CG2, and CG3 and the SL lines WL0, WL1, WL2, and WL3 extend horizontally. The VMM array 1500 is similar to the VMM array 1400, except that the VMM array 1500 implements bidirectional tuning, where each individual cell can be fully erased, partially programmed, and partially erased as needed to achieve a desired charge amount on the floating gate due to the use of separate EG lines. As shown, the reference arrays 1501 and 1502 convert the input current in the terminals BLR0, BLR1, BLR2, and BLR3 into control gate voltages CG0, CG1, CG2, and CG3 to be applied to the memory cells in the row direction (through the action of diode-connected reference cells via the multiplexer 1514). The current output (neuron) is in the bit lines BL0 - BLN, where each bit line sums all the currents from the non-volatile memory cells connected to that specific bit line.
[0164] Table 8 shows the operating voltages for the VMM array 1500. The columns in the table indicate the voltages placed on the word line for the selected cell, the word line for the unselected cell, the bit line for the selected cell, the bit line for the unselected cell, the control gate for the selected cell, the control gate for the unselected cell in the same sector as the selected cell, the control gate for the unselected cell in a different sector from the selected cell, the erase gate for the selected cell, the erase gate for the unselected cell, the source line for the selected cell, and the source line for the unselected cell. The rows indicate read, erase, and program operations.
[0165] Table 8: Fig.15 Operation of the VMM array 1500
[0166]
[0167] Fig.16 Shows the neuron VMM array 1600, which is particularly suitable for Figure 2 the memory cell 210 shown and serves as a synapse and component of a neuron between the input layer and the next layer. In the VMM array 1600, the inputs INPUT0...., INPUT N are received on the bit lines BL0,…BL N respectively, and the outputs OUTPUT1, OUTPUT2, OUTPUT3, and OUTPUT4 are generated on the source lines SL0, SL1, SL2, and SL3 respectively.
[0168] Fig.17 Shows neuron VMM array 1700, which is particularly suitable for Figure 2 the memory cell 210 shown, and serves as the synapse and component of the neuron between the input layer and the next layer. In this example, inputs INPUT0, INPUT1, INPUT2, and INPUT3 are received on source lines SL0, SL1, SL2, and SL3 respectively, and outputs OUTPUT0, … OUTPUT N are generated on bit lines BL0, …, BL N .
[0169] Fig.18 Shows neuron VMM array 1800, which is particularly suitable for Figure 2 the memory cell 210 shown, and serves as the synapse and component of the neuron between the input layer and the next layer. In this example, inputs INPUT0, …, INPUT M are received on word lines WL0, …, WL M respectively, and outputs OUTPUT0, … OUTPUT N are generated on bit lines BL0, …, BL N .
[0170] Fig.19 Shows neuron VMM array 1900, which is particularly suitable for Figure 3 the memory cell 310 shown, and serves as the synapse and component of the neuron between the input layer and the next layer. In this example, inputs INPUT0, …, INPUT M are received on word lines WL0, …, WL M respectively, and outputs OUTPUT0, … OUTPUT N are generated on bit lines BL0, …, BL N .
[0171] Fig. 20 Shows neuron VMM array 2000, which is particularly suitable for Figure 4 the memory cell 410 shown, and serves as the synapse and component of the neuron between the input layer and the next layer. In this example, inputs INPUT 0, …, INPUT n are received on vertical control gate lines CG0, …, CG N respectively, and outputs OUTPUT1 and OUTPUT2 are generated on source lines SL0 and SL1.
[0172] Fig.21Shows a neuron VMM array 2100, which is particularly suitable for Figure 4 the memory cell 410 shown, and serves as the synapse and component of the neuron between the input layer and the next layer. In this example, the inputs INPUT0 to INPUT N are received on the gates of bit line control gates 2901-1, 2901-2 to 2901-(N-1) and 2901-N respectively, and these gates are coupled to bit lines BL0 to BL N respectively. Exemplary outputs OUTPUT1 and OUTPUT2 are generated on source lines SL0 and SL1.
[0173] Fig. 22 Shows a neuron VMM array 2200, which is particularly suitable for Figure 3 the memory cell 310 shown, Figure 5 the memory cell 510 shown, and Figure 7 the memory cell 710 shown, and serves as the synapse and component of the neuron between the input layer and the next layer. In this example, the inputs INPUT0, …, INPUT M are received on word lines WL0, …, WL M and the outputs OUTPUT0, …, OUTPUT N are generated on bit lines BL0, …, BL N respectively.
[0174] Fig.23 Shows a neuron VMM array 2300, which is particularly suitable for Figure 3 the memory cell 310 shown, Figure 5 the memory cell 510 shown, and Figure 7 the memory cell 710 shown, and serves as the synapse and component of the neuron between the input layer and the next layer. In this example, the inputs INPUT0 to INPUT M are received on control gates CG0 to CG M and the outputs OUTPUT0, …, OUTPUT N are generated on vertical source lines SL0, …, SL N respectively, where each source line SL i is coupled to the source lines of all memory cells in column i.
[0175] Fig.24 Shows a neuron VMM array 2400, which is particularly suitable for Figure 3 the memory cell 310 shown, Figure 5 the memory cell 510 shown, and Figure 7The memory cell 710 shown and serves as the synapse and component of the neuron between the input layer and the next layer. In this example, the inputs INPUT0 to INPUT M are received on the control gate lines CG0 to CG M The outputs OUTPUT0, …, OUTPUT N are generated on the vertical bit lines BL0, …, BL N respectively, where each bit line BL i is coupled to the bit lines of all the memory cells in column i.
[0176] Long Short-Term Memory
[0177] The prior art includes the concept known as long short-term memory (LSTM). LSTM is commonly used in artificial neural networks. LSTM allows an artificial neural network to remember information over an arbitrary predetermined time interval and use that information in subsequent operations. Conventional LSTM includes cells, an input gate, an output gate, and a forget gate. The three gates regulate the flow of information into and out of the cell and the time interval during which information is remembered in the LSTM. VMM is particularly useful in LSTM.
[0178] Fig.25 An exemplary LSTM 2500 is shown. The LSTM 2500 in this example includes cells 2501, 2502, 2503, and 2504. Cell 2501 receives the input vector x0 and generates the output vector h0 and the cell state vector c0. Cell 2502 receives the input vector x1, the output vector (hidden state) h0 from cell 2501, and the cell state c0 from cell 2501, and generates the output vector h1 and the cell state vector c1. Cell 2503 receives the input vector x2, the output vector (hidden state) h1 from cell 2502, and the cell state c1 from cell 2502, and generates the output vector h2 and the cell state vector c2. Cell 2504 receives the input vector x3, the output vector (hidden state) h2 from cell 2503, and the cell state c2 from cell 2503, and generates the output vector h3. Additional cells can be used, and the LSTM with four cells is just an example.
[0179] Fig.26 An exemplary embodiment of the LSTM cell 2600 that can be used for Fig.25 the cells 2501, 2502, 2503, and 2504 is shown. The LSTM cell 2600 receives the input vector x(t), the cell state vector c(t−1) from the previous cell, and the output vector h(t−1) from the previous cell, and generates the cell state vector c(t) and the output vector h(t).
[0180] The LSTM cell 2600 includes sigmoid function devices 2601, 2602, and 2603, each of which applies a number between 0 and 1 to control the amount of each component in the input vector that is allowed to pass through to the output vector. The LSTM cell 2600 also includes tanh devices 2604 and 2605 for applying the hyperbolic tangent function to the input vector, multiplier devices 2606, 2607, and 2608 for multiplying two vectors together, and adder device 2609 for adding two vectors together. The output vector h(t) can be provided to the next LSTM cell in the system, or it can be accessed for other purposes.
[0181] Fig. 27 FIG. shows an LSTM cell 2700, which is an example of a specific implementation of the LSTM cell 2600. For the convenience of the reader, the same numbers are used in the LSTM cell 2700 as in the LSTM cell 2600. The sigmoid function devices 2601, 2602, and 2603 and the tanh device 2604 each include a plurality of VMM arrays 2701 and activation circuit blocks 2702. Thus, it can be seen that the VMM array is particularly useful in LSTM cells used in some neural network systems.
[0182] An alternative form of the LSTM cell 2700 (and another example of a specific implementation of the LSTM cell 2600) is shown in Fig.28 In Fig.28 the sigmoid function devices 2601, 2602, and 2603 and the tanh device 2604 share the same physical hardware (VMM array 2801 and activation function block 2802) in a time-division multiplexing manner. The LSTM cell 2800 also includes a multiplier device 2803 for multiplying two vectors together, an adder device 2808 for adding two vectors together, a tanh device 2605 (which includes an activation circuit block 2802), a register 2807 for storing the value i(t) when the value i(t) is output from the sigmoid function block 2802, a register 2804 for storing the value when the value f(t)*c(t-1) is output from the multiplier device 2803 through the multiplexer 2810, a register 2805 for storing the value when the value i(t)*u(t) is output from the multiplier device 2803 through the multiplexer 2810, a register 2806 for storing the value when the value o(t)*c~(t) is output from the multiplier device 2803 through the multiplexer 2810, and a multiplexer 2809.
[0183] The LSTM cell 2700 includes multiple groups of VMM arrays 2701 and corresponding activation function blocks 2702, while the LSTM cell 2800 only includes one group of VMM arrays 2801 and activation function blocks 2802, which are used to represent multiple layers in the implementation of the LSTM cell 2800. The LSTM cell 2800 will require less space than the LSTM 2700 because, compared with the LSTM cell 2700, the LSTM cell 2800 only needs 1 / 4 of its space for the VMM and activation function blocks.
[0184] It can also be understood that an LSTM cell generally includes multiple VMM arrays, and each VMM array requires functions provided by certain circuit blocks outside the VMM array (such as summing and activation circuit blocks and high-voltage generation blocks). Providing separate circuit blocks for each VMM array will require a large amount of space within the semiconductor device and will be somewhat inefficient. Therefore, the embodiments described below attempt to minimize the circuits required outside the VMM array itself.
[0185] Gate Controlled Recursive Unit
[0186] The analog VMM implementation can be used for GRUs (gated recurrent units). A GRU is a gated mechanism in a recurrent artificial neural network. A GRU is similar to an LSTM, except that a GRU cell generally includes fewer components than an LSTM cell.
[0187] Fig.29 An exemplary GRU 2900 is shown. The GRU 2900 in this example includes cells 2901, 2902, 2903, and 2904. The cell 2901 receives the input vector x0 and generates the output vector h0. The cell 2902 receives the input vector x1, the output vector h0 from the cell 2901, and generates the output vector h1. The cell 2903 receives the input vector x2 and the output vector (hidden state) h1 from the cell 2902 and generates the output vector h2. The cell 2904 receives the input vector x3 and the output vector (hidden state) h2 from the cell 2903 and generates the output vector h3. Additional cells can be used, and a GRU with four cells is just an example.
[0188] Fig.30 Shown can be used for Fig.29Exemplary specific implementations of GRU unit 3000 for units 2901, 2902, 2903, and 2904. The GRU unit 3000 receives an input vector x(t) and an output vector h(t - 1) from a previous GRU unit and generates an output vector h(t). The GRU unit 3000 includes sigmoid function devices 3001 and 3002, each of which applies a number between 0 and 1 to components from the output vector h(t - 1) and the input vector x(t). The GRU unit 3000 also includes a tanh device 3003 for applying the hyperbolic tangent function to the input vector, multiple multiplier devices 3004, 3005, and 3006 for multiplying two vectors together, an adder device 3007 for adding two vectors together, and a complementary device 3008 for subtracting the input from 1 to generate the output.
[0189] Fig.31 FIG. shows GRU unit 3100, which is an example of a specific implementation of GRU unit 3000. For the convenience of the reader, the same numbers are used in GRU unit 3100 as in GRU unit 3000. As Fig.31 shown, the sigmoid function devices 3001 and 3002 and the tanh device 3003 each include a plurality of VMM arrays 3101 and activation function blocks 3102. Thus, it can be seen that VMM arrays are particularly useful in GRU units used in certain neural network systems.
[0190] An alternative form of GRU unit 3100 (and another example of a specific implementation of GRU unit 3000) is shown in Fig.32 In Fig.32 In, GRU unit 3200 utilizes a VMM array 3201 and an activation function block 3202, which, when configured as a sigmoid function, applies a number between 0 and 1 to control how much of each component in the input vector is allowed to pass through to the output vector. In Fig.32In it, the sigmoid function devices 3001 and 3002 and the tanh device 3003 share the same physical hardware (the VMM array 3201 and the activation function block 3202) in a time-division multiplexing manner. The GRU unit 3200 also includes a multiplier device 3203 that multiplies two vectors together, an adder device 3205 that adds two vectors together, a complementary device 3209 that subtracts the input from 1 to generate an output, a multiplexer 3204, a register 3206 that holds the value when the value h(t - 1)*r(t) is output from the multiplier device 3203 through the multiplexer 3204, a register 3207 that holds the value when the value h(t - 1)*z(t) is output from the multiplier device 3203 through the multiplexer 3204, and a register 3208 that holds the value when the value h^(t)*(1 - z(t)) is output from the multiplier device 3203 through the multiplexer 3204.
[0191] The GRU unit 3100 includes multiple sets of VMM arrays 3101 and activation function blocks 3102, while the GRU unit 3200 only includes a single set of VMM arrays 3201 and activation function blocks 3202, which are used to represent multiple layers in the implementation of the GRU unit 3200. The GRU unit 3200 will require less space than the GRU unit 3100 because, compared to the GRU unit 3100, the GRU unit 3200 only needs 1 / 3 of its space for the VMM and activation function blocks.
[0192] It can also be understood that a system using GRUs will generally include multiple VMM arrays, and each VMM array requires functions provided by certain circuit blocks outside the VMM array (such as summing and activation circuit blocks and high-voltage generation blocks). Providing separate circuit blocks for each VMM array will require a large amount of space within the semiconductor device and will be somewhat inefficient. Therefore, the embodiments described below attempt to minimize the circuits required outside the VMM array itself.
[0193] The input to the VMM array can be an analog level, a binary level, a timing pulse, or digital bits, and the output can be an analog level, a binary level, a timing pulse, or digital bits (in which case, an output ADC is required to convert the output analog level current or voltage into digital bits).
[0194] For each memory cell in the VMM array, each weight w can be implemented by a single memory cell, or by a differential cell, or by two hybrid memory cells (the average of 2 or more cells). In the case of a differential cell, two memory cells are required to implement the weight w as a differential weight (w = w+ – w-). In two hybrid memory cells, two memory cells are required to implement the weight w as the average of the two cells.
[0195] Implementation for fine tuning units in a VMM
[0196] Fig.33 A block diagram of the VMM system 3300 is shown. The VMM system 3300 includes a VMM array 3301, a row decoder 3302, a high-voltage decoder 3303, a column decoder 3304, a bit-line driver 3305, an input circuit 3306, an output circuit 3307, a control logic component 3308, and a bias generator 3309. The VMM system 3300 further includes a high-voltage generation block 3310, which includes a charge pump 3311, a charge pump regulator 3312, and a high-voltage level generator 3313. The VMM system 3300 further includes an algorithm controller 3314, an analog circuit 3315, a control logic component 3316, and a test control logic component 3317. The systems and methods described below can be implemented in the VMM system 3300.
[0197] The input circuit 3306 can include circuits such as a DAC (digital-to-analog converter), a DPC (digital-to-pulse converter), an AAC (analog-to-analog converter, such as a current-to-voltage converter), a PAC (pulse-to-analog level converter), or any other type of converter. The input circuit 3306 can implement a normalization, a scaling function, or an arithmetic function. The input circuit 3306 can implement a temperature compensation function for the input. The input circuit 3306 can implement an activation function, such as a ReLU or sigmoid function.
[0198] The output circuit 3307 can include circuits such as an ADC (analog-to-digital converter for converting the neuron analog output to digital bits), an AAC (analog-to-analog converter, such as a current-to-voltage converter), an APC (analog-to-pulse converter), or any other type of converter. The output circuit 3307 can implement an activation function, such as a ReLU or sigmoid function. The output circuit 3307 can implement a normalization, a scaling function, or an arithmetic function for the neuron output. The output circuit 3307 can implement a temperature compensation function for the neuron output or the array output (such as the bit-line output), as described below.
[0199] Fig.34Shows a tuning correction method 3400, which can be executed by the algorithm controller 3314 in the VMM system 3300. The tuning correction method 3400 generates an adaptive target based on the final error generated by the cell output and the cell initial target. The method generally starts in response to receiving a tuning command (step 3401). The initial current target (for the programming / verification algorithm) of the selected cell or selected cell group Itargetv(i) is determined using a predictive target model (such as by using a function or a look-up table), and the variable DeltaError is set to 0 (step 3402). The target function (if used) will be based on the I-V programming curve of the selected memory cell or cell group. The target function also depends on various variations caused by the array characteristics, such as the degree of programming interference exhibited by the cells (which depends on the cell address and cell hierarchy within the sector, where if a cell exhibits relatively large interference, the cell undergoes more programming time under the suppression condition, and cells with higher currents generally have more interference), the coupling between cells, and various types of array noise. These variations for silicon can be characterized in terms of PVT (process, voltage, temperature). The look-up table (if used) can be characterized in the same way to simulate the I-V curve and various variations.
[0200] Then, a soft erase is performed on all cells in the VMM, which erases all cells to an intermediate weak erase level such that each cell will consume, for example, approximately 3 μA - 5 μA of current during a read operation (step 3403). For example, the soft erase is performed by applying an incremental erase pulse voltage to the cells until the intermediate cell current is reached. Next, a deep programming operation is performed on all unused cells (step 3404) to reach the <pA current level. Then, a target adjustment (correction) based on the error result is performed. If DeltaError > 0, meaning the cell has experienced overshoot during programming, then Itargetv(i + 1) is set to Itarget + θ * DeltaError, where θ is, for example, 1 or a number close to 1 (step 3405A).
[0201] Itarget(i + 1) can also be adjusted based on the appropriate error target adjustment / correction with respect to the previous Itarget(i). If DeltaError < 0, meaning the cell has experienced undershoot during programming, which means the cell current has not reached the target, then Itargetv(i + 1) is set to the previous target Itargetv(i) (step 3405B).
[0202] Next, perform rough and / or fine programming and verification operations (step 3406). Multiple adaptive rough programming methods can be used to accelerate programming, such as by targeting multiple progressively smaller rough targets before performing the exact (fine) programming step. Adaptive exact programming is done, for example, by programming voltage pulses or constant programming timing pulses in fine (exact) increments. Embodiments of systems and methods for performing rough programming and fine programming are described in U.S. Provisional Patent Application No. 62 / 933,809, filed on November 11, 2019, and titled "Precise Programming Method and Apparatus for Analog Neural Memory in a Deep Learning Artificial Neural Network", which is assigned to the same assignee as the present application and is incorporated herein by reference.
[0203] Measure Icell in the selected cell (step 3407). For example, the cell current can be measured by a galvanometer circuit. For example, the cell current can be measured by an ADC (analog-to-digital converter) circuit, where in this case, the output is represented by digital bits. For example, the cell current can be measured by an I-V (current-voltage converter) circuit, where in this case, the output is represented by an analog voltage. Calculate DeltaError, i.e., Icell - Itarget, which represents the difference between the actual current (Icell) and the target current (Itarget) in the measured cell. If |DeltaError| < DeltaMargin, the cell has reached the target current within a certain tolerance (DeltaMargin), and the method ends (step 3410). |DeltaError| = abs(DeltaError) = the absolute value of DeltaError. If not, the method returns to step 3403 and the steps are executed again in sequence (step 3410).
[0204] Fig.35A and Fig.35B illustrates a tuning correction method 3500, which can be executed by the algorithm controller 3314 in the VMM system 3300. Refer to Fig.35A, the start of the method (step 3501) is typically performed in response to receiving a tuning command. The entire VMM array is erased, such as by a soft erase method (step 3502). A deep programming operation is performed on all unused cells (step 3503) to reach a cell current <pA level. All cells in the VMM array are programmed to an intermediate value, such as 0.5 μA - 1.0 μA, using coarse and / or fine programming loops (step 3504). Embodiments of systems and methods for performing coarse programming and fine programming are described in U.S. Provisional Patent Application No. 62 / 933,809, filed on November 11, 2019, and titled "Precise Programming Method and Apparatus for Analog Neural Memory in a Deep Learning Artificial Neural Network", which is assigned to the same assignee as the present application and is incorporated herein by reference. A prediction target for the used cells is set using the function or look-up table as described above (step 3505). Then, a sector tuning method 3507 is performed on each sector in the VMM (step 3506). A sector typically consists of two or more adjacent rows in the array.
[0205] Fig.35BShows the adaptive target sector tuning method 3507. All cells in the sector are programmed to the final desired value (e.g., 1 nA - 50 nA) using individual or combined programming / verification (P / V) methods such as: (1) rough / fine / constant P / V cycling; (2) CG+ (only CG increment) or EG+ (only EG increment) or complementary CG+ / EG- (CG increment and EG decrement); and (3) first programming the deepest programmed cells (such as progressive grouping, meaning dividing the cells into different groups, and the group of cells with the lowest current is programmed first) (step 3508A). Next, determine whether Icell < Itarget. If so, the method proceeds to step 3509. If not, the method repeats step 3508A. In step 3509, measure DeltaError, which is equal to the measured Icell - Itarget(i + 1) (step 3509). Determine whether |DeltaError| < DeltaMargin (step 3510). If so, the method is complete (step 3511). If not, perform target adjustment. If DeltaError > 0, meaning the cell has experienced overshoot during programming, adjust the target by setting the new target to Itarget + θ * DeltaError, where θ is typically = 1 (step 3512A). Itarget(i + 1) can also be adjusted based on the appropriate error target adjustment / correction of the previous Itarget(i). If DeltaError < 0, meaning the cell has experienced undershoot during programming, which means the cell has not reached the target, adjust the target by maintaining the previous target, i.e., Itargetv(i + 1) = Itargetv(i) (step 3512B). Perform a soft erase on the sector (step 3513). Program all cells in the sector to an intermediate value (step 3514), and return to step 3509.
[0206] A typical neural network can have positive weights w+ and negative weights w- and a combined weight = w+ - w-. w+ and w- are implemented by memory cells (Iw+ and Iw- respectively), and the combined weight (Iw = Iw+ - Iw-, current subtraction) can be performed at the peripheral circuit level (such as at the array bit line output circuit). Thus, the weight tuning implementation for the combined weight can include, for example, tuning the w+ cells and w- cells simultaneously, tuning only the w+ cells, or tuning only the w- cells, as shown in Table 8. Using the previous reference Fig.34 / Fig.35A / Fig.35BExecute tuning using the described programming / verification and error target adjustment methods. Verification can be performed only for the combined weights (e.g., measuring / reading the weight current of the combination rather than the individual positive w+ cell current or w- cell current), only for the w+ cell current, or only for the w- cell current.
[0207] For example, for a combined Iw of 3na, Iw+ can be 3na and Iw- can be 0na; or, Iw+ can be 13na and Iw- can be 10na, meaning that both the positive weight Iw+ and the negative weight Iw- are non-zero (e.g., where zero would represent a deeply programmed cell). Under certain operating conditions, this may be preferred as it will make both Iw+ and Iw- less susceptible to noise.
[0208] Table 8: Weight tuning methods
[0209] I Iw+ Iw- illustrate Initial goal 3na 3na 0na Tuning Iw+ and Iw- Initial goal -2na 0na 2na Tuning Iw+ and Iw- Initial goal 3na 13na 10na Tuning Iw+ and Iw- New goals 2na 12na 10na Tune Iw+ only New goals 2na 11na 9na Tuning Iw+ and Iw- New goals 4na 13na 9na Tune only Iw- New target 4na 12na 8na Tune Iw+ and Iw- New target -2na 8na 10na Tune Iw+ and Iw- New target -2na 7na 9na Tune Iw+ and Iw-
[0210] Figure 36A Show the data behavior (I-V curve) varying with temperature (e.g., in the subthreshold region), Figure 36B Show the problems caused by data drift during the operation of the VMM system, and Figure 36C and Figure 36D Show the block for compensating data drift and regarding Figure 36C show the block for compensating temperature variations.
[0211] Figure 36A Show the known characteristics of the VMM system as the operating temperature increases, where the sense current in any given selected non-volatile memory cell in the VMM array increases in the subthreshold region, decreases in the saturation region, or generally decreases in the linear region.
[0212] Figure 36B Show the array current distribution (data drift) over time, and it shows that the total output from the VMM array (which is the sum of the currents from all bit lines in the VMM array) shifts to the right (or left, depending on the technology used) as the operating time progresses, meaning that the total combined output will drift over the lifetime of the VMM system. This phenomenon is called data drift because the data drifts due to usage conditions and deteriorates due to environmental factors.
[0213] Figure 36C Show the bit line compensation circuit 3600, which can include a compensation current i COMPInject the output of the bitline output circuit 3610 to compensate for data drift. The bitline compensation circuit 3600 may include amplifying or scaling the output by a scaler circuit based on a resistor or capacitor network. The bitline compensation circuit 3600 may include shifting or offsetting the output by a shifter circuit based on its resistor or capacitor network.
[0214] Figure 36D A data drift monitor 3620 that shows the amount of data drift detected is shown. This information is then used as an input to the bitline compensation circuit 3600 so that an appropriate level of i COMP .
[0215] Figure 37 A bitline compensation circuit 3700 is shown, which is an embodiment of the bitline compensation circuit 3600 in FIG. 36. The bitline compensation circuit 3700 includes an adjustable current source 3701 and an adjustable current source 3702, which together generate i COMP , where i COMP is equal to the current generated by the adjustable current source 3701 minus the current generated by the adjustable current source 3701.
[0216] Figure 38 A bitline compensation circuit 3700 is shown, which is an embodiment of the bitline compensation circuit 3600 in FIG. 36. The bitline compensation circuit 3800 includes an operational amplifier 3801, an adjustable resistor 3802, and an adjustable resistor 3803. The operational amplifier 3801 receives a reference voltage VREF at its non-inverting terminal and receives V INPUT , where V INPUT is the voltage received from the bitline output circuit 3610 in Figure 36C , and generates an output V OUTPUT , where V OUTPUT is a scaled version of V INPUT to compensate for data drift based on the ratio of resistors 3803 and 3802. By configuring the values of resistors 3803 and / or 3802, V OUTPUT can be amplified or scaled.
[0217] Figure 39 A bitline compensation circuit 3900 is shown, which is an embodiment of the bitline compensation circuit 3600 in FIG. 36. The bitline compensation circuit 3900 includes an operational amplifier 3901, a current source 3902, a switch 3904, and an adjustable integrating output capacitor 3903. Here, the current source 3902 is actually the output current on a single bitline or a set of multiple bitlines (such as one for summing the positive weight w+ and one for summing the negative weight w-) in the VMM array. The operational amplifier 3901 receives a reference voltage VREF at its non-inverting terminal and receives V INPUT , where VINPUT is the voltage received from the bit line output circuit 3610 in Figure 36C . The bit line compensation circuit 3900 acts as an integrator that integrates the current Ineu through the capacitor 3903 over an adjustable integration time to generate the output voltage V OUTPUT , where V OUTPUT = Ineu * integration time / C 3903 , where C 3903 is the value of the capacitor 3903. Thus, the output voltage V OUTPUT is proportional to the (bit line) output current Ineu, proportional to the integration time, and inversely proportional to the capacitance of the capacitor 3903. The bit line compensation circuit 3900 generates the output V OUTPUT , where the value of V OUTPUT is scaled based on the configured value of the capacitor 3903 and / or the integration time to compensate for data drift.
[0218] Figure 40 shows a bit line compensation circuit 4000, which is an embodiment of the bit line compensation circuit 3600 in FIG. 36. The bit line compensation circuit 4000 includes a current mirror 4010 having an M:N ratio, which means I COMP = (M / N) * i input . The current mirror 4010 receives the current i INPUT , and mirrors the current and optionally scales the current to generate i COMP . Thus, by configuring the M parameter and / or the N parameter, i COMP can be amplified or reduced.
[0219] Figure 41 shows a bit line compensation circuit 4100, which is an embodiment of the bit line compensation circuit 3600 in FIG. 36. The bit line compensation circuit 4100 includes an operational amplifier 4101, an adjustable scaling resistor 4102, an adjustable shift resistor 4103, and an adjustable resistor 4104. The operational amplifier 4101 receives a reference voltage V REF at its non-inverting terminal and receives V IN at its inverting terminal. V IN is generated in response to V INPUT and Vshft, where V INPUT is the voltage received from the bit line output circuit 3610 in Figure 36C , and Vshft is the voltage intended to achieve a shift between V INPUT and V OUTPUT .
[0220] Thus, V OUTPUT is a scaled and shifted version of V INPUT to compensate for data drift.
[0221] Figure 42 Shows a bit line compensation circuit 4200, which is an embodiment of the bit line compensation circuit 3600 in FIG. 36. The bit line compensation circuit 4200 includes an operational amplifier 4201, an input current source Ineu 4202, a current shifter 4203, switches 4205 and 4206, and an adjustable integration output capacitor 4204. Here, the current source 4202 is actually the output current Ineu on a single bit line or multiple bit lines in the VMM array. The operational amplifier 4201 receives a reference voltage VREF at its non-inverting terminal and receives I IN , where I IN is the sum of Ineu and the current output by the current shifter 4203, and generates an output V OUTPUT , where V OUTPUT is scaled (based on the capacitor 4204) and shifted (based on Ishifter 4203) to compensate for data drift.
[0222] Figures 43 to 48 Shows various circuits that can be used to provide the W value to be programmed or read into each selected cell during a programming or read operation.
[0223] Figure 43 Shows a neuron output circuit 4300, which includes an adjustable current source 4301 and an adjustable current source 4302, which together generate I OUT , where I OUT is equal to the current I W+ generated by the adjustable current source 4301 minus the current I W- generated by the adjustable current source 4302. The adjustable current Iw+ 4301 is a scaled current of the cell current or neuron current (such as the bit line current) for implementing a positive weight. The adjustable current Iw- 4302 is a scaled current of the cell current or neuron current (such as the bit line current) for implementing a negative weight. Current scaling is done, for example, through an M:N ratio current mirror circuit, where Iout = (M / N)*Iin.
[0224] Figure 44 Shows a neuron output circuit 4400, which includes an adjustable capacitor 4401, a control transistor 4405, switches 4402, 4403, and an adjustable current source 4404Iw+, which is a scaled output current of the cell current or (bit line) neuron current such as an M:N current mirror circuit. The transistor 4405 is used to apply a fixed bias voltage to the current 4404, for example. The circuit 4404 generates V OUT , where V OUT is inversely proportional to the capacitor 4401, proportional to the adjustable integration time (the time when switch 4403 is closed and switch 4402 is open), and proportional to the adjustable current source 4404IW+ is proportional to the generated current. V OUT equals V +- ((Iw+ * integration time) / C 4401 ), where C 4401 is the value of capacitor 4401. The positive terminal V+ of capacitor 4401 is connected to the positive power supply voltage, and the negative terminal V- of capacitor 4401 is connected to the output voltage V OUT .
[0225] Figure 45 illustrates neuron circuit 4500, which includes capacitor 4401 and adjustable current source 4502, and the current is a scaled current of a unit current such as an M:N current mirror or a (bit line) neuron current. Circuit 4500 generates V OUT , where V OUT is inversely proportional to capacitor 4401, proportional to the adjustable integration time (the time when switch 4501 is open), and proportional to the current generated by adjustable current source 4502I Wi . Capacitor 4401 is reused from neuron output circuit 44 after completing its operation of integrating current Iw+. Then, the positive and negative terminals (V+ and V-) are swapped in neuron output circuit 45, where the positive terminal is connected to output voltage V OUT , and this output voltage is de-integrated by current Iw-. The negative terminal is held at the previous voltage value by a clamping circuit (not shown). In fact, output circuit 44 is used for positive weight implementation, and circuit 45 is used for negative weight implementation, where the final charge on capacitor 4401 effectively represents the combined weight (Qw = Qw+ - Qw-).
[0226] Figure 46 illustrates neuron circuit 4600, which includes adjustable capacitor 4601, switch 4602, control transistor 4604, and adjustable current source 4603. Circuit 4600 generates V OUT , where V OUT is inversely proportional to capacitor 4601, proportional to the adjustable integration time (the time when switch 4602 is open), and proportional to the current generated by adjustable current source 4603I W- . The negative terminal V- of capacitor 4601 is equal to ground, for example. The positive terminal V+ of capacitor 4601 is initially pre-charged to a positive voltage before integrating current Iw-. Neuron circuit 4600 can be used to replace neuron circuit 4500 and neuron circuit 4400 to implement the combined weight (Qw = Qw+ - Qw-).
[0227] Figure 47A neuron circuit 4700 is shown, which includes operational amplifiers 4703 and 4706; adjustable current sources Iw+ 4701 and Iw- 4702; and adjustable resistors 4704, 4705, and 4707. The neuron circuit 4700 generates V OUT , which voltage is equal to R 4707 *(Iw+ - Iw-). The adjustable resistor 4707 implements scaling of the output. The adjustable current sources Iw+ 4701 and Iw- 4702 also implement scaling of the output, for example, through an M:N ratio current mirror circuit (Iout = (M / N)*Iin).
[0228] Figure 48 A neuron circuit 4800 is shown, which includes operational amplifiers 4803 and 4806; switches 4808 and 4809; adjustable current sources Iw- 4802 and Iw+ 4801; adjustable capacitors 4804, 4805, and 4807. The neuron circuit 4800 generates V OUT , which voltage is proportional to (Iw+ - Iw-), proportional to the integration time (the time when switches 4808 and 4809 are open), and inversely proportional to the capacitance of capacitor 4807. The adjustable capacitor 4807 implements scaling of the output. The adjustable current sources Iw+ 4801 and Iw- 4802 also implement scaling of the output, for example, through an M:N ratio current mirror circuit (Iout = (M / N)*Iin). The integration time can also adjust the output scaling.
[0229] Figure 49A , Figure 49B and Figure 49C show a block diagram of an output circuit such as the output circuit 3307 in Figure 33 .
[0230] In Figure 49A , the output circuit 4901 includes an ADC circuit 4911, which is used to directly digitize the analog neuron output 4910 to provide digital output bits 4912.
[0231] In Figure 49B , the output circuit 4902 includes a neuron output circuit 4921 and an ADC 4911. The neuron output circuit 4921 receives the neuron output 4920 and shapes it, and then it is digitized by the ADC circuit 4911 to generate the output 4912. The neuron output circuit 4921 can be used for normalization, scaling, shifting, mapping, arithmetic operations, activation, and / or temperature compensation, as previously described. The ADC circuit can be a serial (ramp or jump or counting) ADC, SAR ADC, pipelined ADC, Σ-Δ ADC, or any type of ADC.
[0232] In Figure 49CIn it, the output circuit includes a neuron output circuit 4921 and a converter circuit 4931. The neuron output circuit receives a neuron output 4930, and the converter circuit is used to convert the output from the neuron output circuit 4921 into an output 4932. The converter 4931 can include an ADC, an AAC (similar to an analog-to-analog converter, such as a current-to-voltage converter), an APC (analog-to-pulse converter), or any other type of converter. The ADC 4911 or the converter 4931 can be used to implement an activation function, for example, through bit mapping (e.g., quantization) or clipping (e.g., clipped ReLU). The ADC 4911 and the converter 4931 can be configurable, such as for lower or higher precision (e.g., lower or higher number of bits), lower or higher performance (e.g., slower or faster speed), etc.
[0233] Another implementation for scaling and shifting is through an ADC (analog-digital) conversion circuit (such as a serial ADC, SAR ADC, pipelined ADC, ramp ADC, etc.) configured to convert an array (bit line) output into digital bits with, for example, lower or higher bit precision, and then manipulating the digital output bits according to a certain function (e.g., linear or non-linear, compression, non-linear activation, etc.), such as through normalization (e.g., 12 bits to 8 bits), shifting, or remapping. An embodiment of the ADC conversion circuit is described in U.S. Provisional Patent Application No. 62 / 933,809, filed on November 11, 2019, and titled "Precise Programming Method and Apparatus for Analog Neural Memory in a Deep Learning Artificial Neural Network", by the same assignee as this application, which is incorporated herein by reference.
[0234] Table 9 shows an alternative method for performing read, erase, and program operations:
[0235] Table 9: Operations of Flash Memory Cells
[0236] <![CDATA SL > <![CDATA BL > <![CDATA WL > <![CDATA CG > <![CDATA EG > <![CDATA P-Sub > <![CDATA Read > 0 0.5 1 0 0 0 <![CDATA Erase > 0 0 0 0 / -8V 10 - 12V / +8V 0 <![CDATA Program1 > 0-5V 0 0 8V -10 to -12V 0 <![CDATA Program2 > 0 0 0 8V 0-5V -10V
[0237] The read and erase operations are similar to the previous table. However, two methods for programming are implemented by the Fowler-Nordheim (FN) tunneling mechanism.
[0238] An implementation for scaling the input can be done, for example, by enabling a certain number of rows of the VMM at a time and then fully combining the results.
[0239] Another implementation is to scale the input voltage and appropriately rescale the output to achieve normalization.
[0240] Another implementation for scaling a pulse-width modulation input is by modulating the timing of the pulse width. An example of this technique is described in U.S. Patent Application No. 16 / 449,201, filed on June 21, 2019, and titled "Configurable Input Blocks and Output Blocks and Physical Layout for Analog Neural Memory in Deep Learning Artificial Neural Network," which is assigned to the same assignee as this application and is incorporated herein by reference.
[0241] Another implementation for scaling an input is by enabling one input bit at a time. For example, for an 8-bit input IN7:0, IN0, IN1, …, IN7 are evaluated sequentially and then the output results are combined with appropriate bit weighting. An example of this technique is described in U.S. Patent Application No. 16 / 449,201, filed on June 21, 2019, and titled "Configurable Input Blocks and Output Blocks and Physical Layout for Analog Neural Memory in Deep Learning Artificial Neural Network," which is assigned to the same assignee as this application and is incorporated herein by reference.
[0242] Optionally, in the above implementations, the cell current can be averaged or measured multiple times, such as 8 to 32 times, for the purpose of verifying or reading the current to reduce the impact of noise (such as RTN or any random noise) and / or detect any outlier bits that are defective and need to be replaced by redundant bits.
[0243] It should be noted that, as used herein, the terms "above" and "on" both inclusively encompass "directly on" (with no intervening material, element, or space therebetween) and "indirectly on" (with intervening material, element, or space therebetween). Similarly, the term "adjacent" includes "directly adjacent" (with no intervening material, element, or space therebetween) and "indirectly adjacent" (with intervening material, element, or space therebetween), "mounted to" includes "directly mounted to" (with no intervening material, element, or space therebetween) and "indirectly mounted to" (with intervening material, element, or space therebetween), and "electrically coupled to" includes "directly electrically coupled to" (with no intervening material or element electrically connecting the elements together) and "indirectly electrically coupled to" (with intervening material or element electrically connecting the elements together). For example, forming an element "above a substrate" may include directly forming the element on the substrate with no intervening material / element therebetween, and indirectly forming the element on the substrate with one or more intervening material / elements therebetween.
Claims
1. A circuit for compensating for drift errors during a read operation in a vector matrix multiplication array, the circuit comprising: A data drift monitoring circuit coupled to the vector matrix multiplication array to monitor a data drift amount in the vector matrix multiplication array and generate an output indicative of the data drift amount; And A bit line compensation circuit for generating a compensation current and injecting the compensation current into an array output current received from one or more bit lines of the vector matrix multiplication array, wherein a level of the compensation current is selected based on the data drift amount from the data drift monitoring circuit.
2. The circuit according to claim 1, wherein the bit line compensation circuit includes a first adjustable current source and a second adjustable current source, and the compensation current is a difference between a current generated by the first adjustable current source and a current generated by the second adjustable current source.
3. The circuit according to claim 1, wherein the bit line compensation circuit includes an operational amplifier, a first adjustable resistor, and a second adjustable resistor.
4. The circuit according to claim 1, wherein the bit line compensation circuit includes an operational amplifier, a current source, and an adjustable capacitor.
5. The circuit according to claim 1, wherein the bit line compensation circuit includes an M:N current mirror.
6. The circuit according to claim 1, wherein the bit line compensation circuit includes an operational amplifier, a first adjustable resistor, a second adjustable resistor, and a third adjustable resistor.
7. The circuit according to claim 1, wherein the bit line compensation circuit includes an operational amplifier, a current source, a current shifter, and an adjustable capacitor.
8. A circuit for compensating for drift errors during a read operation in a vector matrix multiplication array, the circuit comprising: A bit line compensation circuit for scaling an array output current received from one or more bit lines of the array to compensate for drift errors during a read operation of the array, wherein the bit line compensation circuit includes an operational amplifier, a current source, a switch, and an adjustable integration output capacitor; the operational amplifier receives a reference voltage at its non-inverting terminal and receives an input voltage corresponding to the array output current at its inverting terminal and generates an output voltage, the adjustable integration output capacitor is coupled between the input voltage and the output voltage, and a value of the output voltage is scaled based on a configured value of the adjustable integration output capacitor and / or an adjustable integration time to compensate for data drift, or wherein the bit line compensation circuit includes a current mirror having an M:N ratio, and the array output current is scaled by configuring the M and / or N parameters.
9. The circuit according to claim 8, wherein the bit line compensation circuit shifts the output.
10. The circuit according to claim 8, wherein the scaling includes amplification.
11. The circuit according to claim 8, wherein the scaling includes reduction.
12. The circuit according to claim 8, wherein the vector matrix multiplication array is formed of split-gate non-volatile memory cells.
13. The circuit according to claim 8, wherein the circuit further comprises: A bit line compensation circuit for shifting an array output current received from one or more bit lines of the array to compensate for drift errors during a read operation of the array.
14. The circuit according to claim 13, wherein the vector matrix multiplication array is formed of split gate non-volatile memory cells.
15. The circuit according to claim 13, wherein one or more cells in the vector matrix multiplication array are programmed using Fowler-Nordheim tunneling.
16. The circuit according to claim 14, wherein one or more cells in the vector matrix multiplication array are programmed using Fowler-Nordheim tunneling.
17. A method for compensating drift errors during a read operation in a vector matrix multiplication array, the method comprising: Monitoring a data drift amount in the vector matrix multiplication array; Generating a bit line compensation current, wherein a level of the bit line compensation current is selected based on the data drift amount; And Injecting the bit line compensation current into an array output current received from one or more bit lines of the vector matrix multiplication array during a read operation to compensate for the drift error.
18. The method according to claim 17, wherein the vector matrix multiplication array is formed of split gate non-volatile memory cells.
Citation Information
Patent Citations
Deep learning neural network classifier using non-volatile memory array
US11308383B2
Deep Learning Neural Network Classifier Using Non-volatile Memory Array
US20170337466A1
High Precision And Highly Efficient Tuning Mechanisms And Algorithms For Analog Neuromorphic Memory In Artificial Neural Networks
US20190164617A1
Configurable input blocks and output blocks and physical layout for analog neural memory in deep learning artificial neural network
US20200349421A1
Single transistor non-valatile electrically alterable semiconductor memory device
US5029130A