Precision tuning for programming of analog neural memory in deep learning artificial neural network
The precision tuning algorithm and apparatus address the challenge of accurately programming non-volatile memory cells in VMM arrays by enabling precise charge deposition on floating gates, achieving high accuracy in representing multiple weight values and improving the performance of analog neuromorphic memory systems.
Patent Information
- Application Number
- JP2025006935
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2018-12-21
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2039-07-25
AI Technical Summary
Existing technologies face challenges in accurately programming non-volatile memory cells within vector matrix multiplication (VMM) arrays in analog neuromorphic memory systems, particularly in achieving high precision and granularity for representing a large number of different weight values.
A precision tuning algorithm and apparatus are developed to precisely and rapidly deposit an accurate amount of charge on the floating gates of non-volatile memory cells within a VMM array, enabling extremely high accuracy in programming selected cells to hold one of N different values.
The solution enables precise programming of non-volatile memory cells, allowing for accurate representation of a large number of weight values, thereby enhancing the performance and efficiency of analog neuromorphic memory systems in artificial neural networks.
Smart Images

Figure 2025081310000001_ABST
Abstract
Description
Technical Field
[0001] (Claim of Priority) This application claims priority to U.S. Provisional Patent Application No. 62 / 746,470, filed Oct. 16, 2018, entitled “Precision Tuning For the Programming Of Analog Neural Memory In A Deep Learning Artificial Neural Network,” and to U.S. Patent Application No. 16 / 231,231, filed Dec. 21, 2018, entitled “Precision Tuning For the Programming Of Analog Neural Memory In A Deep Learning Artificial Neural Network.”
[0002] (Field of the Invention) Disclosed are numerous embodiments of a precision tuning algorithm and apparatus for precisely and rapidly depositing an accurate amount of charge on a floating gate of a non-volatile memory cell within a vector matrix multiplication (VMM) array within an artificial neural network.
Background Art
[0003] An artificial neural network mimics a biological neural network (the central nervous system of an animal, particularly the brain), may depend on a large number of inputs, and is used to estimate or approximate a generally unknown function. An artificial neural network generally includes layers of interconnected “neurons” that exchange messages.
[0004] Figure 1 shows an artificial neural network, in which the circles represent the input or layers of neurons. The connections (referred to as synapses) are represented by arrows and have numerical weights that can be adjusted based on experience. As a result, the neural network adapts to the input and becomes capable of learning. Typically, a neural network includes multiple input layers. Typically, there is one or more intermediate layers of neurons and an output layer of neurons that provides the output of the neural network. At each level, the neurons make decisions individually or jointly based on the data received from the synapses.
[0005] One of the main challenges in the development of artificial neural networks for high-performance information processing is the lack of appropriate hardware technology. In practice, practical neural networks rely on a very large number of synapses, which enables high connectivity between neurons, that is, a very high degree of parallelization of computational processing. In principle, such complexity can be realized by a digital supercomputer or a dedicated GPU (graphics processing unit) cluster. However, in addition to high costs, these approaches also suffer from poor energy efficiency compared to biological networks that mainly perform low-precision analog calculations and consume far less energy. CMOS analog circuits have been used in artificial neural networks, but most CMOS implementation synapses have been too bulky assuming a large number of neurons and synapses.
[0006] The applicant has previously disclosed, in U.S. Patent Application No. 15 / 594,439, incorporated by reference, an artificial (analog) neural network that utilizes one or more non-volatile memory arrays as synapses. The non-volatile memory arrays operate as analog neuromorphic memories. The neural network device includes a first plurality of synapses configured to receive a first plurality of inputs and then generate a first plurality of outputs, and a first plurality of neurons configured to receive the first plurality of outputs. The first plurality of synapses includes a plurality of memory cells, each memory cell including a spaced-apart source region and drain region formed in a semiconductor substrate with a channel region extending therebetween, a floating gate disposed above a first portion of the channel region and insulated from the first portion of the channel region, and a non-floating gate disposed above a second portion of the channel region and insulated from the second portion of the channel region. Each of the plurality of memory cells is configured to store a weight value corresponding to a number of electrons on the floating gate. The plurality of memory cells is configured to multiply the stored weight values by the first plurality of inputs to generate the first plurality of outputs.
[0007] Each non-volatile memory cell used in an analog neuromorphic memory system must hold a charge, i.e., a number of electrons, in the floating gate in a very specific and accurate amount corresponding to an erase / program. For example, each floating gate must hold one of N different values, where N is the number of different weights that can be represented by each cell. Examples of N include 16, 32, 64, 128, and 256.
[0008] One challenge in a VMM system is the ability to program a selected cell with the accuracy and granularity required for N different values. For example, if a selected cell can include one of 64 different values, a very high accuracy is required in the programming operation.
[0009] What is needed is an improved programming system and method suitable for use with a VMM in an analog neuromorphic memory system.
SUMMARY OF THE INVENTION
[0010] Numerous embodiments are disclosed of a precision tuning algorithm and apparatus for precisely and rapidly depositing an accurate amount of charge on the floating gates of non-volatile memory cells within a vector matrix multiplication (VMM) array in an artificial neural network. Thereby, selected cells can be programmed with extremely high accuracy to hold one of N different values.
[0011]
[0012]
[0013]
[0014]
[0015]
[0016]
[0017]
[0018]
[0019]
[0020]
[0021]
[0022]
[0023]
[0024]
[0025]
[0026]
[0027]
[0028]
[0029]
[0030]
[0031]
[0032]
[0033]
[0034]
[0035]
[0036]
[0037]
[0038]
[0039]
[0040]
[0041]
[0042]
[0043]
[0044]
[0045]
[0046]
[0047]
[0048]
[0049]
[0050]
[0051]
[0052]
[0053]
Brief Description of the Drawings
[0054]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22A
Figure 22B
Figure 23
Figure 24
Figure 25
Figure 26
Figure 27
Figure 28
Figure 29
Figure 30
Figure 31
Figure 32
Figure 33
Figure 34
Figure 35
Figure 36A
Figure 36B
Figure 36C
Figure 36D
Figure 36E
Figure 36F
Figure 37A
Figure 37B
Figure 38
Mode for Carrying Out the Invention
[0055] The artificial neural network of the present invention utilizes a combination of CMOS technology and a non-volatile memory array. Non-volatile memory cell
[0056] Digital non-volatile memories are well known. For example, U.S. Patent No. 5,029,130 (the “’130 patent”), which is incorporated herein by reference, discloses an array of split-gate non-volatile memory cells, which are a type of flash memory cell. Such a memory cell 210 is shown in FIG. 2. Each memory cell 210 includes a source region 14 and a drain region 16 formed in a semiconductor substrate 12, and there is a channel region 18 between the source region 14 and the drain region 16. A floating gate 20 is formed above a first portion of the channel region 18, insulated from (and controlling the conductivity of) the first portion of the channel region 18, and formed over a portion of the source region 14. A word line terminal 22 (typically coupled to a word line) is disposed above a second portion of the channel region 18, insulated from (and controlling the conductivity of) the second portion of the channel region 18, and having a first portion extending upwardly above the floating gate 20 and a second portion extending upwardly over the floating gate 20. The floating gate 20 and the word line terminal 22 are insulated from the substrate 12 by a gate oxide. A bit line 24 is coupled to the drain region 16.
[0057] By applying a high positive voltage to the word line terminal 22, the memory cell 210 is erased (electrons are removed from the floating gate), whereby electrons in the floating gate 20 pass through the insulator therebetween from the floating gate 20 to the word line terminal 22 via Fowler-Nordheim tunneling.
[0058] The memory cell 210 is programmed (electrons are applied to the floating gate) by applying a positive voltage to the word line terminal 22 and a positive voltage to the source region 14. An electron current flows from the source region 14 toward the drain region 16. The electrons are accelerated and heat up when they reach the gap between the word line terminal 22 and the floating gate 20. A portion of the heated electrons is injected into the floating gate 20 through the gate oxide due to the electrostatic attraction from the floating gate 20.
[0059] The memory cell 210 is read by applying a positive read voltage to the drain region 16 and the word line terminal 22 (turning on the portion of the channel region 18 below the word line terminal). When the floating gate 20 is positively charged (i.e., electrons are erased), the portion of the channel region 18 below the floating gate 20 is also turned on, and current flows through the channel region 18, which is detected as the erased state, i.e., the "1" state. When the floating gate 20 is negatively charged (i.e., programmed with electrons), the portion of the channel region below the floating gate 20 becomes almost or completely off, and current does not flow (or hardly flows) through the channel region 18, which is detected as the programmed state, i.e., the "0" state.
[0060] Table 1 shows the typical voltage ranges that can be applied to the terminals of the memory cell 110 to perform read, erase, and program operations. Table 1: Operation of the flash memory cell 210 of FIG. 3
Table 1
[0061] As another type of flash memory cell, other split - gate type memory cell configurations are also known. For example, FIG. 3 shows a four - gate memory cell 310 including a source region 14, a drain region 16, a floating gate 20 above a first portion of the channel region 18, a select gate 22 (typically coupled to a word line, WL) above a second portion of the channel region 18, a control gate 28 above the floating gate 20, and an erase gate 30 above the source region 14. This configuration is described in U.S. Patent No. 6,747,310, which is hereby incorporated by reference for all purposes. Here, all gates are non - floating gates except for the floating gate 20, i.e., they are electrically connected or connectable to a voltage source. Programming is performed by injecting hot electrons from the channel region 18 into the floating gate 20 itself. Erase is performed by electrons tunneling from the floating gate 20 to the erase gate 30.
[0062] Table 2 shows the typical voltage ranges that can be applied to the terminals of the memory cell 310 to perform read, erase, and program operations. Table 2: Operation of the flash memory cell 310 in FIG. 3 [Table 2]
[0063] FIG. 4 shows a 3-gate memory cell 410, which is another type of flash memory cell. The memory cell 410 is identical to the memory cell 310 in FIG. 3, except that the memory cell 410 does not have a separate control gate. The erase operation (erasure occurs through the use of an erase gate) and the read operation are the same as those in FIG. 3, except that no control gate bias is applied. The programming operation is also performed without a control gate bias. As a result, during the programming operation, a higher voltage must be applied to the source line to compensate for the lack of control gate bias.
[0064] Table 3 shows the typical voltage ranges that can be applied to the terminals of the memory cell 410 to perform read, erase, and program operations. Table 3: Operation of the flash memory cell 410 in FIG. 4 [Table 3]
[0065] FIG. 5 shows a stacked gate memory cell 510, which is another type of flash memory cell. The memory cell 510 is the same as the memory cell 210 in FIG. 2, except that the floating gate 20 extends over the entire channel region 18 and the control gate 22 (coupled to the word line) extends over the floating gate 20 separated by an insulating layer (not shown). The erase, programming, and read operations operate in a similar manner as those described above for the memory cell 210.
[0066] Table 4 shows typical voltage ranges that can be applied to the memory cell 510 and the terminals of the substrate 12 to perform read, erase, and program operations. Table 4: Operation of the flash memory cell 510 of FIG. 5 [Table 4]
[0067] To utilize a memory array that includes one of the types of non-volatile memory cells in the artificial neural network described above, two modifications are made. First, the lines are configured so that each memory cell can be individually programmed, erased, and read without adversely affecting the memory states of other memory cells in the array, as further described below. Second, continuous (analog) programming of the memory cells is provided.
[0068] Specifically, the memory state of each memory cell in the array (i.e., the charge on the floating gate) can be changed independently, with minimal interference from other memory cells, continuously from a fully erased state to a fully programmed state. In another embodiment, the memory state of each memory cell in the array (i.e., the charge on the floating gate) can be changed independently, with minimal interference from other memory cells, continuously from a fully programmed state to a fully erased state and vice versa. This means that cell storage is analog or can store at least one of a number of discrete values (such as 16 or 64 different values), which allows all cells in the memory array to be very accurately and individually adjustable, and the memory array to be ideal for storage and allows for fine-tuning of the synaptic weights of the neural network. Neural Network Using a Non-Volatile Memory Cell Array
[0069] Figure 6 conceptually shows a non-limiting example of a neural network utilizing the non-volatile memory array of the present embodiment. This example uses a non-volatile memory array neural network for a face recognition application, but it is also possible to implement other suitable applications using a non-volatile memory array-based neural network.
[0070] S0 is the input layer, and in this example, it is a 32×32 pixel RGB image with 5-bit precision (i.e., three 32×32 pixel arrays, one for each color R, G, and B, and each pixel has 5-bit precision). The synapses CB1 going from the input layer S0 to layer C1 apply different sets of weights to some instances and shared weights to other instances, scanning the input image with an overlapping filter of 3×3 pixels (kernel) and shifting the filter by 1 pixel (or more than 2 pixels depending on the model) at a time. Specifically, the 9 pixel values in the 3×3 portion of the image (i.e., what is called the filter or kernel) are provided to the synapses CB1, where these 9 input values are multiplied by appropriate weights, and after summing the outputs of the multiplications, a single output value is determined and given by the first synapse of CB1 to generate one pixel of the layer of the feature map C1. The 3×3 filter is then shifted 1 pixel to the right within the input layer S0 (i.e., a column of 3 pixels is added on the right and a column of 3 pixels is dropped on the left), and thus the 9 pixel values of this newly positioned filter are provided to the synapses CB1, where they are multiplied by the same weights as above, and a second single output value is determined by the relevant synapses. This process is continued until the 3×3 filter has scanned over the entire 32×32 pixel image of the input layer S0 for all three colors and all bits (precision values). The process is then repeated using different sets of weights until all the feature maps of layer C1 are calculated, generating different feature maps of C1.
[0071] In this example, in layer C1, there are 16 feature maps each having 30×30 pixels. Each pixel is a new feature pixel extracted from the multiplication of the input and the kernel. Thus, each feature map is a two-dimensional array. Therefore, in this example, layer C1 consists of 16 layers of two-dimensional arrays (note that the layers and arrays referred to in this specification are not necessarily in a physical relationship but a logical relationship, that is, the array is not necessarily oriented as a physical two-dimensional array). Each of the 16 feature maps in layer C1 is generated by one of 16 different sets of synaptic weights applied to the filter scan. All of the C1 feature maps can target different aspects of the same image feature, such as edge identification. For example, the first map (generated using a first set of weights shared by all scans used to generate this first map) can identify circular edges, and the second map (generated using a second set of weights different from the first set of weights) can identify square edges or the aspect ratio of a specific feature, etc.
[0072] Before going from layer C1 to layer S1, an activation function P1 (pooling) that pools values from non - overlapping and consecutive 2×2 regions within each feature map is applied. The purpose of the pooling function is to average neighboring positions (or it is also possible to use the max function), for example, to reduce the dependence on edge positions, and to reduce the data size before going to the next stage. In layer S1, there are 16 15×15 feature maps (i.e., 16 different arrays of 15×15 pixels each). The synapses CB2 going from layer S1 to layer C2 scan the maps in S1 with a 4×4 filter with a 1 - pixel filter shift. In layer C2, there are 22 12×12 feature maps. Before going from layer C2 to layer S2, an activation function P2 (pooling) that pools values from non - overlapping and consecutive 2×2 regions within each feature map is applied. In layer S2, there are 22 6×6 feature maps. In the synapses CB3 going from layer S2 to layer C3, an activation function (pooling) is applied, where all neurons in layer C3 are connected to all maps in layer S2 via each synapse of CB3. In layer C3, there are 64 neurons. The synapses CB4 going from layer C3 to the output layer S3 fully connect C3 to S3, i.e., all neurons in layer C3 are connected to all neurons in layer S3. The output in S3 contains 10 neurons, and the neuron with the highest output determines the class. This output can indicate, for example, the identification or classification of the content of the original image.
[0073] Each layer of synapses is implemented using an array or a part of an array of non - volatile memory cells.
[0074] FIG. 7 is a block diagram of an array that can be used for that purpose. A vector-by-matrix multiplication (VMM) array 32 includes non-volatile memory cells and is utilized as synapses (such as CB1, CB2, CB3, and CB4 in FIG. 6) between one layer and the next layer. Specifically, the VMM array 32 includes an array 33 of non-volatile memory cells, an erase gate and word line gate decoder 34, a control gate decoder 35, a bit line decoder 36, and a source line decoder 37, and these decoders decode respective inputs to the non-volatile memory cell array 33. Inputs to the VMM array 32 can be made from the erase gate and word line gate decoder 34 or from the control gate decoder 35. The source line decoder 37 in this example also decodes the output of the non-volatile memory cell array 33. Alternatively, the bit line decoder 36 can decode the output of the non-volatile memory cell array 33.
[0075] The non-volatile memory cell array 33 serves two purposes. First, it stores the weights used by the VMM array 32. Second, the non-volatile memory cell array 33 effectively multiplies the inputs by the weights stored in the non-volatile memory cell array 33 and adds them for each output line (source line or bit line) to generate an output, which becomes the input to the next layer or the input to the last layer. By the non-volatile memory cell array 33 performing the functions of multiplication and addition, the need for separate multiplication and addition logic circuits is eliminated, and the power efficiency is also good due to the calculation within the memory.
[0076] The output of the non-volatile memory cell array 33 is supplied to a differential adder (such as an addition op-amp or an addition current mirror) 38 that sums the outputs of the non-volatile memory cell array 33 to create a single value for convolution. The differential adder 38 is arranged to perform the sum of the positive and negative weights.
[0077] The total output value of the differential adder 38 is then supplied to an activation function circuit 39 that rectifies the output. The activation function circuit 39 can provide a sigmoid, tanh, or ReLU function. The rectified output value of the activation function circuit 39 becomes an element of the feature map as the next layer (e.g., C1 in FIG. 6), and is then applied to the next synapse to generate the next feature map layer or the last layer. Thus, in this example, the non-volatile memory cell array 33 constitutes a plurality of synapses (receiving inputs from the previous layer of neurons or from an input layer such as an image database), and the adder op-amp 38 and the activation function circuit 39 constitute a plurality of neurons.
[0078] The inputs (WLx, EGx, CGx, and optionally BLx and SLx) to the VMM array 32 of FIG. 7 can be at an analog level, binary level, or digital bits (in which case a DAC is provided to convert the digital bits to an appropriate input analog level), and the outputs can be at an analog level, binary level, or digital bits (in which case an output ADC is provided to convert the output analog level to digital bits).
[0079] FIG. 8 is a block diagram showing the use of multiple layers of the VMM array 32, labeled as VMM arrays 32a, 32b, 32c, 32d, and 32e in the figure. As shown in FIG. 8, an input (indicated as Inputx) is converted from digital to analog by a digital-to-analog converter 31 and provided to the input VMM array 32a. The converted analog input can be a voltage or a current. The input D / A conversion of the first layer can be performed by using a function or a LUT (look-up table) that maps the input Inputx to an appropriate analog level of the matrix multiplier of the input VMM array 32a. The input conversion can also be performed by an analog-to-analog (A / A) converter to convert an external analog input to the mapped analog input to the input VMM array 32a.
[0080] The output generated by the input VMM array 32a is then provided as input to the next VMM array (hidden level 1) 32b, which in turn generates an output that is provided as input to the input VMM array (hidden level 2) 32c, and so on. The various layers of the VMM array 32 function as the respective layers of synapses and neurons of a convolutional neural network (CNN). The VMM arrays 32a, 32b, 32c, 32d, and 32e can each be a stand-alone physical non-volatile memory array, or multiple VMM arrays can utilize different portions of the same physical non-volatile memory array, or multiple VMM arrays can utilize overlapping portions of the same physical non-volatile memory array. The example shown in FIG. 8 includes five layers (32a, 32b, 32c, 32d, 32e), namely, one input layer (32a), two hidden layers (32b, 32c), and two fully connected layers (32d, 32e). One of ordinary skill in the art will understand that this is merely exemplary, and alternatively, the system can include more than two hidden layers and more than two fully connected layers. Vector Matrix Multiplication (VMM) array
[0081] FIG. 9 shows a neuron VMM array 900 that is particularly suitable for the memory cell 310 shown in FIG. 3 and is utilized as part of the synapses and neurons between the input layer and the next layer. The VMM array 900 includes a memory array 901 of non-volatile memory cells and a reference array 902 of non-volatile reference memory cells (located at the top of the array). Alternatively, another reference array can be located at the bottom.
[0082] In the VMM array 900, control gate lines such as control gate line 903 extend in the vertical direction (thus, the reference array 902 in the row direction is orthogonal to the control gate line 903), and erasure gate lines such as erasure gate line 904 extend in the horizontal direction. Here, the input to the VMM array 900 is provided to the control gate lines (CG0, CG1, CG2, CG3), and the output of the VMM array 900 appears on the source lines (SL0, SL1). In one embodiment, only even rows are used, and in another embodiment, only odd rows are used. The current applied to each source line (SL0 and SL1 respectively) performs a summation function of all the currents from the memory cells connected to that particular source line.
[0083] As described herein for neural networks, the non-volatile memory cells of the VMM array 900, i.e., the flash memory of the VMM array 900, are preferably configured to operate in the subthreshold region.
[0084] The non-volatile reference memory cells and non-volatile memory cells described herein are biased with weak inversion as follows: Ids = Io * e (Vg-Vth) / kVt = w * Io * e (Vg) / kVt where w = e (-Vth) / kVt is.
[0085] When using an I-V logarithmic converter that uses a memory cell (such as a reference memory cell or a peripheral memory cell) or a transistor to convert an input current into an input voltage: Vg = k * Vt * log[Ids / wp * Io] where wp is the w of the reference or peripheral memory cell.
[0086] Regarding the memory array used as the vector matrix multiplier VMM array, the output current is as follows: Iout = wa * Io* e (Vg) / kVt That is Iout = (wa / wp) * Iin = W * Iin W = e (Vthp-Vtha) / kVt In the formula, wa = w for each memory cell of the memory array.
[0087] The word line or control gate can be used as the input of the memory cell for the input voltage.
[0088] Alternatively, the flash memory cells of the VMM array described in this specification can be configured to operate in the linear region. Ids = β * (Vgs - Vth) * Vds; β = u * Cox * W / L W = α(Vgs - Vth)
[0089] The word line or control gate or bit line or source line can be used as the input of the memory cell operating in the linear region.
[0090] For an I-V linear converter, the input and output current can be linearly converted into the input and output voltage by using a memory cell (such as a reference memory cell or a peripheral memory cell) or a transistor operating in the linear region.
[0091] Another embodiment for the VMM array 32 of FIG. 7 is described in U.S. Patent Application No. 15 / 826,345, which is incorporated herein by reference. As described in the above application, the source line or bit line can be used as the neuron output (current sum output).
[0092] FIG. 10 shows a neuron VMM array 1000 that is particularly suitable for the memory cell 210 shown in FIG. 2 and is used as a synapse between the input layer and the next layer. The VMM array 1000 includes a memory array 1003 of non-volatile memory cells, a reference array 1001 of first non-volatile reference memory cells, and a reference array 1002 of second non-volatile reference memory cells. The reference arrays 1001 and 1002 arranged in the column direction of the array function to convert the current inputs flowing into the terminals BLR0, BLR1, BLR2, and BLR3 into voltage inputs WL0, WL1, WL2, and WL3. In practice, the first and second non-volatile reference memory cells are diode-connected through a multiplexer 1014 (only part shown) in a state where the current inputs flow in. The reference cells are adjusted (e.g., programmed) to a target reference level. The target reference level is provided by a reference mini-array matrix (not shown).
[0093] The memory array 1003 serves two purposes. First, it stores the weights used by the VMM array 1000 in each memory cell. Second, the memory array 1003 effectively multiplies the input (i.e., the current inputs provided to the terminals BLR0, BLR1, BLR2, and BLR3, which the reference arrays 1001 and 1002 convert into voltage inputs and supply to the word lines WL0, WL1, WL2, and WL3) by the weights stored in the memory cell array 1003, and then adds up all the results (memory cell currents) to generate an output on each bit line (BL0 to BLN). This output becomes the input to the next layer or the input to the last layer. By the memory array 1003 performing the functions of multiplication and addition, the need for separate multiplication and addition logic circuits is eliminated, and the power efficiency is also good. Here, the voltage inputs are provided to the word lines WL0, WL1, WL2, and WL3, and the outputs appear on each of the bit lines BL0 to BLN during the read (inference) operation. The currents arranged on each bit line BL0 to BLN perform the total function of the currents from all the non-volatile memory cells connected to that particular bit line.
[0094] Table 5 shows the operating voltages of the VMM array 1000. The columns in the table show the voltages applied to the word line of the selected cell, the word line of the non-selected cell, the bit line of the selected cell, the bit line of the non-selected cell, the source line of the selected cell, and the source line of the non-selected cell. The rows show each of the read, erase, and program operations. Table 5: Operations of the VMM Array 1000 in FIG. 10 [Table 5]
[0095] FIG. 11 shows a neuron VMM array 1100 particularly suitable for the memory cell 210 shown in FIG. 2 and is used as part of the synapses and neurons between the input layer and the next layer. The VMM array 1100 includes a memory array 1103 of non-volatile memory cells, a reference array 1101 of first non-volatile reference memory cells, and a reference array 1102 of second non-volatile reference memory cells. The reference arrays 1101 and 1102 extend in the row direction of the VMM array 1100. The VMM array is similar to the VMM1000 except that the word lines extend vertically in the VMM array 1100. Here, the inputs are provided to the word lines (WLA0, WLB0, WLA1, WLB2, WLA2, WLB2, WLA3, WLB3), and the outputs appear on the source lines (SL0, SL1) during the read operation. The current applied to each source line performs a summation function of all the currents from the memory cells connected to that particular source line.
[0096] Table 6 shows the operating voltages of the VMM array 1100. The columns in the table show the voltages applied to the word line of the selected cell, the word line of the non-selected cell, the bit line of the selected cell, the bit line of the non-selected cell, the source line of the selected cell, and the source line of the non-selected cell. The rows show each of the read, erase, and program operations. Table 6: Operations of the VMM Array 1100 in FIG. 11 [Table 6]
[0097] FIG. 12 shows a neuron VMM array 1200 that is particularly suitable for the memory cell 310 shown in FIG. 3 and is used as part of the synapses and neurons between the input layer and the next layer. The VMM array 1200 includes a memory array 1203 of non-volatile memory cells, a reference array 1201 of first non-volatile reference memory cells, and a reference array 1202 of second non-volatile reference memory cells. The reference arrays 1201 and 1202 function to convert the current inputs flowing into the terminals BLR0, BLR1, BLR2, and BLR3 into voltage inputs CG0, CG1, CG2, and CG3. In practice, the first and second non-volatile reference memory cells are diode-connected through a multiplexer 1212 (only part shown) in a state where the current inputs flow through BLR0, BLR1, BLR2, and BLR3. The multiplexer 1212 includes respective multiplexers 1205 and cascode transistors 1204 to ensure a constant voltage on each bit line (such as BLR0) of the first and second non-volatile reference memory cells during the read operation. The reference cells are adjusted to a target reference level.
[0098] The memory array 1203 serves two purposes. First, it stores the weights used by the VMM array 1200. Second, the memory array 1203 effectively multiplies the input (the current inputs provided to the terminals BLR0, BLR1, BLR2, and BLR3, which the reference arrays 1201 and 1202 convert these current inputs into input voltages and supply to the control gates (CG0, CG1, CG2, and CG3)) by the weights stored in the memory cell array, and then adds up all the results (cell currents) to generate an output, which appears on BL0 to BLN and becomes the input to the next layer or the input to the last layer. By the memory array performing the functions of multiplication and addition, the need for separate multiplication and addition logic circuits is eliminated, and the power efficiency is also good. Here, the input is provided to the control gate lines (CG0, CG1, CG2, and CG3), and the output appears on the bit lines (BL0 to BLN) during the read operation. The current applied to each bit line performs the summation function of all the currents from the memory cells connected to that particular bit line.
[0099] The VMM array 1200 performs a unidirectional adjustment of the non-volatile memory cells within the memory array 1203. That is, each non-volatile memory cell is erased and then partially programmed until the desired charge on the floating gate is reached. This can be performed, for example, using the novel precision programming techniques described below. If too much charge is applied to the floating gate (such as an incorrect value being stored in the cell), the cell must be erased and a series of partial programming operations must be repeated. As shown, two rows sharing the same erase gate (such as EG0 or EG1) need to be erased together (known as page erase), and then each cell is partially programmed until the desired charge on the floating gate is reached.
[0100] Table 7 shows the operating voltages of the VMM array 1200. The columns in the table show the voltages applied to the word line of the selected cell, the word line of the non-selected cell, the bit line of the selected cell, the bit line of the non-selected cell, the control gate of the selected cell, the control gate of the non-selected cell within the same sector as the selected cell, the control gate of the non-selected cell in a different sector from the selected cell, the erase gate of the selected cell, the erase gate of the non-selected cell, the source line of the selected cell, and the source line of the non-selected cell. The rows show each of the read, erase, and program operations. Table 7: Operation of the VMM array 1200 of FIG. 12 [Table 7]
[0101] FIG. 13 shows a neuron VMM array 1300 that is particularly suitable for the memory cell 310 shown in FIG. 3 and is used as part of the synapses and neurons between the input layer and the next layer. The VMM array 1300 includes a memory array 1303 of non-volatile memory cells, a reference array 1301 or a first non-volatile reference memory cell, and a reference array 1302 of second non-volatile reference memory cells. The EG lines EGR0, EG0, EG1, and EGR1 extend vertically, and the CG lines CG0, CG1, CG2, and CG3 and the SL lines WL0, WL1, WL2, and WL3 extend horizontally. The VMM array 1300 is similar to the VMM array 1400 except that the VMM array 1300 implements bidirectional adjustment, and individual cells can be completely erased, partially programmed, and partially erased as needed to reach the desired charge amount on the floating gate by using individual EG lines. As shown, the reference arrays 1301 and 1302 convert the input current in the terminals BLR0, BLR1, BLR2, and BLR3 into control gate voltages CG0, CG1, CG2, and CG3 (through the action of diode-connected reference cells via the multiplexer 1314), and these voltages are applied to the memory cells in the row direction. The current outputs (neurons) are in the bit lines BL0 to BLN, and each bit line sums all the currents from the non-volatile memory cells connected to that particular bit line.
[0102] Table 8 shows the operating voltages of the VMM array 1300. The columns in the table show the voltages applied to the word line of the selected cell, the word line of the non-selected cell, the bit line of the selected cell, the bit line of the non-selected cell, the control gate of the selected cell, the control gate of the non-selected cell in the same sector as the selected cell, the control gate of the non-selected cell in a different sector from the selected cell, the erase gate of the selected cell, the erase gate of the non-selected cell, the source line of the selected cell, and the source line of the non-selected cell. The rows show each operation of read, erase, and program. Table 8: Operations of the VMM Array 1300 in FIG. 13
Table 8
[0103] The prior art includes a concept known as long short-term memory (LSTM). LSTM units are often used within neural networks. With LSTM, a neural network can store information over an arbitrary predetermined period and use that information in subsequent operations. Conventional LSTM units include a cell, an input gate, an output gate, and a forget gate. The three gates regulate the flow of information into and out of the cell and the period during which information is stored within the LSTM. VMM is particularly useful in LSTM units.
[0104] FIG. 14 shows an exemplary LSTM 1400. The LSTM 1400 in this example includes cells 1401, 1402, 1403, and 1404. Cell 1401 receives an input vector x 0 and generates an output vector h 0 and a cell state vector c 0 Cell 1402 receives an input vector x 1 and the output vector (hidden state) h 0 from cell 1401 and the cell state vector c 0 from cell 1401 and generates an output vector h 1 and a cell state vector c 1 Cell 1403 receives an input vector x 2 and the output vector (hidden state) h 1 from cell 1402 and the cell state vector c 1 from cell 1402 and generates an output vector h 2 and a cell state vector c 2 Cell 1404 receives an input vector x 3 and the output vector (hidden state) h 2 from cell 1403 and the cell state vector c 2 from cell 1403 and generates an output vector h 3 Additional cells are also available, and an LSTM with four cells is merely an example.
[0105] FIG. 15 shows an exemplary implementation of an LSTM cell 1500 that can be used for cells 1401, 1402, 1403, and 1404 of FIG. 14. The LSTM cell 1500 receives an input vector x(t), a cell state vector c(t−1) from a preceding cell, and an output vector h(t−1) from a preceding cell, and generates a cell state vector c(t) and an output vector h(t).
[0106] The LSTM cell 1500 includes sigmoid function devices 1501, 1502, and 1503, each of which applies a number between 0 and 1 to control the degree to which each component of the input vector contributes to the output vector. The LSTM cell 1500 also includes tanh devices 1504 and 1505 for applying a hyperbolic tangent function to the input vector, multiplier devices 1506, 1507, and 1508 for multiplying two vectors, and an adder device 1509 for adding two vectors. The output vector h(t) can be provided to the next LSTM cell in the system or accessed for other purposes.
[0107] FIG. 16 shows an LSTM cell 1600, which is an example implementation of the LSTM cell 1500. For the convenience of the reader, the same numbering scheme from the LSTM cell 1500 is used in the LSTM cell 1600. The sigmoid function devices 1501, 1502, and 1503, and the tanh device 1504 each include a plurality of VMM arrays 1601 and activation circuit blocks 1602. Thus, it can be understood that the VMM array is particularly useful in LSTM cells used in a particular neural network system.
[0108] An alternative to the LSTM cell 1600 (and another example of the implementation of the LSTM cell 1500) is shown in FIG. 17. In FIG. 17, the sigmoid function devices 1501, 1502, and 1503, and the tanh device 1504 share the same physical hardware (the VMM array 1701 and the activation function block 1702) in a time-division multiplexed manner. The LSTM cell 1700 also includes a multiplier device 1703 for multiplying two vectors, an adder device 1708 for adding two vectors, a tanh device 1505 (including the activation circuit block 1702), a register 1707 for storing the value i(t) output from the sigmoid function block 1702, and the value f(t) output from the multiplier device 1703 via the multiplexer 1710 * a register 1704 for storing c(t-1), and the value i(t) output from the multiplier device 1703 via the multiplexer 1710 * a register 1705 for storing u(t), and the value o(t) output from the multiplier device 1703 via the multiplexers 1710 and 1709 * and a register 1706 for storing c~(t).
[0109] While the LSTM cell 1600 includes a plurality of sets of the VMM arrays 1601 and their respective activation function blocks 1602, the LSTM cell 1700 includes only one set of the VMM array 1701 and the activation function block 1702 used to represent multiple layers in the embodiment of the LSTM cell 1700. The LSTM cell 1700 requires less space than the LSTM 1600 because it requires only 1 / 4 of the space required for the VMM and the activation function block compared to the LSTM cell 1600.
[0110] An LSTM unit typically includes a plurality of VMM arrays, each of which can be understood to require functions provided by specific circuit blocks outside the VMM array, such as adders, activation circuit blocks, and high-voltage generation blocks. Providing separate circuit blocks for each VMM array would require a significant amount of space within the semiconductor device and would be somewhat inefficient. Thus, in the embodiments described below, an attempt is made to minimize the circuitry required outside the VMM array itself. Gated Recurrent Unit
[0111] The analog VMM implementation can be utilized in a gated recurrent unit (GRU) system. A GRU is a gating mechanism within a recurrent neural network. A GRU is similar to an LSTM, except that a GRU cell generally includes fewer components than an LSTM cell.
[0112] FIG. 18 shows an exemplary GRU 1800. The GRU 1800 in this example includes cells 1801, 1802, 1803, and 1804. Cell 1801 receives an input vector x 0 and generates an output vector h 0 Cell 1802 receives the input vector x 1 and the output vector h 0 from cell 1801 and generates an output vector h 1 Cell 1803 receives the input vector x 2 and the output vector (hidden state) h 1 from cell 1802 and generates an output vector h 2 Cell 1804 receives the input vector x 3 and the output vector (hidden state) h 2 from cell 1803 and generates an output vector h 3 Additional cells are also available, and a GRU with four cells is merely an example.
[0113] FIG. 19 shows an exemplary implementation of GRU cell 1900 that can be used for cells 1801, 1802, 1803, and 1804 of FIG. 18. The GRU cell 1900 receives an input vector x(t) and an output vector h(t−1) from a preceding GRU cell, and generates an output vector h(t). The GRU cell 1900 includes sigmoid function devices 1901 and 1902, each of which applies a number between 0 and 1 to components from the output vector h(t−1) and the input vector x(t). The GRU cell 1900 also includes a tanh device 1903 for applying a hyperbolic tangent function to the input vector, multiplier devices 1904, 1905, and 1906 for multiplying two vectors, an adder device 1907 for adding two vectors, and a complementary device 1908 for subtracting an input from 1 to generate an output.
[0114] FIG. 20 shows a GRU cell 2000, which is an example implementation of the GRU cell 1900. For the convenience of the reader, the same numbering method from the GRU cell 1900 is used for the GRU cell 2000. As can be seen from FIG. 20, the sigmoid function devices 1901 and 1902, and the tanh device 1903 each include a plurality of VMM arrays 2001 and activation function blocks 2002. Thus, it can be understood that the VMM array is particularly used in GRU cells used in a specific neural network system.
[0115] An alternative example of the GRU cell 2000 (and another example of the implementation of the GRU cell 1900) is shown in FIG. 21. In FIG. 21, the GRU cell 2100 uses a VMM array 2101 and an activation function block 2102. When configured as a sigmoid function, by applying a number between 0 and 1, it controls the degree to which each component of the input vector contributes to the output vector. In FIG. 21, the sigmoid function devices 1901 and 1902, and the tanh device 1903 share the same physical hardware (VMM array 2101 and activation function block 2102) in a time-division multiplexed manner. The GRU cell 2100 also includes a multiplier device 2103 for multiplying two vectors, an adder device 2105 for adding two vectors, a complementary device 2109 for subtracting the input from 1 to generate an output, a multiplexer 2104, and the value h(t - 1) output from the multiplier device 2103 via the multiplexer 2104 * A register 2106 that holds r(t), and the value h(t - 1) output from the multiplier device 2103 via the multiplexer 2104 * A register 2107 that holds z(t), and the value h^(t) output from the multiplier device 2103 via the multiplexer 2104 * A register 2108 that holds (1 - z((t)), and.
[0116] While the GRU cell 2000 includes multiple sets of a plurality of VMM arrays 2001 and activation function blocks 2002, the GRU cell 2100 includes only one set of a VMM array 2101 and an activation function block 2102 used to represent multiple layers in an embodiment of the GRU cell 2100. The GRU cell 2100 requires less space than the GRU cell 2000 because it requires only 1 / 3 of the space required for the VMM and the activation function block.
[0117] A GRU system typically includes a plurality of VMM arrays, each of which can be understood to require functions provided by specific circuit blocks outside the VMM array, such as adders, activation circuit blocks, and high-voltage generation blocks. Providing separate circuit blocks for each VMM array would require a significant amount of space within the semiconductor device and would be somewhat inefficient. Therefore, in the embodiments described below, an attempt is made to minimize the circuits required outside the VMM array itself.
[0118] The input to the VMM array can be at an analog level, a binary level, or a digital bit (in which case a DAC is required to convert the digital bit to an appropriate input analog level), and the output can be at an analog level, a binary level, or a digital bit (in which case an output ADC is required to convert the output analog level to a digital bit).
[0119] For each memory cell within the VMM array, each weight w can be implemented by a single memory cell, or by differential cells, or by two blended memory cells (the average of two cells). In the case of differential cells, two memory cells are required to implement the weight w as a differential weight (w = w+ - w-). In the case of two blended memory cells, two memory cells are required to implement the weight w as the average of two cells. Embodiments for precise programming of cells within the VMM
[0120] Figure 22A shows a programming method 2200. First, the method typically starts in response to a received program command (step 2201). Next, a bulk programming operation programs all cells to the "0" state (step 2202). Then, a soft erase operation erases all cells to an intermediate weak erase level so that each cell draws a current of about 3 - 5 μA during a read operation (step 2203). This is in contrast to a deeply erased level where each cell draws a current of about 20 - 30 μA during a read operation. Next, a hard program that adds electrons to the floating gates of the cells to a very deeply programmed state is performed on all non - selected cells (step 2204) to ensure that those cells are truly "off", i.e., those cells draw only a negligible amount of current during a read operation.
[0121] Then, a rough programming method is executed on the selected cells (step 2205), followed by a fine programming method being executed on the selected cells (step 2206) to program the desired precise value into each selected cell.
[0122] Figure 22B shows another programming method 2210 similar to the programming method 2200. However, after the method starts (step 2201), instead of a program operation that programs all cells to the "0" state as in step 2202 of Figure 22A, an erase operation is used to erase all cells to the "1" state (step 2212). Then, a soft program operation (step 2213) is used to program all cells to an intermediate state (level) so that each cell draws a current of about 3 - 5 μA during a read operation. Thereafter, a rough programming method and a fine programming method follow as in the case of Figure 22A. A variation of the embodiment of Figure 22B completely removes the soft programming method (step 2213).
[0123] Figure 23 shows a first embodiment of the rough programming method 2205, which is the search and execution method 2300. First, a lookup table search is performed to determine the rough target current value (I CT ) of the selected cell based on the value intended to be stored in the selected cell (step 2301). Assume that the selected cell can be programmed to store one of N possible values (e.g., 128, 64, 32, etc.). Each of the N values may correspond to a different desired current value (I D ) drawn during the read operation by the selected cell. In one embodiment, the lookup table may include M possible current values to be used as the rough target current value I CT of the selected cell during the implementation of the search and execution method 2300, where M is an integer less than N. For example, when N is 8, M may be 4, which means that there are 8 possible values that the selected cell can store, and one of the 4 rough target current values is selected as the rough target of the search and execution method 2300. That is, the search and execution method 2300 (although repetitive, is an embodiment of the rough programming method 2205) is intended to quickly program the selected cell to a value (I D ) somewhat close to the desired value (I CT ), and then the fine programming method 2206 is intended to more precisely program the selected cell so that it is extremely close to the desired value (I D ).
[0124] Examples of cell values, desired current values, and rough target current values are shown in Tables 9 and 10 for a simple example of N = 8 and M = 4. Table 9: Examples of N Desired Current Values for N = 8
Table 9
Table 10
[0125] Coarse target current value I CT Once selected, the selected cell is programmed by applying a voltage v to the appropriate terminal of the selected cell based on the cell architecture type of the selected cell (e.g., memory cells 210, 310, 410, or 510) (step 2302). If the selected cell is of the type of memory cell 310 in FIG. 3, the voltage v 0 is applied to the control gate terminal 28, and v 0 is, and v 0 is the coarse target current value I CT can be 5 - 7V depending on. The value of v 0 can optionally be determined from a voltage look-up table that stores v CT corresponding to the coarse target current value I 0 .
[0126] Next, the selected cell is programmed by applying a voltage v i =v i-1 +v increment , where i starts at 1 and increments each time this step is repeated, and v increment is a small voltage that causes programming commensurate with the desired granularity of change (step 2303). Thus, the first time step 2303 is executed with i = 1, v 1 is v 0 +v increment . Then, a read operation is performed on the selected cell, and a verification operation is performed where the current (I cell ) drawn through the selected cell is measured (step 2304). If I cell is less than or equal to I CT (here the first threshold value), the search and execution method 2300 is complete and the fine programming method 2206 can be started. If I cell is not less than or equal to I CT , step 2303 is repeated and i is incremented.
[0127] Thus, when the rough programming method 2205 ends and the precise programming method 2206 starts, the voltage v i is the last voltage used to program the selected cell, and the selected cell will store a value associated with the rough target current value I CT . The goal of the precise programming method 2206 is to program the selected cell to a point where the selected cell draws a current I D (plus or minus an acceptable amount of deviation such as 50 pA or less), which is the desired current value associated with the value intended to be stored in the selected cell.
[0128] FIG. 24 shows an example of different voltage progressions that can be applied to the control gate of the selected memory cell during the precise programming method 2206.
[0129] Under the first approach, an increasing voltage is gradually applied to the control gate to further program the selected memory cell. The starting point is v i , which is the last voltage applied during the rough programming method 2205. An increment v p1 is added to v 1 , and then the voltage v 1 +v p1 is used to program the selected cell (shown by the second pulse from the left in progression 2401). v p1 is an increment smaller than v increment (the voltage increment used during the rough programming method 2205). After each programming voltage is applied, a verification step (similar to step 2304) is performed to determine whether Icell is less than or equal to I PT1 (the first precise target current value, which is the second threshold value here), where I PT1 =I D +I PT1OFFSET , and I PT1OFFSET is an offset value added to prevent program overshoot. If the determination is false, another increment vp1 is added to the previously applied programming voltage and the process is repeated. I cell is I PT1 At the point when it is below, this part of the programming sequence stops. Optionally, I PT1 is I D is equal to, or approximately equal to I with sufficient accuracy D to, the selected memory cell is successfully programmed.
[0130] I PT1 is I D is not sufficiently close to, further programming with a smaller granularity can be performed. Here, progression 2402 is used. The starting point of progression 2402 is the last voltage used for programming under progression 2401. Increment V p2 (smaller than v p1 ) is added to that voltage and the combined voltage is applied to program the selected memory cell. After each programming voltage is applied, I cell is I PT2 (the second precision target current value, here the third threshold value) is determined whether it is below or not, a verification step (similar to step 2304) is executed, I PT2 = ID + I PT2OFFSET and I PT2OFFSET is the offset value added to prevent program overshoot. If the determination is false, another increment V p2 is added to the previously applied programming voltage and the process is repeated. I cell is I PT2 At the point when it is below, this part of the programming sequence stops. Here, since the target value is achieved with sufficient accuracy, I PT2 is I D is equal to, or programming can stop so that I DIt is assumed to be sufficiently close. A person skilled in the art can understand that the programming increments used can be gradually reduced and additional progressions can be applied. For example, in FIG. 25, not only two but three progressions (2501, 2502, and 2503) are applied.
[0131] A second approach is shown in progression 2403. Here, instead of increasing the voltage applied during the programming of the selected memory cell, the same voltage is applied for an increasing duration. Instead of adding incremental voltages such as v in progression 2401 p1 and v in progression 2403 p2 each applied pulse is made t p1 longer than the previously applied pulse by adding an additional time increment t p1 to the programming pulse. After each programming pulse is applied, the same verification steps as described above for progression 2401 are executed. Optionally, an additional progression can be applied where the additional time increment added to the programming pulse has a shorter duration than the previously used progression. Although only one temporal progression is shown, a person skilled in the art will understand that any number of different temporal progressions can be applied.
[0132] Here, more details are provided regarding two further embodiments of the rough programming method 2205.
[0133] FIG. 26 shows a second embodiment of the rough programming method 2205, which is the adaptive calibration method 2600. The method starts (step 2601). The cell is programmed with a default starting value v 0 (step 2602). Different from the search and execution method 2300, here v 0Cannot be obtained from the look-up table and can instead be a relatively small initial value. The control gate voltage of the cell is measured at a first current value IR1 (e.g., 100 na) and a second current value IR2 (e.g., 10 na), and the subthreshold slope is determined (e.g., 360 mV / dec) based on those measured values and stored (step 2603).
[0134] The new desired voltage v i is determined. When this step is first executed, i = 1 and v 1 is determined based on the stored subthreshold slope value, as well as the current target and offset values, using the following subthreshold equation. Vi = Vi-1 + Increment, where Increment is proportional to the slope Vg Vg = k * Vt * log[Ids / wa * Io] where wa is the w of the memory cell and Ids is the current target plus the offset value.
[0135] If the stored slope value is relatively steep, a relatively small current offset value can be used. If the stored slope value is relatively flat, a relatively high current offset value can be used. Thus, determining the slope information allows a current offset value customized for the particular cell in question to be selected. This ultimately makes the programming process shorter. When this step is repeated, i is incremented and v i = v i-1 + v increment . The cell is then programmed using vi. v increment can be determined from a look-up table that stores the value of v increment corresponding to the target current value.
[0136] Next, a read operation is performed on the selected cell, and the current (Icell ) is measured, and a verification operation is performed (step 2605). I cell is I CT If it is below (here it is the rough target threshold value) (I CT = I D + I CTOFFSET is set to, and I CTOFFSET is an offset value added to prevent program overshoot), the adaptive calibration method 2600 is completed, and the precise programming method 2206 can be started. I cell If it is not I CT or less, steps 2604 to 2605 are repeated, and i is incremented.
[0137] FIG. 27 shows an aspect of the adaptive calibration method 2600. During step 2603, the current source 2701 is used to apply exemplary current values IR1 and IR2 to the selected cell (here the memory cell 2702), and then the voltages at the control gate of the memory cell 2702 (CGR1 for IR1 and CGR2 for IR2) are measured. The slope is (CGR2 - CGR1) / dec.
[0138] FIG. 28 shows a second embodiment of the rough programming method 2205, which is the absolute calibration method 2800. The method starts (step 2801). The cell is programmed with a default starting value V 0 (step 2802). The control gate voltage of the cell (VCGRx) is measured and stored with the current value Itarget (step 2803). The new desired voltage v 1 is determined based on the stored control gate voltage, as well as the current target and offset value Ioffset + Itarget (step 2804). For example, the new desired voltage v 1 can be calculated as follows: v 1 = v 0 + (VCGBIAS - stored VCGR), where VCGBIAS = ~1.5V, which is the default read control gate voltage at the maximum target current, and the stored VCGR is the measured read control gate voltage of step 2803.
[0139] The cell then i When i=1, the voltage v from step 2804 is programmed using 1 When i>=2, the voltage v i =v i-1 +V increment is used. increment corresponds to the target current value v increment A read operation is then performed on the selected cell, and the current drawn through the selected cell (I cell ) is measured and a verification operation is performed (step 2806). cell I CT If so, the absolute calibration method 2800 is complete and the fine programming method 2206 may begin. cell I CT If not, steps 2805-2806 are repeated and i is incremented.
[0140] FIG. 29 shows a circuit 2900 for implementing step 2803 of the absolute calibration method 2800. A voltage source (not shown) generates a VCGR, which starts at an initial voltage and rises. Here, n + 1 different current sources 2901 (2901-0, 2901-1, 2901-2,..., 2901-n) generate different increasing currents IO0, IO1, IO2,... IOn. Each current source 2901 is connected to an inverter 2902 (2902-0, 2902-1, 2902-2,..., 2902-n) and a memory cell 2903 (2903-0, 2903-1, 2903-2,... 2903-n). As the VCGR rises, each memory cell 2903 draws an increasing amount of current, and the input voltage to each inverter 2902 decreases. Since IO0 < IO1 < IO2 <... < IOn, as the VCGR increases, first the output of inverter 2902-0 switches from low to high. Next, the output of inverter 2902-1 switches from low to high, then inverter 2902-2 switches in the same way, and so on until the output of inverter 2902-n switches from low to high. Each inverter 2902 controls a switch 2904 (2904-0, 2904-1, 2904-2,..., 2904-n), so that when the output of the inverter 2902 is high, the switch 2904 is closed, whereby the VCGR is sampled by the capacitor 2905 (2905-0, 2905-1, 2905-2,..., 2905-n). Thus, the switches 2904 and the capacitors 2905 form a sample-and-hold circuit. The values of IO0, IO1, IO2,... IOn are used as possible values of Itarget, and each sampled voltage is used as the associated value VCGRx in the absolute calibration method 2800 of FIG. 28. Graph 2906 shows the rising VCGR over time, as well as the outputs of inverters 2902-0, 2902-1, and 2902-n that switch from low to high at various times.
[0141] FIG. 30 shows an exemplary progression 3000 for programming selected cells during the adaptive calibration method 2600 or the absolute calibration method 2800. In one embodiment, the voltage Vcgp is applied to the control gate of the memory cells in the selected row. The number of selected memory cells in the selected row is, for example, 32. Thus, up to 32 memory cells in the selected row can be programmed in parallel. Each memory cell can be coupled to the programming current Iprog by a bit line enable signal. When the bit line enable signal is inactive (meaning a positive voltage is applied to the selected bit line), the memory cell is in an inhibited state (not programmed). As shown in FIG. 30, the bit line activation signals En_blx (where x varies from 1 to n, and n is the number of bit lines) are activated at different times at the desired Vcgp voltage level for that bit line (and thus for the selected memory on that bit line). In another embodiment, the voltage applied to the control gate of the selected cell can be controlled using an enable signal on the bit line. Each bit line enable signal applies the desired voltage (such as vi described in FIG. 28) corresponding to that bit line as Vcgp. The bit line enable signal can also control the programming current flowing into the bit line. In this example, each subsequent control gate voltage Vcgp is higher than the previous voltage. Alternatively, each subsequent control gate voltage can be either lower or higher than the previous voltage. Each subsequent increment of Vcgp can be either equal to or not equal to the previous increment.
[0142] FIG. 31 shows an exemplary progression 3100 for programming selected cells during the adaptive calibration method 2600 or the absolute calibration method 2800. In one embodiment, the bit line enable signal enables the selected bit line (meaning the selected memory cell within the bit line) to be programmed at the corresponding Vcgp voltage level. In another embodiment, the voltage applied to the control gate that performs the incremental raise of the selected cell can be controlled using the bit line enable signal. Each bit line enable signal applies the desired voltage (such as vi described in FIG. 28) corresponding to that bit line to the control gate voltage. In this example, each subsequent increment is equal to the previous increment.
[0143] FIG. 32 shows a system for implementing an input and output method for reading or verifying in a VMM array. The input function circuit 3201 receives digital bit values, converts those digital values into analog signals for use, and applies a voltage to the control gate of the selected cell in the array 3204 determined via the control gate decoder 3202. At the same time, the word line decoder 3203 is also used to select the row in which the selected cell is located. The output neuron circuit block 3205 performs the output function of each column (neuron) of cells in the array 3204. The output circuit block 3205 can be implemented using an integrated analog-to-digital converter (ADC), a successive approximation register (SAR) ADC, or a sigma-delta ADC.
[0144] In one embodiment, the digital values provided to the input function circuit 3201 include, by way of example, four bits (DIN3, DIN2, DIN1, and DIN0), and the various bit values correspond to different numbers of input pulses applied to the control gate. The larger the number of pulses, the larger the output value (current) of the cell. Examples of bit values and pulse values are shown in Table 11. Table 11: Digital Bit Inputs and Generated Pulse Numbers
Table 11
[0145] In the above example, there are at most 16 pulses for a 4-bit digital value for reading the cell value. Each pulse is equal to 1 unit of cell value (current). For example, when Icell unit = 1 nA, if DIN[3~0] = 0001, then Icell = 1 * 1 nA = 1 nA, and if DIN[3~0] = 1111, then Icell = 15 * 1 nA = 15 nA.
[0146] In another embodiment, the digital bit input uses digital bit position addition to read the cell value, as shown in Table 12. Here, only 4 pulses are required to evaluate a 4-bit digital value. For example, the first pulse is used to evaluate DIN0, the second pulse is used to evaluate DIN1, the third pulse is used to evaluate DIN2, and the fourth pulse is used to evaluate DIN3. Then, the results from the 4 pulses are added according to the bit position. The realized digital bit addition formula is as follows: Output = 2^0 * DIN0 + 2^1 * DIN1 + 2^2 * DIN2 + 2^3 * DIN3) * Icell unit.
[0147] For example, when Icell unit = 1 nA, if DIN[3~0] = 0001, then Icell total = 0 + 0 + 0 + 1 * 1 nA = 1 nA, and if DIN[3~0] = 1111, then Icell total = 8 * 1 nA + 4 * 1 nA + 2 * 1 nA + 1 * 1 nA = 15 nA. Table 12: Digital Bit Input Addition
Table 12
[0148] FIG. 33 shows an example of a charge adder 3300 that can be used to sum the outputs of the VMM during the verification operation to obtain a single analog value representing the output, and this single analog value can optionally be converted to a digital bit value. The charge adder 3300 includes a current source 3301 and a sample-and-hold circuit including a switch 3302 and a sample-and-hold (S / H) capacitor 3303. As shown by an example of a 4-bit digital value, there are four S / H circuits for holding the values from four evaluation pulses, and these values are summed at the end of the process. The S / H capacitor 3303 is selected at a ratio associated with the 2^n * DINn-bit position, for example, C_DIN3 = x8 Cu, C_DIN2 = x4 Cu, C_DIN1 = x2 Cu, DIN0 = x1 Cu. The current source 3301 is also multiplied by the ratio accordingly.
[0149] FIG. 34 shows a current adder 3400 that can be used to sum the outputs of the VMM during the verification operation. The current adder 3400 includes a current source 3401, a switch 3402, a switch 3403, a switch 3404, and a switch 3405. As shown by an example of a 4-bit digital value, there are current source circuits for holding the values from four evaluation pulses, and these values are summed at the end of the process. The current source is multiplied by a ratio based on the 2^n * DINn-bit position, for example, I_DIN3 = x8 Icell units, _I_DIN2 = x4 Icell units, I_DIN1 = x2 Icell units, I_DIN0 = x1 Icell units.
[0150] FIG. 35 shows a digital adder 3500 that receives a plurality of digital values, sums them together, and generates an output DOUT representing the sum of the inputs. The digital adder 3500 can be used during the verification operation. As shown by an example of a 4-bit digital value, there are digital output bits for holding the values from four evaluation pulses, and these values are summed at the end of the process. The digital output is 2^n *Digitally scaled based on the DINn bit position, for example, DOUT3 = x8 DOUT0, _DOUT2 = x4 DOUT1, I_DOUT1 = x2 DOUT0, I_DOUT0 = DOUT0.
[0151] Figure 36A shows a dual-slope integrating type ADC 3600 applied to the output neuron to convert the cell current into digital output bits. The integrator consisting of the integrating operational amplifier 3601 and the integrating capacitor 3602 integrates the cell current ICELL with respect to the reference current IREF. As shown in Figure 36B, during the fixed time t1, the cell current is integrated upward (Vout rises), and then the reference current is applied so as to be integrated downward over time t2 (Vout drops). The current Icell = t2 / t1 * is IREF. For example, for t1, with a 10-bit digital bit resolution, 1024 cycles are used, and the number of cycles for t2 varies from 0 to 1024 cycles depending on the Icell value.
[0152] Figure 36C shows a single-slope integrating type ADC 3660 applied to the output neuron to convert the cell current into digital output bits. The integrator consisting of the integrating operational amplifier 3661 and the integrating capacitor 3662 integrates the cell current ICELL. As shown in Figure 36D, during the time t1, the cell current is integrated upward (rises until Vout reaches Vref2), and during the time t2, another cell current is integrated upward. The cell current Icell = Cint * is Vref2 / t. The pulse counter is used to count the number of pulses (digital output bits) during the integration time t. For example, as shown, the digital output bits for t1 are fewer than those for t2, which means that the cell current during t1 is larger than the cell current during t2 integration. Initial calibration is performed to calibrate the integrating capacitor value with the reference current and the fixed time, and Cint = Tref * is Iref / Vref2.
[0153] FIG. 36E shows a dual-slope integrating type ADC 3680 applied to an output neuron to convert a cell current into digital output bits. The dual-slope integrating type ADC 3680 does not utilize an integrating operational amplifier. The cell current or the reference current is directly integrated into a capacitor 3682. A pulse counter is used to count the pulses (digital output bits) during the integration time. The current Icell = t2 / t1 * is IREF.
[0154] FIG. 36F shows a single-slope integrating type ADC 3690 applied to an output neuron to convert a cell current into digital output bits. The single-slope integrating type ADC 3680 does not utilize an integrating operational amplifier. The cell current is directly integrated into a capacitor 3692. A pulse counter is used to count the pulses (digital output bits) during the integration time. The cell current Icell = Cint * Vref2 / t.
[0155] FIG. 37A shows a SAR (successive approximation type) ADC applied to an output neuron to convert a cell current into digital output bits. The cell current can be dropped across a resistor to be converted into VCELL. Alternatively, the cell current can charge up an S / H capacitor to be converted into VCELL. A binary search is used to calculate the bits starting from the MSB bit (the most significant bit). A DAC 3702 is used to set an appropriate analog reference voltage to a comparator 3703 based on the digital bits from the SAR 3701. The output of the comparator 3703 is fed back to the SAR 3701 in turn to select the next analog level. As shown in FIG. 37B, in the example of 4-bit digital output bits, there are 4 evaluation periods, a first pulse for evaluating DOUT3 by setting the analog level to the middle, then a second pulse for evaluating DOUT2 by setting the analog level to the middle of the upper half or the middle of the lower half, and so on.
[0156] FIG. 38 shows a sigma-delta type ADC 3800 applied to an output neuron to convert a cell current into digital output bits. An integrator consisting of an operational amplifier 3801 and a capacitor 3805 integrates the sum of the current from a selected cell current and a reference current provided from a 1-bit current DAC 3804. A comparator 3802 compares the integrated output voltage against a reference voltage. A clocked DFF 3803 provides a digital output stream in response to the output of the comparator 3802. The digital output stream typically proceeds to a digital filter before being output to the digital output bits.
[0157] As used herein, it should be noted that both the terms "over" and "on" include both "directly over" (no intervening material, element, or gap disposed therebetween) and "indirectly over" (an intervening material, element, or gap is disposed therebetween). Similarly, the term "adjacent" includes "directly adjacent" (no intervening material, element, or gap disposed therebetween) and "indirectly adjacent" (an intervening material, element, or gap is disposed therebetween), "attached to" includes "directly attached to" (no intervening material, element, or gap disposed therebetween) and "indirectly attached to" (an intervening material, element, or gap is disposed therebetween), and "electrically coupled" includes "directly electrically coupled" (no intervening material or element electrically connecting the elements together therebetween) and "indirectly electrically coupled" (an intervening material or element electrically connecting the elements together is therebetween). For example, forming an element "over a substrate" can include forming the element directly on the substrate without an intervening material / element therebetween and forming the element indirectly over the substrate with one or more intervening materials / elements therebetween.
Claims
1. 1. A method of programming a selected non-volatile memory cell to store one of N possible values, where N is an integer greater than two, and the selected non-volatile memory cell includes a floating gate, the method comprising: performing a coarse programming process, said coarse programming process comprising: selecting one of M different current values as a first threshold current value, where M<N; adding charge to the floating gate; repeating the adding step until a current through the selected non-volatile memory cell during a verify operation is less than or equal to the first threshold current value; performing a precision programming process until a current through the selected non-volatile memory cell during a verify operation is below a second threshold current value.
2. 2. The method of claim 1, further comprising: performing a second precision programming process until a current through the selected non-volatile memory cell during a verify operation is below a third threshold current value.
3. 2. The method of claim 1, wherein the precision programming process comprises applying voltage pulses of increasing magnitude to the control gates of the selected non-volatile memory cells.
4. 2. The method of claim 1, wherein the precision programming process comprises applying voltage pulses of increasing duration to the control gates of the selected non-volatile memory cells.
5. 3. The method of claim 2, wherein the second precision programming process comprises applying voltage pulses of increasing magnitude to the control gates of the selected non-volatile memory cells.
6. 3. The method of claim 2, wherein the second precision programming process comprises applying voltage pulses of increasing duration to the control gates of the selected non-volatile memory cells.
7. The method of claim 1 , wherein the selected non-volatile memory cell comprises a floating gate.
8. 8. The method of claim 7, wherein the selected non-volatile memory cells are split-gate flash memory cells.
9. 2. The method of claim 1, wherein the selected non-volatile memory cell is in a vector matrix multiplication array in an analog memory deep neural network.
10. Prior to performing the coarse programming process, programming the selected non-volatile memory cells to a "0" state; 2. The method of claim 1, further comprising the step of: erasing the selected non-volatile memory cells to a weak erase level.
11. Prior to performing the coarse programming process, erasing the selected non-volatile memory cells to a "1" state; 2. The method of claim 1, further comprising: programming the selected non-volatile memory cells to a weak program level.
12. performing a read operation on the selected non-volatile memory cell; 5. The method of claim 1, further comprising: integrating the current drawn by the selected non-volatile memory cell during the read operation to generate a digital bit using an integrating analog-to-digital converter.
13. performing a read operation on the selected non-volatile memory cell; 2. The method of claim 1, further comprising: converting the current drawn by the selected non-volatile memory cell during the read operation into a digital bit using a sigma-delta analog-to-digital converter.
14. 1. A method of programming a selected non-volatile memory cell to store one of N possible values, where N is an integer greater than two, the selected non-volatile memory cell including a floating gate and a control gate, the method comprising: performing a coarse programming process, said coarse programming process comprising: determining a slope value based on the change in voltage on the control gate of the selected non-volatile memory cell and the change in current drawn by the selected non-volatile memory cell; determining a next programming voltage value based on the slope value; adding an amount of charge from the floating gate of the selected non-volatile memory cell until a current through the selected non-volatile memory cell during a verify operation is below a first threshold current value; performing a precision programming process until a current through the selected non-volatile memory cell during a verify operation is below a second threshold current value.
15. 15. The method of claim 14, further comprising: performing a second precision programming process until a current through the selected non-volatile memory cell during a verify operation is below a third threshold current value.
16. 15. The method of claim 14, wherein the precision programming process comprises applying voltage pulses of increasing magnitude to the control gates of the selected non-volatile memory cells.
17. 15. The method of claim 14, wherein the precision programming process comprises applying voltage pulses of increasing duration to the control gates of the selected non-volatile memory cells.
18. 16. The method of claim 15, wherein the precision programming process comprises applying voltage pulses of increasing magnitude to the control gates of the selected non-volatile memory cells.
19. 16. The method of claim 15, wherein the precision programming process comprises applying voltage pulses of increasing duration to the control gates of the selected non-volatile memory cells.
20. The step of determining the slope value comprises: applying a programming voltage to the control gate of the selected non-volatile memory cell; applying a first current through the selected non-volatile memory cell to determine a first voltage on the control gate; applying a second current through the selected non-volatile memory cell to determine a second voltage on the control gate; and calculating the slope value by dividing the difference between the second voltage and the first voltage by the difference between the second current and the first current.
21. 15. The method of claim 14, wherein the selected non-volatile memory cells are split-gate flash memory cells.
22. 15. The method of claim 14, wherein the selected non-volatile memory cell is in a vector matrix multiplication array in an analog memory deep neural network.
23. Prior to performing the coarse programming process, programming the selected non-volatile memory cells to a "0" state; 15. The method of claim 14, further comprising: erasing the selected non-volatile memory cells to a weak erase level.
24. Prior to performing the coarse programming process, erasing the selected non-volatile memory cells to a "1" state; programming the selected non-volatile memory cells to a weak program level; The method of claim 14 further comprising:
25. performing a read operation on the selected non-volatile memory cell; 15. The method of claim 14, further comprising: integrating the current drawn by the selected non-volatile memory cell during the read operation to generate a digital bit using an integrating analog-to-digital converter.
26. performing a read operation on the selected non-volatile memory cell; 15. The method of claim 14, further comprising: converting the current drawn by the selected non-volatile memory cell during the read operation into a digital bit using a sigma-delta analog-to-digital converter.
27. 1. A method of programming a selected non-volatile memory cell to store one of N possible values, where N is an integer greater than two, the selected non-volatile memory cell including a floating gate and a control gate, the method comprising: performing a coarse programming process, said coarse programming process comprising: applying a programming voltage to the control gate of the selected non-volatile memory cell; repeating the applying step, increasing the programming voltage by an incremental voltage each time the applying step is performed, until a current through the selected non-volatile memory cell during a verify operation is less than or equal to the threshold current value; performing a precision programming process until a current through the selected non-volatile memory cell during a verify operation is below a second threshold current value.
28. 30. The method of claim 27, further comprising: performing a precision programming process until a current through the selected non-volatile memory cell during a verify operation is below a third threshold current value.
29. 28. The method of claim 27, wherein the precision programming process comprises applying voltage pulses of increasing magnitude to the control gates of the selected non-volatile memory cells.
30. 28. The method of claim 27, wherein the predictive programming process comprises applying voltage pulses of increasing duration to the control gates of the selected non-volatile memory cells.
31. 30. The method of claim 28, wherein the precision programming process comprises applying voltage pulses of increasing magnitude to the control gates of the selected non-volatile memory cells.
32. 30. The method of claim 28, wherein the predictive programming process comprises applying voltage pulses of increasing duration to the control gates of the selected non-volatile memory cells.
33. 30. The method of claim 27, wherein the selected non-volatile memory cell comprises a floating gate.
34. 34. The method of claim 33, wherein the selected non-volatile memory cells are split-gate flash memory cells.
35. 28. The method of claim 27, wherein the selected non-volatile memory cell is in a vector matrix multiplication array in an analog memory deep neural network.
36. Prior to performing the coarse programming process, programming the selected non-volatile memory cells to a "0" state; 28. The method of claim 27, further comprising: erasing the selected non-volatile memory cells to a weak erase level.
37. Prior to performing the coarse programming process, erasing the selected non-volatile memory cells to a "1" state; 30. The method of claim 27, further comprising: programming the selected non-volatile memory cells to a weak program level.
38. performing a read operation on the selected non-volatile memory cell; 30. The method of claim 27, further comprising: integrating the current drawn by the selected non-volatile memory cell during the read operation to generate a digital bit using an integrating analog-to-digital converter.
39. performing a read operation on the selected non-volatile memory cell; 28. The method of claim 27, further comprising: converting the current drawn by the selected non-volatile memory cell during the read operation into a digital bit using a sigma-delta analog-to-digital converter.
40. 1. A method of reading a selected non-volatile memory cell that stores one of N possible values, where N is an integer greater than two, the method comprising: applying a digital input pulse to the selected non-volatile memory cell; determining a value stored in the selected non-volatile memory cell based on an output of the selected non-volatile memory cell in response to each of the digital input pulses.
41. 41. The method of claim 40, wherein the number of digital input pulses corresponds to a binary value.
42. 41. The method of claim 40, wherein the number of digital input pulses corresponds to a digital bit position value.
43. 41. The method of claim 40, wherein the determining step comprises receiving output neurons in an integrated analog-to-digital converter and generating a digital bit indicative of the value stored in the non-volatile memory cell.
44. 41. The method of claim 40, wherein the determining step comprises receiving output neurons in a successive approximation analog-to-digital converter and generating a digital bit indicative of the value stored in the non-volatile memory cell.
45. 41. The method of claim 40, wherein the output is a current.
46. 41. The method of claim 40, wherein the output is an electric charge.
47. 41. The method of claim 40, wherein the output is a digital bit.
48. 41. The method of claim 40, wherein the selected non-volatile memory cell comprises a floating gate.
49. 49. The method of claim 48, wherein the selected non-volatile memory cells are split-gate flash memory cells.
50. 41. The method of claim 40, wherein the selected non-volatile memory cell is in a vector matrix multiplication array in an analog memory deep neural network.
51. 1. A method of reading a selected non-volatile memory cell that stores one of N possible values, where N is an integer greater than two, the method comprising: applying an input to the selected non-volatile memory cell; determining a value stored in the selected non-volatile memory cell based on an output of the selected non-volatile memory cell using an analog-to-digital converter circuit in response to the input.
52. 52. The method of claim 51, wherein the input is a digital input.
53. 52. The method of claim 51, wherein the input is an analog input.
54. 52. The method of claim 51 , wherein the determining step comprises receiving output neurons in a single or dual slope integrating analog to digital converter and generating a digital bit indicative of the value stored in the non-volatile memory cell.
55. 52. The method of claim 51, wherein the determining step comprises receiving output neurons in a SAR analog-to-digital converter and generating a digital bit indicative of the value stored in the non-volatile memory cell.
56. 52. The method of claim 51, wherein the determining step comprises receiving output neurons in a sigma-delta analog-to-digital converter and generating a digital bit indicative of the value stored in the non-volatile memory cell.
57. 52. The method of claim 51, wherein the selected non-volatile memory cell comprises a floating gate.
58. 52. The method of claim 51, wherein the selected non-volatile memory cells are split-gate flash memory cells.
59. 52. The method of claim 51, wherein the selected non-volatile memory cell is in a vector matrix multiplication array in an analog memory deep neural network.
60. 52. The method of claim 51, wherein the selected non-volatile memory cells operate in the sub-threshold region.
61. 52. The method of claim 51, wherein the selected non-volatile memory cells operate in a linear region.
Citation Information
Patent Citations
Nonvolatile semiconductor multilevel memory
JP1997091971A
Deep learning neural network classifier using non-volatile memory array
WO2017200883A1