Precise data tuning method and apparatus for analog neural memory in artificial neural network
The precision tuning algorithm for artificial neural networks precisely programs non-volatile memory cells by using differential pairs, achieving high accuracy and efficiency in storing weight values within the VMM array.
Patent Information
- Application Number
- JP2025012856
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-02-25
- Filing Date
- 2025-01-29
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-05-19
AI Technical Summary
Existing technologies face challenges in precisely programming non-volatile memory cells in a vector-by-matrix multiplication (VMM) array within an artificial neural network, requiring high accuracy and granularity to store different weight values.
A precision tuning algorithm and apparatus that store weight values as differential pairs (w+ and w-) in non-volatile memory cells, allowing for precise programming and verification of zero values, and distributing storage evenly among columns in the array.
Enables extremely high accuracy in programming non-volatile memory cells, allowing them to hold one of N different values, thereby improving the performance and efficiency of artificial neural networks.
Smart Images

Figure 2025081343000001_ABST
Abstract
Description
Technical Field
[0001] (Claim of Priority) This application claims the benefit of priority of U.S. Provisional Patent Application No. 62 / 957,013, filed on January 3, 2020, entitled "Precise Data Tuning Method and Apparatus for Analog Neuromorphic Memory in an Artificial Neural Network", and is a continuation-in-part of U.S. Patent Application No. 16 / 829,757, filed on March 25, 2020, entitled "Precise Data Tuning Method and Apparatus for Analog Neuromorphic Memory in an Artificial Neural Network", and claims the benefit of priority of U.S. Patent Application No. 17 / 185,725, filed on February 25, 2021, entitled "Precise Data Tuning Method and Apparatus for Analog Neural Memory in an Artificial Neural Network".
[0002] (Field of the Invention) A number of embodiments are disclosed for precise tuning methods and apparatuses for precisely and rapidly depositing an accurate amount of charge on the floating gates of non-volatile memory cells in a vector-by-matrix multiplication (VMM) array within an artificial neural network.
Background Art
[0003] Artificial neural networks mimic biological neural networks (the central nervous system of animals, especially the brain), can rely on multiple inputs, and are used to estimate or approximate functions that are generally unknown. Artificial neural networks generally include layers of interconnected "neurons" that exchange messages with each other.
[0004] Figure 1 shows an artificial neural network, in which the circles represent the input or layers of neurons. Connections (referred to as synapses) are represented by arrows and have numerical weights that can be tuned based on experience. This enables the artificial neural network to adapt to the input and become learnable. Typically, an artificial neural network includes multiple input layers. Typically, there is one or more intermediate layers of neurons and an output layer of neurons that provides the output of the neural network. At each level, the neurons make decisions individually or collectively based on the data received from the synapses.
[0005] One of the major challenges in the development of artificial neural networks for high-performance information processing is the lack of appropriate hardware technologies. In practice, practical artificial neural networks rely on a very large number of synapses, which enables a high connectivity between neurons, i.e., a very high degree of parallelization of computational processing. In principle, such complexity can be realized by digital supercomputers or dedicated graphics processing unit clusters. However, in addition to high costs, these approaches also suffer from poor energy efficiency compared to biological networks, which mainly perform low-precision analog calculations and consume far less energy. CMOS analog circuits have been used for artificial neural networks, but most CMOS-implemented synapses have been too bulky assuming a large number of neurons and synapses.
[0006] The applicant has previously disclosed, in U.S. Patent Application No. 15 / 594,439, published as U.S. Patent Publication No. 2017 / 0337466, incorporated by reference, an artificial (analog) neural network that utilizes one or more non-volatile memory arrays as synapses. The non-volatile memory arrays operate as analog neuromorphic memories. As used herein, the term neuromorphic means a circuit that implements a model of the nervous system. An analog neuromorphic memory includes a first plurality of synapses configured to receive a first plurality of inputs and then generate a first plurality of outputs, and a first plurality of neurons configured to receive the first plurality of outputs. The first plurality of synapses includes a plurality of memory cells, each memory cell including a source region and a drain region spaced apart within a semiconductor substrate with a channel region extending therebetween, a floating gate insulated and disposed above a first portion of the channel region, and a non-floating gate insulated and disposed above a second portion of the channel region. Each of the plurality of memory cells is configured to store a weight value corresponding to a plurality of electrons in the floating gate. The plurality of memory cells are configured to multiply the stored weight values by the first plurality of inputs to generate the first plurality of outputs. An array of memory cells arranged in this manner may be referred to as a vector matrix multiplication (VMM) array.
[0007] Each non-volatile memory cell used in the VMM array must be erased and programmed to hold a charge, i.e., a number of electrons, in the floating gate in a very specific and precise amount. For example, each floating gate must hold one of N different values, where N is the number of different weights that can be represented by each cell. Examples of N include 16, 32, 64, 128, and 256. One challenge is the ability to program cells selected with the accuracy and granularity required for different values of N. For example, if a selected cell can include one of 64 different values, a very high accuracy is required in the programming operation.
[0008] What is needed is an improved programming system and method suitable for use with a VMM array in an analog neuromorphic memory.
SUMMARY OF THE INVENTION
[0009] Numerous embodiments are disclosed for a precision tuning algorithm and apparatus for precisely and rapidly depositing an accurate amount of charge on the floating gates of non-volatile memory cells in a VMM array within an analog neuromorphic memory system. Thereby, selected cells can be programmed with extremely high accuracy to hold one of N different values.
[0010] In one embodiment, a neural network includes a vector matrix multiplication array of non-volatile memory cells, and the weight value w is stored as a differential pair w+ and w- of a first non-volatile memory cell and a second non-volatile memory cell in the array according to the formula w=(w+)-(w-), where w+ and w- include non-zero offset values.
[0011] In another embodiment, a neural network comprises a vector matrix multiplication array of non-volatile memory cells, the array is organized into rows and columns of non-volatile memory cells, and the weight value w is stored as a differential pair w+ and w- of a first non-volatile memory cell and a second non-volatile memory cell according to the formula w=(w+)-(w-), where the storage of the w+ and w- values is distributed approximately evenly among all columns in the array.
[0012] In another embodiment, a method of programming, verifying, and reading zero values in a differential pair of non-volatile memory cells in a vector matrix multiplication array includes programming a first cell w+ of the differential pair to a first current value, verifying the first cell by applying a voltage equal to the first voltage plus a bias voltage to a control gate terminal of the first cell, programming a second cell w- of the differential pair to the first current value, verifying the second cell by applying a voltage equal to the first voltage plus a bias voltage to a control gate terminal of the second cell, reading the first cell by applying a voltage equal to the first voltage to a control gate terminal of the first cell, reading the second cell by applying a voltage equal to the first voltage to a control gate terminal of the second cell, and calculating a value w according to the formula w=(w+)-(w-). 。
[0013] Separate In an embodiment, the neural network includes a vector matrix multiplication array of non-volatile memory cells, the array is organized into rows and columns of non-volatile memory cells, the weight value w is stored as differential pairs w+ and w- according to the formula w=(w+)-(w-), w+ is stored as a differential pair of a first non-volatile memory cell and a second non-volatile memory cell in the array, w- is stored as a differential pair of a third non-volatile memory cell and a fourth non-volatile memory cell in the array, and the storage of w+ and w- is offset by a bias value.
[0014] In another embodiment, the neural network includes a vector matrix multiplication array of non-volatile memory cells, the array is organized into rows and columns of non-volatile memory cells, the weight value w is stored as differential pairs w+ and w- of a first non-volatile memory cell and a second non-volatile memory cell according to the formula w=(w+)-(w-), the value for w+ is selected from a first range of non-zero values, the value for w- is selected from a second range of non-zero values, and the first range and the second range do not overlap.
[0015] In another embodiment, a method of reading a non-volatile memory cell in a vector matrix multiplication array includes the step of reading the weights stored in a selected cell in the array, and the step of reading includes applying a zero voltage bias to the control gate terminal of the selected cell and detecting a neuron output current including the current output from the selected cell.
[0016] In another embodiment, a method of operating a non-volatile memory cell in a vector matrix multiplication array includes reading the non-volatile memory cell by applying a first bias voltage to the control gate of the non-volatile memory cell, and during one or more of standby operation, deep power down operation, or test operation, reading the non-volatile memory cell by applying a first bias voltage to the control gate of the non-volatile memory cell and applying a second bias voltage to the control gate of the non-volatile memory cell.
[0017]
[0018]
[0019]
[0020]
[0021]
[0022]
[0023]
[0024]
[0025]
[0026]
[0027]
[0028]
[0029]
[0030]
[0031]
[0032]
[0033]
[0034]
[0035]
[0036]
[0037]
[0038]
[0039]
[0040]
[0041]
[0042]
[0043]
[0044]
[0045]
[0046]
[0047]
[0048]
[0049]
[0050]
[0051]
[0052]
[0053]
[0054]
[0055]
[0056]
[0057]
[0058]
[0059]
[0060]
[0061]
[0062]
[0063]
[0064]
[0065]
[0066]
[0067]
[0068]
[0069]
[0070]
Brief Description of the Drawings
[0071]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
Figure 26
Figure 27
Figure 28
Figure 29
Figure 30
Figure 31
Figure 32
Figure 33
Figure 34
Figure 35A
Figure 35B
Figure 36A
Figure 36B
Figure 36C
Figure 36D
Figure 37
Figure 38
Figure 39
Figure 40
Figure 41
Figure 42
Figure 43
Figure 44
Figure 45
Figure 46
Figure 47
Figure 48
Figure 49A
Figure 49B
Figure 49C
Embodiments for Carrying Out the Invention
[0072] The artificial neural network of the present invention utilizes a combination of CMOS technology and a non-volatile memory array. Non-volatile memory cell
[0073] Digital non-volatile memories are well known. For example, U.S. Patent No. 5,029,130 (the " '130 patent"), which is incorporated herein by reference, discloses an array of split-gate non-volatile memory cells, which are a type of flash memory cell. Such a memory cell 210 is shown in FIG. 2. Each memory cell 210 includes a source region 14 and a drain region 16 formed in a semiconductor substrate 12, and there is a channel region 18 between the source region 14 and the drain region 16. The floating gate 20 is formed insulated above a first portion of the channel region 18 (and controls the conductivity of the first portion of the channel region 18) and extends over a portion of the source region 14. The word line terminal 22 (typically coupled to a word line) is disposed insulated above a second portion of the channel region 18 and has a first portion (which controls the conductivity of the second portion of the channel region 18) and a second portion that extends upward above the floating gate 20. The floating gate 20 and the word line terminal 22 are insulated from the substrate 12 by a gate oxide. The bit line terminal 24 is coupled to the drain region 16.
[0074] By applying a high positive voltage to the word line terminal 22, erasure is performed on the memory cell 210 (electrons are removed from the floating gate). As a result, the electrons in the floating gate 20 pass through the insulator between them from the floating gate 20 to the word line terminal 22 via Fowler-Nordheim tunneling.
[0075] The memory cell 210 is programmed (electrons are applied to the floating gate) by applying a positive voltage to the word line terminal 22 and a positive voltage to the source region 14. The electron current flows from the source region 14 (source line terminal) toward the drain region 16. The electrons accelerate and heat up when they reach the gap between the word line terminal 22 and the floating gate 20. A part of the heated electrons is injected into the floating gate 20 through the gate oxide due to the electrostatic attraction from the floating gate 20.
[0076] The memory cell 210 is read by applying a positive read voltage to the drain region 16 and the word line terminal 22 (turning on the portion of the channel region 18 below the word line terminal). When the floating gate 20 is positively charged (i.e., electrons are erased), the portion of the channel region 18 below the floating gate 20 is also turned on, and current flows through the channel region 18, which is detected as the erased state, i.e., the "1" state. When the floating gate 20 is negatively charged (i.e., programmed with electrons), the portion of the channel region below the floating gate 20 is almost or completely turned off, and current does not flow (or hardly flows) through the channel region 18, which is detected as the programmed state, i.e., the "0" state.
[0077] Table 1 shows the typical voltage ranges that can be applied to the terminals of the memory cell 110 to perform read, erase, and program operations. Table 1: Operation of the flash memory cell 210 in FIG. 2
Table 1
[0078] FIG. 3 shows a memory cell 310 similar to the memory cell 210 of FIG. 2 with an additional control gate (CG) terminal 28. The control gate terminal 28 is biased at a high voltage (e.g., 10V) during programming, a low or negative voltage (e.g., 0V / -8V) during erasure, and a low or medium voltage (e.g., 0V / 2.5V) during readout. The other terminals are biased in the same manner as the terminals of FIG. 2.
[0079] FIG. 4 shows a four-gate memory cell 410 comprising a source region 14, a drain region 16, a floating gate 20 above a first portion of the channel region 18, a select gate 22 (typically coupled to a word line, WL) above a second portion of the channel region 18, a control gate 28 above the floating gate 20, and an erase gate 30 above the source region 14. This configuration is described in U.S. Patent No. 6,747,310, which is incorporated herein by reference for all purposes. Here, all gates are non-floating gates except for the floating gate 20, i.e., they are electrically connected or connectable to a voltage source. Programming is performed by injecting hot electrons from the channel region 18 into the floating gate 20 itself. Erasure is performed by tunneling electrons from the floating gate 20 to the erase gate 30.
[0080] Table 2 shows typical voltage ranges that can be applied to the terminals of the memory cell 410 to perform readout, erasure, and program operations. Table 2: Operation of the Flash Memory Cell 410 of FIG. 4 [Table 2] "Readout 1" is a readout mode in which the cell current is output to the bit line. "Readout 2" is a readout mode in which the cell current is output to the source line terminal.
[0081] FIG. 5 shows a memory cell 510 similar to the memory cell 410 of FIG. 4, except that the memory cell 510 does not include an erase gate EG terminal. Erasure is performed by biasing the substrate 18 to a high voltage and biasing the control gate CG terminal 28 to a low voltage or a negative voltage. Alternatively, erasure is performed by biasing the word line terminal 22 to a positive voltage and biasing the control gate terminal 28 to a negative voltage. Programming and reading are the same as those in FIG. 4.
[0082] FIG. 6 shows a three-gate memory cell 610, which is another type of flash memory cell. The memory cell 610 is identical to the memory cell 410 of FIG. 4, except that the memory cell 610 does not have a separate control gate terminal. (Erasure occurs through the use of an erase gate terminal) The erase operation and the read operation are the same as those in FIG. 4, except that no control gate bias is applied. Since the programming operation is also performed without a control gate bias, as a result, a higher voltage must be applied to the source line terminal during the program operation to compensate for the lack of control gate bias.
[0083] Table 3 shows the typical voltage ranges that can be applied to the terminals of the memory cell 610 to perform read, erase, and program operations. Table 3: Operation of the flash memory cell 610 of FIG. 6
Table 3
[0084] FIG. 7 shows a stacked gate memory cell 710, which is another type of flash memory cell. The memory cell 710 is similar to the memory cell 210 of FIG. 2, except that the floating gate 20 extends over the entire channel region 18 and the control gate terminal 22 (coupled to the word line) is separated by an insulating layer (not shown) and extends over the floating gate 20. The erase, programming, and read operations operate in a similar manner as described above for the memory cell 210.
[0085] Table 4 shows typical voltage ranges that can be applied to the terminals of the memory cell 710 and the substrate 12 to perform read, erase, and program operations. Table 4: Operation of the Flash Memory Cell 710 of FIG. 7
Table 4
[0086] "Read 1" is a read mode in which the cell current is output to the bit line. "Read 2" is a read mode in which the cell current is output to the source line terminal. Optionally, in an array including rows and columns of the memory cells 210, 310, 410, 510, 610, or 710, the source line can be coupled to one row of memory cells or two adjacent rows of memory cells. That is, the source line terminal can be shared by adjacent rows of memory cells.
[0087] To utilize a memory array including one of the types of non-volatile memory cells in the above artificial neural network, two modifications are made. First, the lines are configured so that each memory cell can be individually programmed, erased, and read without adversely affecting the memory states of other memory cells in the array, as further described below. Second, continuous (analog) programming of the memory cells is provided.
[0088] Specifically, the memory state of each memory cell in the array (i.e., the charge in the floating gate) can be continuously changed independently, with minimal interference from other memory cells, from a completely erased state to a completely programmed state. In another embodiment, the memory state of each memory cell in the array (i.e., the charge in the floating gate) can be continuously changed independently, with minimal interference from other memory cells, from a completely programmed state to a completely erased state and vice versa. This means that the cell memory is either analog or can store at least one of a number of discrete values (such as 16 or 64 different values), which allows all cells in the memory array to be very precisely and individually tunable, and makes the memory array ideal for memory and fine-tuning adjustments to the synaptic weights of neural networks.
[0089] The methods and means described herein can be applied, without limitation, to other non-volatile memory technologies such as SONOS (silicon-oxide-nitride-oxide-silicon, charge trap in nitride), MONOS (metal-oxide-nitride-oxide-silicon, metal charge trap in nitride), ReRAM (resistive ram), PCM (phase change memory), MRAM (magnetic ram), FeRAM (ferroelectric ram), OTP (bi-level or multi-level one time programmable), and CeRAM (correlated electron ram). The methods and means described herein can be applied, without limitation, to volatile memory technologies used in neural networks such as SRAM, DRAM, and other volatile synaptic cells. Neural network using a non-volatile memory cell array
[0090] FIG. 8 conceptually illustrates a non-limiting example of a neural network utilizing the non-volatile memory array of this embodiment. This example uses a non-volatile memory array neural network for a face recognition application, but it is also possible to implement other suitable applications using a non-volatile memory array-based neural network.
[0091] S0 is the input layer, which in this example is a 32×32 pixel RGB image with 5-bit precision (i.e., three 32×32 pixel arrays, one for each color R, G, and B, and each pixel has 5-bit precision). The synapse CB1 going from the input layer S0 to layer C1 applies a different set of weights to some instances and shared weights to other instances, scans the input image with a 3×3 pixel overlapping filter (kernel), and shifts the filter by 1 pixel (or more than 2 pixels in some models) at a time. Specifically, the 9 pixel values in the 3×3 portion of the image (i.e., what is referred to as the filter or kernel) are provided to the synapse CB1, where these 9 input values are multiplied by appropriate weights, and after summing the output of that multiplication, a single output value is determined and given by the first synapse of CB1 to generate one pixel of the layer of the feature map C1. The 3×3 filter is then shifted 1 pixel to the right within the input layer S0 (i.e., a 3-pixel column is added on the right and a 3-pixel column is dropped on the left), and thus the 9 pixel values of this newly positioned filter are provided to the synapse CB1, where they are multiplied by the same weights as above, and a second single output value is determined by the associated synapse. This process is continued until the 3×3 filter has scanned across the entire 32×32 pixel image of the input layer S0 for all three colors and all bits (precision values). The process is then repeated using different sets of weights until all of the feature maps of layer C1 are calculated, generating different feature maps of C1.
[0092] In this example, in layer C1, there are 16 feature maps each having 30×30 pixels. Each pixel is a new feature pixel extracted from the multiplication of the input and the kernel. Thus, each feature map is a two-dimensional array. Therefore, in this example, layer C1 consists of 16 layers of two-dimensional arrays (note that the layers and arrays referred to in this specification are logical relationships rather than necessarily physical relationships, that is, the arrays are not necessarily oriented in a physical two-dimensional array). Each of the 16 feature maps in layer C1 is generated by one of 16 different sets of synaptic weights applied to the filter scan. All of the C1 feature maps can target different aspects of the same image feature, such as edge identification. For example, the first map (generated using a first set of weights shared by all scans used to generate this first map) can identify circular edges, and the second map (generated using a second set of weights different from the first set of weights) can identify rectangular edges or the aspect ratio of a particular feature, etc.
[0093] Before going from layer C1 to layer S1, an activation function P1 (pooling) is applied that pools values from non - overlapping and consecutive 2×2 regions within each feature map. The purpose of the pooling function is to average neighboring positions (or it is also possible to use the max function), for example, to reduce the dependence on edge positions, and to reduce the data size before going to the next stage. In layer S1, there are 16 15×15 feature maps (i.e., 16 different arrays of 15×15 pixels each). The synapses CB2 going from layer S1 to layer C2 scan the maps in S1 with a 4×4 filter with a 1 - pixel filter shift. In layer C2, there are 22 12×12 feature maps. Before going from layer C2 to layer S2, an activation function P2 (pooling) is applied that pools values from non - overlapping and consecutive 2×2 regions within each feature map. In layer S2, there are 22 6×6 feature maps. In the synapses CB3 going from layer S2 to layer C3, an activation function (pooling) is applied, where all neurons in layer C3 are connected to all maps in layer S2 via their respective synapses of CB3. In layer C3, there are 64 neurons. The synapses CB4 going from layer C3 to the output layer S3 fully connect C3 to S3, i.e., all neurons in layer C3 are connected to all neurons in layer S3. The output in S3 contains 10 neurons, and the neuron with the highest output determines the class. This output can indicate, for example, the identification or classification (categorization) of the content of the original image.
[0094] Each layer of synapses is implemented using an array or a part of an array of non - volatile memory cells.
[0095] Figure 9 is a block diagram of a system that can be used for that purpose. The VMM system 32 includes non - volatile memory cells and synapses between one layer and the next layer (Figure 8It is used as CB1, CB2, CB3, CB4, etc. Specifically, the VMM system 32 includes a VMM array 33 composed of non-volatile memory cells arranged in rows and columns, an erase gate and word line gate decoder 34, a control gate decoder 35, a bit line decoder 36, and a source line decoder 37. These decoders decode their respective inputs to the non-volatile memory cell array 33. The input to the VMM array 33 can be from the erase gate and word line gate decoder 34 or from the control gate decoder 35. The source line decoder 37 in this example also decodes the output of the VMM array 33. Alternatively, the bit line decoder 36 can decode the output of the VMM array 33.
[0096] The VMM array 33 serves two purposes. First, the VMM array 33 stores the weights used by the VMM system 32. Second, the VMM array 33 effectively multiplies the inputs by the weights stored in the VMM array 33 and sums them for each output line (source line or bit line) to generate an output, which becomes the input to the next layer or the input to the last layer. By performing the functions of multiplication and addition, the VMM array 33 eliminates the need for separate multiplication and addition logic circuits and is also power-efficient due to in-memory computing on the spot.
[0097] The output of the VMM array 33 is supplied to a differential adder (such as an addition op-amp or an addition current mirror) 38 that sums the output of the VMM array 33 to create a single value for convolution. The differential adder 38 is arranged to perform the sum of both the positive and negative weight inputs and output a single value.
[0098] The total output value of the differential adder 38 is then supplied to an activation function circuit 39 that rectifies the output. The activation function circuit 39 can provide a sigmoid function, a tanh function, a ReLU function, or any other non-linear function. The rectified output value of the activation function circuit 39 becomes an element of the feature map of the next layer (e.g., C1 in FIG. 8), and is then applied to the next synapse to generate the next feature map layer or the last layer. Thus, in this example, the VMM array 33 constitutes a plurality of synapses (receiving inputs from the previous layer of neurons or from an input layer such as an image database), and the adder 38 and the activation function circuit 39 constitute a plurality of neurons.
[0099] The inputs (WLx, EGx, CGx, and optionally BLx and SLx) to the VMM system 32 of FIG. 9 can be at an analog level, a binary level, a digital pulse (in which case a pulse - analog converter PAC may be required to convert the pulse to an appropriate input analog level), or a digital bit (in which case a DAC is provided to convert the digital bit to an appropriate input analog level), and the output can be at an analog level, a binary level, a digital pulse, or a digital bit (in which case an output ADC is provided to convert the output analog level to a digital bit).
[0100] FIG. 10 is a block diagram showing the use of multiple layers of the VMM system 32, labeled as VMM systems 32a, 32b, 32c, 32d, and 32e in the figure. As shown in FIG. 10, an input (denoted as Inputx) is converted from digital to analog by a digital - analog converter 31 and provided to the input VMM system 32a. The converted analog input can be a voltage or a current. The input D / A conversion of the first layer can be performed by using a function or a LUT (look - up table) that maps the input Inputx to an appropriate analog level of the matrix multiplier of the input VMM system 32a. The input conversion can also be performed by an analog - to - analog (A / A) converter to convert an external analog input to the mapped analog input to the input VMM system 32a. The input conversion can also be performed by a digital - to - digital pules (D / P) converter to convert an external digital input to the mapped digital pulses to the input VMM system 32a.
[0101] The output generated by the input VMM system 32a is then provided as input to the next VMM system (hidden level 1) 32b, which then generates an output that is provided as input to the input VMM system (hidden level 2) 32c, and so on. The various layers of the VMM system 32 function as the layers of synapses and neurons of a convolutional neural network (CNN). Each of the VMM systems 32a, 32b, 32c, 32d, and 32e can be a stand-alone physical system with its corresponding non-volatile memory array, or multiple VMM systems can utilize different portions of the same physical non-volatile memory array, or multiple VMM systems can utilize overlapping portions of the same physical non-volatile memory array. Each of the VMM systems 32a, 32b, 32c, 32d, and 32e can also be time-multiplexed across the various portions of its array or neurons. The example shown in FIG. 10 includes five layers (32a, 32b, 32c, 32d, 32e), namely, one input layer (32a), two hidden layers (32b, 32c), and two fully connected layers (32d, 32e). One skilled in the art will understand that this is merely exemplary and that the system can alternatively include more than two hidden layers and more than two fully connected layers. VMM array
[0102] FIG. 11 shows a neuron VMM array 1100 that is particularly suitable for the memory cell 310 shown in FIG. 3 and is utilized as part of the synapses and neurons between the input layer and the next layer. The VMM array 1100 includes a memory array 1101 of non-volatile memory cells and a reference array 1102 of non-volatile reference memory cells (located at the top of the array). Alternatively, another reference array can be located at the bottom.
[0103] In the VMM array 1100, control gate lines such as control gate line 1103 extend in the vertical direction (thus, the reference array 1102 in the row direction is orthogonal to the control gate line 1103), and erase gate lines such as erase gate line 1104 extend in the horizontal direction. Here, the input to the VMM array 1100 is provided to the control gate lines (CG0, CG1, CG2, CG3), and the output of the VMM array 1100 appears on the source lines (SL0, SL1). In one embodiment, only even rows are used, and in another embodiment, only odd rows are used. The current of each source line (SL0 and SL1 respectively) performs a summation function of all the currents from the memory cells connected to that particular source line.
[0104] As described herein for the neural network, the non-volatile memory cells of the VMM array 1100, i.e., the flash memory of the VMM array 1100, are preferably configured to operate in the subthreshold region.
[0105] The non-volatile reference memory cells and non-volatile memory cells described herein are biased with weak inversion as follows: Ids = Io * e (Vg-Vth) / nVt = w * Io * e (Vg) / nVt where w = e (-Vth) / nVt and where Ids is the drain-source current, Vg is the gate voltage of the memory cell, Vth is the threshold voltage of the memory cell, Vt is the thermal voltage = k * T / q, where k is the Boltzmann constant, T is the Kelvin temperature, q is the electronic charge, n is the slope factor = 1+(Cdep / Cox), Cdep is the capacitance of the depletion layer, and Cox is the capacitance of the gate oxide layer, Io is the memory cell current at a gate voltage equal to the threshold voltage, and Io is (Wt / L) * u * Cox * (n - 1) * Vt 2proportional to, where u is the carrier mobility, and Wt and L are the width and length of the memory cell, respectively.
[0106] When an I-V log converter that converts the input current Ids to the input voltage Vg using a memory cell (such as a reference memory cell or a peripheral memory cell) or a transistor is used, Vg is as follows: Vg = n * Vt * log[Ids / wp * Io] where wp is the w of the reference or peripheral memory cell.
[0107] When an I-V log converter that converts the input current Ids to the input voltage Vg using a memory cell (such as a reference memory cell or a peripheral memory cell) or a transistor is used, Vg is as follows: Vg = n * Vt * log[Ids / wp * Io]
[0108] where wp is the w of the reference or peripheral memory cell.
[0109] For the memory array used as the vector matrix multiplier VMM array, the output current is as follows: Iout = wa * Io * e (Vg) / nVt that is Iout = (wa / wp) * Iin = W * Iin W = e (Vthp-Vtha) / nVt Iin = wp * Io * e (Vg) / nVt where wa is the w of each memory cell of the memory array.
[0110] The word line or control gate can be used as the input of the memory cell for the input voltage.
[0111] Alternatively, the non-volatile memory cells of the VMM array described herein can be configured to operate in the linear region. Ids = beta * (Vgs - Vth) * Vds, beta = u * Cox * Wt / L, W ∝ (Vgs - Vth) That is, the weight W in the linear region is proportional to (Vgs - Vth).
[0112] The word line or control gate or bit line or source line can be used as an input to the memory cell operating in the linear region. The bit line or source line can be used as an output of the memory cell.
[0113] For an I-V linear converter, a memory cell (such as a reference memory cell or a peripheral memory cell) or a transistor or a resistor operating in the linear region can be used to linearly convert the input / output current into the input / output voltage.
[0114] Alternatively, the memory cells of the VMM array described herein can be configured to operate in the saturation region. Ids = 1 / 2 * beta * (Vgs - Vth) 2 , beta = u * Cox * Wt / L W ∝ (Vgs - Vth) 2 , that is, the weight W is (Vgs - Vth) 2 proportional to.
[0115] The word line, control gate, or erase gate can be used as an input to the memory cell operating in the saturation region. The bit line or source line can be used as an output of the output neuron.
[0116] Alternatively, the memory cells of the VMM arrays described herein can be used in all regions or combinations thereof (subthreshold, linear, or saturated).
[0117] Another embodiment for the VMM array 33 of FIG. 9 is described in U.S. Patent Application No. 15 / 826,345, which is incorporated herein by reference. As described in the above application, the source line or bit line can be used as a neuron output (current sum output).
[0118] FIG. 12 shows a neuron VMM array 1200 particularly suitable for the memory cell 210 shown in FIG. 2 and is utilized as a synapse between the input layer and the next layer. The VMM array 1200 includes a memory array 1203 of non-volatile memory cells, a reference array 1201 of first non-volatile reference memory cells, and a reference array 1202 of second non-volatile reference memory cells. The reference arrays 1201 and 1202 arranged in the column direction of the array function to convert the current inputs flowing into the terminals BLR0, BLR1, BLR2, and BLR3 into voltage inputs WL0, WL1, WL2, and WL3. In practice, the first and second non-volatile reference memory cells are diode-connected through a multiplexer 1214 (only part shown) with the current inputs flowing in. The reference cells are tuned (e.g., programmed) to a target reference level. The target reference level is provided by a reference min-array matrix (not shown).
[0119] The memory array 1203 serves two purposes. First, the memory array 1003 stores the weights used by the VMM array 1200 in respective memory cells. Second, the memory array 1203 effectively multiplies the weights stored in the memory array 1203 with the input (i.e., the current inputs provided to terminals BLR0, BLR1, BLR2, and BLR3, which the reference arrays 1201 and 1202 convert into input voltages and supply to word lines WL0, WL1, WL2, and WL3), and then adds up all the results (memory cell currents) to generate the output of each bit line (BL0~BLN), and this output serves as the input to the next layer or the input to the last layer. By performing the functions of multiplication and addition, the memory array 1203 eliminates the need for separate multiplication and addition logic circuits and also has good power efficiency. Here, the voltage input is provided to word lines WL0, WL1, WL2, and WL3, and the output appears on respective bit lines BL0~BLN during the read (inference) operation. The current of each of the bit lines BL0~BLN performs the total function of the currents from all the non-volatile memory cells connected to that specific bit line.
[0120] Table 5 shows the operating voltages of the VMM array 1200. The columns in the table indicate the voltages applied to the word line of the selected cell, the word line of the non-selected cell, the bit line of the selected cell, the bit line of the non-selected cell, the source line of the selected cell, and the source line of the non-selected cell, where FLT indicates floating, i.e., no voltage is applied. The rows indicate the operations of read, erase, and program. Table 5: Operation of the VMM Array 1200 in FIG. 12
Table 5
[0121] FIG. 13 shows a neuron VMM array 1300 that is particularly suitable for the memory cell 210 shown in FIG. 2 and is used as part of synapses and neurons between the input layer and the next layer. The VMM array 1300 includes a memory array 1303 of non-volatile memory cells, a reference array 1301 of first non-volatile reference memory cells, and a reference array 1302 of second non-volatile reference memory cells. The reference arrays 1301 and 1302 extend in the row direction of the VMM array 1300. The VMM array is similar to the VMM1000 except that the word lines extend vertically in the VMM array 1300. Here, the inputs are provided to the word lines (WLA0, WLB0, WLA1, WLB2, WLA2, WLB2, WLA3, WLB3), and the outputs appear on the source lines (SL0, SL1) during a read operation. The current of each source line performs a summation function of all the currents from the memory cells connected to that particular source line.
[0122] Table 6 shows the operating voltages of the VMM array 1300. The columns in the table show the voltages applied to the word line of the selected cell, the word line of the non-selected cell, the bit line of the selected cell, the bit line of the non-selected cell, the source line of the selected cell, and the source line of the non-selected cell. The rows show the operations of read, erase, and program. Table 6: Operations of the VMM array 1300 in FIG. 13
Table 6
[0123] FIG. 14 shows a neuron VMM array 1400 that is particularly suitable for the memory cell 310 shown in FIG. 3 and is used as part of synapses and neurons between the input layer and the next layer. The VMM array 1400 includes a memory array 1403 of non-volatile memory cells, a reference array 1401 of first non-volatile reference memory cells, and a reference array 1402 of second non-volatile reference memory cells. The reference arrays 1401 and 1402 function to convert the current inputs flowing into the terminals BLR0, BLR1, BLR2, and BLR3 into voltage inputs CG0, CG1, CG2, and CG3. In practice, the first and second non-volatile reference memory cells are diode-connected through a multiplexer 1412 (only part shown) in a state where the current inputs flow through BLR0, BLR1, BLR2, and BLR3. The multiplexer 1412 includes corresponding multiplexers 1405 and cascode transistors 1404, respectively, to ensure a constant voltage of each bit line (such as BLR0) of the first and second non-volatile reference memory cells during the read operation. The reference cells are tuned to a target reference level.
[0124] The memory array 1403 serves two purposes. First, the memory array 1203 stores the weights used by the VMM array 1400. Second, the memory array 1403 effectively multiplies the weights stored in the memory array by the inputs (the current inputs provided to the terminals BLR0, BLR1, BLR2, and BLR3, and the reference arrays 1401 and 1402 convert these current inputs into input voltages and supply them to the control gates (CG0, CG1, CG2, and CG3)), and then adds up all the results (cell currents) to generate an output, which appears on BL0~BLN and becomes an input to the next layer or the last layer. By the memory array performing the functions of multiplication and addition, the need for separate multiplication and addition logic circuits is eliminated, and the power efficiency is also good. Here, the inputs are provided to the control gate lines (CG0, CG1, CG2, and CG3), and the outputs appear on the bit lines (BL0~BLN) during the read operation. The current of each bit line performs the total function of all the currents from the memory cells connected to that specific bit line.
[0125] The VMM array 1400 implements unidirectional tuning of non-volatile memory cells within the memory array 1403. That is, each non-volatile memory cell is erased and then partially programmed until the desired charge on the floating gate is reached. This can be performed, for example, using the precision programming techniques described below. If too much charge is applied to the floating gate (resulting in an incorrect value being stored in the cell), the cell must be erased and the series of partial programming operations must be repeated. As shown, two rows sharing the same erase gate (such as EG0 or EG1) need to be erased together (known as page erase), after which each cell is partially programmed until the desired charge on the floating gate is reached.
[0126] Table 7 shows the operating voltages of the VMM array 1400. The columns in the table show the voltages applied to the word line of the selected cell, the word line of the non-selected cell, the bit line of the selected cell, the bit line of the non-selected cell, the control gate of the selected cell, the control gate of the non-selected cell within the same sector as the selected cell, the control gate of the non-selected cell in a different sector from the selected cell, the erase gate of the selected cell, the erase gate of the non-selected cell, the source line of the selected cell, and the source line of the non-selected cell. The rows show the read, erase, and program operations. Table 7: Operation of the VMM Array 1400 in FIG. 14 [Table 7]
[0127] FIG. 15 shows a neuron VMM array 1500 that is particularly suitable for the memory cell 310 shown in FIG. 3 and is used as part of synapses and neurons between the input layer and the next layer. The VMM array 1500 includes a memory array 1503 of non-volatile memory cells, a reference array 1501 or a first non-volatile reference memory cell, and a reference array 1502 of second non-volatile reference memory cells. The EG lines EGR0, EG0, EG1, and EGR1 extend vertically, and the CG lines CG0, CG1, CG2, and CG3 and the SL lines WL0, WL1, WL2, and WL3 extend horizontally. The VMM array 1500 is similar to the VMM array 1400 except that the VMM array 1500 implements bidirectional tuning, and each individual cell can be completely erased, partially programmed, and partially erased as needed to reach the desired charge level of the floating gate by using individual EG lines. As shown, the reference arrays 1501 and 1502 convert the input current at the terminals BLR0, BLR1, BLR2, and BLR3 into control gate voltages CG0, CG1, CG2, and CG3 (through the action of diode-connected reference cells via the multiplexer 1514), and these voltages are applied to the memory cells in the row direction. The current outputs (neurons) are in the bit lines BL0 to BLN, and each bit line sums all the currents from the non-volatile memory cells connected to that particular bit line.
[0128] Table 8 shows the operating voltages of the VMM array 1500. The columns in the table show the voltages applied to the word line of the selected cell, the word line of the non-selected cell, the bit line of the selected cell, the bit line of the non-selected cell, the control gate of the selected cell, the control gate of the non-selected cell in the same sector as the selected cell, the control gate of the non-selected cell in a different sector from the selected cell, the erase gate of the selected cell, the erase gate of the non-selected cell, the source line of the selected cell, and the source line of the non-selected cell. The rows show the read, erase, and program operations. Table 8: Operation of the VMM Array 1500 in FIG. 15
Table 8
[0129] Figure 16 shows a neuron VMM array 1600 that is particularly suitable for the memory cell 210 shown in FIG. 2 and is used as part of synapses and neurons between the input layer and the next layer. In the VMM array 1600, the input INPUT 0 ...., INPUT N is received by bit lines BL 0 ,...BL N respectively, and the outputs OUTPUT 1 , OUTPUT 2 , OUTPUT 3 , and OUTPUT 4 are generated on source lines SL 0 , SL 1 , SL 2 , and SL 3 respectively.
[0130] Figure 17 shows a neuron VMM array 1700 that is particularly suitable for the memory cell 210 shown in FIG. 2 and is used as part of synapses and neurons between the input layer and the next layer. In this example, the inputs INPUT 0 , INPUT 1 , INPUT 2 , and INPUT 3 are received by source lines SL 0 , SL 1 , SL 2 , and SL 3 respectively, and the outputs OUTPUT 0 ,...OUTPUT N are generated on bit lines BL 0 ,...,, BL N .
[0131] Figure 18 shows a neuron VMM array 1800 that is particularly suitable for the memory cell 210 shown in FIG. 2 and is used as part of synapses and neurons between the input layer and the next layer. In this example, the inputs INPUT 0 ,...,, INPUT M are received by word lines WL 0 ,...,, WL M respectively, and the outputs OUTPUT 0 ,...OUTPUTN is generated on bit lines BL 0 ,..., BL N respectively.
[0132] FIG. 19 shows a neuron VMM array 1900 that is particularly suitable for the memory cell 310 shown in FIG. 3 and is used as part of synapses and neurons between the input layer and the next layer. In this example, the inputs INPUT 0 ,..., INPUT M are respectively received by word lines WL 0 ,..., WL M and the outputs OUTPUT 0 ,... OUTPUT N are generated on bit lines BL 0 ,..., BL N respectively.
[0133] FIG. 20 shows a neuron VMM array 2000 that is particularly suitable for the memory cell 410 shown in FIG. 4 and is used as part of synapses and neurons between the input layer and the next layer. In this example, the inputs INPUT 0 ,..., INPUT n are respectively received by vertical control gate lines CG 0 ,..., CG N and the outputs OUTPUT 1 and OUTPUT 2 are generated on source lines SL 0 and SL 1 respectively.
[0134] FIG. 21 shows a neuron VMM array 2100 that is particularly suitable for the memory cell 410 shown in FIG. 4 and is used as part of synapses and neurons between the input layer and the next layer. In this example, the inputs INPUT 0 ,..., INPUT N are respectively received by the gates of bit line control gates 2901-1, 2901-2,..., 2901-(N-1) and 2901-N that are respectively coupled to bit lines BL 0 ,..., BL N . An exemplary output OUTPUT 1and OUTPUT 2 is generated on the source line SL 0 and SL 1 is generated.
[0135] FIG. 22 shows a neuron VMM array 2200 that is particularly suitable for the memory cell 310 shown in FIG. 3, the memory cell 510 shown in FIG. 5, and the memory cell 710 shown in FIG. 7, and is used as part of synapses and neurons between the input layer and the next layer. In this example, the input INPUT 0 ,..., INPUT M is received on the word lines WL 0 ,..., WL M and the output OUTPUT 0 ,..., OUTPUT N is respectively generated on the bit lines BL 0 ,..., BL N is respectively generated.
[0136] FIG. 23 shows a neuron VMM array 2300 that is particularly suitable for the memory cell 310 shown in FIG. 3, the memory cell 510 shown in FIG. 5, and the memory cell 710 shown in FIG. 7, and is used as part of synapses and neurons between the input layer and the next layer. In this example, the input INPUT 0 ,..., INPUT M is received on the control gate lines CG 0 ,..., CG M . The output OUTPUT 0 ,..., OUTPUT N is respectively generated on the vertical source lines SL 0 ,..., SL N and each source line SL i is connected to the source lines of all the memory cells in column i.
[0137] FIG. 24 shows a neuron VMM array 2400 that is particularly suitable for the memory cell 310 shown in FIG. 3, the memory cell 510 shown in FIG. 5, and the memory cell 710 shown in FIG. 7, and is used as part of synapses and neurons between the input layer and the next layer. In this example, the input INPUT 0,..., INPUT M is the control gate line CG 0 ,..., CG M is received at. The output OUTPUT 0 ,..., OUTPUT N is the vertical bit line BL 0 ,..., BL N is respectively generated at each bit line BL i is connected to the bit lines of all memory cells within column i. Long Short-Term Memory
[0138] The prior art includes a concept known as long short-term memory (LSTM). LSTM is often used in artificial neural networks. LSTM enables an artificial neural network to remember information over an arbitrary given period and use that information in subsequent operations. Conventional LSTM includes cells, an input gate, an output gate, and a forget gate. The three gates regulate the flow of information into and out of the cell and the period during which information is stored within the LSTM. VMM is particularly useful in LSTM.
[0139] FIG. 25 shows an exemplary LSTM2500. The LSTM2500 in this example includes cells 2501, 2502, 2503, and 2504. Cell 2501 receives the input vector x 0 and generates the output vector h 0 and the cell state vector c 0 . Cell 2502 receives the input vector x 1 , the output vector (hidden state) h 0 from cell 2501, and the cell state c 0 from cell 2501, and generates the output vector h 1 and the cell state vector c 1 . Cell 2503 receives the input vector x 2 , the output vector (hidden state) h 1 from cell 2502, and the cell state c 1 from cell 2502, and generates the output vector h 2and the cell state vector c 2 and generates. Cell 2504 receives the input vector x 3 and the output vector (hidden state) h from cell 2503 2 and the cell state c from cell 2503 2 and generates the output vector h 3 Additional cells are available, and an LSTM with four cells is merely an example.
[0140] FIG. 26 shows an exemplary implementation of an LSTM cell 2600 that can be used for cells 2501, 2502, 2503, and 2504 of FIG. 25. The LSTM cell 2600 receives the input vector x(t), the cell state vector c(t−1) from the preceding cell, and the output vector h(t−1) from the preceding cell, and generates the cell state vector c(t) and the output vector h(t).
[0141] The LSTM cell 2600 includes sigmoid function devices 2601, 2602, and 2603, each of which applies a number between 0 and 1 to control the degree to which each component of the input vector contributes to the output vector. The LSTM cell 2600 also includes tanh devices 2604 and 2605 for applying the hyperbolic tangent function to the input vector, multiplier devices 2606, 2607, and 2608 for multiplying two vectors, and an adder device 2609 for adding two vectors. The output vector h(t) can be provided to the next LSTM cell in the system or accessed for other purposes.
[0142] FIG. 27 shows an LSTM cell 2700 that is an example of an implementation of the LSTM cell 2600. For the convenience of the reader, the same numbering scheme from the LSTM cell 2600 is used for the LSTM cell 2700. Each of the sigmoid function devices 2601, 2602, and 2603, and the tanh device 2604 includes a plurality of VMM arrays 2701 and activation circuit blocks 2702. Thus, it can be seen that the VMM array is particularly useful in LSTM cells used in a particular neural network system.
[0143] An alternative to the LSTM cell 2700 (and another example of an implementation of the LSTM cell 2600) is shown in FIG. 28. In FIG. 28, the sigmoid function devices 2601, 2602, and 2603, and the tanh device 2604 can share the same physical hardware (the VMM array 2801 and the activation function block 2802) in a time-division multiplexed manner. The LSTM cell 2800 also includes a multiplier device 2803 for multiplying two vectors, an adder device 2808 for adding two vectors, a tanh device 2605 (including the activation circuit block 2802), a register 2807 for storing the value i(t) when i(t) is output from the sigmoid function block 2802, and the value f(t) * c(t - 1), a register 2804 for storing the value when its value is output from the multiplier device 2803 via the multiplexer 2810, and the value i(t) * u(t), a register 2805 for storing the value when its value is output from the multiplier device 2803 via the multiplexer 2810, and the value o(t) * ĉ(t), a register 2806 for storing the value when its value is output from the multiplier device 2803 via the multiplexer 2810, and the multiplexer 2809.
[0144] Whereas the LSTM cell 2700 includes multiple sets of the VMM array 2701 and their respective activation function blocks 2702, the LSTM cell 2800 includes only one set of the VMM array 2801 and the activation function block 2802, which is used to represent multiple layers in an embodiment of the LSTM cell 2800. The LSTM cell 2800 requires 1 / 4 of the space needed for the VMM and the activation function block compared to the LSTM cell 2700, so the LSTM cell 2800 requires less space than the LSTM 2700.
[0145] The LSTM unit typically includes a plurality of VMM arrays, and it can be further understood that each of these requires functions provided by specific circuit blocks outside the VMM array, such as adders, activation circuit blocks, and high-voltage generation blocks. Providing separate circuit blocks for each VMM array would require a significant amount of space within the semiconductor device and would be somewhat inefficient. Therefore, in the embodiments described below, an attempt is made to minimize the circuits required outside the VMM array itself. Gated Recurrent Unit
[0146] The analog VMM implementation can be utilized in a gated recurrent unit (GRU). A GRU is a gate mechanism within a recurrent artificial neural network. A GRU is similar to an LSTM, except that a GRU cell generally includes fewer components than an LSTM cell.
[0147] FIG. 29 shows an exemplary GRU 2900. The GRU 2900 in this example includes cells 2901, 2902, 2903, and 2904. Cell 2901 receives an input vector x 0 and generates an output vector h 0 Cell 2902 receives an input vector x 1 and the output vector h 0 from cell 2901 and generates an output vector h 1 Cell 2903 receives an input vector x 2 and the output vector (hidden state) h 1 from cell 2902 and generates an output vector h 2 Cell 2904 receives an input vector x 3 and the output vector (hidden state) h 2 from cell 2903 and generates an output vector h 3 Additional cells are also available, and a GRU with four cells is merely an example.
[0148] FIG. 30 shows an exemplary implementation of a GRU cell 3000 that can be used for the cells 2901, 2902, 2903, and 2904 of FIG. 29. The GRU cell 3000 receives an input vector x(t) and an output vector h(t−1) from a preceding GRU cell and generates an output vector h(t). The GRU cell 3000 includes sigmoid function devices 3001 and 3002, each of which applies a number between 0 and 1 to components from the output vector h(t−1) and the input vector x(t). The GRU cell 3000 also includes a tanh device 3003 for applying a hyperbolic tangent function to the input vector, a plurality of multiplier devices 3004, 3005, and 3006 for multiplying two vectors, an adder device 3007 for adding two vectors, and a complement device 3008 for subtracting the input from 1 to generate an output.
[0149] FIG. 31 shows a GRU cell 3100 that is an example of an implementation of the GRU cell 3000. For the convenience of the reader, the same numbering scheme from the GRU cell 3000 is used for the GRU cell 3100. As can be seen from FIG. 31, the sigmoid function devices 3001 and 3002, and the tanh device 3003 each include a plurality of VMM arrays 3101 and activation function blocks 3102. Thus, it can be seen that the VMM array is particularly used in GRU cells used in a specific neural network system.
[0150] An alternative example of GRU cell 3100 (and another example of an implementation of GRU cell 3000) is shown in FIG. 32. In FIG. 32, GRU cell 3200 utilizes VMM array 3201 and activation function block 3202. When configured as a sigmoid function, it applies a number between 0 and 1 to control the degree to which each component of the input vector contributes to the output vector. In FIG. 32, sigmoid function devices 3001 and 3002, and tanh device 3003 share the same physical hardware (VMM array 3201 and activation function block 3202) in a time-division multiplexed manner. GRU cell 3200 also includes a multiplier device 3203 for multiplying two vectors, an adder device 3205 for adding two vectors, a complement device 3209 for subtracting the input from 1 to generate an output, a multiplexer 3204, and the value h(t - 1) * r(t), a register 3206 for holding its value when the value is output from the multiplier device 3203 via the multiplexer 3204, and the value h(t - 1) * z(t), a register 3207 for holding its value when the value is output from the multiplier device 3203 via the multiplexer 3204, and the value h^(t) * (1 - z((t)), a register 3208 for holding its value when the value is output from the multiplier device 3203 via the multiplexer 3204, are provided.
[0151] While GRU cell 3100 includes multiple sets of VMM array 3101 and activation function block 3102, GRU cell 3200 includes only one set of VMM array 3201 and activation function block 3202, which is used to represent multiple layers in an embodiment of GRU cell 3200. GRU cell 3200 requires only 1 / 3 of the space needed for the VMM and activation function blocks compared to GRU cell 3100, so GRU cell 3200 requires less space than GRU cell 3100.
[0152] A system using GRUs typically includes a plurality of VMM arrays, and it can be further understood that each of these requires functions provided by specific circuit blocks outside the VMM array, such as adders, activation circuit blocks, and high-voltage generation blocks. Providing separate circuit blocks for each VMM array would require a significant amount of space within the semiconductor device and would be somewhat inefficient. Thus, in the embodiments described below, an attempt is made to minimize the circuits required outside the VMM array itself.
[0153] Inputs to the VMM array can be at an analog level, binary level, timing pulse, or digital bit, and the outputs can be at an analog level, binary level, timing pulse, or digital bit (in which case an output ADC is required to convert the output analog level current or voltage to a digital bit).
[0154] For each memory cell within the VMM array, each weight w can be implemented by a single memory cell, or by differential cells, or by two blended memory cells (the average of two or more cells). In the case of differential cells, two memory cells are required to implement the weight w as a differential weight (w = w+ - w-). In the case of two blended memory cells, two memory cells are required to implement the weight w as the average of two cells. Embodiments for precise tuning of cells within the VMM
[0155] FIG. 33 shows a block diagram of a VMM system 3300. The VMM system 3300 includes a VMM array 3301, a row decoder 3302, a high voltage decoder 3303, a column decoder 3304, a bit line driver 3305, an input circuit 3306, an output circuit 3307, control logic 3308, and a bias generator 3309. The VMM system 3300 further includes a high voltage generation block 3310 that includes a charge pump 3311, a charge pump regulator 3312, and a high voltage level generator 3313. The VMM system 3300 further includes an algorithm controller 3314, an analog circuit 3315, control logic 3316, and test control logic 3317. The systems and methods described below may be implemented in the VMM system 3300.
[0156] The input circuit 3306 may include circuits such as a DAC (digital to analog converter), a DPC (digital to pulses converter), an AAC (analog to analog converter) such as a current-voltage converter, a PAC (pulse to analog level converter), or any other type of converter. The input circuit 3306 may implement a normalization function, a scaling function, or an arithmetic function. The input circuit 3306 may implement a temperature compensation function for the input. The input circuit 3306 may implement an activation function such as a ReLU or sigmoid function.
[0157] The output circuit 3307 may include circuits such as an ADC (analog-to-digital converter for converting neuron analog outputs to digital bits), an AAC (analog converter such as a current-voltage converter), an APC (analog-pulse converter), or any other type of converter. The output circuit 3307 may implement an activation function such as a ReLU or sigmoid function. The output circuit 3307 may implement a normalization function, a scaling function, or an arithmetic function for neuron outputs. The output circuit 3307 may implement a temperature compensation function for neuron outputs or array outputs (such as bit line outputs) as described below.
[0158] FIG. 34 shows a tuning (programming or erasing a memory cell to a target) correction method 3400 that can be executed by an algorithm controller 3314 within a VMM system 3300. The tuning correction method 3400 generates an adaptive target based on a cell output and a final error resulting from an original target of the cell. This method typically begins in response to receiving a tuning command (step 3401). An initial current target Itargetv(i) (used in a program / verify algorithm) for a selected cell, or a group of selected cells, is determined using a prediction target model, such as by using a function or a look-up table, and a variable DeltaError is set to 0 (step 3402). The target function, if used, will be based on the I-V program curve of the selected memory cell or group of cells. The target function also depends on various variations caused by array characteristics such as the degree of programming anomalies exhibited by cells (where cells that exhibit more anomalies are typically subjected to more programming time under a suppression condition, and cells with higher currents typically have more anomalies), cell-to-cell coupling, and various types of array noise. These variations can be characterized on silicon in terms of PVT (process, voltage, temperature). The look-up table, if used, can be characterized in the same way to emulate the I-V curve and various variations.
[0159] Next, soft erasure is performed on all cells in the VMM. This soft erasure erases all cells to an intermediate weak erasure level such that each cell draws a current of, for example, about 3 to 5 μA during a read operation (step 3403). The soft erasure is performed, for example, by applying an incremental erasure pulse voltage to the cells until an intermediate cell current is reached. Next, a programming operation (such as deep programming to a target or rough / fine programming) is performed on all unused cells such that a <pA current level or an equivalent zero weight is reached (step 3404). Then, a target adjustment (correction) based on an error result is performed. If DeltaError > 0, which means the cell has received an overshoot during programming, Itargetv(i+1) is set to Itarget + theta * where theta is set to DeltaError and theta is, for example, 1 or a number close to 1 (step 3405A).
[0160] Itarget(i+1) can also be adjusted based on the previous Itarget(i) using an appropriate error target adjustment / correction. If DeltaError < 0, which means the cell has received an undershoot during programming, this means the cell current has not yet reached the target, and then Itargetv(i+1) is set to the previous target Itargetv(i) (step 3405B).
[0161] Next, a coarse and / or fine program and a verification operation are executed (step 3406). Before performing the precise (fine) programming step, programming can be accelerated by using a plurality of adaptive coarse programming methods, such as aiming at a plurality of gradually smaller coarse targets. Adaptive precise programming is performed, for example, with a fine (precise) incremental program voltage pulse or a steady program timing pulse. Examples of systems and methods for performing coarse programming and fine programming are described in U.S. Provisional Patent Application No. 62 / 933,809, entitled "Precise Programming Method and Apparatus for Analog Neural Memory in a Deep Learning Artificial Neural Network", filed on November 11, 2019, by the same assignee as this application, which is incorporated herein by reference.
[0162] Measure Icell at the selected cell (step 3407). For example, the cell current can be measured by an ammeter circuit. For example, the cell current can be measured by an ADC (analog-to-digital converter) circuit, in which case the output is represented by digital bits. For example, the cell current can be measured by an I-V (current-voltage converter) circuit, in which case the output is represented by an analog voltage. DeltaError is calculated, which is Icell - Itarget, representing the difference between the actual current (Icell) and the target current (Itarget) in the measured cell. If |DeltaError| < DeltaMargin, the cell has achieved the target current within a specific tolerance (DeltaMargin) and the method ends (step 3410). |DeltaError| = abs(DeltaError) = the absolute value of DeltaError. Otherwise, the method returns to step 3403 and sequentially executes the steps again (step 3410).
[0163] Figures 35A and 35B show a tuning correction method 3500 that can be executed by an algorithm controller 3314 in the VMM system 3300. Referring to Figure 35A, the method starts (step 3501), which is typically done in response to receiving a tuning command. The entire VMM array is erased, such as by soft erasure (step 3502). To obtain cell currents at the <pA level or equivalent zero weights, a programming operation (such as deep programming towards a target or rough / fine programming) is executed for all unused cells (step 3503). All cells in the VMM array are programmed to an intermediate value, such as 0.5 to 1.0 μA, using rough and / or fine programming cycles (step 3504). Examples of systems and methods for performing rough and fine programming are described in U.S. Provisional Patent Application No. 62 / 933,809, entitled "Precise Programming Method and Apparatus for Analog Neural Memory in a Deep Learning Artificial Neural Network," filed on November 11, 2019, by the same assignee as this application, which is incorporated herein by reference. The prediction target is set for the cells to be used using a function or look-up table as described above (step 3505). Next, a sector tuning method 3507 is executed for each sector of the VMM (step 3506). A sector typically consists of two or more adjacent rows in the array.
[0164] Figure 35B shows the adaptive target sector tuning method 3507. All cells within the sector are programmed to the final desired value (e.g., 1 nA to 50 nA) using individual or combined program / verify (P / V) methods such as (1) coarse / fine / steady P / V cycles, (2) CG + (CG increment only) or EG + (EG increment only), or complementary CG + / EG - (CG increment and EG decrement), and (3) the deepest programmed cell first (which means grouping cells into groups, such as progressive grouping where the group has the cell programmed with the lowest current first) (step 3508A). Next, a determination is made as to whether Icell < Itarget. If yes, the method proceeds to step 3509. If no, the method repeats step 3508A. In step 3509, DeltaError is measured, which is equal to the measured Icell - Itarget(i + 1) (step 3509). A determination is made as to whether |DeltaError| < DeltaMargin (step 3510). If yes, the method is complete (step 3511). If no, target adjustment is performed. If DeltaError > 0, which means the cell has received an overshoot during programming, the target is adjusted by setting the new target to Itarget + theta * DeltaError, where theta is typically = 1 (step 3512A). Itarget(i + 1) can also be adjusted based on the previous Itarget(i) using appropriate error target adjustment / correction. If DeltaError < 0, which means the cell has received an undershoot during programming, but this means the cell has not yet reached the target, the target is adjusted by maintaining the previous target, which means Itargetv(i + 1) = Itargetv(i) (step 3512B). The sector is softly erased (step 3513). All cells within the sector are programmed to an intermediate value (step 3514), and the method returns to step 3509.
[0165] A typical neural network can have positive weights w+ and negative weights w-, with a total weight = w+ - w-. w+ and w- are each implemented by memory cells (Iw+ and Iw- respectively), and the total weight (Iw = Iw+ - Iw-, current subtraction) can be performed at the peripheral circuit level (by using, for example, an array bit line output circuit). Thus, embodiments of total weight weight tuning can, in an example as shown in Table 9 include tuning both the w+ cell and the w- cell simultaneously, tuning only the w+ cell, or tuning only the w- cell. The tuning operation is performed using the program / verification and error target adjustment methods described above with respect to FIGS. 34 / 35A / 35B. The verification operation can be performed for only the total weight (e.g., measuring / reading the total weight current but not measuring / reading the individual positive w+ cell current or w- cell current), only the w+ cell current, or only the w- cell current.
[0166] For example, for a total Iw of 3 nA, Iw+ can be 3 nA and Iw- can be 0 nA. Alternatively, Iw+ can be 13 nA and Iw- can be 10 nA, which means that both the positive weight Iw+ and the negative weight Iw- are not zero (e.g., zero can indicate a deeply programmed cell). This may be preferred under certain operating conditions because both Iw+ and Iw- should be less susceptible to the influence of noise. Table 9 : Weight Tuning Method (Values are current values in nA units) [Table 9]
[0167] Thus, differential weight mapping according to the formula w = (w+) - (w-) can be used to store tuning value w for use in a neural network. The mapping of w+ and w- can be optimized to address specific problems that occur in the VMM array within the neural network, for example, by including offset values in w+ and w- where each value is stored.
[0168] In one embodiment, the weight mapping is optimized to reduce RTN noise. For example, if the desired value of w is 0 nA (zero w or unused memory cell), one possible mapping is (w+) = 0 nA and (w-) = 0 nA, and another possible mapping is (w+) = 30 nA and (w-) = 30 nA, where 30 nA is an example of a non-zero offset value that is added to each w+ and w- value before storage. Including such non-zero offset values ultimately consumes more power and introduces greater inaccuracy during the tuning process, but minimizes the impact of any noise. Similarly, if the desired value of w is 1 nA, one possible mapping is (w+) = 1 nA and (w-) = 0 nA, and another possible mapping is (w+) = 30 nA and (w-) = 29 nA. The latter may consume more power and introduce greater inaccuracy during the tuning process, but minimizes the impact of any noise.
[0169] In another embodiment, for a zero weight w, both w+ and w- can be tuned to be approximately zero, such as by using the method of bias offset voltage for the verification operation described herein. In this method, both w+ and w- are tuned to 5 nA at a control gate voltage (VCG) higher than the normal CG voltage used for inference (readout), for example. For instance, when a value of VCG = 1.5 V is used in inference for dVCG / Ir = 2 mV / 1 nA, in the tuning algorithm (program verification algorithm for weight tuning), a value of VCG = 1.510 V is used to verify a zero weight cell reaching the target of 5 nA. In the inference operation, since VCG = 1.5 V is used, the cell current is shifted down by 2 mV per 1 nA, and thus 5 nA in the verification operation becomes approximately 0 nA in the inference operation.
[0170] In another embodiment, for a zero weight w, both w+ and w- can be tuned to have a negative current, such as by using the negative current tuning method of bias offset voltage for the verification operation, which will be described below. In this case, both w+ and w- are tuned to -10 nA at a control gate voltage (VCG) higher than the normal CG voltage used for inference (readout), for example. For instance, when a value of VCG = 1.5 V is used in inference for dVCG / Ir = 2 mV / 1 nA, in the tuning algorithm (program verification algorithm for weight tuning), VCG = 1.530 V is used to verify a zero weight cell reaching the target of 5 nA. In the inference operation, since VCG = 1.5 V is used, the cell current is shifted down by 2 mV per 1 nA, and thus 5 nA in the verification operation becomes approximately 10 nA in the inference operation.
[0171] A similar method of offset bias conditions or bias (voltage and / or current and / or timing and / or temperature) sequences can be used for the purpose of test screening to detect abnormal bits such as cells that are susceptible to significant noise (such as RTN noise, thermal noise, or any other noise source). Basically, there are bias conditions or bias sequences that can be used to better detect noise by significantly attenuating the noise from the memory cell compared to other bias conditions or bias sequences (noise attenuation test). For example, for a 20nA memory cell, by modulating the bias conditions for this cell, such as changing the control gate bias voltage in test screening, placing this cell under different conditions can be advantageous for detecting undesirable behavior. For example, due to tester limitations or circuit limitations, modulating the bias conditions can be advantageous for detecting bits / cells that are susceptible to the influence of noise (such as RTN noise) at higher current levels.
[0172] Another way to screen or verify the noise level, such as RTN noise screening for a memory cell, is to sample the memory cell output (by measuring the output multiple times, such as 4 / 8 / ... / 1024 times, etc.). The screening criteria are such that the value of any sample instance is greater than the average of the samples by a certain amount. Another screening criterion is that the value of one sample is greater than the next sample by a certain amount. These techniques are described in U.S. Provisional Patent Application No. 62 / 933,809, titled "Precise Programming Method and Apparatus for Analog Neural Memory in a Deep Learning Artificial Neural Network", filed by the applicant on November 11, 2019, which is incorporated herein by reference.
[0173] A method of tuning the weights of memory cells incorporating some of the above weight assignments (programming or erasing the cells) can include soft erasing the cells, then programming zero-weight cells (such as those above), and then performing coarse, fine, and / or ultra-fine tuning algorithms using noise screening as described above. Techniques for coarse, fine, and ultra-fine tuning algorithms were previously disclosed by the application in U.S. Patent Application No. 16 / 829,757, filed on March 25, 2020, entitled "Precise Data Tuning Method and Apparatus for Analog Neural Memory in an Artificial Neural Network", which is hereby incorporated by reference.
[0174] In another embodiment, the noise contribution can be reduced by using a method of readout (inference) or verification that includes bias condition sequencing. For example, an accumulation condition is performed on the memory cells before the readout or verification operation is executed.
[0175] In another embodiment, the noise contribution can be reduced by applying a negative voltage range to the control gate. In another embodiment, the background data of the array for zero-weight and unused cells can be a specific pattern for reducing variations. For example, a high current level background may be desirable for noise such as RTN noise reduction. For example, a low current level background may be desirable for noise such as data drift. In standby or deep power down, the array is placed in the correct state by modulating the control gate voltage, for example, by using a specific voltage level compared to the control gate voltage used during the verification operation, which means that the control gate level can be set to lower the current or increase the current level during standby or deep power down operations.
[0176] In another embodiment, the read or verify method is performed by applying 0V, approximately 0V, or a low bias voltage to the control gate during the read or verify operation. The word line is used instead of the control gate line to receive the row data input (activation value) via, for example, a pulse width modulated input or an analog voltage applied to the word line.
[0177] In another embodiment, the "0" value for w (zero w or unused cell) can be defined as <10nA or another predetermined threshold. That is, when (w+) - (w-) < 10nA, w is given a value of "0". This provides a larger tolerance each time w = 0 and is more robust to inaccuracies caused by noise, temperature variations, or other forces.
[0178] In another embodiment, the weight mapping is optimized to reduce temperature variations. For example, if the desired value of w is 5nA, one possible mapping is (w+) = 5nA and (w-) = 0nA, and another possible mapping is (w+) = 30nA and (w-) = 25nA. The latter consumes more power and may introduce a larger inaccuracy during the tuning process but minimizes temperature variations.
[0179] In another embodiment, the weight mapping is optimized to reduce the overall noise or temperature variations of the neurons. For example, the number of stored w+ and w- values per bit line (which can be implemented as, for example, 30nA - 25nA or 50nA - 45nA or 80nA - 75nA for 5nA) can be mapped to balance among all bit lines such that the number of stored values per bit line is approximately the same for all bit lines.
[0180] In another embodiment, the weight mapping is optimized to reduce the total noise of the neurons (bit lines). For example, the number of stored w+ and w- values (cells) per bit line can be balanced among all bit lines such that the total noise contribution of all weights (cells) within a neuron (bit line) is optimal (has the minimum noise). Table 9A: Weight Tuning Method [Table 9A] Table 9B: Weight Tuning Method [Table 9B]
[0181] Tables 9A and 9B show exemplary embodiments of 16 levels (states) at nA. That is, the memory cells can have 16 levels as shown here. Table 9A shows a situation where w can be one of 16 different positive values according to the formula Iw = Iw+ - Iw-, and Table 9B shows a situation where w can be one of 16 different negative values. The current ranges shown in the table are 0 to 80 nA. Table 10A: Weight Tuning Method [Table 10A] Table 10B: Weight Tuning Method [Table 10B]
[0182] Tables 10A and 10B show embodiments that compress the dynamic full current range of the levels from 0 to 80 nA to 40 nA to 85 nA.
[0183] Regarding Tables 10A / 11A / 12A, Iw- can be tuned using coarse or fine or ultra-fine tuning steps (e.g., programming and verification of the Iw- cell), and Iw+ can be tuned using fine or ultra-fine steps (e.g., programming and verification of Iw = (Iw+ - Iw-), or verification of only the Iw+ cell). Regarding Tables 10B / 11B / 12B, Iw+ can be tuned using coarse or fine or ultra-fine tuning steps (e.g., programming and verification of the Iw+ cell), and Iw- can be tuned using fine or ultra-fine steps (e.g., programming and verification of Iw = (Iw+ - Iw-), or verification of only the Iw- cell).
[0184] This is advantageous in terms of reducing variations and inconsistencies due to process, temperature, noise, operating stress, or operating conditions, which is similar to the concept shown in FIGS. 11 - 12.
[0185] As shown in Tables 10A and 10B, the values in the lower half of the table are shifted up by a positive amount (offset bias) such that the entire range of those values is approximately the same as the upper half of the table. The offset is approximately equal to half of the maximum current (level). Table 11A: Weight Tuning Method
Table 11A
Table 11B
[0186] Tables 11A and 11B show embodiments similar to those of 10A and 10B having zero weight (w = 0) equal to the offset bias value. By way of example, one range of weight values for one offset bias value and two sub-ranges of weights for two offset values are shown. Table 10A and 10B show embodiments similar to those of 10A and 10B having zero weight (w = 0) equal to the offset bias value. By way of example, one range of weight values for one offset bias value and two sub-ranges of weights for two offset values are shown. Table 12A: Weight Tuning Method
Table 12A
Table 12B
Table 13
Table 14
Table 15
Table 16
[0187] Table 12A and Table 12B show embodiments similar to those of FIGS. 10A and 10B, where each w+ or w+ value is implemented by two memory cells, further reducing the total dynamic range by about half.
[0188] It should be understood that the values for w+ and w- provided in the above embodiments are merely examples, and other values can be used according to the disclosed concept. For example, the offset bias value may be any value that shifts the values of each level, or may be a fixed value for all levels. In practice, each w is implemented as a differential cell, which can be effective in minimizing variations or mismatches from process, temperature, noise (such as RTN or power supply noise), stress, or operating conditions.
[0189] FIG. 36A shows data behavior (I-V curve) over temperature (in the subthreshold region as an example), FIG. 36B shows problems caused by data drift during operation of the VMM system, FIGS. 36C and 36D show blocks for compensating data drift, and with respect to FIG. 36C, shows a block for compensating temperature change.
[0190] FIG. 36A shows known characteristics of the VMM system, where the characteristics are such that as the operating temperature increases, the sense current in any given selected non-volatile memory cell within the VMM array increases within the subthreshold region, decreases within the saturation region, or generally decreases within the linear region.
[0191] FIG. 36B shows the array current distribution over time of use (data drift), which shows that the aggregate output from the VMM array (which is the sum of the currents from all bit lines within the VMM array) shifts to the right (or, depending on the technology used, to the left) over the operating time of use, which means that the total aggregate output drifts over the lifetime of use of the VMM system. This phenomenon is known as data drift when data drifts due to usage conditions and degradation by environmental factors.
[0192] FIG. 36C shows a compensation current i at the output of the bit line output circuit 3610 to compensate for data drift COMPA bit line compensation circuit 3600 that may include injecting is shown. The bit line compensation circuit 3600 may include scaling the output up or down by a scaling circuit based on a resistor or capacitor network. The bit line compensation circuit 3600 may include shifting or offsetting the output by a shift circuit based on its resistor or capacitor network.
[0193] FIG. 36D shows a data drift monitor 3620 that detects the amount of data drift. Then, that information is used as an input to the bit line compensation circuit 3600, and as a result, COMP an appropriate level of i may be selected.
[0194] FIG. 37 shows a bit line compensation circuit 3700 that is an embodiment of the bit line compensation circuit 3600 of FIG. 36. The bit line compensation circuit 3700 includes an adjustable current source 3701 and an adjustable current source 3702, which together COMP generate i, and COMP i is equal to the current generated from the current generated by the adjustable current source 3701 minus the current generated by the adjustable current source 37 2 02.
[0195] FIG. 38 shows a bit line compensation circuit 3700 that is an embodiment of the bit line compensation circuit 3600 of FIG. 36. The bit line compensation circuit 3800 includes an operational amplifier 3801, an adjustable resistor 3802, and an adjustable resistor 3803. The operational amplifier 3801 receives a reference voltage VREF on its non-inverting terminal and receives V INPUT on its inverting terminal, where INPUT V is the voltage received from the bit line output circuit 3610 of FIG. 36C, and generates an output of V OUTPUT , where OUTPUT V is a scaled version of V for compensating data drift based on the ratio of resistor 3803 to resistor 3802. By configuring the values of resistor 3803 and / or 3802, V INPUT can be scaled up or down. OUTPUT
[0196] FIG. 39 shows a bit line compensation circuit 3900, which is an embodiment of the bit line compensation circuit 3600 of FIG. 36. The bit line compensation circuit 3900 includes an operational amplifier 3901, a current source 3902, a switch 3904, and an adjustable integration output capacitor 3903. Here, the current source 3902 is actually the output current of a single bit line or a set of multiple bit lines (such as one for summing the positive weight w+ and one for summing the negative weight w−) in the VMM array. The operational amplifier 3901 receives the reference voltage VREF on its non-inverting terminal and receives V INPUT there, where V INPUT is the voltage received from the bit line output circuit 3610 of FIG. 36C. The bit line compensation circuit 3900 functions as an integrator that integrates the current Ineu across the entire capacitor 3903 at an adjustable integration time to generate an output voltage V OUTPUT , where V OUTPUT = Ineu * integration time / C 3903 , where C 3903 is the value of the capacitor 3903. Therefore, the output voltage V OUTPUT is proportional to the (bit line) output current Ineu, proportional to the integration time, and inversely proportional to the capacitance of the capacitor 3903. The bit line compensation circuit 3900 generates the output of V OUTPUT , where the value of V OUTPUT is scaled based on the configured value of the capacitor 3903 and / or the integration time for compensating data drift.
[0197] FIG. 40 shows a bit line compensation circuit 4000, which is an embodiment of the bit line compensation circuit 3600 of FIG. 36. The bit line compensation circuit 4000 includes a current mirror 4010 having an M:N ratio, which means I COMP =(M / N) * i input . The current mirror 4010 receives the current i INPUT , mirrors that current, and optionally scales that current to generate i COMP . Therefore, by configuring the M and / or N parameters, iCOMP can be scaled up or down.
[0198] FIG. 41 shows a bit line compensation circuit 4100 which is an embodiment of the bit line compensation circuit 3600 of FIG. 36. The bit line compensation circuit 4100 includes an operational amplifier 4101, an adjustable scaling resistor 4102, an adjustable shift resistor 4103, and an adjustable resistor 4104. The operational amplifier 4101 receives a reference voltage V REF at its non-inverting terminal and receives V IN at its inverting terminal. V IN is generated in response to V INPUT and Vshft, where V INPUT is the voltage received from the bit line output circuit 3610 of FIG. 36C and Vshft is a voltage intended to implement a shift between V INPUT and V OUTPUT . Thus, V OUTPUT is a scaled and shifted version of V INPUT for compensating data drift.
[0199] FIG. 42 shows a bit line compensation circuit 4200 which is an embodiment of the bit line compensation circuit 3600 of FIG. 36. The bit line compensation circuit 4200 includes an operational amplifier 4201, an input current source Ineu4202, a current shifter 4203, switches 4205 and 4206, and an adjustable integrating output capacitor 4204. Here, the current source 4202 is actually the output current Ineu on a single or multiple bit lines within the VMM array. The operational amplifier 4201 receives a reference voltage VREF at its non-inverting terminal and receives I IN at its inverting terminal, where I IN is the sum of Ineu and the current output by the current shifter 4203, and generates the output of V OUTPUT , where V OUTPUT is scaled (based on capacitor 4204) and shifted (based on Ishifter4203) to compensate for data drift.
[0200] Figures 43 to 48 show various circuits that can be used to provide the W value to be programmed or read to each selected cell during the programming or read operation.
[0201] Figure 43 shows a neuron output circuit 4300 including an adjustable current source 4301 and an adjustable current source 4302, which together generate an I OUT and I OUT is equal to the current obtained by subtracting the current I W+ generated by the adjustable current source 4302 from the current I W- generated by the adjustable current source 4301. The adjustable current Iw+4301 is a scaled current of the cell current or neuron current (such as the bit line current) for implementing a positive weight. The adjustable current Iw-4302 is a scaled current of the cell current or neuron current (such as the bit line current) for implementing a negative weight. The current scaling is performed by an M:N ratio current mirror circuit, Iout=(M / N) * Iin, etc.
[0202] Figure 44 shows an adjustable capacitor 4401, a control transistor 4405, a switch 4402, a switch 4403, and and M :N current mirror circuit, etc., which is a scaled output current of the cell current or (bit line) neuron current Showing neuron output circuit 4400 with an adjustable current source 4404 for generating current Iw+ . The transistor 4405 is used, for example, to impose a fixed bias voltage on the current 4404. The circuit 4404 generates a V OUT where V OUT is inversely proportional to the capacitor 4401, proportional to the adjustable integration time (the time when the switch 4403 is closed and the switch 4402 is open), and proportional to the current I W+ generated by the adjustable current source 4404. V OUT is equal to V+-((Iw+ * integration time) / C 4401 ), where C 4401is the value of capacitor 4401. The positive terminal V+ of capacitor 4401 is connected to the positive supply voltage, and the negative terminal V- of capacitor 4401 is connected to the output voltage V OUT is connected.
[0203] Figure 45 shows capacitor 4401 and and M :N current mirror Circuit and so on Therefore a scaled current of the cell current or (bit line) neuron current Showing neuron circuit 4500 with an adjustable current source 4502 for generating Iw- . Circuit 4500 generates V OUT , where V OUT is inversely proportional to capacitor 4401, proportional to the adjustable integration time (the time switch 4501 is open), and proportional to the current I Wi generated by the adjustable current source 4502. After capacitor 4401 completes the operation of integrating current Iw+, it is reused from neuron output circuit 44. Then, the positive and negative terminals (V+ and V-) are exchanged within neuron output circuit 45, where the positive terminal is connected to the output voltage V OUT , which is discharged by current Iw-. The negative terminal is held at the previous voltage value by a clamp circuit (not shown). In practice, output circuit 44 is used for positive weight implementation, circuit 45 is used for negative weight implementation, and the final charge on capacitor 4401 effectively represents the total weight (Qw = Qw+ - Qw-).
[0204] Figure 46 shows neuron circuit 4600 including adjustable capacitor 4601, switch 4602, control transistor 4604, and adjustable current source 4603. Circuit 4600 generates V OUT , where V OUT is inversely proportional to capacitor 4601, proportional to the adjustable integration time (the time switch 4602 is open), and proportional to the current I W-is proportional to. The negative terminal V- of the capacitor 4601 is equal to, for example, ground. The positive terminal V+ of the capacitor 4601 is initially pre-charged to a positive voltage, for example, before integrating the current Iw-. The neuron circuit 4600 can be used instead of the neuron circuit 4500 together with the neuron circuit 4400 to implement the total weight (Qw = Qw+ - Qw-).
[0205] Figure 47 shows a neuron circuit 4700 comprising operational amplifiers 4703 and 4706, adjustable current sources Iw+ 4701 and Iw- 4702, and adjustable resistors 4704, 4705, and 4707. The neuron circuit 4700 generates V OUT which is equal to R 4707 * (Iw+ - Iw-). The adjustable resistor 4707 implements output scaling. The adjustable current sources Iw+ 4701 and Iw- 4702 also implement output scaling by, for example, an M:N ratio current mirror circuit (Iout = (M / N) * Iin).
[0206] Figure 48 shows a neuron circuit 4800 comprising operational amplifiers 4803 and 4806, switches 4808 and 4809, adjustable current sources Iw- 4802 and Iw+ 4801, and adjustable capacitors 4804, 4805, and 4807. The neuron circuit 4800 generates V OUT which is proportional to (Iw+ - Iw-), proportional to the integration time (the time the switches 4808 and 4809 were open), and inversely proportional to the capacitance of the capacitor 4807. The adjustable capacitor 4807 implements output scaling. The adjustable current sources Iw+ 4801 and Iw- 4802 also implement output scaling by, for example, an M:N ratio current mirror circuit (Iout = (M / N) * Iin). The integration time can also adjust the output scaling.
[0207] Figures 49A, 49B, and 49C show block diagrams of output circuits such as the output circuit 3307 of FIG. 33.
[0208] In FIG. 49A, output circuit 4901 includes an ADC circuit 4911, which is used to directly digitize the analog neuron output 4910 to provide a digital output bit 4912.
[0209] In FIG. 49B, output circuit 4902 includes a neuron output circuit 4921 and an ADC 4911. The neuron output circuit 4921 receives the neuron output 4920, shapes it, and then it is digitized by the ADC circuit 4911 to generate an output 4912. The neuron output circuit 4921 can be used for normalization, scaling, shifting, mapping, arithmetic operations, activation, and / or temperature compensation as described above. The ADC circuit can be a serial (ramp or ramp or count) ADC, SAR ADC, pipeline ADC, sigma-delta ADC, or any type of ADC.
[0210] In FIG. 49C, the output circuit includes a neuron output circuit 4921 that receives the neuron output 4930, and the converter circuit 4931 is for converting the output from the neuron output circuit 4921 to an output 4932. The converter 4931 can include an ADC, an AAC (analog-to-analog converter such as a current-voltage converter), an APC (analog-to-pulse converter), or any other type of converter. The ADC 4911 or the converter 4931 can be used to implement an activation function, for example, by bit mapping (e.g., quantization) or clipping (e.g., clipped ReLU). The ADC 4911 and the converter 4931 can be configured to have lower or higher accuracy (e.g., a lower or higher number of bits), lower or higher performance (e.g., a slower or faster speed), etc.
[0211] Another embodiment for scaling and shifting is to convert the array (bit line) output to digital bits, such as those having less or more bit precision, and then operate on the digital output bits by normalization (e.g., from 12 bits to 8 bits), shifting, or remapping, etc., according to a specific function (e.g., linear or non - linear, compression, non - linear activation, etc.), which is achieved by constructing an ADC (analog - to - digital) conversion circuit (such as a serial ADC, SAR ADC, pipeline ADC, ramp ADC, etc.). An example of the ADC conversion circuit is described in U.S. Patent Provisional Application No. 62 / 933,809, entitled "Precise Programming Method and Apparatus for Analog Neural Memory in a Deep Learning Artificial Neural Network", filed on November 11, 2019 by the same assignee as this application, which is incorporated herein by reference.
[0212] Table 17 shows an alternative approach for performing read, erase, and program operations. Table 17 : Operation of Flash Memory Cells
Table 17
[0213] An embodiment for scaling the input can enable a specific number of rows of the VMM at a time and then perform it, for example, by fully combining the results.
[0214] Another embodiment is to scale the input voltage and appropriately rescale the output for normalization.
[0215] Another embodiment for scaling a pulse-width modulation input is by modulating the timing of the pulse width. An example of this technique is described in U.S. Patent Application No. 16 / 449,201, entitled "Configurable Input Blocks and Output Blocks and Physical Layout for Analog Neural Memory in Deep Learning Artificial Neural Network", filed on June 21, 2019, by the same assignee as this application, which is incorporated herein by reference.
[0216] Another embodiment for scaling an input is, for example, for an 8-bit input IN7:0, enabling the input binary bits one at a time, evaluating IN0, IN1,..., IN7 in order, and then combining the output results with appropriate binary bit weighting. An example of this technique is described in U.S. Patent Application No. 16 / 449,201, entitled "Configurable Input Blocks and Output Blocks and Physical Layout for Analog Neural Memory in Deep Learning Artificial Neural Network", filed on June 21, 2019, by the same assignee as this application, which is incorporated herein by reference.
[0217] Optionally, in the above embodiments, measuring the cell current for the purpose of verifying or reading out the current can be, for example, taking 8 to 32 average or multiple measurements to reduce the influence of noise (such as RTN or any random noise) and / or detecting any outlier bits that are defective and need to be replaced by redundant bits.
[0218] As used herein, it should be noted that both the terms "over" and "on" include both "directly on" (with no intervening material, element, or gap therebetween) and "indirectly on" (with an intervening material, element, or gap therebetween). Similarly, the term "adjacent" includes "directly adjacent" (with no intervening material, element, or gap therebetween) and "indirectly adjacent" (with an intervening material, element, or gap therebetween), "attached to" includes "directly attached to" (with no intervening material, element, or gap therebetween) and "indirectly attached to" (with an intervening material, element, or gap therebetween), and "electrically coupled" includes "directly electrically coupled" (with no intervening material or element electrically connecting the elements together therebetween) and "indirectly electrically coupled" (with an intervening material or element electrically connecting the elements together therebetween). For example, forming an element "over a substrate" can include forming the element directly on the substrate without an intervening material / element therebetween, and forming the element indirectly on the substrate with one or more intervening materials / elements therebetween.
Claims
1. A neural network comprising: A neural network comprising a vector matrix multiplication array of non-volatile memory cells, wherein weight values w are stored as differential pairs w+ and w− of first and second non-volatile memory cells in the array according to a formula w=(w+)-(w−), where w+ and w− include non-zero offset values.
2. 2. The neural network of claim 1, wherein the non-zero offset value is a positive value.
3. 2. The neural network of claim 1, wherein the non-zero offset value is a negative value.
4. 2. The neural network of claim 1, wherein when a value of 0 is desired for w, values equal to the non-zero offset value are stored for w+ and w-.
5. 2. The neural network of claim 1, wherein a value of 0 for w is indicated by a stored value for w that is less than a predetermined threshold.
6. 6. The neural network of claim 5, wherein the predetermined threshold is 5 nA.
7. 6. The neural network of claim 5, wherein the predetermined threshold is 10 nA.
8. 2. The neural network of claim 1, wherein the non-volatile memory cells are split-gate flash memory cells.
9. 2. The neural network of claim 1, wherein the non-volatile memory cells are stacked gate flash memory cells.
10. 2. The neural network of claim 1, wherein the first non-volatile memory cell is tuned to w+ by one or more of a coarse, fine, or ultra-fine tuning algorithm, and the second non-volatile memory cell is tuned to w- by one or more of a fine or ultra-fine tuning algorithm.
11. A neural network comprising:
1. A neural network comprising: a vector matrix multiplication array of non-volatile memory cells, said array being organized into rows and columns of non-volatile memory cells, wherein weight values w are stored as differential pairs w+ and w− of first and second non-volatile memory cells according to a formula w=(w+)-(w−), and wherein storage of w+ and w− values is approximately evenly distributed among all columns in said array.
12. 12. The neural network of claim 11, wherein the non-zero offset value is a positive value.
13. 12. The neural network of claim 11, wherein the non-zero offset value is a negative value.
14. 12. The neural network of claim 11, wherein when a value of 0 is desired for w, values equal to the non-zero offset value are stored for w+ and w-.
15. 12. The neural network of claim 11, wherein a value of 0 for w is indicated by a stored value for w that is less than a predetermined threshold.
16. 16. The neural network of claim 15, wherein the predetermined threshold is 5 nA.
17. 16. The neural network of claim 15, wherein the predetermined threshold is 10 nA.
18. 12. The neural network of claim 11, wherein the non-volatile memory cells are split-gate flash memory cells.
19. 12. The neural network of claim 11, wherein the non-volatile memory cells are stacked gate flash memory cells.
20. 12. The neural network of claim 11, wherein the first non-volatile memory cell is tuned to w+ by one or more of a coarse, fine, or ultra-fine tuning algorithm, and the second non-volatile memory cell is tuned by one or more of a fine or ultra-fine tuning algorithm.
21. 1. A method for programming, verifying, and reading zero values in differential pairs of non-volatile memory cells in a vector matrix multiplication array, the method comprising: programming a first cell w+ in the differential pair to a first current value; verifying the first cell by applying a voltage equal to a first voltage plus a bias voltage to a control gate terminal of the first cell; programming a second cell w− in the differential pair to the first current value; verifying the second cell by applying a voltage equal to the first voltage plus the bias voltage to a control gate terminal of the second cell; reading the first cell by applying a voltage equal to the first voltage to the control gate terminal of the first cell; reading the second cell by applying a voltage equal to the first voltage to the control gate terminal of the second cell; and calculating a value w according to the formula w=(w+)-(w-).
22. 22. The method of claim 21, wherein the first cell is tuned by one or more of a coarse, fine, or ultra-fine tuning algorithm and the second cell is tuned by one or more of a fine or ultra-fine tuning algorithm.
23. 1. A method for programming, verifying, and reading zero values in differential pairs of non-volatile memory cells in a vector matrix multiplication array, the method comprising: programming a first cell w+ in the differential pair to a first current value; verifying the first cell by applying a voltage equal to a first voltage plus a bias voltage to a control gate terminal of the first cell; programming a second cell w− in the differential pair to the first current value; verifying the second cell by applying a voltage equal to the first voltage plus the bias voltage to a control gate terminal of the second cell; reading the first cell by applying a voltage equal to the first voltage to the control gate terminal of the first cell; reading the second cell by applying a voltage equal to the first voltage to the control gate terminal of the second cell; and calculating a value w according to the formula w=(w+)-(w-).
24. 24. The method of claim 23, wherein the first current value is a positive value.
25. 24. The method of claim 23, wherein the second current value is a negative value.
26. 24. The method of claim 23, wherein the first cell is tuned to w+ by one or more of a coarse, fine, or ultra-fine tuning algorithm, and the second cell is tuned to w- by one or more of a fine or ultra-fine tuning algorithm.
27. A neural network comprising:
1. A neural network comprising: a vector matrix multiplication array of non-volatile memory cells, the array being organized into rows and columns of non-volatile memory cells, weight values w being stored as differential pairs w+ and w− according to a formula w=(w+)-(w−), w+ being stored as a differential pair of a first non-volatile memory cell and a second non-volatile memory cell in the array, and w− being stored as a differential pair of a third non-volatile memory cell and a fourth non-volatile memory cell in the array, and the storage of the w+ and w− values being offset by a bias value.
28. 28. The neural network of claim 27, wherein the current ranges of all said memory cells are reduced by approximately a memory cell current value representing a bias value.
29. 28. The neural network of claim 27, wherein the memory cells have several current levels.
30. 30. The neural network of claim 28, wherein bias values are shared across several memory cell levels.
31. 28. The neural network of claim 27, wherein the bias value is a positive value.
32. 28. The neural network of claim 27, wherein the bias value is a negative value.
33. 28. The neural network of claim 27, wherein when a value of 0 is desired for w, values equal to the bias value are stored for w+ and w-.
34. 28. The neural network of claim 27, wherein a value of 0 for w is indicated by a stored value for w that is less than a predetermined threshold.
35. 35. The neural network of claim 34, wherein the predetermined threshold is 5 nA.
36. 35. The neural network of claim 34, wherein the predetermined threshold is 10 nA.
37. 28. The neural network of claim 27, wherein the non-volatile memory cells are split-gate flash memory cells.
38. 28. The neural network of claim 27, wherein the non-volatile memory cells are stacked gate flash memory cells.
39. 28. The neural network of claim 27, wherein the first and third non-volatile memory cells are tuned by one or more of a coarse, fine, or ultra-fine tuning algorithm, and the second and fourth non-volatile memory cells are tuned by one or more of a fine or ultra-fine tuning algorithm.
40. A neural network comprising:
1. A neural network comprising: a vector matrix multiplication array of non-volatile memory cells, the array being organized into rows and columns of non-volatile memory cells, and weight values w being stored as differential pairs w+ and w− of first and second non-volatile memory cells according to a formula w=(w+)-(w−), a value for w+ being selected from a first range of non-zero values and a value for w− being selected from a second range of non-zero values, the first range and the second ranges not overlapping.
41. 41. The neural network of claim 40, wherein a value of 0 for w is indicated by a stored value for w less than a predetermined threshold.
42. 42. The neural network of claim 41, wherein the predetermined threshold is 5 nA.
43. 42. The neural network of claim 41, wherein the predetermined threshold is 10 nA.
44. 41. The neural network of claim 40, wherein the non-volatile memory cells are split-gate flash memory cells.
45. 41. The neural network of claim 40, wherein the non-volatile memory cells are stacked gate flash memory cells.
46. 41. The neural network of claim 40, wherein the first non-volatile memory cell is tuned to w+ by one or more of a coarse, fine, or ultra-fine tuning algorithm, and the second non-volatile memory cell is tuned to w- by one or more of a fine or ultra-fine tuning algorithm.
47. 1. A method for reading non-volatile memory cells in a vector matrix multiplication array, the method comprising: reading weights stored in selected cells in the array, said reading step comprising: applying a zero voltage bias to a control gate terminal of the selected cell; and sensing a neuronal output current comprising a current output from the selected cell.
48. The reading step includes:
48. The method of claim 47, further comprising applying a voltage to a word line terminal of the selected cell.
49. 1. A method of operating non-volatile memory cells in a vector matrix multiplication array, the method comprising: reading the non-volatile memory cell by applying a first bias voltage to a control gate of the non-volatile memory cell; applying a second bias voltage to the control gate of the non-volatile memory cell during one or more of a standby operation, a deep power down operation, or a test operation.
50. 50. The method of claim 49, further comprising the step of modulating a background data pattern or zero weight or non-user cells in the array.
Citation Information
Patent Citations
Non-volatile semiconductor memory
JP2000268593A
Semiconductor memory device
JP2008084499A
Vector-by-matrix multiplier modules based on non-volatile 2d and 3D memory arrays
US20190213234A1
Precision tuning for the programming of analog neural memory in a deep learning artificial neural network
WO2020081140A1