Verification method and system in artificial neural network arrays
Patent Information
- Application Number
- KR1020257006207
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-12-13
- Filing Date
- 2022-12-15
- Publication Date
- 2026-09-21
- Estimated Expiration
- 2042-12-15
Smart Images

Figure 112025021553939-PCT00016_ABST
Abstract
Description
Technology Field
[0001] Claim of priority
[0002] This application claims priority to U.S. Patent Application No. 18 / 080,545, filed December 13, 2022, with the title "Verification Method and System in Artificial Neural Network Array," and U.S. Provisional Patent Application No. 63 / 409,142, filed September 22, 2022, with the title "Verification Method and System in Artificial Neural Network Array."
[0003] Technology field
[0004] A number of examples of verification circuits and associated methods in artificial neural networks are disclosed. Background Technology
[0005] Artificial neural networks mimic biological neural networks (the central nervous system of animals, particularly the brain) and are used to estimate or approximate functions that may rely on multiple inputs and are generally unknown. Artificial neural networks generally consist of layers of interconnected "neurons" that exchange messages with one another.
[0006] Figure 1 illustrates an artificial neural network, where circles represent layers of neurons or inputs. Connections (referred to as synapses) are represented by arrows and have numerical weights that can be tuned based on experience. This enables the neural network to adapt to inputs and learn. Typically, a neural network includes multiple layers of inputs. Typically, there are one or more intermediate layers of neurons and an output layer of neurons that provides the output of the neural network. Neurons at each level make decisions individually or collectively based on data received from synapses.
[0007] One of the major challenges in the development of artificial neural networks for high-performance information processing is the lack of suitable hardware technology. In fact, real neural networks rely on a very large number of synapses to enable high connectivity between neurons—that is, very high computational parallelism. In principle, such complexity can be achieved with digital supercomputers or clusters of specialized graphics processing units. However, in addition to high costs, these approaches are also concerned about their moderate energy efficiency compared to biological networks, which consume much less energy, because they primarily perform low-precision analog calculations. Although CMOS analog circuits have been used in artificial neural networks, most CMOS-implemented synapses were too bulky considering the large number of neurons and synapses.
[0008] The applicant previously disclosed an artificial (analog) neural network utilizing one or more non-volatile memory arrays as synapses in U.S. Patent Application Publication No. 2017 / 0337466A1, incorporated by reference. The non-volatile memory array operates as an analog neural memory and comprises non-volatile memory cells arranged in rows and columns. The neural network comprises a first plurality of synapses configured to receive a first plurality of inputs and generate a first plurality of outputs therefrom, and a first plurality of neurons configured to receive the first plurality of outputs. The first plurality of synapses comprises a plurality of memory cells, wherein each memory cell comprises a separated source region and a drain region formed within a semiconductor substrate and having a channel region extending between them, a floating gate disposed over a first portion of the channel region and isolated therefrom, and a non-floating gate disposed over a second portion of the channel region and isolated therefrom. Each of the plurality of memory cells stores a weight value corresponding to the number of electrons on the floating gate. The plurality of memory cells generate a first plurality of outputs by multiplying the first plurality of inputs by the stored weight value.
[0009] non-volatile memory cells
[0010] Non-volatile memory is well known. For example, U.S. Patent No. 5,029,130 ("130 patent") incorporated herein by reference discloses an array of separated gate non-volatile memory cells, which is a type of flash memory cell. Such memory cells (210) are illustrated in FIG. 2. Each memory cell (210) comprises a source region (14) and a drain region (16) formed within a semiconductor substrate (12), with a channel region (18) between these regions. A floating gate (20) is formed over a portion of the source region (14) and over a first portion of the channel region (18) and is insulated from there (and controls its conductivity). A word line terminal (22) (typically coupled to a word line) has a first portion disposed over a second portion of the channel region (18) and is insulated from there (and controls its conductivity), and a second portion extending upward and over the floating gate (20). The floating gate (20) and the word line terminal (22) are insulated from the substrate (12) by the gate oxide. The bit line (24) is coupled to the drain region (16).
[0011] The memory cell (210) is erased by applying a high positive voltage to the word line terminal (22) (where electrons are removed from the floating gate), which causes electrons on the floating gate (20) to tunnel from the floating gate (20) to the word line terminal (22) by passing through an intermediate insulator via Fowler-Nordheim (FN) tunneling.
[0012] The memory cell (210) is programmed by source-side injection (SSI) with hot electrons (where the electrons are placed in the floating gate) by applying a positive voltage to the word line terminal (22) and a positive voltage to the source region (14). An electron current will flow from the drain region (16) toward the source region (14). The electrons will be accelerated and heated when they reach the gap between the word line terminal (22) and the floating gate (20). Some of the heated electrons will pass through the gate oxide due to electrostatic attraction from the floating gate (20) and be injected onto the floating gate (20).
[0013] The memory cell (210) is read by placing a positive read voltage on the drain region (16) and the word line terminal (22) (which turns on the portion of the channel region (18) below the word line terminal). When the floating gate (20) is positively charged (i.e., when electrons are erased), the portion of the channel region (18) below the floating gate (20) is also turned on, and current will flow across the channel region (18), which is detected as an erased state or a "1" state. When the floating gate (20) is negatively charged (i.e. programmed with electrons), the portion of the channel region below the floating gate (20) is mostly or completely turned off, and current will not flow across the channel region (18) (or there will be almost no flow), which is detected as a programmed state or a "0" state.
[0014] Table 1 shows typical voltage and current ranges that can be applied to the terminals of the memory cell (210) to perform read, erase, and program operations:
[0015] [Table 1]
[0016]
[0017] Other types of flash memory cells, such as other isolated gate memory cell configurations, are known. For example, FIG. 3 illustrates a 4-gate memory cell (310) comprising a source region (14), a drain region (16), a floating gate (20) over a first portion of a channel region (18), a select gate (22) (typically connected to a word line (WL)) over a second portion of the channel region (18), a control gate (28) over the floating gate (20), and an erase gate (30) over the source region (14). Such a configuration is described in U.S. Patent No. 6,747,310, incorporated herein by reference for all purposes. Here, all gates are non-floating gates except for the floating gate (20), which means that they are electrically connected or can be connected to a voltage source. Programming is performed by heated electrons from the channel region (18) injecting heated electrons onto the floating gate (20). Erasing is performed by electrons tunneling from the floating gate (20) to the erase gate (30).
[0018] Table 2 shows typical voltage and current ranges that can be applied to the terminals of the memory cell (310) to perform read, erase, and program operations:
[0019] [Table 2]
[0020]
[0021] FIG. 4 illustrates a 3-gate memory cell (410), which is another type of flash memory cell. The memory cell (410) is identical to the memory cell (310) of FIG. 3, except that the memory cell (410) does not have a separate control gate. The erase operation (whereby erasure occurs through the use of the erase gate) and the read operation are similar to those of FIG. 3, except that no control gate bias is applied. The programming operation is also performed without control gate bias, and as a result, a higher voltage is applied to the source line during the programming operation to compensate for the lack of control gate bias.
[0022] Table 3 shows typical voltage and current ranges that can be applied to the terminals of the memory cell (410) to perform read, erase, and program operations:
[0023] [Table 3]
[0024]
[0025] FIG. 5 illustrates a stacked gate memory cell (510), which is another type of flash memory cell. The memory cell (510) is similar to the memory cell (210) of FIG. 2, except that a floating gate (20) extends over the entire channel region (18) and a control gate (22) (which will be connected to the word line) extends over the floating gate (20) and they are separated by an insulating layer (not shown). Erasing is performed by FN tunneling of electrons from the FG to the substrate, programming is performed by channel hot electron (CHE) injection in the region between the channel (18) and the drain region (16) by electrons flowing from the source region (14) toward the drain region (16), and reading operations are similar to those of the memory cell (210) but are performed at a higher control gate voltage.
[0026] Table 4 shows typical voltage ranges that can be applied to the terminals of the substrate (12) and the memory cell (510) to perform read, erase, and program operations:
[0027] [Table 4]
[0028]
[0029] The methods and means described herein may be applied, without limitation, to other non-volatile memory technologies such as FINFET isolated gate flash or stacked gate flash memory, NAND flash, SONOS (silicon-oxide-nitride-oxide-silicon, charge trap in nitride), MONOS (metal-oxide-nitride-oxide-silicon, metal charge trap in nitride), ReRAM (resistive ram), PCM (phase change memory), MRAM (magnetic ram), FeRAM (ferroelectric ram), CT (charge trap) memory, CN (carbon-tube) memory, OTP (bi-level or multi-level one time programmable), and CeRAM (correlated electron ram).
[0030] To utilize a memory array containing one of the types of non-volatile memory cells described above in an artificial neural network, two modifications are made. First, the line is configured so that each memory cell can be programmed, erased, and read individually without adversely affecting the memory state of other memory cells within the array, as further described below. Second, continuous (analog) programming of the memory cells is provided.
[0031] Specifically, the memory state of each memory cell in the array (i.e., the charge on the floating gate) can be continuously changed from a completely erased state to a fully programmed state, independently and with minimal disturbance to other memory cells, and vice versa. This means that the cell storage is actually analog or at least can store one of many individual values (e.g., 16 or 64 different values), which allows for very precise and individual tuning of all memory cells in the memory array, making the memory array ideal for storing synaptic weights of a neural network and performing fine-tuning adjustments on them.
[0032] Neural network employing a non-volatile memory cell array
[0033] FIG. 6 conceptually illustrates a non-limiting example of a neural network utilizing a non-volatile memory array of the present example. While this example uses a non-volatile memory array neural network for a facial recognition application, any other suitable application may be implemented using a non-volatile memory array-based neural network.
[0034] S0 is an input layer that is a 32x32 pixel RGB image with 5-bit precision in this example (i.e., three 32x32 pixel arrays, one for each color R, G, and B, each pixel having 5-bit precision). The synapse (CB1) from the input layer (S0) to the layer (C1) applies different sets of weights in some cases and shared weights in others, scans the input image with a 3x3 pixel overlapping filter (kernel), and shifts the filter by 1 pixel (or more than 1 pixel as indicated by the model). Specifically, values for nine pixels (i.e., referred to as filters or kernels) within a 3x3 portion of the image are provided to a synapse (CB1), where these nine input values are multiplied by appropriate weights and the outputs of the multiplication are summed, and a single output value is determined and provided by the first synapse of CB1 to generate one pixel of the feature map of layer (C1). Then, the 3x3 filter is shifted one pixel to the right within the input layer (S0) (i.e., a column of three pixels is added to the right and a column of three pixels is subtracted from the left), thereby providing the nine pixel values in this newly positioned filter to the synapse (CB1), where they are multiplied by the same weights and a second single output value is determined by the associated synapse. This process continues for all three colors and for all bits (precision values) until a 3x3 filter scans across the entire 32x32 pixel image of the input layer (S0). Then, the process is repeated using different sets of weights to generate different feature maps of layer (C1) until all feature maps of layer (C1) are calculated.
[0035] In this example, layer (C1) has 16 feature maps, each having 30x30 pixels. Each pixel is a new feature pixel extracted by multiplying the input and the kernel, and thus each feature map is a two-dimensional array, and thus, in this example, layer (C1) constitutes 16 layers of a two-dimensional array (note that the layers and arrays mentioned herein are not necessarily physical but logical—that is, the array is not necessarily physically oriented as a two-dimensional array). Each of the 16 feature maps in layer (C1) is generated by one of 16 different sets of synaptic weights applied to the filter scan. All of the C1 feature maps may relate to different aspects of the same image feature, such as boundary identification. For example, a first map (generated using a first set of weights shared for all scans used to generate this first map) can identify circular edges, and a second map (generated using a second set of weights different from the first set of weights) can identify rectangular edges, or the aspect ratio of a specific feature, etc.
[0036] An activation function (P1) (pooling) is applied before moving from layer (C1) to layer (S1), which pools values from consecutive non-overlapping 2x2 regions within each feature map. The purpose of the pooling function (P1) is to average nearby locations (or a maximum function may be used) to reduce edge location dependency and decrease data size before moving to the next stage, for example. In layer (S1), there are 16 15x15 feature maps (i.e., 16 different arrays of 15x15 pixels each). A synapse (CB2) moving from layer (S1) to layer (C2) scans the maps within layer (S1) with a 4x4 filter and a filter shift of 1 pixel. In layer (C2), there are 22 12x12 feature maps. An activation function (P2) (pooling) is applied before going from layer (C2) to layer (S2), which pools values from consecutive non-overlapping 2x2 regions within each feature map. In layer (S2), there are 22 6x6 feature maps. An activation function (pooling) is applied at the synapse (CB3) going from layer (S2) to layer (C3), where all neurons in layer (C3) are connected to all maps in layer (S2) through their respective synapses in CB3. In layer (C3), there are 64 neurons. The synapse (CB4) going from layer (C3) to the output layer (S3) completely connects C3 to S3, meaning that all neurons in layer (C3) are connected to all neurons in layer (S3). The output in S3 contains 10 neurons, where the highest output neuron determines the class. This output may, for example, indicate the identification or classification of the content of the original image.
[0037] Each layer of synapses is implemented using an array of non-volatile memory cells or a part of an array thereof.
[0038] FIG. 7 is a block diagram of an array that can be used for that purpose. A vector-by-matrix multiplication (VMM) array (32) comprises non-volatile memory cells and is used as a synapse (e.g., CB1, CB2, CB3, and CB4 in FIG. 6) between one layer and the next. Specifically, the VMM array (32) comprises an array of non-volatile memory cells (33), an erase gate and word line gate decoder (34), a control gate decoder (35), a bit line decoder (36), and a source line decoder (37), which each decode their respective inputs to the non-volatile memory cell array (33). The input to the VMM array (32) may be from the erase gate and word line gate decoder (34) or from the control gate decoder (35). In this example, the source line decoder (37) also decodes the output of the non-volatile memory cell array (33). Alternatively, the bit line decoder (36) can decode the output of the non-volatile memory cell array (33).
[0039] The non-volatile memory cell array (33) serves two purposes. First, it stores weights to be used by the VMM array (32). Second, the non-volatile memory cell array (33) effectively multiplies the input with the weights stored in the non-volatile memory cell array (33) and adds them for each output line (source line or bit line) to produce an output, which will be an input to the next layer or an input to the final layer. By performing the multiplication and addition functions, the non-volatile memory cell array (33) eliminates the need for separate multiplication and addition logic circuits and is also power efficient due to its in-situ memory calculation.
[0040] The output of the non-volatile memory cell array (33) is supplied to a differential summer (e.g., summing operation amplifier or summing current mirror) (38) that sums the outputs of the non-volatile memory cell array (33) to generate a single value for the corresponding convolution. The differential summer (38) is arranged to perform the summing of positive weights and negative weights.
[0041] Next, the summed output value of the differential summer (38) is supplied to the activation function block (39), which rectifies the output. The activation function block (39) may provide a sigmoid, tanh, or ReLU function. The rectified output value of the activation function block (39) becomes an element of the feature map as the next layer (e.g., C1 in FIG. 6) and is then applied to the next synapse to create the next feature map layer or the final layer. Thus, in this example, the non-volatile memory cell array (33) constitutes multiple synapses (which receive their inputs from the previous neural layer or from an input layer such as an image database), and the summing operation amplifier (38) and the activation function block (39) constitute multiple neurons.
[0042] The inputs (WLx, EGx, CGx, and optionally BLx and SLx) to the VMM array (32) of FIG. 7 may be analog levels, binary levels, or digital bits (in which case a DAC is provided to convert the digital bits to an appropriate input analog level), and the outputs may be analog levels, binary levels, or digital bits (in which case an output ADC is provided to convert the output analog levels to digital bits).
[0043] FIG. 8 is a block diagram illustrating the use of multiple layers of a VMM array (32), denoted herein as VMM arrays (32a, 32b, 32c, 32d, and 32e). As illustrated in FIG. 8, an input denoted as Inputx is converted from digital to analog by a digital-to-analog converter (31) and provided to an input VMM array (32a). The converted analog inputs may be voltage or current. Input D / A conversion for the first layer may be performed by using a function or lookup table (LUT) that maps the input (Inputx) to an appropriate analog level for a matrix multiplier of the input VMM array (32a). Input conversion may also be performed by an analog-to-analog (A / A) converter to convert an external analog input into a mapped analog input to the input VMM array (32a).
[0044] The output generated by the input VMM array (32a) is provided as input to the next VMM array (hidden level 1) (32b), which then generates an output provided as input to the next VMM array (hidden level 2) (32c), and so on. The various layers of the VMM array (32) function as different layers of synapses and neurons of a convolutional neural network (CNN). Each VMM array (32a, 32b, 32c, 32d, and 32e) may be a standalone physical non-volatile memory array, multiple VMM arrays may utilize different parts of the same physical non-volatile memory array, or multiple VMM arrays may utilize overlapping parts of the same physical non-volatile memory array. The example illustrated in FIG. 8 includes the following five layers (32a, 32b, 32c, 32d, 32e): one input layer (32a), two hidden layers (32b, 32c), and two fully connected layers (32d, 32e). A person skilled in the art will recognize that this is merely an example and that the system may instead include more than two hidden layers and more than two fully connected layers.
[0045] Vector-Matrix Multiplication (VMM) Array
[0046] FIG. 9 illustrates a neural VMM array (900) that is particularly suitable for memory cells (310) as illustrated in FIG. 3 and is used as parts of neurons and synapses between an input layer and a next layer. The VMM array (900) includes a memory array (901) of non-volatile memory cells and a reference array (902) of non-volatile reference memory cells (located at the top of the array). Alternatively, another reference array may be placed at the bottom.
[0047] In the VMM array (900), control gate lines such as control gate line (903) are connected in a vertical direction (thus, the reference array (902) in the row direction is orthogonal to the control gate line (903)), and erase gate lines such as erase gate line (904) are connected in a horizontal direction. Here, input to the VMM array (900) is provided to the control gate lines (CG0, CG1, CG2, CG3), and output of the VMM array (900) appears on the source lines (SL0, SL1). In one example, only even rows are used, and in another example, only odd rows are used. The current placed in each source line (SL0, SL1, each) performs a summation function of all currents from the memory cells connected to that specific source line.
[0048] As described herein with respect to the neural network, the non-volatile memory cell of the VMM array (900), that is, the memory cell (310) of the VMM array (900), can be configured to operate in a sub-threshold region.
[0049] The non-volatile reference memory cell and the non-volatile memory cell described herein are biased with weak inversion as follows (lower threshold region):
[0050] Ids = Io * e (Vg- Vth) / nVt = w * Io * e (Vg) / nVt ,
[0051] Here w = e (- Vth) / nVt
[0052] Here, Ids is the drain-source current; Vg is the gate voltage on the memory cell; Vth is the threshold voltage of the memory cell; Vt is the thermal voltage = k*T / q, where k is the Boltzmann constant, T is the temperature in Kelvin, and q is the electron charge; n is the gradient coefficient = 1 + (Cdep / Cox), where Cdep is the capacitance of the depletion layer and Cox is the capacitance of the gate oxide layer; Io is the memory cell current at a gate voltage equal to the threshold voltage, and Io is proportional to (Wt / L)*u*Cox* (n-1)*Vt2, where u is the carrier mobility and Wt and L are the width and length of the memory cell, respectively.
[0053] For an IV log converter using a memory cell (e.g., a reference memory cell or a peripheral memory cell) or a transistor for converting input current into input voltage:
[0054] Vg = n * Vt * log [Ids / wp * Io]
[0055] Here, wp is the w of the reference or surrounding memory cell.
[0056] For a memory array used as a vector matrix multiplier (VMM) array with current input, the output current is as follows:
[0057] Iout = wa * Io * e (Vg) / nVt, in other words
[0058] Iout = (wa / wp) * Iin = W * Iin
[0059] W = e (Vthp - Vtha) / nVt
[0060] Here, wa = w of each memory cell in the memory array.
[0061] Vthp is the effective threshold voltage of the peripheral memory cells, and Vtha is the effective threshold voltage of the main (data) memory cells. Note that the transistor's threshold voltage is a function of the substrate body bias voltage, and that the substrate body bias voltage, denoted as Vsb, can be modulated to compensate for various conditions at such temperatures. The threshold voltage Vth can be expressed as follows:
[0062] Vth = Vth0 + gamma (SQRT |Vsb - 2*φF) - SQRT |2*φF|)
[0063] Here, Vth0 is the threshold voltage at zero substrate bias, φF is the surface potential, and gamma is the body effect parameter.
[0064] A word line or control gate can be used as an input to the memory cell for the input voltage.
[0065] Alternatively, the flash memory cells of the VMM arrays described herein may be configured to operate in a linear region:
[0066] Ids = Beta * (Vgs - Vth) * Vds; Beta = u * Cox * Wt / L
[0067] W = α (Vgs-Vth)
[0068] This means that the weight W in the linear region is proportional to (Vgs-Vth).
[0069] A word line or control gate or bit line or source line can be used as an input to a memory cell operating in a linear region. A bit line or source line can be used as an output to a memory cell.
[0070] For an IV linear converter, a memory cell (e.g., a reference memory cell or a peripheral memory cell) or a transistor operating in the linear region can be used to linearly convert an input / output current into an input / output voltage.
[0071] Alternatively, the memory cells of the VMM arrays described herein may be configured to operate in a saturation region:
[0072] Ids = ½ * β * (Vgs-Vth) 2 ; Beta = u*Cox*Wt / L
[0073] Wα (Vgs-Vth) 2 , this means that weight W is (Vgs-Vth) 2 It means that it is proportional to
[0074] A word line, control gate, or erase gate can be used as an input to a memory cell operating in the saturation region. A bit line or source line can be used as an output to an output neuron.
[0075] Alternatively, the memory cells of the VMM array described herein may be used in all regions of each layer or multiple layers of the neural network or a combination thereof (sub-threshold, linear, or saturated).
[0076] Another example of the VMM array (32) of FIG. 7 is described in U.S. Patent No. 10,748,630, which is incorporated herein by reference. As described in that application, a source line or a bit line can be used as a neural output (current summing output).
[0077] FIG. 10 illustrates a neural VMM array (1000), which is particularly suitable for memory cells (210) as illustrated in FIG. 2 and is utilized as a synapse between an input layer and a next layer. The VMM array (1000) includes a memory array (1003) of non-volatile memory cells, a reference array (1001) of first non-volatile reference memory cells, and a reference array (1002) of second non-volatile reference memory cells. The reference arrays (1001 and 1002), arranged in the column direction of the array, serve to convert current inputs flowing into terminals (BLR0, BLR1, BLR2, and BLR3) into voltage inputs (WL0, WL1, WL2, and WL3). In practice, the first and second non-volatile reference memory cells are diode-connected to the current inputs flowing into them through a multiplexer (1014) (only partially illustrated). The reference cell is tuned (e.g., programmed) to the target reference level. The target reference level is provided by a reference mini-array matrix (not shown).
[0078] The memory array (1003) serves two purposes. First, the memory array stores weights to be used by the VMM array (1000) in their respective memory cells. Second, the memory array (1003) effectively multiplies the input (i.e., current inputs provided to terminals (BLR0, BLR1, BLR2, and BLR3), which are converted by reference arrays (1001 and 1002) into input voltages to be supplied to word lines (WL0, WL1, WL2, and WL3)) with the weights stored in the memory array (1003), and then adds all results (memory cell currents) to generate outputs in their respective bit lines (BL0–BLN), which will be inputs to the next layer or inputs to the final layer. By performing the multiplication and addition functions, the memory array (1003) eliminates the need for separate multiplication and addition logic circuits and is also power efficient. Here, voltage inputs are provided to word lines (WL0, WL1, WL2, and WL3), and outputs appear in their respective bit lines (BL0–BLN) during a read (inference) operation. The current placed in each bit line (BL0–BLN) performs a summation function of the currents from all non-volatile memory cells connected to that specific bit line.
[0079] Table 5 shows the operating voltages and currents for the VMM array (1000). The columns in the table represent the voltages placed on the word line for the selected cell, the word line for the unselected cell, the bit line for the selected cell, the bit line for the unselected cell, the source line for the selected cell, and the source line for the unselected cell. The rows represent the read operation, the erase operation, and the program operation.
[0080] [Table 5]
[0081]
[0082] FIG. 11 illustrates a neural VMM array (1100) that is particularly suitable for memory cells (210) as illustrated in FIG. 2 and is used as parts of neurons and synapses between an input layer and a next layer. The VMM array (1100) includes a memory array (1103) of non-volatile memory cells, a reference array (1101) of first non-volatile reference memory cells, and a reference array (1102) of second non-volatile reference memory cells. The reference arrays (1101 and 1102) are connected in the row direction of the VMM array (1100). The VMM array is similar to the VMM (1000) except that the word lines in the VMM array (1100) are connected in the vertical direction. Here, input is provided to word lines (WLA0, WLB0, WLA1, WLB2, WLA2, WLB2, WLA3, WLB3), and output appears in source lines (SL0, SL1) during a read operation. The current placed in each source line performs a summation function of all currents from the memory cells connected to that specific source line.
[0083] Table 6 shows the operating voltages and currents for the VMM array (1100). The columns in the table represent the voltages placed on the word line for the selected cell, the word line for the unselected cell, the bit line for the selected cell, the bit line for the unselected cell, the source line for the selected cell, and the source line for the unselected cell. The rows represent the read operation, the erase operation, and the program operation.
[0084] [Table 6]
[0085]
[0086] FIG. 12 illustrates a neural VMM array (1200) that is particularly suitable for memory cells (310) as illustrated in FIG. 3 and is used as parts of neurons and synapses between an input layer and a next layer. The VMM array (1200) includes a memory array (1203) composed of non-volatile memory cells, a reference array (1201) composed of a first non-volatile reference memory cell, and a reference array (1202) composed of a second non-volatile reference memory cell. The reference arrays (1201 and 1202) serve to convert current inputs flowing into terminals (BLR0, BLR1, BLR2, and BLR3) into voltage inputs (CG0, CG1, CG2, and CG3). In practice, the first and second non-volatile reference memory cells are diode-connected through a multiplexer (1212) (only partially illustrated) with current inputs flowing into them via BLR0, BLR1, BLR2, and BLR3. Each multiplexer (1212) includes its own multiplexer (1205) and cascoding transistor (1204) to ensure a constant voltage on the bit line (e.g., BLR0) of each of the first and second non-volatile reference memory cells during a read operation. The reference cells are tuned to a target reference level.
[0087] The memory array (1203) serves two purposes. First, it stores weights to be used by the VMM array (1200). Second, the memory array (1203) effectively multiplies the inputs (current inputs provided to the terminals (BLR0, BLR1, BLR2, and BLR3), for which reference arrays (1201 and 1202) convert these current inputs into input voltages to be supplied to the control gates (CG0, CG1, CG2, and CG3)) with the weights stored in the memory array, and then adds all results (cell currents) to produce an output, which appears on BL0 - BLN and will be an input to the next layer or an input to the final layer. By performing the multiplication and addition functions, the memory array eliminates the need for separate multiplication and addition logic circuits and is also power efficient. Here, inputs are provided on control gate lines (CG0, CG1, CG2, and CG3), and outputs appear on bit lines (BL0 - BLN) during a read operation. The current placed on each bit line performs a summation function of all currents from memory cells connected to that specific bit line.
[0088] The VMM array (1200) implements unidirectional tuning for non-volatile memory cells within the memory array (1203). That is, each non-volatile memory cell is erased and then partially programmed until a desired charge is reached on the floating gate. If too much charge is placed on the floating gate (so that an incorrect value is stored in the cell), the cell is erased, and the sequence of partial programming operations is restarted. As illustrated, two rows sharing the same erase gate (e.g., EG0 or EG1) are erased together (known as page erasure), and then each cell is partially programmed until a desired charge is reached on the floating gate.
[0089] Table 7 shows the operating voltages and currents for the VMM array (1200). The columns in the table represent the word lines for selected cells, word lines for unselected cells, bit lines for selected cells, bit lines for unselected cells, control gates for selected cells, control gates for unselected cells within the same sector as the selected cells, control gates for unselected cells within a different sector from the selected cells, erase gates for selected cells, erase gates for unselected cells, source lines for selected cells, and voltages placed on the source lines for unselected cells. The rows represent read operations, erase operations, and program operations.
[0090] [Table 7]
[0091]
[0092] FIG. 13 illustrates a neural VMM array (1300) that is particularly suitable for memory cells (310) as illustrated in FIG. 3 and is used as parts of neurons and synapses between an input layer and a next layer. The VMM array (1300) includes a memory array (1303) of non-volatile memory cells, a reference array (1301) of a first non-volatile reference memory cell, and a reference array (1302) of a second non-volatile reference memory cell. The EG lines (EGR0, EG0, EG1, and EGR1) are connected vertically, while the CG lines (CG0, CG1, CG2, and CG3) and SL lines (WL0, WL1, WL2, and WL3) are connected horizontally. The VMM array (1300) is similar to the VMM array (1400) except that the VMM array (1300) implements bidirectional tuning, where each individual cell can be completely erased, partially programmed, and partially erased as needed to reach a desired charge amount on the floating gate due to the use of separate EG lines. As illustrated, the reference arrays (1301 and 1302) convert the input current in the terminals (BLR0, BLR1, BLR2, and BLR3) (through the action of the diode-connected reference cells via the multiplexer (1314)) into control gate voltages (CG0, CG1, CG2, and CG3) to be applied to the memory cells in the row direction. The current output (neural) is in the bit lines (BL0–BLN), where each bit line sums all currents from the non-volatile memory cells connected to that specific bit line.
[0093] Table 8 shows the operating voltages and currents for the VMM array (1300). The columns in the table represent the word lines for selected cells, word lines for unselected cells, bit lines for selected cells, bit lines for unselected cells, control gates for selected cells, control gates for unselected cells within the same sector as the selected cells, control gates for unselected cells within a different sector from the selected cells, erase gates for selected cells, erase gates for unselected cells, source lines for selected cells, and voltages placed on the source lines for unselected cells. The rows represent read operations, erase operations, and program operations.
[0094] [Table 8]
[0095]
[0096] FIG. 22 illustrates a neural VMM array (2200) particularly suitable for memory cells (210) as illustrated in FIG. 2, which is used as parts of neurons and synapses between an input layer and a next layer. In the VMM array (2200), inputs (INPUT0..., INPUT N ) are each beat lines (BL0, ..., BL N It is received on ), and outputs (OUTPUT1, OUTPUT2, OUTPUT3, and OUTPUT4) are generated on source lines (SL0, SL1, SL2, SL3), respectively.
[0097] FIG. 23 illustrates a neural VMM array (2300) particularly suitable for the memory cell (210) illustrated in FIG. 2 and utilized as a portion of the neural and synapses between the input layer and the next layer. In this example, the input (INPUT 0, INPUT 1, INPUT2 and INPUT3) are received on the source lines (SL0, SL1, SL2, and SL3) respectively, and outputs (OUTPUT0, ..., OUTPUT N) is a bit line(BL0, ..., BL N It is generated on ).
[0098] FIG. 24 illustrates a neural VMM array (2400) that is particularly suitable for memory cells (210) as illustrated in FIG. 2 and is used as parts of neurons and synapses between an input layer and a next layer. In this example, inputs (INPUT0, ..., INPUT M ) are each word line(WL0, ..., WL M Received on ), and output(OUTPUT0, ..., OUTPUT N ) is a bit line(BL0, ..., BL N It is generated on ).
[0099] FIG. 25 illustrates a neural VMM array (2500) that is particularly suitable for memory cells (310) as illustrated in FIG. 3 and is used as parts of neurons and synapses between an input layer and a next layer. In this example, the input (INPUT 0, ..., INPUT M ) are each word line(WL0, ..., WL M Received on ), and output(OUTPUT0, ..., OUTPUT N ) is a bit line(BL0, ..., BL N It is generated on ).
[0100] FIG. 26 illustrates a neural VMM array (2600) particularly suitable for a memory cell (410) as illustrated in FIG. 4 and utilized as a portion of a neuron and a synapse between an input layer and a next layer. In this example, the input (INPUT 0, … , INPUT n ) are each vertical control gate lines (CG0, …, CG N It is received from ), and the outputs (OUTPUT1 and OUTPUT2) are generated from source lines (SL0 and SL1).
[0101] FIG. 27 illustrates a neural VMM array (2700) that is particularly suitable for memory cells (410) as illustrated in FIG. 4 and is used as parts of neurons and synapses between an input layer and a next layer. In this example, the input (INPUT 0, … , INPUT N ) are each beat lines (BL0, …, BL N It is received at the gates of bit line control gates (2701-1, 2701-2, …, 2701-(N-1), 2701-N) each connected to ). Exemplary outputs (OUTPUT1 and OUTPUT2) are generated from source lines (SL0 and SL1).
[0102] FIG. 28 illustrates a neural VMM array (2800) particularly suitable for memory cells (310) as shown in FIG. 3, memory cells (510) as shown in FIG. 5, and memory cells (710) as shown in FIG. 7, which are used as parts of neurons and synapses between an input layer and a next layer. In this example, inputs (INPUT0, …, INPUT M ) are each word line(WL0, …, WL M Received from ), and output(OUTPUT0, …, OUTPUT N ) are bit lines (BL0, …, BL) respectively. N It is generated in ).
[0103] FIG. 29 illustrates a neural VMM array (2900) particularly suitable for a memory cell (310) as shown in FIG. 3, a memory cell (510) as shown in FIG. 5, and a memory cell (710) as shown in FIG. 7, which is used as parts of neurons and synapses between an input layer and a next layer. In this example, the input (INPUT 0, … , INPUT M ) is the control gate line (CG0, …, CG M Received from ). Output(OUTPUT0, …, OUTPUT N) are each vertical source lines (SL0, …, SL N It is generated from ), where each source line (SL i ) is connected to the source line of all memory cells within column i.
[0104] FIG. 30 illustrates a neural VMM array (3000) that is particularly suitable for memory cells (310) as shown in FIG. 3, memory cells (510) as shown in FIG. 5, and memory cells (710) as shown in FIG. 7, and is utilized as a portion of a neuron and a synapse between an input layer and a next layer. In this example, the input (INPUT 0, … , INPUT M ) is the control gate line (CG0, …, CG M Received from ). Output(OUTPUT 0, … , OUTPUT N Each ) is a vertical bit line (BL0, …, BL N It is generated in ), where each bit line (BLi) is connected to the bit line of every memory cell in column i.
[0105] Short-term and long-term memory
[0106] Conventional technology includes a concept known as long short-term memory (LSTM). LSTM units are often used in neural networks. LSTMs allow neural networks to remember information over a predetermined, arbitrary time interval and use that information in subsequent operations. Conventional LSTM units include a cell, an input gate, an output gate, and a forget gate. The three gates control the flow of information into and out of the cell and the time interval during which information is recalled in the LSTM. VMMs are particularly useful in LSTM units.
[0107] FIG. 14 illustrates an exemplary LSTM (1400). The LSTM (1400) in this example includes cells (1401, 1402, 1403, and 1404). Cell (1401) receives an input vector (x0) and generates an output vector (h0) and a cell state vector (c0). Cell (1402) receives an input vector (x1), an output vector (hidden state) (h0) from cell (1401), and a cell state (c0) from cell (1401), and generates an output vector (h1) and a cell state vector (c1). Cell (1403) receives an input vector (x2), an output vector (hidden state) (h1) from cell (1402), and a cell state (c1) from cell (1402), and generates an output vector (h2) and a cell state vector (c2). Cell (1404) receives an input vector (x3), an output vector (hidden state) (h2) from cell (1403), and a cell state (c2) from cell (1403), and generates an output vector (h3). Additional cells may be used, and an LSTM with four cells is just an example.
[0108] FIG. 15 illustrates an exemplary embodiment of an LSTM cell (1500) that can be used for the cells (1401, 1402, 1403, and 1404) of FIG. 14. The LSTM cell (1500) receives an input vector (x(t)), a cell state vector (c(t-1)) from a preceding cell, and an output vector (h(t-1)) from a preceding cell, and generates a cell state vector (c(t)) and an output vector (h(t)).
[0109] The LSTM cell (1500) includes sigmoid function devices (1501, 1502, and 1503), each of which applies a number between 0 and 1 to control how many of each component within the input vector is allowed to pass to the output vector. The LSTM cell (1500) also includes tanh devices (1504 and 1505) for applying a hyperbolic tangent function to the input vector, multiplier devices (1506, 1507, and 1508) for multiplying two vectors together, and an adder device (1509) for adding two vectors together. The output vector (h(t)) may be provided to the next LSTM cell in the system, or it may be accessed for other purposes.
[0110] FIG. 16 illustrates an LSTM cell (1600) which is an example of an implementation of an LSTM cell (1500). For the convenience of the reader, the same numbering in the LSTM cell (1500) is used in the LSTM cell (1600). Each of the sigmoid function devices (1501, 1502, and 1503) and the tanh device (1504) includes a plurality of VMM arrays (1601) and activation function blocks (1602). Thus, it can be seen that VMM arrays are particularly useful for LSTM cells used in a given neural network system. The multiplier devices (1506, 1507, and 1508) and the adder device (1509) are implemented digitally or analogously. The activation function blocks (1602) can be implemented digitally or analogously.
[0111] An alternative to the LSTM cell (1600) (and another example of an implementation of the LSTM cell (1500)) is shown in FIG. 17. In FIG. 17, the sigmoid function devices (1501, 1502, and 1503) and the tanh device (1504) share the same physical hardware (VMM array (1701) and activation function block (1702)) in a time-multiplexed manner. The LSTM cell (1700) also includes a multiplier device (1703) for multiplying two vectors together, an adder device (1708) for adding two vectors together, a tanh device (1505) (including an activation function block (1702), a register (1707) for storing a value i(t) when i(t) is output from the sigmoid function block (1702), a register (1704) for storing a value f(t)*c(t-1) when this value is output from the multiplier device (1703) through the multiplexer (1710), a register (1705) for storing a value i(t)*u(t) when this value is output from the multiplier device (1703) through the multiplexer (1710), and a value o(t)*c~(t) when this value is output from the multiplier device (1703) through the multiplexer (1710). It includes a register (1706) and a multiplexer (1709) for storing when output from the device (1703).
[0112] The LSTM cell (1600) contains multiple sets of VMM arrays (1601) and their respective activation function blocks (1602), whereas the LSTM cell (1700) contains only one set of VMM arrays (1701) and activation function blocks (1702), which is used to represent multiple layers in the example of the LSTM cell (1700). The LSTM cell (1700) will require less space than the LSTM (1600) because the LSTM cell (1700) will require 1 / 4 of the space for the VMM and activation function blocks compared to the LSTM cell (1600).
[0113] It can be further understood that an LSTM unit typically includes multiple VMM arrays, each of which requires functions provided by specific circuit blocks outside the VMM array, such as summers, activation function blocks, and high voltage generation blocks. Providing separate circuit blocks for each VMM array would require a significant amount of space within the semiconductor device and would be somewhat inefficient. Therefore, the example described below reduces the circuitry required outside the VMM array itself.
[0114] Gated Recurrent Unit
[0115] An example of an analog VMM implementation can be utilized in a GRU (gated recurrent unit) system. A GRU is a gating mechanism in a recurrent neural network. A GRU is similar to an LSTM, except that GRU cells generally contain fewer components than LSTM cells.
[0116] FIG. 18 illustrates an exemplary GRU (1800). The GRU (1800) in this example includes cells (1801, 1802, 1803, and 1804). Cell (1801) receives an input vector (x0) and generates an output vector (h0). Cell (1802) receives an input vector (x1) and an output vector (h0) from cell (1801), and generates an output vector (h1). Cell (1803) receives an input vector (x2) and an output vector (hidden state) (h1) from cell (1802), and generates an output vector (h2). Cell (1804) receives an input vector (x3) and an output vector (hidden state) (h2) from cell (1803), and generates an output vector (h3). Additional cells may be used, and a GRU having four cells is just an example.
[0117] FIG. 19 illustrates an exemplary embodiment of a GRU cell (1900) that can be used in the cells (1801, 1802, 1803, and 1804) of FIG. 18. The GRU cell (1900) receives an input vector (x(t)) and an output vector (h(t-1)) from a preceding GRU cell and generates an output vector (h(t)). The GRU cell (1900) includes sigmoid function devices (1901 and 1902), each of which applies a number between 0 and 1 to a component from the output vector (h(t-1)) and the input vector (x(t)). The GRU cell (1900) also includes a tanh device (1903) for applying a hyperbolic tangent function to an input vector, a plurality of multiplier devices (1904, 1905, and 1906) for multiplying two vectors together, an adder device (1907) for adding two vectors together, and a complementary device (1908) for subtracting the input from 1 to produce an output.
[0118] FIG. 20 illustrates a GRU cell (2000), which is an example of an implementation of a GRU cell (1900). For the convenience of the reader, the same reference numerals used in the GRU cell (1900) are used for the GRU cell (2000). As can be seen in FIG. 20, the sigmoid function devices (1901 and 1902) and the tanh device (1903) each include a plurality of VMM arrays (2001) and activation function blocks (2002). Thus, it can be seen that VMM arrays are particularly useful in GRU cells used in a given neural network system. The multiplier devices (1904, 1905, 1906), the adder device (1907), and the complementary device (1908) are implemented digitally or analogously. The activation function blocks (2002) can be implemented digitally or analogously.
[0119] An alternative to the GRU cell (2000) (and another example of an implementation of the GRU cell (1900)) is illustrated in FIG. 21. In FIG. 21, the GRU cell (2100) utilizes a VMM array (2101) and an activation function block (2102), and the activation function block, when configured as a sigmoid function, applies a number between 0 and 1 to control how many of each component within the input vector is allowed to pass to the output vector. In FIG. 21, the sigmoid function devices (1901 and 1902) and the tanh device (1903) share the same physical hardware (VMM array (2101) and activation function block (2102)) in a time-multiplexed manner. The GRU cell (2100) also includes a multiplier device (2103) for multiplying two vectors together, an adder device (2105) for adding two vectors together, a complement device (2109) for subtracting an input from 1 to generate an output, a multiplexer (2104), a register (2106) for holding a value h(t-1)*r(t) when this value is output from the multiplier device (2103) through the multiplexer (2104), a register (2107) for holding a value h(t-1)*z(t) when this value is output from the multiplier device (2103) through the multiplexer (2104), and a register (2108) for holding a value h^(t)*(1-z(t)) when this value is output from the multiplier device (2103) through the multiplexer (2104).
[0120] The GRU cell (2000) contains multiple sets of VMM arrays (2001) and activation function blocks (2002), whereas the GRU cell (2100) contains only one set of VMM arrays (2101) and activation function blocks (2102), which is used to represent multiple layers in the example of the GRU cell (2100). The GRU cell (2100) will require less space than the GRU cell (2000) because the GRU cell (2100) will require 1 / 3 of the space for the VMM and activation function blocks compared to the GRU cell (2000).
[0121] It can be further understood that a GRU system typically includes multiple VMM arrays, each of which requires functions provided by specific circuit blocks outside the VMM array, such as summers, activation function blocks, and high voltage generation blocks. Providing separate circuit blocks for each VMM array would require a significant amount of space within the semiconductor device and would be somewhat inefficient. Therefore, the example described below reduces the circuitry required outside the VMM array itself.
[0122] The input to the VMM array can be an analog level, a binary level, a pulse, a time-modulated pulse, or a digital bit (in this case, a DAC is required to convert the digital bit to an appropriate input analog level), and the output can be an analog level, a binary level, a timing pulse, a pulse, or a digital bit (in this case, an output ADC is required to convert the output analog level to a digital bit).
[0123] Generally, for each memory cell within a VMM array, each weight (W) can be implemented by a single memory cell, by a differential cell, or by two blend memory cells (average of two cells). In the case of a differential cell, two memory cells are required to implement the weight (W) as a differential weight (W = W+ - W-). In the case of two blend memory cells, two memory cells are required to implement the weight (W) as the average of two cells.
[0124] FIG. 31 illustrates a VMM system (3100). In some examples, the weights (W) stored in the VMM array are stored as differential pairs, namely W+ (positive weight) and W- (negative weight), where W = (W+) - (W-). In the VMM system (3100), half of the bit lines are designated as W+ lines, i.e., bit lines connected to memory cells to store positive weights (W+), and the other half of the bit lines are designated as W- lines, i.e., bit lines connected to memory cells to implement negative weights (W-). W- lines are alternately interspersed between the W+ lines. Subtraction operations are performed by summing circuits, such as summing circuits (3101 and 3102), which receive current from the W+ lines and W- lines. The outputs of the W+ lines and the W- lines are combined to effectively provide W = W+ - W- for each pair of (W+, W-) cells for every pair of (W+, W-) lines. Although the above was described in relation to W- lines alternately interspersed between W+ lines, in other examples, W+ lines and W- lines can be arbitrarily located anywhere within the array.
[0125] FIG. 32 illustrates another example. In a VMM system (3210), positive weights (W+) are implemented in a first array (3211) and negative weights (W-) are implemented in a second array (3212), the second array (3212) is separate from the first array, and the resulting weights are appropriately combined together by a summing circuit (3213).
[0126] FIG. 33 illustrates a VMM system (3300), wherein weights (W) stored in a VMM array are stored as differential pairs W+ (positive weights) and W- (negative weights), where W = (W+) - (W-). The VMM system (3300) includes an array (3301) and an array (3302). Half of the bit lines within each array (3301 and 3302) are designated as W+ lines, i.e., bit lines connected to memory cells to store positive weights (W+), and the other half of the bit lines within each array (3301 and 3302) are designated as W- lines, i.e., bit lines connected to memory cells to implement negative weights (W-). W- lines are alternately interspersed between the W+ lines. The subtraction operation is performed by a summing circuit that receives current from the W+ line and the W- line, such as summing circuits (3303, 3304, 3305, and 3306). The outputs of the W+ line and the W- line from each array (3301, 3302) are combined together to effectively provide W = W+ - W- for each pair of (W+, W-) cells for every pair of (W+, W-) lines. Additionally, the W values from each array (3301 and 3302) can be further combined through summing circuits (3307 and 3308), so that each W value is the result of W value from array (3301) - W value from array (3302), which means that the final result from the summing circuits (3307 and 3308) is the differential value of two differential values.
[0127] Each non-volatile memory cell used in an analog neural memory system must be erased and programmed to hold a very specific and precise amount of charge, i.e., the number of electrons, in a floating gate. For example, each floating gate must hold one of N different values, where N is the number of different weights that can be represented by each cell. Examples of N include 16, 32, 64, 128, and 256.
[0128] It is important to be able to accurately verify the programming operation after it has been performed.
[0129] A number of examples of verification circuits and associated methods in artificial neural networks are disclosed. Brief explanation of the drawing
[0130] Figure 1 is a diagram illustrating an artificial neural network. FIG. 2 illustrates a conventional separated gate flash memory cell. FIG. 3 illustrates a separate gate flash memory cell of another prior art. FIG. 4 illustrates a separate gate flash memory cell of another prior art. FIG. 5 illustrates a separate gate flash memory cell of another prior art. Figure 6 is a diagram illustrating exemplary artificial neural networks of different levels utilizing one or more non-volatile memory arrays. Figure 7 is a block diagram illustrating a VMM system. Figure 8 is a block diagram illustrating an exemplary artificial neural network utilizing one or more VMM systems. Figure 9 illustrates another example of a VMM system. Figure 10 illustrates another example of a VMM system. Figure 11 illustrates another example of a VMM system. Figure 12 illustrates another example of a VMM system. Figure 13 illustrates another example of a VMM system. FIG. 14 illustrates a conventional short-term and long-term memory system. FIG. 15 illustrates an exemplary cell for use in a short-term and long-term memory system. FIG. 16 illustrates an exemplary implementation of the cell of FIG. 15. FIG. 17 illustrates another exemplary embodiment of the cell of FIG. 15. FIG. 18 illustrates a gated circulation unit system of the prior art. FIG. 19 illustrates an exemplary cell for use in a gated circulation unit system. FIG. 20 illustrates an exemplary implementation of the cell of FIG. 19. FIG. 21 illustrates another exemplary embodiment of the cell of FIG. 19. Figure 22 illustrates another example of a VMM system. Figure 23 illustrates another example of a VMM system. Figure 24 illustrates another example of a VMM system. Figure 25 illustrates another example of a VMM system. Figure 26 illustrates another example of a VMM system. Figure 27 illustrates another example of a VMM system. Figure 28 illustrates another example of a VMM system. Figure 29 illustrates another example of a VMM system. Figure 30 illustrates another example of a VMM system. Figure 31 illustrates another example of a VMM system. Figure 32 illustrates another example of a VMM system. Figure 33 illustrates another example of a VMM system. Figure 34 illustrates another example of a VMM system. FIGS. 35a and FIGS. 35b illustrate the respective programming methods. Figure 36 illustrates a search and execution method. Fig. 37 illustrates a precision programming method. Fig. 38 illustrates a precision programming method. Figure 39 illustrates an adaptive correction method. Figure 40 illustrates a calibration circuit. Figure 41 illustrates an adaptive correction method. Fig. 42 illustrates an absolute correction method. FIG. 43 illustrates a VMM system including a verification circuit. FIG. 44a illustrates an exemplary verification circuit. FIG. 44b illustrates an exemplary comparator circuit with offset compensation. Figure 45 illustrates a reference voltage generator. FIG. 46 illustrates a physical array including reference arrays. FIG. 47 illustrates a physical array including reference arrays. FIG. 48 illustrates a physical array including a VMM array and another physical array including a reference array. FIG. 49 illustrates a reference array including a plurality of reference sub-arrays. FIG. 50 illustrates a reference array including a plurality of reference sub-arrays. Specific details for implementing the invention
[0131] VMM System Architecture
[0132] FIG. 34 illustrates a block diagram of a VMM system (3400). The VMM system (3400) includes a VMM array (3401), a row decoder (3402), a high voltage decoder (3403), a column decoder (3404), a bit line driver (3405), an input circuit (3406), an output circuit (3407), control logic (3408), and a bias generator (3409). The VMM system (3400) further includes a high voltage generation block (3410) comprising a charge pump (3411), a charge pump controller (3412), and a high voltage level generator (3413). The VMM system (3400) further comprises an algorithm controller (3414) (program / erase, or weight tuning), an analog circuit section (3415), a control engine (3416) (which may include, without limitation, special functions such as arithmetic functions, activation functions, and embedded microcontroller logic), test control logic (3417), and a static random access memory (SRAM) block (3418) for storing intermediate data or data for programming (e.g., data for an entire row or multiple rows) for input circuits (e.g., activation data) or output circuits (neural output data).
[0133] The input circuit (3406) may include circuits such as a DAC (digital-to-analog converter), a DPC (digital-to-pulse converter, digital-time modulated pulse converter), an AAC (analog-to-analog converter, such as a current-to-voltage converter, log converter), a PAC (pulse-to-analog level converter), or any other type of converter. The input circuit (3406) may implement one or more of a normalization, linear or non-linear up / down scaling function, or arithmetic function. The input circuit (3406) may implement a temperature compensation function for the input level. The input circuit (3406) may implement an activation function such as ReLU or sigmoid. The input circuit (3406) may store digital activation data to be applied as an input signal or combined with an input signal during a program or read operation. The digital activation data may be stored in registers. The input circuit (3406) may include circuits for driving array terminals, such as CG, WL, EG, and SL lines, and may include sample-and-hold circuits and buffers. A DAC may be used to convert digital enable data into an analog input voltage to be applied to the array.
[0134] The output circuit (3407) may include circuits such as an ITV (current-voltage circuit), an ADC (analog-to-digital converter for converting neural analog outputs into digital bits), an AAC (analog-to-analog converter such as a current-voltage converter or a log converter), an APC (analog-to-pulse(s) converter, an analog-to-time modulated pulse converter), or any other type of converter. The output circuit (3407) may convert array outputs into activation data. The output circuit (3407) may implement an activation function such as a rectified linear activation function (ReLU) or a sigmoid. The output circuit (3407) may implement one or more of statistical normalization, regularization, up / down scaling / gain functions, statistical rounding, or arithmetic functions (e.g., addition, subtraction, division, multiplication, shift, log) for neural outputs. The output circuit (3407) may implement a temperature compensation function for a neural output section or an array output section (e.g., a bit line output section) to, for example, keep the power consumption of the array approximately constant with respect to temperature, or to, for example, improve the precision of the array (neural) output by keeping the IV slope approximately equal with respect to temperature. The output circuit (3407) may include registers for storing output data.
[0135] FIG. 35a illustrates a programming method (3500). First, a method that occurs in response to a typically received program command is initiated (step 3501). Next, a large program operation programs all cells to a '0' state (step 3502). Then, a weak erase operation erases all cells to a weakly erased level so that each cell will draw, for example, approximately 1 to 5 μA of current during a read operation (step 3503). This is in contrast to a deep erase level where each cell will draw, for example, approximately 20 to 30 μA of current during a read operation. Then, a hard program operation is performed on all unselected cells to a very deep programmed state to add electrons to the floating gates of the cells (step 3504), ensuring that those cells are actually "off," which means that those cells will draw a negligible amount of current during a read operation.
[0136] Then, a coarse programming operation is performed to program the selected cells to a level much closer to the target, for example, a 2X-100X target. Then, a coarse programming operation is performed on the selected cells (step 3505), followed by a precise programming operation on the selected cells (step 3506) to program the required precise value for each selected cell.
[0137] The non-precision programming operation (3505) may consist of a number of non-precision verification / program cycles. In each non-precision verification / program cycle, a verification operation is performed to verify whether the cell output satisfies the non-precision target; if not, a programming operation is performed again for the cell. The verification / program cycles are repeated until the cell outputs of all target cells satisfy the non-precision target.
[0138] The precision programming operation (3506) may consist of a number of precision verification / program verification / program cycles. In each precision verification / program cycle, a verification operation is performed to verify whether the cell output satisfies the precision target. If not, a programming operation is performed again for the cell. The verification / program cycles are repeated until the cell outputs of all target cells satisfy the precision target.
[0139] FIG. 35b illustrates a different programming method (3510) similar to the programming method (3500). However, instead of a programming operation to program all cells to a '0' state as in step 3502 of FIG. 35a, after the method is started (step 3501), an erasure operation is used to erase all cells to a '1' state (step 3512). Then, a soft programming operation (step 3513) is used to program all cells to a weakly programmed state (level) such that each cell draws, for example, approximately 3-5 μA of current during a read operation. After that, a hard programming operation is performed on all unselected cells to a very deeply programmed state (step 3504), and then non-precision and precision programming (3505-3506) are performed as in FIG. 35a. A variation of the example in FIG. 35b will eliminate the weak programming method (step 3513).
[0140] FIG. 36 illustrates a first example of a non-precision programming operation (3505), which is a search and execution method (3600). First, a lookup table search is performed to determine a non-precision target current value (ICT) for a selected cell based on a value intended to be stored in such a selected cell (step 3601). Such a table is generated, for example, by silicon characterization or from calibration from wafer testing. It is assumed that the selected cell can be programmed to store one of N possible values (e.g., 128, 64, 32 without limitation). Each of the N values will correspond to a different desired current value (ID) drawn by the selected cell during a read operation. In one example, the lookup table may contain M possible current values to be used as the non-precision target current value (ICT) for the selected cell during the search and execution method (3600), where M is an integer less than N. For example, if N is 8, M can be 4, which means there are 8 possible values that the selected cell can store, and one of the 4 non-precision target current values will be selected as a non-precision target for the search and execution method (3600). That is, the search and execution method (3600) (which is an example of the non-precision programming method (3505) as shown above) is intended to quickly program the selected cell to a value (ICT) that is somewhat close to the desired value (ID), and then the precision programming method (3506) is intended to program the selected cell more precisely to be close to the desired value (ID).
[0141] Examples of cell values, desired current values, and non-precision target current values are shown in Tables 9 and 10 for simple examples of N=8 and M=4:
[0142] [Table 9]
[0143]
[0144] [Table 10]
[0145]
[0146] Offset value I CTOFFSETx It is used to prevent the desired current value from overshooting during non-precision adjustment.
[0147] First, the non-precision target current value I CT When selected, the selected cell is programmed by applying a voltage v0 to the appropriate terminal of the selected cell based on the cell architecture type of the selected cell (e.g., memory cell (210, 310, 410, or 510)) (step 3602). If the selected cell is a memory cell (310) of the type of FIG. 3, the voltage V0 will be applied to the control gate terminal (28), and V0 is a non-precision target current value I CT It can be 5 to 7 V depending on. The value of v0 is optionally v0 to non-precision target current value (I CT It can be determined from a voltage lookup table that stores values for ).
[0148] Next, the selected cell is voltage v i = v i-1 +v increment It is programmed by applying, where i starts at 1 and is incremented each time this step is repeated, where v increment is a small, micro voltage that causes an appropriate degree of programming for the granularity of the desired change (step 3603). Accordingly, when step 3603 is first performed, i=1, and v1 is v0 + v increment It will be. Then, a verification operation occurs (step 3604), a read operation is performed on the selected cell, and the current (I) flowing through the selected cell cell ) is a non-precision target threshold (I CT is compared with ). I cell This I CT(Here, the first threshold) If it is less than or equal to, the search and execution method (3600) is completed, and the precision programming method (3506) can be started. cell This I CT If it is not less than or equal to, i is incremented and step 3603 is repeated.
[0149] Accordingly, at the point when the non-precision programming method (3505) ends and the precision programming method (3506) begins, voltage v i will be the final voltage used to program the selected cell, and the selected cell is a non-precision target current value (I CT )(Here I cell >= I CT The value associated with the cell will be stored. The precision programming method (3506) programs the selected cell during the read operation to a point where it retrieves a current ID (plus or minus acceptable deviation amount, e.g., + / -30% or less, e.g., + / -50 pA), and the current ID is a desired current value associated with the value intended to be stored in the selected cell.
[0150] FIG. 37 illustrates examples of different voltage progressions that may be applied to the control gate of a selected memory cell during a non-precision programming operation (3505) and / or a precision programming operation (3506). This consists of a number of verification / program cycles.
[0151] Under the first approach, an increasing voltage is applied to the control gate during the process to program the selected memory cell. The starting point is vi, which is the last voltage applied during the non-precision programming method (3505) during the precision programming operation (3506). v p1 The increment of is added to v1, and then, the voltage (v1 + v p1 ) is used to program the selected cell (indicated by the second pulse from the left in progress (3701)). v p1 eu v incrementIt is an increment smaller than (the voltage increment used during the non-precision programming method (3505)). After each programming voltage is applied, a verification operation (similar to step 3404) is performed, and Icell is I PT1 (This is the first precision target current value, and here it is the second threshold value) A determination is made whether it is less than or equal to, and I PT1 = I D + I PT1OFFSET, and, I PT1OFFSET is an offset value added to prevent program overshooting. If not, another increment v p1 This is added to the previously applied programming voltage, and the process is repeated. I cell This I PT1 At the point below, then, this part of the programming sequence is stopped. Optionally, I PT1 This I D I is equal to, or has sufficient precision, i.e., a plus or minus acceptable deviation amount. D If it is nearly identical to, the selected memory cell has been successfully programmed.
[0152] I PT1 This I D If it is not sufficiently close to, i.e., I D If it is not nearly identical with sufficient precision, additional programming of a smaller particle size occurs. Here, the progression (3702) is now used. The starting point for the progression (3702) is the last voltage used for programming under the progression (3701). (v p1 (smaller than) V p2 An increment of is added to the corresponding voltage, and the combined voltage is applied to program the selected memory cell. After each programming voltage is applied, a verification operation (similar to step 3404) is performed, and I cell This I PT2 A determination is made whether it is less than or equal to (the second precision target current value, which is the third threshold value here), and I PT2 = ID + IPT2OFFSET (Here I PT2OFFSET is an offset value added to prevent program overshooting). If not, another increment V p2 is added to the previously applied programming voltage, and the process is repeated. I cell This I PT2 , at the point where the amount of acceptable deviation (plus or minus) is less than or equal to, then, this part of the programming sequence stops. Here, because the target value has been achieved with sufficient precision, I PT2 is I D It is assumed that programming can be stopped when it is equal to or sufficiently close to the ID. A person skilled in the art may recognize that additional progress can be applied by using increasingly smaller programming increments. For example, in FIG. 38, three progresses (3801, 3802, and 3803) are applied instead of just two.
[0153] A second approach is illustrated in the 3703 progression of FIG. 37 and the 3803 progression of FIG. 38. Instead of increasing the voltage applied during programming of the selected memory cell, the same voltage is applied during the duration of the increasing period. That is, each applied pulse is t compared to the previously applied pulse. p1 So that it is longer by that amount, time (t p1 An additional increment of time is added to the programming pulse. After each programming pulse is applied, the same verification operation as previously described for the progression (3701) is performed. Optionally, additional progressions may be applied, and the additional increment of time added to the programming pulse has a shorter duration than the previous progression used. Although only one time progression is shown, those skilled in the art will recognize that any number of different time progressions may be applied.
[0154] Now, additional details regarding three examples of non-precision programming methods (3505) will be provided.
[0155] FIG. 39 illustrates another example of a non-precision programming method (3505) which is an adaptive calibration method (3900). The adaptive calibration method is initiated (step 3901). The cell is programmed with a default starting value v0 (step 3902). Unlike in the search and run method (3600), v0 here is not derived from a lookup table, but instead can be a relatively small initial value. The control gate voltage of the cell is measured at a first current value IR1 (e.g., 100 na) and a second current value IR2 (e.g., 10 na), and the sub-threshold slope is determined and stored based on these measurements (e.g., 360 mV / dec) (step 3903).
[0156] New program voltage v i is determined. When this step is first performed, i=1, and v1 is determined using the following subcritical equation based on the stored subcritical slope value, current target, and offset value:
[0157] Vi = V i-1 + V increment ,
[0158] Here, V increment is proportional to the slope of Vg.
[0159] Vg = n * Vt * log [Ids / wa * Io]
[0160] Here, wa is the w of the memory cell, and Ids is the current target plus offset value.
[0161] If the stored slope value is relatively steep, a relatively small current offset value can be used. If the stored slope value is relatively flat, a relatively high current offset value can be used. Accordingly, determining the slope information allows a customized current offset value to be selected for the specific cell in question. This will ultimately make the programming process shorter. As these steps are repeated, i is incremented, and v i = vi-1 + v increment is. Then, the cell is v i It is programmed using v increment ne, v increment It can be determined from a lookup table that stores the value versus the target current value.
[0162] Next, a verification operation occurs, a read operation is performed on the selected cell, and the current (I) flowing through the selected cell cell ) is a non-precision target threshold (I CT It is compared with ) (Step 3905). I CT is I D + I CTOFFSET, It is set to, and I CTOFFSET is an offset value added to prevent program overshooting, and I cell If this ICT is less than or equal to this, the adaptive correction method (3900) is completed, and the precision programming operation (3506) can be started. cell This I CT If not less than or equal to, steps 3904 through 3905 are repeated, and i is incremented. Then, the precision programming method (3506) starts with voltage vi, which is the last voltage used in the program of the selected cell.
[0163] FIG. 40 illustrates an aspect of an adaptive calibration operation (3900). During step 3903, a current source (4001) applies current values (IR1 and IR2) to a selected cell (here, memory cell (4002)), and then the voltage at the control gate of the memory cell (4002) (CGR1 for IR1 and CGR2 for IR2) is measured. The slope is determined as (CGR2-CGR1) / dec of the current, which is the slope (I) of VCG versus LOG.
[0164] FIG. 41 illustrates another example of a non-precision programming operation (3505) which is an adaptive calibration method (4100). The adaptive calibration method is initiated (step 4101). The cell is programmed with a default starting value v0 (step 4102). v0 is derived from a lookup table such as one generated from silicon characterization, and the table value is an offset such as one that does not overshoot the target program value.
[0165] In the next step 4103, the IV slope parameter is generated and used to predict the next programming voltage, and the first control gate read voltage, V CGR1 This is applied to the selected cell, and the generated cell current, IR1, is measured. Then, the second control gate read voltage, V CGR2 It is applied to the selected cell, and the generated cell current, IR2, is measured. The slope is determined based on these measurements and stored, for example, according to the equation of the sub-critical region (cell operating at a sub-critical):
[0166] Slope = (V CGR1 - V CGR2 ) / (LOG(IR1) - LOG(IR2))
[0167] (Step 4103). V CGR1 and V CGR2 Examples of values for are, for example, 1.5 V and 1.3 V.
[0168] Therefore, determining the gradient information is V customized for that specific cell increment It allows values to be selected. This will ultimately make the programming process shorter.
[0169] When step 4104 is repeated, i is incremented, and the new desired programming voltage, Vi, is determined based on the stored slope value and current target and offset values using the following mathematical formula:
[0170] V i = V i-1+ V increment ,
[0171] Here, for i-1, V increment = alpha * gradient * (LOG (IR1) - LOG (I CT )),
[0172] Here I CT ε is the target current, and alpha is a predetermined constant less than 1 (programming offset value) to prevent overshooting, e.g., 0.9. For example, V i is the VSLP or VCGP, source line or control gate programming voltage.
[0173] Next, the cell is programmed using Vi (step 4104).
[0174] Next, a verification operation occurs, where a read operation is performed on the selected cell, and the current (I) drawn through the selected cell cell ) is I CT It is compared with (step 4106). I cell This I CT If it is less than or equal to (this is the non-precision target threshold here) - where, I CT is I D + I CTOFFSET It is set to, and I CTOFFSET is an offset value added to prevent program overshooting -, the process proceeds to step 4107. Otherwise, the process returns to step 4104, and i is incremented.
[0175] In step 4107, I cell I CT Smaller threshold, CT2 It is compared with ,. The purpose of this is to determine whether overshooting has occurred. In other words, the goal is I cell This I CT It is less than that, but it is excessively I CT If it is less than, overshooting has occurred, and the stored value may actually correspond to an incorrect value. cell This ICT2 If it is not less than or equal to the above, no overshooting occurred, the adaptive correction method (4100) was completed, and at that point, the process proceeds to the precision programming operation (3506). I cell This I CT2 If it is less than or equal to, overshooting has occurred. Subsequently, the selected cell is erased (step 4108), the programming process restarts from step 4102, and i is reset to 0. Optionally, if step 4108 is performed more than a predetermined number of times, the selected cell may be considered a defective cell that should not be used.
[0176] The precision program operation (3506) consists of a number of verification and program (V / P) cycles, wherein the program voltage has a fixed pulse width and is incremented by a constant minute voltage, or wherein the program voltage is fixed and the program pulse width is varied.
[0177] Optionally, the step of determining whether the current through a selected non-volatile memory cell during a read or verification operation is less than or equal to a non-precision target threshold can be performed by applying a fixed bias to the terminals of the non-volatile memory cell, measuring and digitizing the current drawn by the selected non-volatile memory cell to generate digital output bits, and comparing the digital output bits with digital bits representing a first threshold current.
[0178] Optionally, the step of determining whether the current through a selected non-volatile memory cell during a read or verification operation is less than or equal to a non-precision target threshold can be performed by applying a fixed bias to the terminals of the non-volatile memory cell, measuring and digitizing the current drawn by the selected non-volatile memory cell to generate digital output bits, and comparing the digital output bits with digital bits representing a first threshold current.
[0179] Optionally, the step of determining whether the current through a selected non-volatile memory cell during a read or verification operation is less than or equal to a non-precision target threshold can be performed by applying an input to a terminal of the non-volatile memory cell, modulating the current drawn by the selected non-volatile memory cell into an output pulse to generate a modulated output, digitizing the modulated output to generate digital output bits, and comparing the digital output bits with digital bits representing a first threshold current.
[0180] FIG. 42 illustrates a third example of a non-precision programming operation (3505) which is an absolute calibration method (4200). The absolute calibration method is initiated (step 4201). The cell is programmed at a default starting value (v0) (step 4202). The control gate voltage (VCGRx) of the cell is the current value I target It is measured and stored (step 4203). The stored control gate voltage, and current target and offset values, I offset +I target A new desired voltage v1 is determined based on this (step 4204). For example, the new desired voltage v1 can be calculated as follows: v1 = v0 + (VCGBIAS - stored VCGR), where VCGBIAS is the default read control gate voltage at the maximum target current, for example, = ~1.5V, and stored VCGR is the measured read control gate voltage from step 4203.
[0181] Then, the cell is programmed using vi. When i=1, the voltage (v1) from step 4204 is used. When i>1, voltage v i = v i-1 + v increment is used. v increment ne, v incrementIt can be determined from a lookup table storing the value versus the target current value. Next, a verification operation is performed, where a read operation is performed on the selected cell, and the current (I) drawn through the selected cell cell ) is I CT It is compared with (step 4206). I cell This I CT (Here, the threshold is) If it is less than or equal to, the absolute calibration method (4200) is completed and the precision programming method (3506) can be started. cell This I CT If not less than or equal to, steps 4205 through 4206 are repeated, and i is incremented.
[0182] A single weight verification method may be used to verify whether the cell has achieved a weight target as a result of a programming operation. A memory cell is selected for the verification operation, and then the output of the memory cell is verified by a verification mechanism described below with reference to FIGS. 43 through 49.
[0183] As a result of a programming operation, a differential weight verification method may be used to determine whether a differential cell (formed by two cells, where the stored value is the difference between the values stored in the two cells) has achieved a weight target. Two cells associated with the differential weight are selected for the verification operation. Then, the difference in output between the cells is verified by a verification mechanism described below with reference to FIGS. 43 through 49. For example, if cell1 stores a w+ value and cell2 stores a w- value, w = w+ - w- is the differential weight. Cell1 and cell2 are the two cells selected for the verification operation to determine whether the weight w has achieved a target. Alternatively, the w- value of cell2 is verified first, and then the differential weight w is verified. Alternatively, the w+ value of cell1 is verified first, and then the differential weight w is verified. Alternatively, in the first operation, the w+ and w- values of cell1 and cell2 are verified, and then the differential weight w is verified in the second operation.
[0184] FIG. 43 illustrates a VMM system (4300). Current-voltage converter and analog-to-digital converter blocks (4301) receive current from the VMM array (3401), typically from bit lines or source lines in the VMM array (3401), and provide outputs to verification circuits (4302). The current-voltage converter and analog-to-digital converter blocks (4301) and the verification circuits (4302) together are output blocks coupled to the VMM array (3401) to generate voltages during the verification operation of the VMM array (3401) and to generate digital outputs during the read operation of the VMM array (3401). Each current-voltage converter in block (4301) converts current into voltage. The analog-to-digital converters in block (4301) are reconfigured (e.g., in the manner described below with reference to FIG. 44a) to be used during the verification operation (also known as the read-verification operation). During the verification operation, the reference array (4304) is used to generate all N possible current targets (e.g., 32 values in the range of 3-96nA in 3nA increments). Each of the N possible current targets corresponds to one of the N weight targets stored in each reference memory cell in the reference array (4304). Alternatively, the main reference current generator (4305) is used to generate all N current targets. Reference current(s) from either the reference array (4304) or the main reference current generator (4305) are fed to the reference voltage generator (4303), which includes a voltage DAC, which converts the reference currents into reference voltages using the voltage DAC, where there are N possible voltages corresponding to the N possible values. For example, for a 5-bit cell, there are 32 reference voltages corresponding to the 32 current targets.Among N voltage values, an appropriate voltage value is selected for comparison by a verification circuit (4302) by comparing it with the voltage provided by the corresponding current-voltage converter (4301). This comparison is a verification operation performed by the verification circuit (4302). Accordingly, weights can be verified after being programmed into the VMM array (3401). In this approach, verification is performed by comparing the output voltage from the memory cell with the reference voltage from the voltage reference generator (4303).
[0185] Alternatively, a reference current digital-to-analog converter (IDAC) is used to directly verify the cell current without using a current-to-voltage converter, which means the cell current is compared to a reference current. Under this approach, latency and variation will typically be greater due to the settling time of low-current (e.g., several nA) circuits.
[0186] FIG. 44a illustrates a neural output ITV+ADC+rjawmd circuit (4488) comprising a current-voltage converter (ITV) (4401), a successive approximation register (SAR) analog-to-digital converter (ADC) (4402), and a verification circuit (4403). The verification circuit (4403) includes a comparator (4404) (which is also used to generate digital outputs during a read or read neural operation), a reference voltage selection circuit (4406), and verification registers (4405). The verification registers (4405) are the data-out registers of the SAR ADC (4402), which are used here to perform a verification function but can also be used during a read or read neural operation to generate digital outputs. The current-voltage converter (4401) and the SAR analog-to-digital converter (4402) are examples of implementations of the current-voltage converter and the analog-to-digital converter (4301) in FIG. 43, and the verification circuit (4403) is an example of implementation of the verification circuit (4302) in FIG. 43.
[0187] The current-voltage converter (4401) and the SAR analog-to-digital converter (4301) may be used during a read or neural read operation. However, the current-voltage converter (4401) and the SAR analog-to-digital converter (4301) may also be used during a verification operation in which a weight programmed into a non-volatile memory cell in a VMM array (intended to be one of N possible weight values) is verified.
[0188] A current-to-voltage converter (4401) receives current from a VMM array from a single selected cell and converts this current into voltage. Current-to-voltage conversion can be performed by a plurality of resistors ITV (RITV) (4490R) or a plurality of capacitors ITV (CITV) (4490C). One of N possible reference voltages is provided to the verification circuit (4403) through verification registers (4405) and a verification reference voltage selection circuit (4406). The verification registers may be, for example, 8-bit registers used to select one of 256 voltage reference levels in the verification reference voltage selection circuit (4406). As inputs (via verification reference voltage lines) to the verification reference voltage selection circuit (4406), the verification reference voltages are provided by a global verification reference voltage generator such as in FIG. 45. Then, the comparator (4404) compares the voltage from the current-voltage converter (4401) with one of N possible reference voltages to indicate whether the cell is storing an accurate value. During the verification operation, the SAR logic and capacitors in the SAR ADC (4402) are not used. In one example, the control circuit closes the switches (S1A) in the SAR ADC (4402) to provide the positive output of the ITV (4401, Vinp) directly to the non-inverting input of the comparator (4404), opens the switch (NOT NAMED) to provide the output of the verification reference voltage selection circuit (4406) to the inverting input of the comparator (4404), and closes the switch (NOT NAMED).
[0189] The ITV+ADC+verification circuit (4488) can be used for both single-weight verification operations and differential-weight verification operations. For a single-weight verification operation, only one input from a single cell is required. The output voltage of the ITV (4401) is proportional to the value of the cell current and is verified by comparing it to a reference voltage level provided by the verification reference voltage selection circuit (4406). For a differential-weight verification operation, two inputs from two cells are two inputs to the ITV+ADC+verification circuit (4488), and the output voltage of the ITV (4401) (e.g., Vinp) is proportional to the difference between the two cell currents and is verified by comparing it to a reference voltage level provided by the verification reference voltage selection circuit (4406).
[0190] The total offset compensation of the ITV+ADC+verification circuit (4488) can be trimmed using the offset trimming of the comparator (4404). This offset can be further trimmed by trimming the resistor (4490R) or capacitor (4490C) of the 4401 ITV circuit.
[0191] In another example, the total gain compensation of the ITV+ADC+verification circuit (4488) can be trimmed by trimming the resistor (4490R) or capacitor (4490C) of the 4401 ITV circuit.
[0192] An alternative method of offset compensation can be performed in the time domain by using an ITV (4401) having a capacitor (4490C). To enable integration of the capacitor (4490C), a variable-width pulse with a reference current input is used to generate an output voltage by the ITV (4401). The output voltage from the ITV (4401) is compared with the reference voltage by a comparator (4404). The parameters of the variable-width pulse are stored, for example, by a counter (not shown) in the digital domain or by a table (not shown) that stores the analog voltage in the analog domain for each ITV (the analog voltage is converted from the variable pulse input). This information is used by a controller (not shown) to enable the ITV using enable signals (not shown) for verification operations.
[0193] FIG. 44b illustrates a comparator and offset circuit (4490) that can be used instead of the comparator (4404) in FIG. 44a to add an additional function of offset compensation. The comparator and offset circuit (4490) includes a comparator (4491) which compares inputs (VINP and VINN) (which may be the same signals shown in FIG. 44a) to produce an output (COMPOUT) and its complement (COMPOUTB). A calibration circuit (4492) can be adjusted to provide an offset voltage (VON) to the comparator (4491), and a calibration circuit (4493) can be adjusted to provide an offset voltage (VOP) to the comparator (4492).
[0194] FIG. 45 illustrates a reference voltage generator (4500). The reference voltage generator (4500) is an exemplary embodiment of the reference voltage generator (4303). In one example, a reference array (4304) (not illustrated) provides a maximum current from a reference cell, where the current represents the highest possible weight (of N possible weights) that can be stored in a non-volatile memory cell. The maximum current provided from the reference cell is converted into a high-verification reference voltage by a current-voltage converter (4501), which can be done using a resistive current-voltage converter (RITV) (4511) or a capacitor current-voltage converter (CITV) (4510).
[0195] Then, these voltages are used to generate 32 verification reference voltages for a 5-bit cell, for example, by a resistor string (4504) containing N-1 resistors in series. In another example, the reference array provides a reference current that is converted into reference voltages, for example, the reference current may be an intermediate range value, which is appropriately converted into all N reference voltages (for example, by a current proportional mirror, by a trimmed resistor value, or by a trimmed capacitor value through an ITV circuit). The ITV (4501) as illustrated uses a differential operational amplifier. This ITV is a replica of the subcircuit ITV (4301) in the VMM system (4300) and is also illustrated as the ITV (4401) in the ITV+ADC+verification circuit (4488). The differential operational amplifier (op amp) is a replica of the local differential op amp (4480) in FIG. 44, and N global references can track local voltages from the ITV (4301) across PVT (process, power supply, or temperature) variations. Alternatively, the ITV (4501) can be based on a single-ended operational amplifier. Resistors (4511x) and capacitors (4510x) can be trimmed to adjust the range and compensate for mismatch or offset variations. The resistor string (4504) can be trimmed to adjust the range of VN to V1, shift the range VN up / down to V1, or adjust the local value VN to V1.
[0196] In another example, a constant current bias (e.g., from an IDAC) is used instead of the reference current from the reference array.
[0197] The current bias or reference current from the reference array can be adjusted to obtain a target value. These are also compensated for PVT (process, power supply, or temperature) variations.
[0198] The total global offset and mismatch compensation of the ITV+ADC+verification circuit (4488) can be trimmed by using the offset and mismatch trimming of the reference generator (4500). This can be achieved by adjusting the current bias (4512) or the trimming resistors (4511a and 4511b), capacitors (4510a and 4510b), or resistor string (4504).
[0199] Buffer (4502) is used to buffer a high verification reference voltage, for example, representing the 32nd level (L31) of the 32 reference voltage levels (L0-31) for a 5-bit cell, to drive the resistor string (4504) at one end of the resistor string (4504). In another example, buffer (4503) is provided to buffer a low verification reference voltage (VREF2), corresponding to the first level (L0) of the 32 levels (L0-31) for a 5-bit cell, and providing this voltage to one end of the resistor string (4504). For example, the high verification reference voltage may be 900 mV and the low verification reference voltage may be 300 mV. The voltage ladder (resistor string) (4504) generates N voltages ranging from V0 to VN, representing N possible values that can be stored in the VMM array. These reference voltages are then used by the verification circuit (4403) in FIG. 44. In another example, the input voltages and / or output voltages of the buffers (4502 and 4503) are trimmed to adjust the range and correct for any offset, such as the buffer offset. Alternatively, the ITV (4501) can directly drive the resistor string (4504) to provide 32 reference levels (L0-L31).
[0200] In one example, K verification reference voltage lines may be used to provide N different voltages. For example, for a 5-bit cell, 32 verification reference voltage lines are required to supply to the verification reference voltage selection circuit (4406), while for a 6-bit cell, 64 verification reference voltage lines are required. The 32 reference voltage lines may each be used twice to provide 64 verification reference voltages by time-multiplexing their use so that the 32 lines provide a first set of 32 voltages during a first verification period and a second set of 32 voltages during a second verification period. This approach may be extended without limitation by using 4 verification periods to provide 128 voltages for 7-bit cells or 8 verification periods to provide 256 voltages for 8-bit cells.
[0201] In one example, bias voltages for the inputs of the array for verification operation (e.g., CG bias and EG bias) are generated from a reference array so that these biases adapt to temperature to keep the array current as constant as possible.
[0202] FIGS. 46 to 50 illustrate examples of reference arrays that can be used in the reference array (4304) of FIG. 43.
[0203] FIG. 46 illustrates a physical array (4600). The physical array (4600) includes an array of non-volatile memory cells. The non-volatile memory cells may optionally include stacked gated flash memory cells or separate gated flash memory cells. The physical array (4600) is divided into two types of arrays: a VMM array (3401) (see FIG. 34) and a reference array (4304). In one example, the VMM array (3401) and the reference array (4304) share the same bit line. In another example, the VMM array (3401) and the reference array (4304) use separate sets of bit lines, with two separate sets of bit lines.
[0204] FIG. 47 illustrates a physical array (4700) divided into two arrays, namely a VMM array (3401) and a reference array (4304). In one example, the VMM array (3401) and the reference array (4304) share one or more sets of horizontal lines, such as word lines, control gate lines, and erase lines. In another example, the VMM array (3401) and the reference array (4304) do not share any sets of horizontal lines.
[0205] FIG. 48 illustrates an example in which a reference array (4304) and a VMM array (3401) are located in separate physical arrays. For example, there may be substrate separation or active diffusion separation between the two arrays. Physical array (4801) includes the VMM array (3401), and physical array (4802) includes the reference array (4304). The VMM array (3401) and the reference array (4304) do not share bit lines, word lines, control gate lines, or erase lines.
[0206] FIG. 49 illustrates an example of a reference array (4304). Here, the reference array (4304) comprises a plurality of reference sub-arrays, such as sub-reference arrays (4901-0, 4901-1, …, 4901-(n-1), and 4901-n). Accordingly, the reference array (4304) comprises n+1 different reference sub-arrays. Different reference sub-arrays may have different characteristics, and due to these different characteristics, each reference array features an IV curve different from that of other reference sub-arrays. For example, each reference sub-array may differ in one or more of the following dimensions: (1) the width of the control gate line of the transistors of each reference array; (2) the width of the word line of the transistors of each reference array; (3) the width of the floating gate of the transistors of each reference array; (4) the overall width of the non-volatile memory cell of each reference array; (5) the shallow trench isolation (STI) spacing within each reference array; or (6) other characteristics. Additionally, the reference sub-arrays may differ in one or more device implant states or doping characteristics (e.g., well implant state, source implant state, drain implant state, etc., but not limited thereto).
[0207] FIG. 50 illustrates another example of a reference array (4304). Here, the reference array (4303) includes multiple reference arrays such as reference sub-arrays (5001-0, 5001-1, …, 5001-(n-1), and 5001-n and 5001-0, 5002-1, …, 5002-(n-1), and 5002-n). Accordingly, the reference array (4304) includes 2*(n+1) different reference sub-arrays, which means twice as many as in FIG. 49. The different reference sub-arrays in FIG. 50 may have different characteristics as in FIG. 49, and due to these different characteristics, each reference sub-array features an IV curve different from that of the other reference arrays. For example, each reference sub-array may vary in dimensions of one or more of the control gate width, word line width, floating gate width, the total width of the array's non-volatile memory cells, STI spacing, and device implant state—but not limited to these.
[0208] It should be noted that, as used herein, both the terms "on" and "on" encompass "directly on" (without any intermediate material, element, or space placed between them) and "indirectly on" (with intermediate material, element, or space placed between them). Likewise, the term "adjacent" includes "directly adjacent" (without any intermediate material, element, or space placed between them) and "indirectly adjacent" (with intermediate material, element, or space placed between them); "embedded in" includes "directly embedded in" (without any intermediate material, element, or space placed between them) and "indirectly embedded in" (with intermediate material, element, or space placed between them); and "electrically connected" includes "directly electrically connected to" (without any intermediate material or element electrically connecting the elements together between them) and "indirectly electrically connected to" (with an intermediate material or element electrically connecting the elements together between them). For example, forming an element "on a substrate" may include not only forming an element directly on a substrate without any intermediate materials / elements between them, but also forming an element indirectly on a substrate with one or more intermediate materials / elements between them.
Claims
Claim 1 A system comprising a vector x matrix multiplication array including a plurality of non-volatile memory cells arranged in rows and columns, wherein each of the non-volatile memory cells is capable of storing one of N possible levels corresponding to one of N possible currents; and a plurality of output blocks comprising receiving current from each column of the vector x matrix multiplication array, generating analog voltages during a verification operation of the vector x matrix multiplication, comparing the analog voltages with one or more of N reference voltages, and generating digital outputs during a read operation of the vector x matrix multiplication array. Claim 2 A system according to claim 1, wherein the plurality of output blocks convert current from columns of the array into voltages using a plurality of resistors or a plurality of capacitors. Claim 3 A system according to claim 1, comprising a reference voltage generator that generates one or more of N reference voltages during the verification operation. Claim 4 A system according to paragraph 3, comprising a verification circuit that compares the voltage from the reference voltage generator with the voltage from one of the plurality of output blocks. Claim 5 In paragraph 4, the system wherein the verification circuit generates a digital output representing the result of the comparison. Claim 6 In paragraph 3, the system wherein the reference voltage generator generates one or more of the N reference voltages in response to a current received from a reference array. Claim 7 In paragraph 3, the system wherein the reference voltage generator generates one or more of the N reference voltages in response to a current received from a main reference current generator. Claim 8 In paragraph 3, the reference voltage generator comprises: a current-voltage converter for converting the maximum current among the N possible currents into a maximum voltage; and a resistor string for generating N voltages ranging from the maximum voltage to the minimum voltage. Claim 9 In paragraph 8, the system comprises (N-1) resistors in series, wherein the resistor string comprises (N-1) resistors. Claim 10 A system according to claim 8, comprising a first buffer for providing the maximum voltage to the first end of the resistor string. Claim 11 A system according to claim 10, comprising a second buffer for providing the minimum voltage to the second end of the resistor string. Claim 12 A system comprising: a current-voltage converter that converts a current from a vector x matrix array into a voltage; a successive approximation register analog-to-digital converter that receives the voltage from the current-voltage converter and generates a digital output during a read operation; and a verification circuit that receives the voltage from the current-voltage converter and compares it with one or more N reference voltages during a verification operation to verify that the voltage corresponds to an appropriate voltage among N possible voltages. Claim 13 In paragraph 12, the reference voltage is provided by a reference voltage generator, the reference voltage generator comprising: a current-voltage converter for converting a maximum current among N possible currents into a maximum voltage; and a resistor string for generating N voltages ranging from the maximum voltage to the minimum voltage. Claim 14 In paragraph 13, the system comprises (N-1) resistors in series, wherein the resistor string comprises (N-1) resistors. Claim 15 A system according to claim 13, comprising a first buffer for providing the maximum voltage to the first end of the resistor string. Claim 16 A system according to claim 15, comprising a second buffer for providing the minimum voltage to the second end of the resistor string. Claim 17 A system comprising: a vector x matrix multiplication array including a plurality of non-volatile memory cells arranged in rows and columns, wherein each of the non-volatile memory cells is capable of storing one of N possible voltages corresponding to one of N possible currents; and a plurality of current-voltage converters and a plurality of verification circuits that receive current from the columns of the array during a verification operation of the vector x matrix multiplication array to generate analog voltages and compare the analog voltages with one or more of N reference voltages, and generate digital outputs during a read operation of the vector x matrix multiplication array. Claim 18 In claim 17, a system comprising a reference voltage generator that generates one of N voltages during the verification operation of the vector x matrix multiplication array. Claim 19 A system according to claim 18, comprising a comparator that compares a voltage from the reference voltage generator with a voltage from one of the plurality of current-voltage converters. Claim 20 In paragraph 19, the above comparator is a system that performs offset correction. Claim 21 In paragraph 20, the system wherein the offset correction is performed in the time domain. Claim 22 In paragraph 17, the current-voltage converter is a system that converts currents from columns of the array into voltages using a plurality of resistors or a plurality of capacitors. Claim 23 delete Claim 24 delete Claim 25 delete Claim 26 delete Claim 27 delete Claim 28 delete Claim 29 delete Claim 30 delete Claim 31 delete Claim 32 delete
Citation Information
Patent Citations
Resistive memory device for matrix-vector multiplications
US20200013462A1
Precise data tuning method and apparatus for analog neural memory in an artificial neural network
WO2022182378A1