Split-array architecture for analog neural memory in deep learning artificial neural network
By using non-volatile memory arrays as synapses in artificial neural networks and integrating vector matrix multiplication systems, the inefficiencies of existing hardware technologies are addressed, achieving improved energy efficiency and precise synaptic weight tuning.
Patent Information
- Application Number
- JP2025034766
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-08-30
- Filing Date
- 2025-03-05
- Publication Date
- 2025-07-23
- Estimated Expiration
- 2041-09-02
AI Technical Summary
Existing artificial neural networks face challenges in high-performance information processing due to the lack of appropriate hardware technologies, particularly in terms of energy efficiency and scalability, as they rely on bulky CMOS-implemented synapses and digital supercomputers, which are costly and inefficient compared to biological networks.
The implementation of a non-volatile memory array as synapses in an artificial neural network, where each memory cell can be individually programmed and read without affecting others, allowing for continuous and precise tuning of synaptic weights, and utilizing a vector matrix multiplication (VMM) system that integrates non-volatile memory cells to perform multiplication and addition within the memory array, eliminating the need for separate logic circuits.
This approach enhances energy efficiency and reduces the need for separate multiplication and addition logic circuits, enabling precise tuning of synaptic weights and improving the performance of neural networks.
Smart Images

Figure 2025108410000001_ABST
Abstract
Description
Technical Field
[0001] (Claims of Priority) This application claims the benefit of priority of U.S. Provisional Patent Application No. 63 / 190,228, filed May 18, 2021, and entitled "Split Array Architecture for Analog Neural Memory in a Deep Learning Artificial Neural Network", and U.S. Patent Application No. 17 / 461,901, filed Aug. 30, 2021, and entitled "Split Array Architecture for Analog Neural Memory in a Deep Learning Artificial Neural Network", which are hereby incorporated by reference in their entirety.
[0002] (Field of the Invention) A number of embodiments are disclosed for splitting an array into multiple parts in an analog neural memory within a deep learning artificial neural network, each part interacting with a specific circuit dedicated to that part and other circuits shared with one or more other parts.
Background Art
[0003] An artificial neural network mimics a biological neural network (the central nervous system of an animal, particularly the brain), can depend on a large number of inputs, and is used to estimate or approximate a generally unknown function. An artificial neural network generally includes layers of interconnected "neurons" that exchange messages with each other.
[0004] Figure 1 shows an artificial neural network, in which the circles represent layers of inputs or neurons. The connections (referred to as synapses) are represented by arrows and have numerical weights that can be tuned based on experience. This enables the neural network to adapt to the input and become learnable. Typically, a neural network includes multiple input layers. Typically, there is one or more intermediate layers of neurons and an output layer of neurons that provides the output of the neural network. At each level, the neurons make decisions individually or collectively based on the data received from the synapses.
[0005] One of the major challenges in the development of artificial neural networks for high-performance information processing is the lack of appropriate hardware technologies. In practice, practical neural networks rely on a very large number of synapses, which enables high connectivity between neurons, i.e., a very high degree of parallelization of computational processing. In principle, such complexity can be realized by digital supercomputers or dedicated graphics processing unit clusters. However, in addition to high costs, these approaches also suffer from poor energy efficiency compared to biological networks, which mainly perform low-precision analog calculations and consume far less energy. CMOS analog circuits have been used in artificial neural networks, but most CMOS-implemented synapses have been too bulky assuming a large number of neurons and synapses.
[0006] The applicant has previously disclosed in U.S. Patent Application No. 15 / 594,439, incorporated by reference, an artificial (analog) neural network that utilizes one or more non-volatile memory arrays as synapses. The non-volatile memory arrays operate as analog neuromorphic memory. The neural network device includes a first plurality of synapses configured to receive a first plurality of inputs and then generate a first plurality of outputs, and a first plurality of neurons configured to receive the first plurality of outputs. The first plurality of synapses includes a plurality of memory cells, each of the memory cells including a spaced source region and drain region formed in a semiconductor substrate with a channel region extending therebetween, a floating gate insulated and disposed above a first portion of the channel region, and a non-floating gate insulated and disposed above a second portion of the channel region. Each of the plurality of memory cells is configured to store a weight value corresponding to a plurality of electrons in the floating gate. The plurality of memory cells are configured to multiply the stored weight values by the first plurality of inputs to generate the first plurality of outputs. Non-volatile memory cell
[0007] Non-volatile memories are well known. For example, U.S. Patent No. 5,029,130 (the “’130 patent”), which is incorporated herein by reference, discloses an array of split-gate non-volatile memory cells, which are a type of flash memory cell. Such a memory cell 210 is shown in FIG. 2. Each memory cell 210 includes a source region 14 and a drain region 16 formed in a semiconductor substrate 12, with a channel region 18 therebetween. A floating gate 20 is formed insulated above a first portion of the channel region 18 (and controls the conductivity of the first portion of the channel region 18), and extends over a portion of the source region 14. A word line terminal 22 (typically coupled to a word line) is disposed insulated above a second portion of the channel region 18, having a first portion (which controls the conductivity of the second portion of the channel region 18) and a second portion that extends upwardly above the floating gate 20. The floating gate 20 and the word line terminal 22 are insulated from the substrate 12 by a gate oxide. A bit line 24 is coupled to the drain region 16.
[0008] By applying a high positive voltage to the word line terminal 22, the memory cell 210 is erased (electrons are removed from the floating gate), whereby the electrons in the floating gate 20 pass through the insulator therebetween from the floating gate 20 to the word line terminal 22 via Fowler-Nordheim (FN) tunneling.
[0009] The memory cell 210 is programmed (electrons are added to the floating gate) by source side injection (SSI) by hot electrons by applying a positive voltage to the word line terminal 22 and a positive voltage to the source region 14. An electron current flows from the drain region 16 toward the source region 14. The electrons are accelerated and heat up when they reach the gap between the word line terminal 22 and the floating gate 20. A portion of the heated electrons is injected through the gate oxide into the floating gate 20 due to the electrostatic attraction from the floating gate 20.
[0010] The memory cell 210 is read by applying a positive read voltage to the drain region 16 and the word line terminal 22 (turning on the portion of the channel region 18 below the word line terminal). When the floating gate 20 is positively charged (i.e., electrons are erased), the portion of the channel region 18 below the floating gate 20 is also turned on, and current flows through the channel region 18, which is detected as the erased state, i.e., the "1" state. When the floating gate 20 is negatively charged (i.e., programmed with electrons), the portion of the channel region below the floating gate 20 becomes almost or completely off, and current does not flow (or hardly flows) through the channel region 18, which is detected as the programmed state, i.e., the "0" state.
[0011] Table 1 shows typical voltage / current ranges that can be applied to the terminals of the memory cell 110 to perform read, erase, and program operations.
Table 1
[0012] As another type of flash memory cell, other split gate type memory cell configurations are also known. For example, FIG. 3 shows a four-gate memory cell 310 including a source region 14, a drain region 16, a floating gate 20 above a first portion of the channel region 18, a select gate 22 (typically coupled to a word line, WL) above a second portion of the channel region 18, a control gate 28 above the floating gate 20, and an erase gate 30 above the source region 14. This configuration is described in U.S. Patent No. 6,747,310, which is hereby incorporated by reference for all purposes. Here, all gates are non-floating gates except for the floating gate 20, i.e., they are electrically connected or connectable to a voltage source. Programming is performed by injecting hot electrons from the channel region 18 into the floating gate 20 itself. Erase is performed by electrons tunneling from the floating gate 20 to the erase gate 30.
[0013] Table 2 shows typical voltage / current ranges that can be applied to the terminals of memory cell 310 to perform read, erase, and program operations. [Table 2]
[0014] FIG. 4 shows a 3-gate memory cell 410, which is another type of flash memory cell. Memory cell 410 is identical to memory cell 310 of FIG. 3, except that memory cell 410 does not have a separate control gate. (Erasure occurs through the use of an erase gate) The erase operation and the read operation are the same as those of FIG. 3, except that no control gate bias is applied. The programming operation is also performed without a control gate bias. As a result, during the programming operation, a higher voltage must be applied to the source line to compensate for the lack of control gate bias.
[0015] Table 3 shows typical voltage / current ranges that can be applied to the terminals of memory cell 410 to perform read, erase, and program operations. [Table 3]
[0016] FIG. 5 shows a stacked gate memory cell 510, which is another type of flash memory cell. Memory cell 510 is similar to memory cell 210 of FIG. 2, except that floating gate 20 extends over the entire channel region 18, and control gate 22 (coupled to the word line here) is separated by an insulating layer (not shown) and extends over floating gate 20. Erasure is performed by FN tunneling of electrons from the FG to the substrate, and programming is performed by channel hot electron (CHE) injection in the region between channel 18 and drain region 16, by electrons flowing from source region 14 to drain region 16, and by a read operation similar to the read operation of memory cell 210 having a higher control gate voltage.
[0017] Table 4 shows typical voltage ranges that can be applied to the terminals of memory cell 510 and substrate 12 to perform read, erase, and program operations.
Table 4
[0018] The methods and means described herein can be applied to other non-volatile memory technologies such as, but not limited to, FINFET split gate flash or stacked gate flash memory, NAND flash, SONOS (silicon-oxide-nitride-oxide-silicon, charge trap in nitride), MONOS (metal-oxide-nitride-oxide-silicon, metal charge trap in nitride), ReRAM (resistive change memory), PCM (phase change memory), MRAM (magnetoresistive memory), FeRAM (ferroelectric memory), CT (charge trap) memory, CN (carbon nanotube) memory, OTP (one-time programmable with bi-level or multi-level), and CeRAM (strongly correlated electron memory).
[0019] To utilize a memory array that includes one of the types of non-volatile memory cells in the artificial neural network described above, two modifications are made. First, the lines are configured so that each memory cell can be individually programmed, erased, and read without adversely affecting the memory states of other memory cells in the array, as further described below. Second, continuous (analog) programming of the memory cells is provided.
[0020] Specifically, the memory state of each memory cell in the array (i.e., the charge on the floating gate) can be continuously varied independently, with minimal disruption to other memory cells, from a completely erased state to a completely programmed state. In another embodiment, the memory state of each memory cell in the array (i.e., the charge on the floating gate) can be continuously varied independently, with minimal disruption to other memory cells, from a completely programmed state to a completely erased state and vice versa. This means that the cell memory is either analog or can store at least one of a number of discrete values (such as 16 or 64 different values), which allows all cells in the memory array to be very precisely and individually tunable, and makes the memory array ideal for fine-tuning adjustments to memory and to the synaptic weights of neural networks. Neural Network Using a Non-Volatile Memory Cell Array
[0021] FIG. 6 conceptually illustrates a non-limiting example of a neural network that utilizes the non-volatile memory array of this embodiment. This example uses a non-volatile memory array neural network for a face recognition application, but it is also possible to implement other suitable applications using a non-volatile memory array-based neural network.
[0022] S0 is the input layer, and in this example, it is a 32×32 pixel RGB image with 5-bit precision (i.e., three 32×32 pixel arrays, one for each of the colors R, G, and B, and each pixel has 5-bit precision). The synapses CB1 going from the input layer S0 to layer C1 apply a different set of weights to some instances and shared weights to other instances, scanning the input image with an overlapping 3×3 pixel filter (kernel) and shifting the filter by 1 pixel (or more than 2 pixels in some models) at a time. Specifically, the 9 pixel values in the 3×3 portion of the image (i.e., what is referred to as the filter or kernel) are provided to the synapses CB1, where these 9 input values are multiplied by appropriate weights, and after summing the outputs of the multiplications, a single output value is determined and given by the first synapse of CB1 to generate one pixel of the feature map of layer C1. The 3×3 filter is then shifted 1 pixel to the right within the input layer S0 (i.e., a 3-pixel column is added on the right and a 3-pixel column is dropped on the left), and the 9 pixel values of this newly positioned filter are provided to the synapses CB1, where they are multiplied by the same weights as above, and a second single output value is determined by the relevant synapse. This process is continued until the 3×3 filter has scanned over the entire 32×32 pixel image of the input layer S0 for all three colors and all bits (precision values). The process is then repeated using different sets of weights until all of the feature maps of layer C1 are calculated, generating different feature maps of layer C1.
[0023] In this example, in layer C1, there are 16 feature maps each having 30×30 pixels. Each pixel is a new feature pixel extracted from the multiplication of the input and the kernel. Thus, each feature map is a two-dimensional array. Therefore, in this example, layer C1 consists of 16 layers of two-dimensional arrays (note that the layers and arrays referred to in this specification are logical relationships rather than necessarily physical relationships, that is, the array is not necessarily oriented as a physical two-dimensional array). Each of the 16 feature maps in layer C1 is generated by one of 16 different sets of synaptic weights applied to the filter scan. All of the C1 feature maps can target different aspects of the same image feature, such as edge identification. For example, the first map (generated using a first set of weights shared by all scans used to generate this first map) can identify circular edges, and the second map (generated using a second set of weights different from the first set of weights) can identify rectangular edges or the aspect ratio of a particular feature, etc.
[0024] Before going from layer C1 to layer S1, an activation function P1 (pooling) that pools values from non - overlapping and consecutive 2×2 regions within each feature map is applied. The purpose of the pooling function P1 is to average neighboring positions (or it is also possible to use the max function), for example, to reduce the dependence on edge positions, and to reduce the data size before going to the next stage. In layer S1, there are 16 15×15 feature maps (i.e., 16 different arrays of 15×15 pixels each). The synapse CB2 going from layer S1 to layer C2 scans the maps within layer S1 with a 4×4 filter with a 1 - pixel filter shift. In layer C2, there are 22 12×12 feature maps. Before going from layer C2 to layer S2, an activation function P2 (pooling) that pools values from non - overlapping and consecutive 2×2 regions within each feature map is applied. In layer S2, there are 22 6×6 feature maps. In the synapse CB3 going from layer S2 to layer C3, an activation function (pooling) is applied, where all neurons within layer C3 are connected to all maps within layer S2 via each synapse of CB3. In layer C3, there are 64 neurons. The synapse CB4 going from layer C3 to the output layer S3 fully connects C3 to S3, that is, all neurons within layer C3 are connected to all neurons within layer S3. The output in S3 contains 10 neurons, and the neuron with the highest output determines the class. This output can, for example, indicate the identification or classification (class) of the content of the original image.
[0025] Each layer of the synapse is implemented using an array or a part of an array of non - volatile memory cells.
[0026] FIG. 7 is a block diagram of an array that can be used for that purpose. The vector-by-matrix multiplication (VMM) array 32 includes non-volatile memory cells and is utilized as synapses (such as CB1, CB2, CB3, and CB4 in FIG. 6) between one layer and the next layer. Specifically, the VMM array 32 includes an array 33 of non-volatile memory cells, an erase gate and word line gate decoder 34, a control gate decoder 35, a bit line decoder 36, and a source line decoder 37, and those decoders decode respective inputs to the non-volatile memory cell array 33. Inputs to the VMM array 32 can be made from the erase gate and word line gate decoder 34 or from the control gate decoder 35. The source line decoder 37 in this example also decodes the output of the non-volatile memory cell array 33. Alternatively, the bit line decoder 36 can decode the output of the non-volatile memory cell array 33.
[0027] The non-volatile memory cell array 33 serves two purposes. First, the non-volatile memory cell array 33 stores the weights used by the VMM array 32. Second, the non-volatile memory cell array 33 effectively multiplies the weights stored in the non-volatile memory cell array 33 by the inputs and adds them for each output line (source line or bit line) to generate an output, and this output becomes the input to the next layer or the input to the last layer. By the non-volatile memory cell array 33 performing the functions of multiplication and addition, the need for separate multiplication and addition logic circuits is eliminated, and the power efficiency is also good due to the calculation within the memory.
[0028] The output of the non-volatile memory cell array 33 is supplied to a differential summing device (such as a summing operational amplifier or a summing current mirror) 38 that sums the output of the non-volatile memory cell array 33 to create a single value for convolution. The differential summing device 38 is arranged to perform the sum of the positive and negative weights.
[0029] The total output value of the differential aggregator 38 is then supplied to an activation function block 39 that rectifies the output. The activation function block 39 can provide a sigmoid, tanh, or ReLU function. The rectified output value of the activation function block 39 becomes an element of the feature map as the next layer (e.g., C1 in FIG. 6), and is then applied to the next synapse to generate the next feature map layer or the last layer. Thus, in this example, the non-volatile memory cell array 33 constitutes a plurality of synapses (receiving inputs from the previous layer of neurons or from an input layer such as an image database), and the summing operational amplifier 38 and the activation function block 39 constitute a plurality of neurons.
[0030] The inputs (WLx, EGx, CGx, and optionally BLx and SLx) to the VMM array 32 of FIG. 7 can be at an analog level, a binary level, or digital bits (in which case, a DAC is provided to convert the digital bits to an appropriate input analog level), and the outputs can be at an analog level, a binary level, or digital bits (in which case, an output ADC is provided to convert the output analog level to digital bits).
[0031] FIG. 8 is a block diagram showing the use of multiple layers of the VMM array 32, labeled as VMM arrays 32a, 32b, 32c, 32d, and 32e in the figure. As shown in FIG. 8, an input (denoted as Inputx) is converted from digital to analog by a digital-to-analog converter 31 and provided to the input VMM array 32a. The converted analog input can be a voltage or a current. The input D / A conversion of the first layer can be performed by using a function or a look-up table (LUT) that maps the input Inputx to an appropriate analog level of the matrix multiplier of the input VMM array 32a. The input conversion can also be performed by an analog-to-analog (A / A) converter to convert an external analog input to the mapped analog input to the input VMM array 32a.
[0032] The output generated by the input VMM array 32a is then provided as input to the next VMM array (hidden level 1) 32b, which then generates an output that is provided as input to the input VMM array (hidden level 2) 32c, and so on. The various layers of the VMM array 32 function as the layers of synapses and neurons of a convolutional neural network (CNN). Each of the VMM arrays 32a, 32b, 32c, 32d, and 32e can be a stand-alone physical non-volatile memory array, or multiple VMM arrays can utilize different portions of the same physical non-volatile memory array, or multiple VMM arrays can utilize overlapping portions of the same physical non-volatile memory array. The example shown in FIG. 8 includes five layers (32a, 32b, 32c, 32d, 32e): one input layer (32a), two hidden layers (32b, 32c), and two fully connected layers (32d, 32e). One of ordinary skill in the art will understand that this is merely exemplary and that the system can alternatively include more than two hidden layers and more than two fully connected layers. Vector Matrix Multiplication (VMM) Array
[0033] FIG. 9 shows a neuron VMM array 900 that is particularly suitable for the memory cell 310 shown in FIG. 3 and is utilized as part of the synapses and neurons between the input layer and the next layer. The VMM array 900 includes a memory array 901 of non-volatile memory cells and a reference array 902 of non-volatile reference memory cells (located at the top of the array). Alternatively, another reference array can be located at the bottom.
[0034] In the VMM array 900, control gate lines such as control gate line 903 extend in the vertical direction (therefore, the reference array 902 in the row direction is orthogonal to the control gate line 903), and erase gate lines such as erase gate line 904 extend in the horizontal direction. Here, the input to the VMM array 900 is provided to the control gate lines (CG0, CG1, CG2, CG3), and the output of the VMM array 900 appears on the source lines (SL0, SL1). In one embodiment, only even rows are used, and in another embodiment, only odd rows are used. The current of each source line (SL0 and SL1 respectively) performs a summation function of all the currents from the memory cells connected to that particular source line.
[0035] As described herein for neural networks, the non-volatile memory cells of the VMM array 900, i.e., the memory cells 310 of the VMM array 900, are preferably configured to operate in the subthreshold region.
[0036] The non-volatile reference memory cells and non-volatile memory cells described herein are biased in weak inversion (subthreshold region) as follows: Ids = Io * e (Vg-Vth) / nVt = w * Io * e (Vg) / nVt where w = e (-Vth) / nVt and where Ids is the drain-source current, Vg is the gate voltage of the memory cell, Vth is the threshold voltage of the memory cell, Vt is the thermal voltage = k * T / q, where k is the Boltzmann constant, T is the Kelvin temperature, q is the electron charge, n is the slope factor = 1+(Cdep / Cox), Cdep is the capacitance of the depletion layer, and Cox is the capacitance of the gate oxide layer, Io is the memory cell current at a gate voltage equal to the threshold voltage, and Io is (Wt / L) * u * Cox * (n - 1) * Vt 2Proportional to, where u is the carrier mobility, and Wt and L are the width and length of the memory cell, respectively.
[0037] When using an I-V log converter that converts an input current into an input voltage using a memory cell (such as a reference memory cell or a peripheral memory cell) or a transistor: Vg=n * Vt * log[Ids / wp * Io] Where wp is the w of the reference or peripheral memory cell.
[0038] For a memory array used as a vector matrix multiplier VMM array with a current input, the output current is as follows: Iout=wa * Io * e (Vg) / nVt That is Iout=(wa / wp) * Iin=W * Iin W=e (Vthp-Vtha) / nVt Where wa is the w of each memory cell of the memory array. Vthp is the effective threshold voltage of the peripheral memory cell, and Vtha is the effective threshold voltage of the main (data) memory cell. Note that the threshold voltage of the transistor is a function of the substrate body bias voltage, and the substrate body bias voltage represented as Vsb can be modulated to compensate for various conditions at such a temperature. The threshold voltage Vth can be expressed as follows. Vth=Vth0+gamma(SQRT|Vsb-2 * φF)-SQRT|2 * φF|) Where Vth0 is the threshold voltage with zero substrate bias, φF is the surface potential, and gamma is the body effect parameter.
[0039] The word line or control gate can be used as the input of the memory cell for the input voltage.
[0040] Alternatively, the flash memory cells of the VMM array described herein can be configured to operate in the linear region. Ids = beta * (Vgs - Vth) * Vds; beta = u * Cox * Wt / L W = α(Vgs - Vth) That is, the weight W in the linear region is proportional to (Vgs - Vth).
[0041] The word line or control gate or bit line or source line can be used as an input to the memory cell operating in the linear region. The bit line or source line can be used as an output of the memory cell.
[0042] For an I-V linear converter, a memory cell (such as a reference memory cell or a peripheral memory cell) or a transistor operating in the linear region can be used to linearly convert the input / output current into an input / output voltage.
[0043] Alternatively, the memory cells of the VMM array described herein can be configured to operate in the saturation region. Ids = 1 / 2 * beta * (Vgs - Vth) 2 ; beta = u * Cox * Wt / L W ∝ (Vgs - Vth) 2 , that is, the weight W is proportional to (Vgs - Vth) 2 proportional.
[0044] The word line, control gate, or erase gate can be used as an input to the memory cell operating in the saturation region. The bit line or source line can be used as an output of the output neuron.
[0045] Alternatively, the memory cells of the VMM array described herein can be used for all regions or combinations thereof (sub-threshold, linear, or saturated) for each layer or multiple layers of a neural network.
[0046] Another embodiment for the VMM array 32 of FIG. 7 is described in U.S. Patent Application No. 15 / 826,345, which is incorporated herein by reference. As described in the above application, the source line or bit line can be used as a neuron output (sum of currents output).
[0047] FIG. 10 shows a neuron VMM array 1000 that is particularly suitable for the memory cell 210 shown in FIG. 2 and is utilized as a synapse between the input layer and the next layer. The VMM array 1000 includes a memory array 1003 of non-volatile memory cells, a reference array 1001 of first non-volatile reference memory cells, and a reference array 1002 of second non-volatile reference memory cells. The reference arrays 1001 and 1002 arranged in the column direction of the array function to convert the current inputs flowing into terminals BLR0, BLR1, BLR2, and BLR3 into voltage inputs WL0, WL1, WL2, and WL3. In practice, the first and second non-volatile reference memory cells are diode-connected through a multiplexer 1014 (only part shown) in a state where the current input flows in. The reference cells are tuned (e.g., programmed) to a target reference level. The target reference level is provided by a reference mini-array matrix (not shown).
[0048] The memory array 1003 serves two purposes. First, the memory array 1003 stores the weights used by the VMM array 1000 in respective memory cells. Second, the memory array 1003 effectively multiplies the weights stored in the memory array 1003 with the input (i.e., the current inputs provided to terminals BLR0, BLR1, BLR2, and BLR3, which are converted to input voltages by reference arrays 1001 and 1002 and supplied to word lines WL0, WL1, WL2, and WL3), and then adds up all the results (memory cell currents) to generate the output of each bit line (BL0~BLN), and this output becomes the input to the next layer or the input to the last layer. By performing the multiplication and addition functions, the memory array 1003 eliminates the need for separate multiplication and addition logic circuits and is also power-efficient. Here, the voltage input is provided to word lines WL0, WL1, WL2, and WL3, and the output appears on respective bit lines BL0~BLN during the read (inference) operation. The current of each of the bit lines BL0~BLN performs the total function of the currents from all the non-volatile memory cells connected to that specific bit line.
[0049] Table 5 shows the operating voltages and currents of the VMM array 1000. The columns in the table indicate the voltages applied to the word line of the selected cell, the word line of the non-selected cell, the bit line of the selected cell, the bit line of the non-selected cell, the source line of the selected cell, and the source line of the non-selected cell. The rows indicate the read, erase, and program operations.
Table 5
[0050] FIG. 11 shows a neuron VMM array 1100 that is particularly suitable for the memory cell 210 shown in FIG. 2 and is used as part of synapses and neurons between the input layer and the next layer. The VMM array 1100 includes a memory array 1103 of non-volatile memory cells, a reference array 1101 of first non-volatile reference memory cells, and a reference array 1102 of second non-volatile reference memory cells. The reference arrays 1101 and 1102 extend in the row direction of the VMM array 1100. The VMM array is similar to the VMM 1000 except that the word lines extend vertically in the VMM array 1100. Here, the input is provided to the word lines (WLA0, WLB0, WLA1, WLB2, WLA2, WLB2, WLA3, WLB3), and the output appears on the source lines (SL0, SL1) during a read operation. The current of each source line performs a summation function of all the currents from the memory cells connected to that particular source line.
[0051] Table 6 shows the operating voltages and currents of the VMM array 1100. The columns in the table show the voltages applied to the word line of the selected cell, the word line of the non-selected cell, the bit line of the selected cell, the bit line of the non-selected cell, the source line of the selected cell, and the source line of the non-selected cell. The rows show the read, erase, and program operations.
Table 6
[0052] FIG. 12 shows a neuron VMM array 1200 that is particularly suitable for the memory cell 310 shown in FIG. 3 and is used as part of synapses and neurons between the input layer and the next layer. The VMM array 1200 includes a memory array 1203 of non-volatile memory cells, a reference array 1201 of first non-volatile reference memory cells, and a reference array 1202 of second non-volatile reference memory cells. The reference arrays 1201 and 1202 function to convert the current inputs flowing into the terminals BLR0, BLR1, BLR2, and BLR3 into voltage inputs CG0, CG1, CG2, and CG3. In practice, the first and second non-volatile reference memory cells are diode-connected through a multiplexer 1212 (only part shown) in a state where the current inputs flow through BLR0, BLR1, BLR2, and BLR3. The multiplexer 1212 includes a corresponding multiplexer 1205 and a cascode transistor 1204, respectively, to ensure a constant voltage of each bit line (such as BLR0) of the first and second non-volatile reference memory cells during the read operation. The reference cells are tuned to a target reference level.
[0053] The memory array 1203 serves two purposes. First, the memory array 1203 stores the weights used by the VMM array 1200. Second, the memory array 1203 effectively multiplies the weights stored in the memory array by the inputs (the current inputs provided to the terminals BLR0, BLR1, BLR2, and BLR3, and the reference arrays 1201 and 1202 convert these current inputs into input voltages and supply them to the control gates (CG0, CG1, CG2, and CG3)), and then adds up all the results (cell currents) to generate an output, which appears on BL0~BLN and becomes the input to the next layer or the input to the last layer. By the memory array performing the functions of multiplication and addition, the need for separate multiplication and addition logic circuits is eliminated, and the power efficiency is also good. Here, the inputs are provided to the control gate lines (CG0, CG1, CG2, and CG3), and the outputs appear on the bit lines (BL0~BLN) during the read operation. The current of each bit line performs the total function of all the currents from the memory cells connected to that specific bit line.
[0054] The VMM array 1200 implements unidirectional tuning of non-volatile memory cells within the memory array 1203. That is, each non-volatile memory cell is erased and then partially programmed until the desired charge on the floating gate is reached. If too much charge is applied to the floating gate (such as when an incorrect value is stored in the cell), the cell is erased and a series of partial programming operations are restarted from the beginning. As shown, two rows sharing the same erase gate (such as EG0 or EG1) are erased together (known as page erase), and then each cell is partially programmed until the desired charge on the floating gate is reached.
[0055] Table 7 shows the operating voltages and currents of the VMM array 1200. The columns in the table show the word lines of the selected cells, the word lines of the non-selected cells, the bit lines of the selected cells, the bit lines of the non-selected cells, the control gates of the selected cells, the control gates of the non-selected cells within the same sector as the selected cells, the control gates of the non-selected cells in a different sector from the selected cells, the erase gates of the selected cells, the erase gates of the non-selected cells, the source lines of the selected cells, and the voltages applied to the source lines of the non-selected cells. The rows show the read, erase, and program operations. [Table 7]
[0056] FIG. 13 shows a neuron VMM array 1300 that is particularly suitable for the memory cell 310 shown in FIG. 3 and is used as part of the synapses and neurons between the input layer and the next layer. The VMM array 1300 includes a memory array 1303 of non-volatile memory cells, a reference array 1301 or a first non-volatile reference memory cell, and a reference array 1302 of second non-volatile reference memory cells. The EG lines EGR0, EG0, EG1, and EGR1 extend vertically, and the CG lines CG0, CG1, CG2, and CG3 and the SL lines WL0, WL1, WL2, and WL3 extend horizontally. The VMM array 1300 is similar to the VMM array 1400 except that the VMM array 1300 implements bidirectional tuning, and each individual cell can be completely erased, partially programmed, and partially erased as needed to reach the desired charge level of the floating gate by using individual EG lines. As shown, the reference arrays 1301 and 1302 convert the input current at the terminals BLR0, BLR1, BLR2, and BLR3 into control gate voltages CG0, CG1, CG2, and CG3 (through the action of the diode-connected reference cells via the multiplexer 1314), and these voltages are applied to the memory cells in the row direction. The current outputs (neurons) are in the bit lines BL0 to BLN, and each bit line sums all the currents from the non-volatile memory cells connected to that particular bit line.
[0057] Table 8 shows the operating voltages and currents of the VMM array 1300. The columns in the table show the word lines of the selected cells, the word lines of the non-selected cells, the bit lines of the selected cells, the bit lines of the non-selected cells, the control gates of the selected cells, the control gates of the non-selected cells in the same sector as the selected cells, the control gates of the non-selected cells in a different sector from the selected cells, the erase gates of the selected cells, the erase gates of the non-selected cells, the source lines of the selected cells, and the voltages applied to the source lines of the non-selected cells. The rows show the operations of read, erase, and program.
Table 8
[0058] FIG. 22 shows a neuron VMM array 2200 that is particularly suitable for the memory cell 210 shown in FIG. 2 and is used as part of synapses and neurons between the input layer and the next layer. In the VMM array 2200, the inputs INPUT0...., INPUT N are received by bit lines BL0,... BL N respectively, and the outputs OUTPUT1, OUTPUT2, OUTPUT3, and OUTPUT4 are generated by source lines SL0, SL1, SL2, and SL3 respectively.
[0059] FIG. 23 shows a neuron VMM array 2300 that is particularly suitable for the memory cell 210 shown in FIG. 2 and is used as part of synapses and neurons between the input layer and the next layer. In this example, the inputs INPUT0, INPUT1, INPUT2, and INPUT3 are received by source lines SL0, SL1, SL2, and SL3 respectively, and the outputs OUTPUT0,... OUTPUT N are generated by bit lines BL0,... BL N respectively.
[0060] FIG. 24 shows a neuron VMM array 2400 that is particularly suitable for the memory cell 210 shown in FIG. 2 and is used as part of synapses and neurons between the input layer and the next layer. In this example, the inputs INPUT0,... INPUT M are received by word lines WL0,... WL M respectively, and the outputs OUTPUT0,... OUTPUT N are generated by bit lines BL0,... BL N respectively.
[0061] FIG. 25 shows a neuron VMM array 2500 that is particularly suitable for the memory cell 310 shown in FIG. 3 and is used as part of synapses and neurons between the input layer and the next layer. In this example, the inputs INPUT0,... INPUT M are received by word lines WL0,... WL M respectively, and the outputs OUTPUT0,... OUTPUT Nare bit lines BL0, ..., BL N and are generated. Alternatively, the inputs can be received at control gates CG0, ..., CG M and can be received at control gates CG0, ..., CG
[0062] FIG. 26A shows a neuron VMM array 2600 that is particularly suitable for the memory cell 410 shown in FIG. 4 and is utilized as part of synapses and neurons between an input layer and a next layer. In this example, the inputs INPUT0, ..., INPUT n are received at vertical control gate lines CG0, ..., CG N respectively, and the outputs OUTPUT1 and OUTPUT2 are generated at source lines SL0 and SL1.
[0063] FIG. 26B shows a neuron VMM array 2620 which is an alternative design of the VMM array 2600 where the word lines are vertical instead of horizontal. In this instance, the inputs can be received at vertical word lines WL0, WL1, and the outputs OUTPUT1 and OUTPUT2 are generated at horizontal source lines SL0 and SL1.
[0064] FIG. 26C shows a neuron VMM array 2640 which is an alternative design of the VMM array 2600 where the erase gate lines are vertical instead of horizontal. In this instance, the inputs can be received at vertical erase gate lines EG0, EG1, and the outputs OUTPUT1 and OUTPUT2 are generated at horizontal source lines SL0 and SL1. And the outputs OUTPUT1 and OUTPUT2 are generated at horizontal source lines SL0 and SL1.
[0065] FIG. 27 shows a neuron VMM array 2700 that is particularly suitable for the memory cell 410 shown in FIG. 4 and is utilized as part of synapses and neurons between an input layer and a next layer. In this example, the inputs INPUT0, ..., INPUT N are the bit lines BL0, ..., BL NThey are respectively received by the gates of bit line control gates 2701-1, 2701-2, ..., 2701-(N-1) and 2701-N that are respectively coupled thereto. Exemplary outputs OUTPUT1 and OUTPUT2 are generated on source lines SL0 and SL1.
[0066] FIG. 28 shows a neuron VMM array 2800 that is particularly suitable for the memory cells 310 shown in FIG. 3, the memory cells 510 shown in FIG. 5, and the memory cells 710 shown in FIG. 7, and that is utilized as part of synapses and neurons between an input layer and a next layer. In this example, inputs INPUT0, ..., INPUT M are received on word lines WL0, ..., WL M and outputs OUTPUT0, ..., OUTPUT N are respectively generated on bit lines BL0, ..., BL N Alternatively, the inputs can be received on control gates CG0, ..., CG M .
[0067] FIG. 29 shows a neuron VMM array 2900 that is particularly suitable for the memory cells 310 shown in FIG. 3, the memory cells 510 shown in FIG. 5, and the memory cells 710 shown in FIG. 7, and that is utilized as part of synapses and neurons between an input layer and a next layer. In this example, inputs INPUT0, ..., INPUT M are received on control gate lines CG0, ..., CG M and outputs OUTPUT0, ..., OUTPUT N are respectively generated on vertical source lines SL0, ..., SL N where each source line SL i is coupled to the source lines of all the memory cells within column i. Alternatively, the inputs can be received on word lines WL0, ..., WL M .
[0068] FIG. 30 shows a neuron VMM array 3000 that is particularly suitable for the memory cell 310 shown in FIG. 3, the memory cell 510 shown in FIG. 5, and the memory cell 710 shown in FIG. 7, and is used as part of synapses and neurons between an input layer and a next layer. In this example, the inputs INPUT0, ..., INPUT M are received on the control gate lines CG0, ..., CG M Outputs OUTPUT0, ..., OUTPUT N are respectively generated on the vertical bit lines BL0, ..., BL N and each bit line BL i is coupled to the bit lines of all the memory cells within column i. Long Short-Term Memory
[0069] The prior art includes a concept known as long short-term memory (LSTM). LSTM units are often used within neural networks. With LSTM, a neural network can store information over an arbitrary predetermined period and use that information in subsequent operations. Conventional LSTM units include a cell, an input gate, an output gate, and a forget gate. The three gates regulate the flow of information into and out of the cell and the period for which information is stored within the LSTM. VMM is particularly useful in LSTM units.
[0070] Figure 14 shows an exemplary LSTM 1400. The LSTM 1400 in this example includes cells 1401, 1402, 1403, and 1404. Cell 1401 receives the input vector x0 and generates the output vector h0 and the cell state vector c0. Cell 1402 receives the input vector x1, the output vector (hidden state) h0 from cell 1401, and the cell state c0 from cell 1401, and generates the output vector h1 and the cell state vector c1. Cell 1403 receives the input vector x2, the output vector (hidden state) h1 from cell 1402, and the cell state c1 from cell 1402, and generates the output vector h2 and the cell state vector c2. Cell 1404 receives the input vector x3, the output vector (hidden state) h2 from cell 1403, and the cell state c2 from cell 1403, and generates the output vector h3. Additional cells can also be used, and the LSTM with four cells is merely an example.
[0071] Figure 15 shows an exemplary implementation of an LSTM cell 1500 that can be used for cells 1401, 1402, 1403, and 1404 of Figure 14. The LSTM cell 1500 receives the input vector x(t), the cell state vector c(t - 1) from the preceding cell, and the output vector h(t - 1) from the preceding cell, and generates the cell state vector c(t) and the output vector h(t).
[0072] The LSTM cell 1500 includes sigmoid function devices 1501, 1502, and 1503, each of which applies a number between 0 and 1 to control the degree to which each component of the input vector contributes to the output vector. The LSTM cell 1500 also includes tanh devices 1504 and 1505 for applying the hyperbolic tangent function to the input vector, multiplier devices 1506, 1507, and 1508 for multiplying two vectors, and an adder device 1509 for adding two vectors. The output vector h(t) can be provided to the next LSTM cell in the system or accessed for other purposes.
[0073] FIG. 16 shows an LSTM cell 1600 which is an example of an implementation of the LSTM cell 1500. For the convenience of the reader, the same numbering method from the LSTM cell 1500 is used in the LSTM cell 1600. The sigmoid function devices 1501, 1502, and 1503, and the tanh device 1504 each include a plurality of VMM arrays 1601 and activation function circuit blocks 1602. Thus, it can be seen that the VMM array is particularly useful in the LSTM cell used in a specific neural network system. The multiplier devices 1506, 1507, and 1508, and the adder device 1509 are implemented in a digital or analog manner. The activation function block 1602 can be implemented in a digital or analog manner.
[0074] An alternative example of the LSTM cell 1600 (and another example of an implementation of the LSTM cell 1500) is shown in FIG. 17. In FIG. 17, the sigmoid function devices 1501, 1502 and 1503, and the tanh device 1504 share the same physical hardware (VMM array 1701 and activation function block 1702) in a time-division multiplexed manner. The LSTM cell 1700 also includes a multiplier device 1703 for multiplying two vectors, an adder device 1708 for adding two vectors, a tanh device 1505 (including the activation function block 1702), a register 1707 for storing the value i(t) when i(t) is output from the sigmoid function block 1702, and the value f(t) * c(t - 1), a register 1704 for storing the value when its value is output from the multiplier device 1703 via the multiplexer 1710, and the value i(t) * u(t), a register 1705 for storing the value when its value is output from the multiplier device 1703 via the multiplexer 1710, and the value o(t) * c~(t), a register 1706 for storing the value when its value is output from the multiplier device 1703 via the multiplexer 1710, and the multiplexer 1709.
[0075] While the LSTM cell 1600 includes a plurality of sets of the VMM array 1601 and respective activation function blocks 1602, the LSTM cell 1700 includes only one set of the VMM array 1701 and the activation function block 1702, which is used to represent a plurality of layers in an embodiment of the LSTM cell 1700. The LSTM cell 1700 requires less space than the LSTM 1600 because it requires only 1 / 4 of the space required for the VMM and the activation function block as compared with the LSTM cell 1600.
[0076] It can be further understood that the LSTM unit typically includes a plurality of VMM arrays, each of which requires the functions provided by specific circuit blocks outside the VMM array, such as an adder, an activation function block, and a high-voltage generation block. Providing separate circuit blocks for each VMM array would require a significant amount of space within the semiconductor device and would be somewhat inefficient. Therefore, in the embodiments described below, an attempt is made to minimize the circuits required outside the VMM array itself. Gated Recurrent Unit
[0077] The analog VMM implementation can be utilized in a gated recurrent unit (GRU) system. The GRU is a gate mechanism within a recurrent neural network. The GRU is similar to the LSTM, except that the GRU cell generally includes fewer components than the LSTM cell.
[0078] Figure 18 shows an exemplary GRU 1800. The GRU 1800 in this example includes cells 1801, 1802, 1803, and 1804. Cell 1801 receives an input vector x0 and generates an output vector h0. Cell 1802 receives an input vector x1 and the output vector h0 from cell 1801, and generates an output vector h1. Cell 1803 receives an input vector x2 and the output vector (hidden state) h1 from cell 1802, and generates an output vector h2. Cell 1804 receives an input vector x3 and the output vector (hidden state) h2 from cell 1803, and generates an output vector h3. Additional cells can also be used, and the GRU with four cells is merely an example.
[0079] Figure 19 shows an exemplary implementation of a GRU cell 1900 that can be used for cells 1801, 1802, 1803, and 1804 in Figure 18. The GRU cell 1900 receives an input vector x(t) and an output vector h(t - 1) from a preceding GRU cell, and generates an output vector h(t). The GRU cell 1900 includes sigmoid function devices 1901 and 1902, each of which applies a number between 0 and 1 to components from the output vector h(t - 1) and the input vector x(t). The GRU cell 1900 also includes a tanh device 1903 for applying the hyperbolic tangent function to the input vector, a plurality of multiplier devices 1904, 1905, and 1906 for multiplying two vectors, an adder device 1907 for adding two vectors, and a complement device 1908 for subtracting the input from 1 to generate an output.
[0080] Figure 20 shows a GRU cell 2000, which is an example of an implementation of the GRU cell 1900. For the convenience of the reader, the same numbering method from the GRU cell 1900 is used in the GRU cell 2000. As can be seen from Figure 20, the sigmoid function devices 1901 and 1902, and the tanh device 1903 each include a plurality of VMM arrays 2001 and activation function blocks 2002. Therefore, it can be seen that the VMM array is particularly used in the GRU cell used in a specific neural network system. The multiplier devices 1904, 1905, 1906, the adder device 1907, and the complement device 1908 are implemented in a digital or analog manner. The activation function block 2002 can be implemented in a digital or analog manner.
[0081] An alternative example of the GRU cell 2000 (and another example of an implementation of the GRU cell 1900) is shown in Figure 21. In Figure 21, the GRU cell 2100 utilizes a VMM array 2101 and an activation function block 2102, and when configured as a sigmoid function, it applies a number from 0 to 1 to control the degree to which each component of the input vector contributes to the output vector. In Figure 21, the sigmoid function devices 1901 and 1902, and the tanh device 1903 share the same physical hardware (VMM array 2101 and activation function block 2102) in a time-division multiplexed manner. The GRU cell 2100 also includes a multiplier device 2103 for multiplying two vectors, an adder device 2105 for adding two vectors, a complement device 2109 for subtracting the input from 1 to generate an output, a multiplexer 2104, and the value h(t - 1) * r(t), a register 2106 for holding the value when it is output from the multiplier device 2103 via the multiplexer 2104, and the value h(t - 1) * z(t), a register 2107 for holding the value when it is output from the multiplier device 2103 via the multiplexer 2104, and the value h^(t) *It includes a register 2108 for holding (1 - z((t)) when its value is output from the multiplier device 2103 via the multiplexer 2104.
[0082] While the GRU cell 2000 includes a plurality of sets of the VMM array 2001 and the activation function block 2002, the GRU cell 2100 includes only one set of the VMM array 2101 and the activation function block 2102, which is used to represent multiple layers in the embodiment of the GRU cell 2100. The GRU cell 2100 requires less space than the GRU cell 2000 because it only needs one-third of the space required for the VMM and the activation function block compared to the GRU cell 2000.
[0083] It can be further understood that the GRU system typically includes a plurality of VMM arrays, each of which requires the functions provided by specific circuit blocks outside the VMM array, such as an integrator, an activation function block, and a high-voltage generation block. Providing separate circuit blocks for each VMM array would require a significant amount of space within the semiconductor device and would be somewhat inefficient. Therefore, in the embodiments described below, an attempt is made to minimize the circuits required outside the VMM array itself.
[0084] The input to the VMM array can be at an analog level, binary level, pulse, time-modulated pulse, or digital bit (in which case a DAC is required to convert the digital bit to an appropriate input analog level), and the output can be at an analog level, binary level, timing pulse, pulse, or digital bit (in which case an output ADC is required to convert the output analog level to a digital bit).
[0085] For each memory cell in the VMM array, each weight W can be implemented by a single memory cell, or by a differential cell, or by two blended memory cells (the average of two cells). In the case of a differential cell, two memory cells are required to implement the weight W as a differential weight (W = W+ - W-). In the case of two blended memory cells, two memory cells are required to implement the weight W as the average of two cells.
[0086] Each non-volatile memory cell used in an analog neuromorphic memory system must hold a very specific and precise amount of charge, i.e., the number of electrons, in the floating gate in response to erase / program. For example, each floating gate must hold one of N different values, where N is the number of different weights that can be represented by each cell. Examples of N include 16, 32, 64, 128, and 256.
[0087] In a VMM system, it is necessary to increase throughput and reduce latency as much as possible while reducing the total amount of space required for memory cells and support circuits.
Summary of the Invention
[0088] Numerous embodiments are disclosed for dividing an array in an analog neural memory within a deep learning artificial neural network into multiple parts, where each part interacts with a specific circuit dedicated to that part and other circuits shared with one or more other parts.
[0089]
[0090]
[0091]
[0092]
[0093]
[0094]
[0095]
[0096]
[0097]
[0098]
[0099]
[0100]
[0101]
[0102]
[0103]
[0104]
[0105]
[0106]
[0107]
[0108]
[0109]
[0110]
[0111]
[0112]
[0113]
[0114]
[0115]
[0116]
[0117]
[0118]
[0119]
[0120]
[0121]
[0122]
[0123]
[0124]
[0125]
[0126]
[0127]
[0128]
Brief Description of the Drawings
[0129]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
Figure 26A
Figure 26B
Figure 26C
Figure 27
Figure 28
Figure 29
Figure 30
Figure 31
Figure 32
Figure 33
Figure 34
Figure 35
Figure 36
Figure 37
Figure 38
Figure 39
Mode for Carrying Out the Invention
[0130] The artificial neural network of the present invention utilizes a combination of CMOS technology and a non-volatile memory array. Overview of the VMM System
[0131] FIG. 31 shows a block diagram of a VMM system 3100. The VMM system 3100 includes a VMM array 3101, a row decoder 3102, a high voltage decoder 3103, a column decoder 3104, a bit line driver 3105, an input circuit 3106, an output circuit 3107, control logic 3108, and a bias generator 3109. The VMM system 3100 further includes a high voltage generation block 3110 that includes a charge pump 3111, a charge pump regulator 3112, and a high voltage level generator 3113. The VMM system 3100 further includes a (program / erase, or alias weight adjustment) algorithm controller 3114, an analog circuit 3115, a control engine 3116 (which may include special functions such as arithmetic functions, startup functions, embedded microcontroller logic, etc.), and test control logic 3117. The systems and methods described below may be implemented in the VMM system 3100.
[0132] The input circuit 3106 may include circuits such as a DAC (Digital-to-Analog Converter), DPC (Digital-to-Pulse Converter, Digital-to-Time Modulated Pulse Converter), AAC (Analog-to-Analog Converter such as a Current-to-Voltage Converter, Logarithmic Converter), PAC (Pulse-to-Analog Level Converter), or any other type of converter. The input circuit 3106 may implement a normalization, linear or non-linear up / down scaling function, or arithmetic function. The input circuit 3106 may implement a temperature compensation function for the input level. The input circuit 3106 may implement an activation function such as ReLU or sigmoid. The output circuit 3107 may include circuits such as an ADC (Analog-to-Digital Converter for converting neuron analog output to digital bits), AAC (Analog-to-Analog Converter such as a Current-to-Voltage Converter, Logarithmic Converter), APC (Analog-to-Pulse Converter, Analog-to-Time Modulated Pulse Converter), Current-to-Voltage Converter, or any other type of converter. The output circuit 3107 may implement an activation function such as ReLU or sigmoid. The output circuit 3107 may implement a statistical normalization, regularization, up / down scaling / gain function, statistical rounding, or arithmetic function (e.g., addition, subtraction, division, multiplication, shift, log) of the neuron output. The output circuit 3107 may implement a temperature compensation function for the neuron output or array output (such as bit line output) in order to keep the power consumption of the array substantially constant or to improve the accuracy of the array (neuron) output by keeping the slope of the IV substantially the same, etc.
[0133] Figures 32 to 36 show embodiments of a VMM system that include some commonalities with the VMM system 3100 but also include some modifications.
[0134] FIG. 32 shows a VMM system 3200. The VMM system 3200 includes an array 3201, a shared row decoder 3202, a shared high voltage decoder 3203, column decoders 3204 and 3205, (row) input circuits 3220, output circuits 3206 and 3207, and a shared bit line driver 3208. The shared row decoder 3202 is coupled to all the rows in the array 3201 and applies a voltage to the selected row. The shared high voltage decoder 3203 can be selectively coupled to all the rows in the array 3201. The shared high voltage decoder 3203 optionally includes a control gate high voltage decoder 3231 that can be selectively coupled to all the rows in the array, and a shared erase gate high voltage decoder 3232 that can be selectively coupled to all the rows in the array. The input circuit 3220 is, for example, similar to the input circuit 3106 of FIG. 31. The circuits and functions of the output circuits 3206 and 3207 are each, for example, similar to the circuits and functions of the output circuit 3107 of FIG. 31. Different from the VMM system 3100, in the VMM system 3200, specific operations are divided among different sets of circuits. Specifically, half of the columns in the array 3201 (e.g., all odd columns) are operated by the column decoder 3204 and the output circuit 3206, and the other half of the columns in the array 3201 (e.g., all even columns) are operated by the column decoder 3205 and the output circuit 3207. Thus, the output circuit 3206 is coupled to the column decoder 3204 to generate a first output from one or more columns in the first half of the columns during a read operation, and the output circuit 3207 is coupled to the column decoder 3207 to generate a second output from one or more columns in the second half of the columns during a read operation. In this embodiment, during a program operation or an erase operation, all the columns are coupled to a plurality of shared bit line drivers 3208. This enables a plurality of bit lines to be read in parallel. That is, the plurality of bit lines coupled to the column decoder 3204 and the output circuit 3206 and the plurality of bit lines coupled to the column decoder 3205 and the output circuit 3207 are simultaneously enabled by the plurality of shared bit line drivers 3208 for a read operation. Thus, this increases the throughput for reading the array 3201. Alternatively, the read operations do not need to be simultaneous.
[0135] Optionally, referring further to FIG. 39, continuous diffusion can be implemented between the upper and lower halves of the array.
[0136] FIG. 33 shows a VMM system 3300. The VMM system 3300 includes arrays 3301a and 3301b, a row decoder 3302, a shared high voltage decoder 3303, column decoders 3304 and 3305, an input circuit 3320, current-voltage converter circuits 3306 and 3307, a shared analog-to-digital converter (ADC) 3308, and a shared bit line driver 3309. The current-voltage converter circuit 3306 or 3307 and the shared ADC circuit 3308 are part of the output circuit 3207 of FIG. 32.
[0137] Unlike the VMM system 3100, in the VMM system 3300, specific operations are divided among different sets of circuits. Specifically, the array 3301a is operated by the column decoder 3304 and the current-voltage converter 3306, and the array 3301b is operated by the column decoder 3305 and the current-voltage converter 3307. This allows multiple read operations and / or program operations to be executed simultaneously, and the read operation or program operation can be executed in parallel for one or more cells in the array 3301a and one or more cells in the array 3301b.
[0138] Both current-voltage converter circuits 3306 and 3307 are coupled to a shared analog-to-digital converter 3308 that is used in a time-division multiplexing manner during the read operation and a shared bit line driver 3309 that is used during the program operation and the erase operation. For example, during the read operation, array 3301a is enabled and coupled to column decoder 3304 and current-voltage converter circuit 3306, while at the same time, array 3301b is enabled and coupled to column decoder 3305 and current-voltage converter circuit 3307. The output voltages from current-voltage converter circuits 3306 and 3307 are sample-and-held (S / H) by, for example, the S / H capacitors in the shared ADC 3308, and these array output voltages are digitized (converted) by the time-division multiplexed shared ADC 3308 (since it is shared between current-voltage converter circuits 3306 and 3307). For example, two sets of S / H capacitors are used for one ADC shared between two current-voltage converter circuits. In another embodiment, one ADC can be used for N current-voltage converter circuits, in which case N sets of S / H capacitors are used.
[0139] The use of the shared ADC between two current-voltage converter circuits can also be applied to FIGS. 34 / 35 / 36.
[0140] FIG. 34 shows a VMM system 3400. The VMM system 3400 includes arrays 3401a and 3401b, a shared row decoder 3402, a shared high voltage decoder 3403, column decoders 3404 and 3405, an input circuit 3420, output circuits 3406 and 3407, and a shared bit line driver 3408. Different from the VMM system 3100, in the VMM system 3400, certain operations are divided among different sets of circuits. Specifically, array 3401a is operated by column decoder 3404 and output circuit 3406, and array 3401b is operated by column decoder 3405 and output circuit 3407. This allows multiple read operations / and or program operations to be executed simultaneously, and the read operation or program operation can be executed in parallel for one or more cells in array 3401a and one or more cells in array 3401b. Both arrays 3401a and 3401b are coupled to a shared bit line driver 3408 that is used during program and erase operations.
[0141] FIG. 35 shows a VMM system 3500. The VMM system 3500 includes arrays 3501a, 3501b, 3501c, and 3501d, row decoders 3502 and 3503, a shared high voltage decoder 3504, column decoders 3505, 3506, 3507, and 3508, an input circuit 3520, output circuits 3509, 3510, 3511, and 3512, and shared bit line drivers 3513 and 3514. The shared high voltage decoder 3504 can be selectively coupled to all rows in arrays 3501a, 3501b, 3501c, and 3501d. The row decoder 3502 is shared by arrays 3501a and 3501b, is coupled to all rows in these arrays, and applies a voltage to the selected row. The row decoder 3503 is shared by arrays 3501c and 3501d, is coupled to all rows in these arrays, and applies a voltage to the selected row.
[0142] In the VMM system 3500, certain operations are divided among different sets of circuits. Specifically, array 3501a is operated by column decoder 3505 and output circuit 3509, array 3501b is operated by column decoder 3507 and output circuit 3511, array 3501c is operated by column decoder 3506 and output circuit 3510, and array 3501d is operated by column decoder 3508 and output circuit 3512. This allows multiple read operations / and or program operations to be executed simultaneously at once in all four arrays, and the read operation or program operation can be executed in parallel for one or more cells in array 3501a, one or more cells in array 3501b, one or more cells in array 3501c, and one or more cells in array 3501d. Both arrays 3501a and 3501b are selectively coupled to shared bit line driver 3513 during program and erase operations. Both arrays 3501c and 3501d are selectively coupled to shared bit line driver 3514 during program and erase operations.
[0143] For example, column decoder 3505 and output circuit 3509 can execute a first read operation that generates a first output from one or more rows in array 3501a, column decoder 3506 and output circuit 3510 can execute a second read operation that generates a second output from one or more rows in array 3501c, column decoder 3507 and output circuit 3511 can execute a third read operation that generates a third output from one or more rows in array 3501b, and column decoder 3508 and output circuit 3512 can execute a fourth read operation that generates a fourth output from one or more rows in array 3501d. Optionally, the first and third read operations can be performed in parallel. Optionally, the second and fourth read operations can be performed in parallel.
[0144] FIG. 36 shows a VMM system 3600. The VMM system 3600 includes arrays 3601a, 3601b, 3601c, and 3601d, a row decoder 3621, control gate decoders 3602 and 3603, a shared high voltage decoder 3604, column decoders 3605, 3606, 3607, and 3608, output circuits 3609, 3610, 3611, and 3612, and shared bit line drivers 3613 and 3614. In the VMM system 3600, certain operations are divided among different sets of circuits. Specifically, array 3601a is operated by column decoder 3605 and output circuit 3609, array 3601b is operated by column decoder 3607 and output circuit 3611, array 3601c is operated by column decoder 3606 and output circuit 3610, and array 3601d is operated by column decoder 3608 and output circuit 3612. This allows multiple read operations and / or program operations to be executed simultaneously at once in all four arrays, and the read operation or program operation can be executed in parallel for one or more cells in array 3601a, one or more cells in array 3601b, one or more cells in array 3601c, and one or more cells in array 3601d. Both arrays 3601a and 3601b are selectively coupled to shared bit line driver 3613 during program and erase operations. Both arrays 3601c and 3601d are selectively coupled to shared bit line driver 3614 during program and erase operations.
[0145] FIGS. 32 to 36 show that reading is performed by the row input of the control gate. Alternatively, reading can be performed with a word line or an erase gate. The input circuits 3220 in FIG. 32, 3320 in FIG. 33, 3420 in FIG. 34, 3520 in FIG. 35, and 3620 in FIG. 36 are the same as the input circuit 3106 in FIG. 31. The output circuits 3206 / 3207 in FIG. 32, 3406 / 4307 in FIG. 34, 3507 / 3508 / 3509 / 3510 in FIG. 35, and 3607 / 3608 / 3609 / 3610 in FIG. 36 are the same as the output circuit 3107 in FIG. 31.
[0146] FIG. 37 shows a part of the VMM array 3700. The VMM array 3700 includes rows 3701, 3702, 3703, 3704, 3705, 3706, 3707, and 3708. Rows 3701, 3702, 3705, and 3706 share an erase gate line (EG0) and a source line (SL0), and rows 3703, 3704, 3707, and 3708 share an erase gate line (EG1) and a source line (SL1). In addition, rows 3701 and 3703 share a control gate line (CG0 / CG2), rows 3702 and 3704 share a control gate line (CG1 / CG3), rows 3705 and 3707 share a control gate line (CG4 / CG6), and rows 3706 and 3708 share a control gate line (CG5 / CG7). These connections enable different rows to share a decoder circuit. Array terminals are shared such that program or erase disturb is reduced by reducing the amount of erase or program voltage stress on unselected cells.
[0147] In the arrays of FIGS. 37 and 38 (described below), the row inputs for the VMM arrays 3700 and 3800 for neural read operations (where multiple rows and multiple bit lines are on simultaneously) are on the word lines. If the input for neural read is on the control gate, the control gate cannot be shared between multiple rows within the same sub-array or array bank.
[0148] FIG. 38 shows a part of the array 3800. The array 3800 includes sectors 3809 and 3819. Sector 3809 includes rows 3801, 3802, 3803, 3804, 3805, 3806, 3807, and 3808. Sector 3819 includes rows 3811, 3812, 3813, 3814, 3815, 3816, 3817, and 3818.
[0149] Row 3801 (the first row) and row 3811 (the second row) share a control gate line (CG0) (i.e., the control gate terminals of each cell in these rows are coupled to the same control gate line), row 3802 and 3812 share a control gate line (CG1) (i.e., the control gate terminals of each cell in these rows are coupled to the same control gate line), row 3803 and 3813 share a control gate line (CG2) (i.e., the control gate terminals of each cell in these rows are coupled to the same control gate line), row 3804 and 3814 share a control gate line (CG3) (i.e., the control gate terminals of each cell in these rows are coupled to the same control gate line), row 3805 and 3815 share a control gate line (CG4) (i.e., the control gate terminals of each cell in these rows are coupled to the same control gate line), row 3806 and 3816 share a control gate line (CG5) (i.e., the control gate terminals of each cell in these rows are coupled to the same control gate line), row 3807 and 3817 share a control gate line (CG6) (i.e., the control gate terminals of each cell in these rows are coupled to the same control gate line), row 3808 and 3818 share a control gate line (CG7) (i.e., the control gate terminals of each cell in these rows are coupled to the same control gate line). That is, the control gate is shared between sectors. These couplings enable different rows to share a decoder circuit. The array terminals are shared such that program or erase disturb is reduced by reducing the amount of erase or program voltage stress on unselected cells.
[0150] Row 3801 (the first row), row 3802 (the third row), row 3805, and row 3806 share an erase gate line (EG0) (i.e., the erase gate terminals of each cell in these rows are connected to the same erase gate line) and a source line (SL0) (i.e., the source line terminals of each cell in these rows are connected to the same source line), row 3803, 3084, 3807, and 3808 share an erase gate line (EG1) (i.e., the erase gate terminals of each cell in these rows are connected to the same erase gate line) and a source line (SL1) (i.e., the source line terminals of each cell in these rows are connected to the same source line), row 3811, 3812, 3815, and 3816 share an erase gate line (EG0) (i.e., the erase gate terminals of each cell in these rows are connected to the same erase gate line) and a source line (SL0) (i.e., the source line terminals of each cell in these rows are connected to the same source line), and row 3813, 3114, 3817, and row 3818 share an erase gate line (EG1) (i.e., the erase gate terminals of each cell in these rows are connected to the same erase gate line) and a source line (SL1) (i.e., the source line terminals of each cell in these rows are connected to the same source line).
[0151] FIG. 39 shows an exemplary layout of a portion of a single array 3901 (such as array 3101 of FIG. 31 and array 3201 of FIG. 32) and a split array 3902 (such as arrays 3301a and 3301b of FIG. 33, arrays 3401a and 3401b of FIG. 34, arrays 3501a, 3501b, 3501c, and 3501d of FIG. 35, and arrays 3601a, 3601b, 3601c, and 3601d of FIG. 36). The split array 3902 follows the same design as the array 3901, except that certain contacts and metal connections 3904 are removed (or not formed) to create sub-arrays 3903a and 3903b. A few dummy rows at the interface are disabled, such as by grounding the word lines and control gates. This maintains process uniformity by the front-end layer (i.e., continuous column diffusion within a column and continuous row diffusion within a source line), and the polysilicon is continuous and uniform between two arrays of non-volatile memory cells (between electrically separated arrays). This also results in a reduction of area overhead compared to the physical separation of different arrays.
[0152] As used herein, it should be noted that both the terms "over" and "on" include both "directly on" (with no intervening material, element, or gap therebetween) and "indirectly on" (with an intervening material, element, or gap therebetween). Similarly, the term "adjacent" includes "directly adjacent" (with no intervening material, element, or gap therebetween) and "indirectly adjacent" (with an intervening material, element, or gap therebetween), "attached to" includes "directly attached to" (with no intervening material, element, or gap therebetween) and "indirectly attached to" (with an intervening material, element, or gap therebetween), and "electrically coupled" includes "directly electrically coupled" (with no intervening material or element electrically connecting the elements together therebetween) and "indirectly electrically coupled" (with an intervening material or element electrically connecting the elements together therebetween). For example, forming an element "over a substrate" can include forming the element directly on the substrate without any intervening material / element therebetween and forming the element indirectly on the substrate with one or more intervening material / elements therebetween.
Claims
1. An analog neural memory, comprising: An array of a plurality of non-volatile memory cells arranged in a plurality of rows and a plurality of columns; A first column decoder coupled to a first half of the plurality of columns in the array; A second column decoder coupled to a second half of the plurality of columns in the array; A first output circuit coupled to the first column decoder for generating a first output from one or more columns in the first half of the columns during a first read operation; A second output circuit coupled to the second column decoder for generating a second output from one or more columns in the second half of the columns during a second read operation.
2. The analog neural memory according to claim 1, wherein the first read operation and the second read operation are performed in parallel.
3. A shared bit line driver coupled to the first column decoder and the second column decoder during a program operation. The analog neural memory according to claim 1, further comprising the shared bit line driver.
4. The analog neural memory according to claim 1, wherein a shared high voltage decoder is selectively coupled to all rows in the array.
5. The analog neural memory according to claim 1, wherein a shared control gate high voltage decoder is selectively coupled to all rows in the array.
6. The analog neural memory according to claim 1, wherein a shared erase gate high voltage decoder is selectively coupled to all rows in the array.
7. The analog neural memory according to claim 1, wherein a shared row decoder is coupled to all rows in the array.
8. The analog neural memory according to claim 1, wherein continuous column diffusion is performed between columns in the first half of the plurality of columns and the second half of the plurality of columns.
9. An analog neural memory, comprising: A first array of a plurality of non-volatile memory cells arranged in a plurality of rows and a plurality of columns; A second array of a plurality of non-volatile memory cells arranged in a plurality of rows and a plurality of columns; A third array of a plurality of non-volatile memory cells arranged in a plurality of rows and a plurality of columns; A fourth array of a plurality of non-volatile memory cells arranged in a plurality of rows and a plurality of columns; A first row decoder coupled to the plurality of rows of the first array and the second array. A second row decoder coupled to a plurality of rows of the third array and the fourth array; A first column decoder coupled to the first array; A second column decoder coupled to the second array; A third column decoder coupled to the third array; A fourth column decoder coupled to the fourth array; A first output circuit coupled to the first column decoder for generating a first output from one or more rows in the first array during a first read operation; A second output circuit coupled to the second column decoder for generating a second output from one or more rows in the second array during a first read operation; A third output circuit coupled to the third column decoder for generating a third output from one or more rows in the third array during a second read operation; A fourth output circuit coupled to the fourth column decoder for generating a fourth output from one or more rows in the fourth array during the second read operation, an analog-to-analog memory.
10. The analog-to-analog memory according to claim 9, wherein the first read operation and the third read operation are performed in parallel.
11. The analog-to-analog memory according to claim 9, wherein the second read operation and the fourth read operation are performed in parallel.
12. A first shared bit line driver coupled to the first column decoder and the second column decoder during a program operation; A second shared bit line driver coupled to the third column decoder and the fourth column decoder during a program operation; The analog-to-analog memory according to claim 9, further comprising.
13. The analog-to-analog memory according to claim 9, wherein each of the first output circuit, the second output circuit, the third output circuit, and the fourth output circuit comprises a current-to-voltage converter.
14. The analog-to-analog memory according to claim 13, wherein each of the first output circuit, the second output circuit, the third output circuit, and the fourth output circuit further comprises an analog-to-digital converter coupled to the current-to-voltage converter.
15. The analog-to-analog memory according to claim 9, wherein a shared high voltage decoder is selectively coupled to all rows in the array.
16. The analog neuromorphic memory according to claim 9, wherein the first array, the second array, the third array, and the fourth array each include continuous column diffusion between columns.
17. The analog neuromorphic memory according to claim 9, wherein the first array, the second array, the third array, and the fourth array are formed from one physical array and are divided from each other by a part of the physical array without metal contacts.
18. An analog neuromorphic memory, comprising: an array of a plurality of non-volatile memory cells arranged in a plurality of rows and a plurality of columns; a first output circuit coupled to a first half of the plurality of columns in the array for generating a first output from one or more columns in the first half of the plurality of columns during a first read operation; a second output circuit coupled to a second half of the plurality of columns for generating a second output from one or more columns in the second half of the plurality of columns during a second read operation.
19. The analog neuromorphic memory according to claim 18, wherein the first read operation and the second read operation are performed in parallel.
20. The analog neuromorphic memory according to claim 18, wherein a shared high voltage decoder is selectively coupled to all rows in the array.
21. The analog neuromorphic memory according to claim 18, wherein a shared control gate high voltage decoder is selectively coupled to all rows in the array.
22. The analog neuromorphic memory according to claim 18, wherein a shared erase gate high voltage decoder is selectively coupled to all rows in the array.
23. The analog neuromorphic memory according to claim 18, wherein a shared word line decoder is selectively coupled to all rows in the array.
24. The analog neuromorphic memory according to claim 18, wherein the array includes continuous column diffusion between columns in the first half of the plurality of columns and the second half of the plurality of columns.
25. An analog neuromorphic memory, comprising: an array of a plurality of non-volatile memory cells arranged in a plurality of rows and a plurality of columns, each non-volatile memory cell including a control gate terminal, a word line terminal, a source line terminal, and an erase gate terminal. A plurality of control gate lines, each control gate line being coupled to a plurality of control gate terminals of a plurality of rows of non-volatile memory cells, the plurality of control gate lines; A plurality of word lines, each word line being coupled to a plurality of word line terminals of a plurality of rows of non-volatile memory cells, the plurality of word lines; A plurality of source lines, each source line being coupled to source line terminals of two adjacent rows of non-volatile memory cells, the plurality of source lines; A plurality of erase gate lines, each erase gate line being coupled to a plurality of erase gate terminals of a plurality of rows of non-volatile memory cells, the plurality of erase gate lines, comprising; An analog neuromorphic memory in which the control gate line of the first row is coupled to the control gate line of the second row, the erase gate line of the first row is coupled to the erase gate line of the third row, and the source line of the first row is coupled to the source line of the third row.
26. The analog neuromorphic memory according to claim 25, wherein the first row and the second row are in different sectors.
27. The analog neuromorphic memory according to claim 25, wherein the first row and the third row are in different sectors.
Citation Information
Patent Citations
Efficient verification for coarse / fine programming of non-volatile memory
JP2007520845A
Non-volatile memory device with configurable page size
JP2011511391A
Decoders for analog neural memory in deep learning artificial neural network
WO2019177698A1
Output array neuron conversion and calibration for analog neural memory in deep learning artificial neural network
WO2020222869A1