Artificial neural networks including three-dimensional integrated circuits.

By employing nonvolatile memory arrays for synaptic weights in artificial neural networks, the challenges of energy efficiency and scalability are addressed, enabling precise tuning and efficient information processing.

JP2025514878AActive Publication Date: 2025-05-12SILICON STORAGE TECHNOLOGY INC
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2024559235
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-06-23
Filing Date
2022-07-19
Publication Date
2025-05-12
Estimated Expiration
2042-07-19

AI Technical Summary

Technical Problem

Current artificial neural networks face challenges in achieving high performance information processing due to the lack of suitable hardware technology, specifically in terms of energy efficiency and scalability for large numbers of synapses.

Method used

The use of nonvolatile memory arrays, such as those employing split gate or stacked gate memory cells, as synapses in artificial neural networks, allowing for continuous analog programming and individual tuning of memory cells to store synaptic weights efficiently.

Benefits of technology

This approach enables precise tuning of synaptic weights, improves energy efficiency, and enhances the scalability of artificial neural networks, making them more suitable for complex information processing tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025514878000001_ABST
    Figure 2025514878000001_ABST
Patent Text Reader

Abstract

Numerous examples of artificial neural networks including three-dimensional integrated circuits are disclosed. In one embodiment, a three-dimensional integrated circuit for use in an artificial neural network comprises a first die including a first vector matrix multiplication array and a first input multiplexer, located in a first vertical layer, a second die including an input circuit, located in a second vertical layer different from the first vertical layer, and one or more vertical interfaces coupling the first and second dies, wherein during a read operation, the input circuit provides an input signal to the first input multiplexer through at least one of the one or more vertical interfaces, and the first input multiplexer applies the input signal to one or more rows in the first vector matrix multiplication array, and the first vector matrix multiplication array generates an output.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] (Priority Claim) This application claims priority to U.S. Provisional Patent Application No. 63 / 328,126, filed April 6, 2022, entitled "Artificial Neural Network Comprising a Three-Dimensional Integrated Circuit," and U.S. Patent Application No. 17 / 848,371, filed June 23, 2022, entitled "Artificial Neural Network Comprising a Three-Dimensional Integrated Circuit."

[0002] FIELD OF THEINVENTION Numerous embodiments of artificial neural networks including three-dimensional integrated circuits are disclosed. [Background technology]

[0003] Artificial neural networks mimic biological neural networks (the central nervous systems of animals, particularly the brain) and are used to estimate or approximate functions that may depend on multiple inputs and are generally unknown. Artificial neural networks typically contain layers of interconnected "neurons" that exchange messages between each other.

[0004] FIG. 1 shows an artificial neural network, where the circles represent layers of inputs or neurons. The connections (called synapses) are represented by arrows and have numerical weights that can be tuned based on experience. This allows the neural network to adapt to the inputs and learn. Typically, a neural network contains multiple layers of inputs. Typically, there are one or more hidden layers of neurons, and an output layer of neurons that provide the output of the neural network. Neurons at each level make decisions, individually or collectively, based on the data they receive from the synapses.

[0005] One of the main challenges in the development of artificial neural networks for high-performance information processing is the lack of suitable hardware technology. Indeed, practical neural networks rely on a very large number of synapses, which allows high connectivity between neurons, i.e., a very high degree of parallelization of computation. In principle, such complexity could be achieved by digital supercomputers or clusters of dedicated graphics processing units. However, in addition to high costs, these approaches also suffer from poor energy efficiency compared to biological networks, which mainly perform low-precision analog calculations and therefore consume much less energy. CMOS analog circuits have been used for artificial neural networks, but most CMOS-implemented synapses are too bulky given the large number of neurons and synapses.

[0006] Applicant previously disclosed in U.S. Patent Application Publication No. 2017 / 0337466(A1), incorporated by reference, an artificial (analog) neural network utilizing one or more non-volatile memory arrays as synapses. The non-volatile memory array operates as an analog neural memory and includes non-volatile memory cells arranged in rows and columns. The neural network includes a first plurality of synapses configured to receive a first plurality of inputs and generate a first plurality of outputs therefrom, and a first plurality of neurons configured to receive the first plurality of outputs. The first plurality of synapses includes a plurality of memory cells, each of which includes spaced apart source and drain regions formed in a semiconductor substrate with a channel region extending therebetween, a floating gate disposed above and insulated from a first portion of the channel region, and a non-floating gate disposed above and insulated from a second portion of the channel region. Each of the plurality of memory cells stores a weight value corresponding to a number of electrons in the floating gate. The plurality of memory cells multiply the first plurality of inputs by the stored weight values ​​to generate a first plurality of outputs. <Non-volatile memory cells>

[0007] Non-volatile memories are well known. For example, U.S. Pat. No. 5,029,130 ​​("the '130 patent"), incorporated herein by reference, discloses an array of split-gate non-volatile memory cells, which are a type of flash memory cell. Such a memory cell 210 is shown in FIG. 2. Each memory cell 210 includes a source region 14 and a drain region 16 formed in a semiconductor substrate 12, with a channel region 18 between the source region 14 and the drain region 16. A floating gate 20 is formed above and insulated from (and controls the conductivity of) a first portion of the channel region 18, and is formed above a portion of the source region 14. A wordline terminal 22 (typically coupled to a wordline) has a first portion disposed above and insulated from (and controls the conductivity of) a second portion of the channel region 18, and a second portion extending upwardly above the floating gate 20. A floating gate 20 and a wordline terminal 22 are insulated from the substrate 12 by a gate oxide. A bitline 24 is coupled to the drain region 16.

[0008] The memory cell 210 is erased (electrons are removed from the floating gate) by applying a high positive voltage to the wordline terminal 22, which causes the electrons in the floating gate 20 to pass through the intermediate insulator from the floating gate 20 to the wordline terminal 22 via Fowler-Nordheim (FN) tunneling.

[0009] The memory cell 210 is programmed by hot electron Source Side Injection (SSI) by applying a positive voltage to the word line terminal 22 and a positive voltage to the source region 14 (electrons are applied to the floating gate). An electronic current flows from the drain region 16 towards the source region 14. The electrons accelerate and generate heat when they reach the gap between the word line terminal 22 and the floating gate 20. Some of the heated electrons are injected into the floating gate 20 through the gate oxide due to electrostatic attraction from the floating gate 20.

[0010] The memory cell 210 is read by applying a positive read voltage to the drain region 16 and the word line terminal 22 (turning on the portion of the channel region 18 below the word line terminal). When the floating gate 20 is positively charged (i.e., erased with electrons), the portion of the channel region 18 below the floating gate 20 is also turned on and current will flow through the channel region 18, which is sensed as the erased or "1" state. When the floating gate 20 is negatively charged (i.e., programmed with electrons), the portion of the channel region below the floating gate 20 is mostly or completely off and no (or very little) current will flow through the channel region 18, which is sensed as the programmed or "0" state.

[0011] Table 1 shows typical voltage / current ranges that may be applied to the terminals of memory cell 210 to perform read, erase, and program operations. Table 1: Operation of the Flash Memory Cell 210 of FIG. 2 [Table 1]

[0012] Other split-gate memory cell configurations are also known, as are other types of flash memory cells. For example, FIG. 3 shows a four-gate memory cell 310 including a source region 14, a drain region 16, a floating gate 20 above a first portion of a channel region 18, a select gate 22 (typically coupled to a word line (WL)) above a second portion of the channel region 18, a control gate 28 above the floating gate 20, and an erase gate 30 above the source region 14. This configuration is described in U.S. Pat. No. 6,747,310, which is incorporated herein by reference for all purposes. Here, all gates are non-floating gates, except for the floating gate 20, i.e., they are electrically connected or connectable to a voltage source. Programming is performed by heated electrons injecting themselves from the channel region 18 into the floating gate 20. Erasing is performed by electrons tunneling from the floating gate 20 to the erase gate 30.

[0013] Table 2 shows typical voltage / current ranges that may be applied to the terminals of memory cell 310 to perform read, erase, and program operations. Table 2: Operation of the Flash Memory Cell 310 of FIG. [Table 2]

[0014] Figure 4 shows another type of flash memory cell, a three-gate memory cell 410. Memory cell 410 is identical to memory cell 310 of Figure 3, except that memory cell 410 does not have a separate control gate. Erase and read operations (where erasure occurs through the use of an erase gate) are similar to those of Figure 3, except that no control gate bias is applied. Programming operations are also performed without a control gate bias, and as a result, a higher voltage is applied to the source line during a program operation to compensate for the lack of control gate bias.

[0015] Table 3 shows typical voltage / current ranges that may be applied to the terminals of memory cell 410 to perform read, erase, and program operations. Table 3: Operation of Flash Memory Cell 410 of FIG. [Table 3]

[0016] Figure 5 shows another type of flash memory cell, a stacked gate memory cell 510. The memory cell 510 is similar to the memory cell 210 of Figure 2, except that the floating gate 20 extends over the entire channel region 18, and the control gate 22 (where it is coupled to a word line) extends over the floating gate 20, separated by an insulating layer (not shown). Erasing is done by FN tunneling of electrons from the FG (Frame Ground) to the substrate, and programming is done by electrons flowing from the source region 14 towards the drain region 16 by Channel Hot Electron (CHE) injection in the region between the channel 18 and the drain region 16, and by a read operation that is similar to the read operation of the memory cell 210 with a higher control gate voltage.

[0017] Table 4 shows typical voltage ranges that may be applied to the terminals of memory cell 510 and substrate 12 to perform read, erase, and program operations. Table 4: Operation of Flash Memory Cell 510 of FIG. [Table 4]

[0018] The methods and means described herein may be used in a wide variety of applications, including but not limited to, FINFET (Fin Field Effect Transistor) split gate flash memory or stack gate flash memory, NAND flash, SONOS (Silicon-Oxide-Nitride-Oxide-Silicon, charge traps in nitride), MONOS (Metal-Oxide-Nitride-Oxide-Silicon, metal charge traps in nitride), ReRAM (Resistive Random Access Memory), PCM (Phase Change Memory), MRAM (magnetic ram), FeRAM (ferroelectric ram), CT (charge trap) memory, CN (Carbon-Tube) memory, OTP (One Time Programmable Transformer) memory, and the like. The present invention may also be applied to other non-volatile memory technologies such as EEPROM (Programmable, bi-level or multi-level one-time programmable) and CeRAM (correlated electron ram).

[0019] In order to utilize a memory array containing one of the above non-volatile memory cell types in an artificial neural network, two modifications are made: First, as explained further below, the lines are constructed so that each memory cell can be individually programmed, erased, and read without adversely affecting the memory state of the other memory cells in the array; Second, a continuous (analog) programming of the memory cells is provided.

[0020] Specifically, the memory state (i.e., the charge on the floating gate) of each memory cell in the array can be changed continuously from a fully erased state to a fully programmed state, and vice versa, independently and with minimal disturbance to the other memory cells. This means that the cell storage is essentially analog, or can at a minimum store one of a number of discrete values ​​(such as 16 or 64 different values), making every memory cell in the memory array very precisely and individually tunable, making the memory array ideal for storing the synaptic weights of a neural network and for fine tuning adjustments to the synaptic weights. <Neural network using non-volatile memory cell array>

[0021] 6 conceptually illustrates a non-limiting example of a neural network utilizing the non-volatile memory array of the example embodiment. This example uses the non-volatile memory array neural network for a face recognition application, although the non-volatile memory array based neural network can also be used to implement any other suitable application.

[0022] S0 is the input layer, which in this example is a 32x32 pixel RGB image with 5-bit precision (i.e., three 32x32 pixel arrays, one for each color R, G, and B, each pixel with 5-bit precision). Synapse CB1 going from input layer S0 to layer C1 scans the input image with overlapping filters (kernels) of 3x3 pixels, sometimes applying different sets of weights and sometimes applying shared weights, and shifts the filters by one pixel (or by two or more pixels depending on the model). Specifically, the values ​​of the nine pixels in the 3x3 portion of the image (i.e., referred to as filters or kernels) are provided to synapse CB1, which multiplies these nine input values ​​by appropriate weights and, after summing the outputs of the multiplication, determines a single output value, which is provided by the first synapse of CB1 to generate one pixel of the feature map of layer C1. The 3x3 filter is then shifted by one pixel to the right in the input layer S0 (i.e., add a column of 3 pixels to the right and drop a column of 3 pixels on the left), so that the 9 pixel values ​​of this newly positioned filter are provided to synapse CB1, where they are multiplied by the same weights to determine a second single output value by the associated synapse. This process continues until the 3x3 filter has scanned the entire 32x32 pixel image of the input layer S0 for all three colors and all bits (precision values). The process is then repeated using different sets of weights to generate different feature maps of layer C1 until all feature maps of layer C1 have been calculated.

[0023] In this example, there are 16 feature maps in layer C1, each with 30x30 pixels. Each pixel is a new feature pixel extracted from the multiplication of the input with a kernel, and thus each feature map is a two-dimensional array, and thus in this example, layer C1 constitutes 16 layers of two-dimensional arrays (note that layers and arrays referred to herein are logical, not necessarily physical, relationships, i.e., arrays are not necessarily oriented in a physical two-dimensional array). Each of the 16 feature maps in layer C1 is generated by one of 16 different synaptic weight sets applied to the filter scans. The C1 feature maps may all target different aspects of the same image feature, such as boundary identification. For example, a first map (generated using a first weight set shared by all scans used to generate this first map) may identify circular edges, a second map (generated using a second weight set different from the first weight set) may identify rectangular edges or the aspect ratio of a particular feature, etc.

[0024] Before going from layer C1 to layer S1, an activation function P1 (pooling) is applied that pools values ​​from non-overlapping, consecutive 2x2 regions in each feature map. The purpose of the pooling function P1 is to average nearby positions (or a max function can be used), e.g., to reduce edge position dependency, and to reduce data size before going to the next stage. In layer S1, there are 16 15x15 feature maps (i.e., 16 different arrays of 15x15 pixels each). The synapse CB2 going from layer S1 to layer C2 scans the maps in layer S1 with a 4x4 filter with a filter shift of 1 pixel. In layer C2, there are 22 12x12 feature maps. Before going from layer C2 to layer S2, an activation function P2 (pooling) is applied that pools values ​​from non-overlapping, consecutive 2x2 regions in each feature map. In layer S2, there are 22 6x6 feature maps. At the synapse CB3 going from layer S2 to layer C3, an activation function (pooling) is applied, where every neuron in layer C3 connects to every map in layer S2 through a respective synapse in CB3. There are 64 neurons in layer C3. The synapse CB4 going from layer C3 to the output layer S3 fully connects C3 to S3, i.e. every neuron in layer C3 connects to every neuron in layer S3. The output at S3 contains 10 neurons, where the neuron with the highest output determines the class. This output may indicate, for example, the identification or classification of the content of the original image.

[0025] Each layer of the synapse is implemented using an array or a portion of an array of non-volatile memory cells.

[0026] FIG. 7 is a block diagram of an array that can be used for that purpose. A Vector-by-Matrix Multiplication (VMM) array 32 contains non-volatile memory cells and is utilized as a synapse between one layer and the next (such as CB1, CB2, CB3, and CB4 in FIG. 6). Specifically, the VMM array 32 includes an array of non-volatile memory cells 33, an erase gate and word line gate decoder 34, a control gate decoder 35, a bit line decoder 36, and a source line decoder 37, which decode the respective inputs to the non-volatile memory cell array 33. Inputs to the VMM array 32 can come from the erase gate and word line gate decoder 34 or from the control gate decoder 35. The source line decoder 37 in this embodiment also decodes the output of the non-volatile memory cell array 33. Alternatively, the bit line decoder 36 may decode the output of the non-volatile memory cell array 33.

[0027] The non-volatile memory cell array 33 serves two purposes. First, the non-volatile memory cell array 33 stores the weights used by the VMM array 32. Second, the non-volatile memory cell array 33 effectively multiplies the inputs by the weights stored in the non-volatile memory cell array 33 and adds them for each output line (source line or bit line) to generate an output that becomes the input to the next layer or the input to the last layer. With the non-volatile memory cell array 33 performing the multiplication and addition functions, the need for separate multiplication and addition logic is eliminated and in-memory computation is also more power efficient.

[0028] The outputs of the non-volatile memory cell array 33 are provided to a differential summer (such as a summing op-amp or summing current mirror) 38 that sums the outputs of the non-volatile memory cell array 33 to create a single value for the convolution. The differential summer 38 is arranged to perform summation of the positive and negative weights.

[0029] The summed output values ​​of the differential summers 38 are then fed to an activation function block 39 that normalizes the output. The activation function block 39 may provide a sigmoid, tanh, or ReLU function. The normalized output values ​​of the activation function block 39 become elements of a feature map as the next layer (e.g., C1 in FIG. 6) and are then applied to the next synapse to generate the next feature map layer or the last layer. Thus, in this embodiment, the non-volatile memory cell array 33 constitutes a number of synapses (receiving inputs from a previous layer of neurons or from an input layer such as an image database), and the summing op-amps 38 and the activation function block 39 constitute a number of neurons.

[0030] The inputs to the VMM array 32 of FIG. 7 (WLx, EGx, CGx, and optionally BLx and SLx) may be analog levels, binary levels, or digital bits (in which case a Digital-to-Analog Converter (DAC) is provided to convert the digital bits to appropriate input analog levels), and the outputs may be analog levels, binary levels, or digital bits (in which case an output Analog-to-Digital Converter (ADC) is provided to convert the output analog levels to digital bits).

[0031] FIG. 8 is a block diagram illustrating the use of multiple layers of VMM array 32, labeled in the figure as VMM arrays 32a, 32b, 32c, 32d, and 32e. As shown in FIG. 8, the input (represented by Inputx) is converted from digital to analog by a digital-to-analog converter 31 and fed to the input VMM array 32a. The converted analog input can be a voltage or a current. The first layer input D / A (Digital / Analog) conversion can be done by using a function or a LUT (Look Up Table) that maps the input Inputx to the appropriate analog level of the matrix multiplier of the input VMM array 32a. The input conversion can also be done by an analog to analog (A / A) converter to convert an external analog input to a mapped analog input to the input VMM array 32a.

[0032] The output generated by the input VMM array 32a is provided as an input to the next VMM array (hidden level 1) 32b, which in turn generates an output that is provided as an input to the next VMM array (hidden level 2) 32c, and so on. The various layers of the VMM arrays 32 function as different layers of synapses and neurons of a Convolutional Neural Network (CNN). Each VMM array 32a, 32b, 32c, 32d, and 32e may be a stand-alone physical non-volatile memory array, or the multiple VMM arrays may utilize different portions of the same physical non-volatile memory array, or the multiple VMM arrays may utilize overlapping portions of the same physical non-volatile memory array. 8 includes five layers (32a, 32b, 32c, 32d, 32e): one input layer (32a), two hidden layers (32b, 32c), and two fully connected layers (32d, 32e). Those skilled in the art will appreciate that this is merely an example, and that the system may alternatively include more than two hidden layers and more than two fully connected layers. <Vector-by-Matrix Multiplication (VMM) Array>

[0033] 9 shows a neuron VMM array 900 that is particularly suited for the memory cells 310 shown in FIG. 3 and is utilized as part of the synapses and neurons between the input layer and the next layer. The VMM array 900 includes a memory array 901 of non-volatile memory cells and a reference array 902 of non-volatile reference memory cells (at the top of the array). Alternatively, a separate reference array could be located at the bottom.

[0034] In the VMM array 900, the control gate lines, such as control gate line 903, run vertically (so that the row-wise reference array 902 is orthogonal to the control gate line 903), and the erase gate lines, such as erase gate line 904, run horizontally. Here, the input to the VMM array 900 is provided on the control gate lines (CG0, CG1, CG2, CG3), and the output of the VMM array 900 appears on the source lines (SL0, SL1). In one embodiment, only the even rows are used, and in another embodiment, only the odd rows are used. The current in each source line (SL0, SL1, respectively) performs a function of the sum of all the currents from the memory cells connected to that particular source line.

[0035] As described herein for neural networks, the non-volatile memory cells of the VMM array 900, i.e., the memory cells 310 of the VMM array 900, may be configured to operate in the sub-threshold region.

[0036] The non-volatile reference memory cells and non-volatile memory cells described herein are biased in weak inversion (sub-threshold region) as follows. Ids=Io * e (Vg-Vth) / nVt =w * Io * e (Vg) / nVt , In the formula, w=e (-Vth) / nVt and where Ids is the drain-source current, Vg is the gate voltage of the memory cell, Vth is the threshold voltage of the memory cell, and Vt is the thermal voltage = k * where T / q, k is the Boltzmann constant, T is temperature in Kelvin, q is the electron charge, n is the slope coefficient = 1 + (Cdep / Cox), Cdep = capacitance of the depletion layer, and Cox is the capacitance of the gate oxide layer, Io is the memory cell current at a gate voltage equal to the threshold voltage, and Io is (Wt / L) * u * Cox * (n-1) * Vt 2 where u is the carrier mobility, and Wt and L are the width and length of the memory cell, respectively.

[0037] In the case of an IV log converter that uses memory cells (such as reference or peripheral memory cells) or transistors to convert an input current to an input voltage, Vg=n * Vt * log[Ids / wp * Io] where wp is the w of the reference or periphery memory cell.

[0038] For a memory array used as a vector matrix multiplier VMM array with current inputs, the output current is: Iout=wa * Io * e (Vg) / nVt , i.e. Iout = (wa / wp) * Iin=W * Iin W=e (Vthp-Vtha) / nVt where wa is the w of each memory cell in the memory array. Vthp is the effective threshold voltage of the peripheral memory cells and Vtha is the effective threshold voltage of the main (data) memory cells. Note that the threshold voltage of a transistor is a function of the substrate body bias voltage, which is designated Vsb, and can be modulated to compensate for various conditions at such temperatures. The threshold voltage Vth can be expressed as: Vth = Vth0 + gamma (SQRT | Vsb - 2 * φF)-SQRT|2 * φF|) where Vth0 is the threshold voltage with zero substrate bias, φF is the surface potential, and gamma is the body effect parameter.

[0039] The word line or control gate can be used as an input to the memory cell for an input voltage.

[0040] Alternatively, the flash memory cells of the VMM arrays described herein may be configured to operate in the linear region. Ids=Beta * (Vgs-Vth) * Vds, beta = u * Cox * Wt / L W = α(Vgs-Vth) That is, the weight W in the linear region is proportional to (Vgs-Vth).

[0041] The word lines or control gates or bit lines or source lines can be used as inputs to memory cells operating in the linear region. The bit lines or source lines can be used as outputs of the memory cells.

[0042] For an IV linear converter, memory cells (such as reference or peripheral memory cells) or transistors operating in the linear region may be used to linearly convert input and output currents to input and output voltages.

[0043] Alternatively, the memory cells of the VMM arrays described herein may be configured to operate in the saturation region. Ids=1 / 2* beta * (Vgs-Vth) 2 , beta = u * Cox * Wt / L Wα(Vgs-Vth) 2 , that is, the weight W is (Vgs-Vth) 2 is proportional to

[0044] The word lines, control gates, or erase gates can be used as inputs to memory cells operating in the saturation region, and the bit lines or source lines can be used as outputs to output neurons.

[0045] Alternatively, the memory cells of the VMM arrays described herein may be used in all regions or combinations thereof (subthreshold, linear, or saturation) for each layer or layers of a neural network.

[0046] Other examples of the VMM array 32 of Figure 7 are described in U.S. Patent No. 10,748,630, which is incorporated herein by reference. As described in that application, the source lines or bit lines can be used as neuron outputs (current sum outputs).

[0047] FIG. 10 shows a neuron VMM array 1000, which is particularly suitable for the memory cells 210 shown in FIG. 2 and is used as a synapse between an input layer and the next layer. The VMM array 1000 includes a memory array 1003 of non-volatile memory cells, a reference array 1001 of a first non-volatile reference memory cell, and a reference array 1002 of a second non-volatile reference memory cell. The reference arrays 1001 and 1002 arranged in the columns of the array function to convert the current inputs flowing into terminals BLR0, BLR1, BLR2, and BLR3 into voltage inputs WL0, WL1, WL2, and WL3. In practice, the first and second non-volatile reference memory cells are diode-connected through a multiplexer 1014 (only partially shown) with the current inputs flowing into them. The reference cells are tuned (e.g., programmed) to a target reference level. The target reference level is provided by a reference mini-array matrix (not shown).

[0048] The memory array 1003 serves two purposes. First, it stores the weights in each memory cell that are used by the VMM array 1000. Second, it effectively multiplies the inputs (i.e., the current inputs provided to terminals BLR0, BLR1, BLR2, and BLR3, which the reference arrays 1001 and 1002 convert to input voltages that they provide to the word lines WL0, WL1, WL2, and WL3) by the weights stored in the memory array 1003, and then adds all the results (memory cell currents) to generate an output on each bit line (BL0-BLN) that is the input to the next layer or the input to the last layer. By performing the multiplication and addition functions, the memory array 1003 eliminates the need for separate multiplication and addition logic and is also power efficient. Here, voltage inputs are provided to word lines WL0, WL1, WL2, and WL3, and outputs appear on respective bit lines BL0-BLN during a read (inference) operation, with the current in each of the bit lines BL0-BLN performing a function of the sum of the currents from all the non-volatile memory cells connected to that particular bit line.

[0049] Table 5 shows the operating voltages and currents for the VMM array 1000. The columns in the table show the voltages applied to the word line of the selected cell, the word lines of the unselected cells, the bit lines of the selected cells, the bit lines of the unselected cells, the source lines of the selected cells, and the source lines of the unselected cells. The rows show the operations of read, erase, and program. Table 5: Operation of VMM Array 1000 in Figure 10 [Table 5]

[0050] FIG. 11 shows a neuron VMM array 1100 that is particularly suited for the memory cells 210 shown in FIG. 2 and is utilized as part of the synapses and neurons between the input layer and the next layer. The VMM array 1100 includes a memory array 1103 of non-volatile memory cells, a reference array 1101 of first non-volatile reference memory cells, and a reference array 1102 of second non-volatile reference memory cells. The reference arrays 1101 and 1102 extend in the row direction of the VMM array 1100. The VMM array is similar to the VMM 1000, except that the word lines extend vertically in the VMM array 1100. Here, the inputs are provided to the word lines (WLA0, WLB0, WLA1, WLB2, WLA2, WLB2, WLA3, WLB3) and the outputs appear on the source lines (SL0, SL1) during a read operation. The current of each source line performs a function of the sum of all the currents from the memory cells connected to that particular source line.

[0051] Table 6 shows the operating voltages and currents for VMM array 1100. The columns in the table show the voltages applied to the word line of the selected cell, the word lines of the unselected cells, the bit lines of the selected cells, the bit lines of the unselected cells, the source lines of the selected cells, and the source lines of the unselected cells. The rows show the read, erase, and program operations. Table 6: Operation of VMM Array 1100 in FIG. 11 [Table 6]

[0052] 12 shows a neuron VMM array 1200 that is particularly suited for the memory cells 310 shown in FIG. 3 and is used as part of synapses and neurons between the input layer and the next layer. The VMM array 1200 includes a memory array 1203 of non-volatile memory cells, a reference array 1201 of a first non-volatile reference memory cell, and a reference array 1202 of a second non-volatile reference memory cell. The reference arrays 1201 and 1202 function to convert current inputs flowing into terminals BLR0, BLR1, BLR2, and BLR3 into voltage inputs CG0, CG1, CG2, and CG3. In practice, the first and second non-volatile reference memory cells are diode-connected through multiplexers 1212 (only a portion of which is shown), with the current inputs flowing through BLR0, BLR1, BLR2, and BLR3. Each multiplexer 1212 includes a corresponding multiplexer 1205 and cascoding transistor 1204 to ensure a constant voltage on the bit line (e.g., BLR0) of each of the first and second non-volatile reference memory cells during a read operation. The reference cells are tuned to a target reference level.

[0053] The memory array 1203 serves two purposes. First, it stores the weights used by the VMM array 1200. Second, it effectively multiplies the weights stored in the memory array by the inputs (current inputs provided to terminals BLR0, BLR1, BLR2, and BLR3, which the reference arrays 1201 and 1202 convert to input voltages and provide to the control gates (CG0, CG1, CG2, and CG3)) and then adds all the results (cell currents) to produce an output that appears on BL0-BLN and is the input to the next layer or the input to the last layer. Having the memory array perform the multiplication and addition functions eliminates the need for separate multiplication and addition logic and is also power efficient, where the inputs are provided to the control gate lines (CG0, CG1, CG2, and CG3) and the outputs appear on the bit lines (BL0-BLN) during read operations. The current in each bit line is a function of the sum of all the currents from the memory cells connected to that particular bit line.

[0054] VMM array 1200 implements one-way tuning of non-volatile memory cells in memory array 1203. That is, each non-volatile memory cell is erased and then partially programmed until the desired charge on the floating gate is reached. If too much charge is added to the floating gate (such as an incorrect value being stored in the cell), the cell is erased and the series of partial programming operations starts over. As shown, two rows that share the same erase gate (such as EG0 or EG1) are erased together (known as a page erase), and then each cell is partially programmed until the desired charge on the floating gate is reached.

[0055] Table 7 shows the operating voltages and currents for VMM array 1200. The columns in the table show the voltages applied to the word line of the selected cell, the word lines of the unselected cells, the bit line of the selected cell, the bit lines of the unselected cells, the control gate of the selected cell, the control gates of the unselected cells in the same sector as the selected cell, the control gates of the unselected cells in a different sector than the selected cell, the erase gate of the selected cell, the erase gates of the unselected cells, the source line of the selected cell, and the source lines of the unselected cells. The rows show the read, erase, and program operations. Table 7: Operation of VMM Array 1200 in FIG. 12 [Table 7]

[0056] FIG. 13 shows a neuron VMM array 1300 that is particularly suited for the memory cells 310 shown in FIG. 3 and is utilized as part of synapses and neurons between the input layer and the next layer. The VMM array 1300 includes a memory array 1303 of non-volatile memory cells, a reference array 1301 or a first non-volatile reference memory cell, and a reference array 1302 of a second non-volatile reference memory cell. The EG lines EGR0, EG0, EG1, and EGR1 run vertically, and the CG lines CG0, CG1, CG2, and CG3 and the SL lines WL0, WL1, WL2, and WL3 run horizontally. The VMM array 1300 is similar to the VMM array 1400, except that the VMM array 1300 implements bidirectional tuning, and each individual cell can be fully erased, partially programmed, and partially erased as needed to reach a desired amount of charge on the floating gate through the use of individual EG lines. As shown, reference arrays 1301 and 1302 convert input currents at terminals BLR0, BLR1, BLR2 and BLR3 (through the operation of diode-connected reference cells via multiplexer 1314) into control gate voltages CG0, CG1, CG2 and CG3 that are applied to the memory cells in a row direction. The current outputs (neurons) are in bit lines BL0-BLN, each of which sums all the currents from the non-volatile memory cells connected to that particular bit line.

[0057] Table 8 shows the operating voltages and currents for VMM array 1300. The columns in the table show the voltages applied to the word line of the selected cell, the word lines of the unselected cells, the bit line of the selected cell, the bit lines of the unselected cells, the control gate of the selected cell, the control gates of the unselected cells in the same sector as the selected cell, the control gates of the unselected cells in a different sector than the selected cell, the erase gate of the selected cell, the erase gates of the unselected cells, the source line of the selected cell, and the source lines of the unselected cells. The rows show the read, erase, and program operations. Table 8: Operation of VMM Array 1300 in FIG. 13 [Table 8]

[0058] 22 shows a neuron VMM array 2200 that is particularly suited for the memory cells 210 shown in FIG. 2 and is used as part of the synapses and neurons between the input layer and the next layer. In the VMM array 2200, inputs INPUT0, ..., INPUT N are bit lines BL0, ...BL N and outputs OUTPUT1, OUTPUT2, OUTPUT3, and OUTPUT4 are generated on source lines SL0, SL1, SL2, and SL3, respectively.

[0059] 23 illustrates a neuron VMM array 2300 that is particularly suited for memory cells 210 shown in FIG. 2 and that is utilized as part of the synapses and neurons between an input layer and the next layer. In this embodiment, inputs INPUT0, INPUT1, INPUT2, and INPUT3 are received on source lines SL0, SL1, SL2, and SL3, respectively, and outputs OUTPUT0, ...OUTPUT N are bit lines BL0, ..., BL N is generated.

[0060] 24 shows a neuron VMM array 2400 that is particularly suited for the memory cells 210 shown in FIG. 2 and is used as part of the synapses and neurons between the input layer and the next layer. In this embodiment, inputs INPUT0, ..., INPUT M are the word lines WL0, ..., WL M Received at OUTPUT0, ...OUTPUT N are bit lines BL0, ..., BL N is generated.

[0061] 25 shows a neuron VMM array 2500 that is particularly suited for the memory cells 310 shown in FIG. 3 and is used as part of the synapses and neurons between the input layer and the next layer. In this embodiment, inputs INPUT0, ..., INPUT M are the word lines WL0, ..., WL M Received at OUTPUT0, ...OUTPUT N are bit lines BL0, ..., BL N is generated.

[0062] 26 shows a neuron VMM array 2600 that is particularly suited for the memory cells 410 shown in FIG. 4 and is used as part of the synapses and neurons between the input layer and the next layer. In this embodiment, inputs INPUT0, ..., INPUT n are the vertical control gate lines CG0, ..., CG N and outputs OUTPUT1 and OUTPUT2 are generated on source lines SL0 and SL1.

[0063] 27 shows a neuron VMM array 2700 that is particularly suited for the memory cells 410 shown in FIG. 4 and is used as part of the synapses and neurons between the input layer and the next layer. In this embodiment, inputs INPUT0, ..., INPUT N are bit lines BL0, ..., BL N, 2701-(N-1) and 2701-N, which are respectively coupled to bit line control gates 2701-1, 2701-2, ..., 2701-(N-1) and 2701-N. Exemplary outputs OUTPUT1 and OUTPUT2 are generated on source lines SL0 and SL1.

[0064] 28 shows a neuron VMM array 2800 that is particularly suitable for memory cells 310 shown in FIG. 3, memory cells 510 shown in FIG. 5, and memory cells 710 shown in FIG. 7, and is used as part of synapses and neurons between the input layer and the next layer. In this embodiment, inputs INPUT0, ..., INPUT M are the word lines WL0, ..., WL M Received at OUTPUT0, ..., OUTPUT N are bit lines BL0, ..., BL N is generated.

[0065] 29 shows a neuron VMM array 2900 that is particularly suitable for memory cells 310 shown in FIG. 3, memory cells 510 shown in FIG. 5, and memory cells 710 shown in FIG. 7, and is used as part of synapses and neurons between the input layer and the next layer. In this embodiment, inputs INPUT0, ..., INPUT M are the control gate lines CG0, ..., CG M The outputs are OUTPUT0, ..., OUTPUT N are the vertical source lines SL0, ..., SL N , and each source line SL i is coupled to the source lines of all memory cells in column i.

[0066] 30 shows a neuron VMM array 3000 that is particularly suitable for memory cells 310 shown in FIG. 3, memory cells 510 shown in FIG. 5, and memory cells 710 shown in FIG. 7, and is used as part of synapses and neurons between the input layer and the next layer. In this embodiment, inputs INPUT0, ..., INPUT M are the control gate lines CG0, ..., CGM The outputs are OUTPUT0, ..., OUTPUT N are vertical bit lines BL0, ..., BL N , and each bit line BL i is coupled to the bit lines of all memory cells in column i. <Long-term and short-term memory>

[0067] The prior art includes a concept known as Long Short-Term Memory (LSTM). LSTM units are often used within neural networks. LSTM allows neural networks to store information for any predefined period of time and use that information in subsequent operations. A traditional LSTM unit includes a cell, an input gate, an output gate, and a forget gate. The three gates regulate the flow of information into and out of the cell, as well as the period of time information is stored within the LSTM. VMMs are particularly useful in LSTM units.

[0068] FIG. 14 illustrates an exemplary LSTM 1400. The LSTM 1400 in this embodiment includes cells 1401, 1402, 1403, and 1404. Cell 1401 receives an input vector x0 and generates an output vector h0 and a cell state vector c0. Cell 1402 receives an input vector x1, an output vector (hidden state) h0 from cell 1401, and a cell state c0 from cell 1401, and generates an output vector h1 and a cell state vector c1. Cell 1403 receives an input vector x2, an output vector (hidden state) h1 from cell 1402, and a cell state c1 from cell 1402, and generates an output vector h2 and a cell state vector c2. Cell 1404 receives an input vector x3, an output vector (hidden state) h2 from cell 1403, and a cell state c2 from cell 1403, and generates an output vector h3. Additional cells can be used, the LSTM with four cells is just an example.

[0069] Figure 15 shows an example implementation of an LSTM cell 1500 that can be used for cells 1401, 1402, 1403, and 1404 in Figure 14. The LSTM cell 1500 receives an input vector x(t), a cell state vector c(t-1) from a previous cell, and an output vector h(t-1) from a previous cell, and produces a cell state vector c(t) and an output vector h(t).

[0070] LSTM cell 1500 includes sigmoid function devices 1501, 1502, and 1503, each of which applies a number between 0 and 1 to control the degree to which each component of the input vector contributes to the output vector. LSTM cell 1500 also includes tanh devices 1504 and 1505 for applying a hyperbolic tangent function to the input vector, multiplier devices 1506, 1507, and 1508 for multiplying two vectors, and adder device 1509 for adding the two vectors. The output vector h(t) may be provided to the next LSTM cell in the system or may be accessed for other purposes.

[0071] FIG. 16 shows an example of an implementation of LSTM cell 1500, LSTM cell 1600. For the convenience of the reader, the same numbering scheme from LSTM cell 1500 is used in LSTM cell 1600. Sigmoid function devices 1501, 1502, and 1503, and tanh device 1504 each include multiple VMM arrays 1601 and activation function blocks 1602. Thus, VMM arrays prove particularly useful in LSTM cells used in certain neural network systems. Multiplier devices 1506, 1507, and 1508, and adder device 1509 are implemented in a digital or analog manner. Activation function block 1602 can be implemented in a digital or analog manner.

[0072] An alternative embodiment of LSTM cell 1600 (and another embodiment of one implementation of LSTM cell 1500) is shown in Figure 17. In Figure 17, sigmoid function devices 1501, 1502, and 1503, and tanh device 1504 share the same physical hardware (VMM array 1701 and activation function block 1702) in a time-multiplexed manner. LSTM cell 1700 also includes a multiplier device 1703 for multiplying two vectors, an adder device 1708 for adding two vectors, a tanh device 1505 (which includes activation function block 1702), a register 1707 for storing the value i(t) when i(t) is output from sigmoid function block 1702, and a register 1708 for storing the value f(t). * A register 1704 for storing c(t-1) as its value is output from the multiplier device 1703 via a multiplexer 1710; * A register 1705 for storing u(t) as its value is output from the multiplier device 1703 via a multiplexer 1710; * It includes a register 1706 for storing {tilde over (c)}(t) as its value is output from the multiplier device 1703 via a multiplexer 1710, and a multiplexer 1709.

[0073] Whereas LSTM cell 1600 includes multiple sets of VMM arrays 1601 and respective activation function blocks 1602, LSTM cell 1700 includes only one set of VMM arrays 1701 and activation function blocks 1702, which are used to represent multiple layers in an embodiment of LSTM cell 1700. LSTM cell 1700 requires less space than LSTM cell 1600 because LSTM cell 1700 requires ¼ the space for the VMMs and activation function blocks compared to LSTM cell 1600.

[0074] It can be further appreciated that an LSTM unit typically includes multiple VMM arrays, each of which requires functionality provided by specific circuit blocks outside the VMM array, such as summer and activation function blocks and high voltage generation blocks. Providing separate circuit blocks for each VMM array requires a significant amount of space within a semiconductor device and is somewhat inefficient. Thus, the embodiments described below reduce the circuitry required outside the VMM array itself. <Gated Recurrent Unit>

[0075] Analog VMM implementations can be used for GRU (Gated Recurrent Unit) systems. GRUs are the gating mechanism in recurrent neural networks. GRUs are similar to LSTMs, except that GRU cells generally contain fewer components than LSTM cells.

[0076] 18 illustrates an exemplary GRU 1800. The GRU 1800 in this example includes cells 1801, 1802, 1803, and 1804. Cell 1801 receives input vector x0 and generates output vector h0. Cell 1802 receives input vector x1 and output vector h0 from cell 1801 and generates output vector h1. Cell 1803 receives input vector x2 and output vector (hidden state) h1 from cell 1802 and generates output vector h2. Cell 1804 receives input vector x3 and output vector (hidden state) h2 from cell 1803 and generates output vector h3. Additional cells may be used, and a GRU with four cells is merely an example.

[0077] FIG. 19 illustrates an example implementation of a GRU cell 1900 that may be used for cells 1801, 1802, 1803, and 1804 of FIG. 18. GRU cell 1900 receives an input vector x(t) and an output vector h(t-1) from a preceding GRU cell and generates an output vector h(t). GRU cell 1900 includes sigmoid function devices 1901 and 1902, each of which applies a number between 0 and 1 to the output vector h(t-1) and components from the input vector x(t). GRU cell 1900 also includes a tanh device 1903 for applying a hyperbolic tangent function to the input vector, a number of multiplier devices 1904, 1905, and 1906 for multiplying two vectors, an adder device 1907 for adding the two vectors, and a complementary device 1908 for subtracting the input from 1 to generate an output.

[0078] FIG. 20 shows a GRU cell 2000, which is an example of one implementation of the GRU cell 1900. For the convenience of the reader, the same numbering method from the GRU cell 1900 is used in the GRU cell 2000. As can be seen from FIG. 20, the sigmoid function devices 1901 and 1902 and the tanh device 1903 each include a plurality of VMM arrays 2001 and activation function blocks 2002. Therefore, it can be seen that the VMM array is particularly used in the GRU cell used in a specific neural network system. The multiplier devices 1904, 1905, 1906, the adder device 1907, and the complementary device 1908 are implemented in a digital or analog manner. The activation function block 2002 can be implemented in a digital or analog manner.

[0079] An alternative embodiment of GRU cell 2000 (and another embodiment of one implementation of GRU cell 1900) is shown in FIG. 21. In FIG. 21, GRU cell 2100 utilizes VMM array 2101 and activation function block 2102, which, when configured as a sigmoid function, applies a number between 0 and 1 to control how much each component of the input vector contributes to the output vector. In FIG. 21, sigmoid function devices 1901 and 1902, and tanh device 1903 share the same physical hardware (VMM array 2101 and activation function block 2102) in a time-multiplexed manner. GRU cell 2100 also includes a multiplier device 2103 for multiplying two vectors together, an adder device 2105 for adding two vectors together, a complementary device 2109 for subtracting the input from 1 to generate the output, a multiplexer 2104, and a value h(t-1) * A register 2106 for holding r(t) as its value is output from the multiplier device 2103 via a multiplexer 2104, and a value h(t-1) * A register 2107 for holding z(t) as its value is output from the multiplier device 2103 via a multiplexer 2104, and a value ĥ(t) * and a register 2108 for holding (1-z((t)) as that value is output from the multiplier device 2103 via multiplexer 2104.

[0080] Whereas GRU cell 2000 includes multiple sets of VMM arrays 2001 and activation function blocks 2002, GRU cell 2100 includes only one set of VMM arrays 2101 and activation function blocks 2102, which are used to represent multiple tiers in an embodiment of GRU cell 2100. GRU cell 2100 requires less space than GRU cell 2000 because GRU cell 2100 requires one-third the space for the VMM and activation function blocks compared to GRU cell 2000.

[0081] It can be further appreciated that a GRU system typically includes multiple VMM arrays, each of which requires functionality provided by specific circuit blocks outside the VMM array, such as summer and activation function blocks and high voltage generation blocks. Providing separate circuit blocks for each VMM array requires a significant amount of space within a semiconductor device and is somewhat inefficient. Thus, the embodiments described below reduce the circuitry required outside the VMM array itself.

[0082] The input to the VMM array can be an analog level, a binary level, a pulse, a time modulated pulse, or a digital bit (in which case a DAC is required to convert the digital bit to the appropriate input analog level), and the output can be an analog level, a binary level, a timing pulse, a pulse, or a digital bit (in which case an output ADC is required to convert the output analog level to a digital bit).

[0083] In general, for each memory cell in the VMM array, each weight W can be implemented by a single memory cell, or by a differential cell, or by two blended memory cells (average of two cells). In the case of a differential cell, two memory cells are required to implement weight W as a differential weight (W=W+-W-). In the case of two blended memory cells, two memory cells are required to implement weight W as an average of two cells.

[0084] FIG. 31 illustrates a VMM system 3100. In some embodiments, the weights W stored in the VMM array are stored as a differential pair, W+ (positive weight) and W- (negative weight), where W=(W+)-(W-). In the VMM system 3100, half of the bit lines are designated as W+ lines, i.e., bit lines that connect to memory cells that will store a positive weight W+, and the other half of the bit lines are designated as W- lines, i.e., bit lines that connect to memory cells that implement a negative weight W-. The W- lines are interspersed alternately among the W+ lines. The subtraction operation is performed by summing circuits, such as summing circuits 3101 and 3102, that receive currents from the W+ and W- lines. The output of the W+ lines and the output of the W- lines are combined together to effectively yield W=W+-W- for each pair of (W+, W-) cells of every pair of (W+, W-) lines. Although described thus far with respect to W- lines interspersed alternatingly among W+ lines, in other embodiments the W+ and W- lines may be arbitrarily located anywhere within the array.

[0085] 32 shows another embodiment. In a VMM system 3210, the positive weights W+ are implemented in a first array 3211 and the negative weights W- are implemented in a second array 3212 that is separate from the first array, and the resulting weights are suitably combined together by a summing circuit 3213.

[0086] FIG. 33 illustrates a VMM system 3300. The weights W stored in the VMM array are stored as a differential pair, W+ (positive weight) and W- (negative weight), where W=(W+)-(W-). The VMM system 3300 includes an array 3301 and an array 3302. Half of the bit lines in each of the arrays 3301 and 3302 are designated as W+ lines, i.e., bit lines that connect to memory cells that store a positive weight W+, and the other half of the bit lines in each of the arrays 3301 and 3302 are designated as W- lines, i.e., bit lines that connect to memory cells that implement a negative weight W-. The W- lines are interspersed alternately among the W+ lines. Subtraction operations are performed by adder circuits, such as adder circuits 3303, 3304, 3305, and 3306, that receive current from the W+ and W- lines. The outputs on the W+ and W- lines from each array 3301, 3302, respectively, are combined together to effectively yield W=W+-W- for each pair of (W+,W-) cells of every pair of (W+,W-) lines. Note that the W values ​​from each array 3301 and 3302 may be further combined via summing circuits 3307 and 3308, meaning that each W value is the result of subtracting the W value from array 3302 from the W value from array 3301, and the final result from summing circuits 3307 and 3308 is one difference value of the two difference values.

[0087] Each non-volatile memory cell used in an analog neural memory system will be erased and programmed to hold a very specific and precise amount of charge, i.e., number of electrons, on its floating gate. For example, each floating gate should hold one of N different values, where N is the number of different weights that can be represented by each cell. Examples of N include 16, 32, 64, 128, and 256.

[0088] 34 shows a block diagram of a VMM system 3400. The VMM system 3400 includes a VMM array 3401, a row decoder 3402, a high voltage decoder 3403, a column decoder 3404, a bit line driver 3405, an input circuit 3406, an output circuit 3407, a control logic 3408, and a bias generator 3409. The VMM system 3400 further includes a high voltage generation block 3410 including a charge pump 3411, a charge pump regulator 3412, and a high voltage analog precision level generator 3413. The VMM system 3400 further includes a (program / erase, or weight tuning) algorithm controller 3414, an analog circuit 3415, a control engine 3416 (which may include special functions such as, but not limited to, arithmetic functions, activation functions, embedded microcontroller logic, etc.), and a test control logic 3417.

[0089] The input circuit 3406 may include circuits such as a DAC (Digital to Analog Converter), a DPC (Digital to Pulses Converter), an AAC (Analog to Analog Converter, such as a current-to-voltage converter, a logarithmic converter, etc.), a PAC (Pulse to Analog level Converter), or any other type of converter. The input circuit 3406 may implement one or more of normalization, linear or nonlinear upscaling / downscaling functions, or arithmetic functions. The input circuit 3406 may implement a temperature compensation function for the input level. The input circuit 3406 may implement an activation function such as ReLU or sigmoid.

[0090] The output circuit 3407 may include circuits such as an ADC (Analog to Digital Converter, for converting neuron analog output to digital bits), an AAC (Analog to Analog Converter, such as a current-to-voltage converter, a logarithmic converter), an APC (Analog to Pulse Converter, an Analog to Time Modulated Pulse Converter), or any other type of converter. The output circuit 3407 may implement activation functions such as a Rectified Linear Activation Function (ReLU) or a sigmoid. The output circuit 3407 may implement one or more of statistical normalization, regularization, upscaling / downscaling / gain functions, statistical rounding, or arithmetic functions (e.g., addition, subtraction, division, multiplication, shift, log) of the neuron output. The output circuit 3407 may implement a temperature compensation function for the neuron output or array output (such as the bit line output) to keep the power consumption of the array approximately constant, or to increase the accuracy of the array (neuron) output, such as by keeping the IV slope approximately the same.

[0091] As artificial neural network applications become more complex, the need for larger VMM arrays increases, while at the same time efficiently using the space within the packaged integrated circuit and conserving power wherever possible while maintaining accuracy so that each of the N distinct weights can still be properly stored and read out.

[0092] . Summary of the Invention

[0093] A number of embodiments are described that provide an artificial neural network system that comprises a three-dimensional integrated circuit that comprises one or more arrays of VMMs.

[0094]

[0095]

[0096]

[0097]

[0098]

[0099]

[0100]

[0101]

[0102]

[0103]

[0104]

[0105]

[0106]

[0107]

[0108]

[0109]

[0110]

[0111]

[0112]

[0113]

[0114]

[0115]

[0116]

[0117]

[0118]

[0119]

[0120]

[0121]

[0122]

[0123]

[0124]

[0125]

[0126]

[0127]

[0128]

[0129]

[0130]

[0131]

[0132]

[0133]

[0134]

[0135]

[0136]

[0137]

[0138]

[0139]

[0140]

[0141]

[0142]

[0143]

[0144]

[0145]

[0146]

[0147] [Brief description of the drawings]

[0148] [Figure 1] FIG. 1 illustrates an artificial neural network. [Diagram 2] 1 shows a prior art split-gate flash memory cell. [Diagram 3] 1 illustrates another prior art split-gate flash memory cell. [Figure 4] 1 illustrates another prior art split-gate flash memory cell. [Diagram 5] 1 illustrates another prior art split-gate flash memory cell. [Figure 6] FIG. 1 illustrates various levels of an exemplary artificial neural network that utilizes one or more non-volatile memory arrays. [Figure 7] FIG. 2 is a block diagram showing a VMM system. [Figure 8] FIG. 1 is a block diagram illustrating an example artificial neural network utilizing one or more VMM systems. [Figure 9] 2 illustrates another embodiment of a VMM system. [Figure 10] 2 illustrates another embodiment of a VMM system. [Figure 11] 2 illustrates another embodiment of a VMM system. [Figure 12] 2 illustrates another embodiment of a VMM system. [Figure 13] 2 illustrates another embodiment of a VMM system. [Figure 14] 1 illustrates a prior art long and short term memory system. [Figure 15] 1 illustrates an exemplary cell for use in a long-short-term memory system. [Figure 16] 16 illustrates an example implementation of the cell of FIG. 15. [Figure 17] 16 illustrates another exemplary implementation of the cell of FIG. 15. [Figure 18]1 shows a prior art gated recurrent unit system. [Figure 19] 1 shows an exemplary cell for use in a gated recurrent unit system. [Figure 20] 20 illustrates an example implementation of the cell of FIG. 19. [Figure 21] 20 illustrates another exemplary implementation of the cell of FIG. 19. [Figure 22] 2 illustrates another embodiment of a VMM system. [Figure 23] 2 illustrates another embodiment of a VMM system. [Figure 24] 2 illustrates another embodiment of a VMM system. [Diagram 25] 2 illustrates another embodiment of a VMM system. [Figure 26] 2 illustrates another embodiment of a VMM system. [Figure 27] 2 illustrates another embodiment of a VMM system. [Figure 28] 2 illustrates another embodiment of a VMM system. [Figure 29] 2 illustrates another embodiment of a VMM system. [Diagram 30] 2 illustrates another embodiment of a VMM system. [Diagram 31] 2 illustrates another embodiment of a VMM system. [Diagram 32] 2 illustrates another embodiment of a VMM system. [Diagram 33] 2 illustrates another embodiment of a VMM system. [Diagram 34] 1 illustrates an embodiment of a 2D VMM system. [Diagram 35] 1 illustrates an embodiment of a 3D VMM system. [Diagram 36] 1 illustrates an embodiment of a 3D VMM system. [Figure 37A] 1 illustrates an embodiment of a 3D VMM system. [Figure 37B] 1 illustrates an embodiment of a 3D VMM system. [Figure 38] 1 illustrates an embodiment of a 3D VMM system. [Figure 39A]1 illustrates an embodiment of a 3D VMM system. [Figure 39B] 1 illustrates an embodiment of a 3D VMM system. [Figure 39C] 1 illustrates an embodiment of a 3D VMM system. [Diagram 40] 1 illustrates an embodiment of a 3D VMM system. [Diagram 41] 1 illustrates an embodiment of a 3D VMM system. [Diagram 42] 1 illustrates an embodiment of a 3D VMM system. [Diagram 43] 1 illustrates an embodiment of a 3D VMM system. [Diagram 44] 1 illustrates an embodiment of a 3D VMM system. [Diagram 45] 1 illustrates an embodiment of a 3D VMM system. [Diagram 46] 1 illustrates an embodiment of a 3D VMM system. [Figure 47] 1 illustrates an embodiment of a 3D VMM system. [Figure 48] 1 illustrates an embodiment of a 3D VMM system. [Figure 49] 1 shows an embodiment of a neuron circuit. [Figure 50] 1 shows an embodiment of a neuron circuit. [Figure 51] 2 shows an embodiment of a driving circuit. [Figure 52] 1 shows an example of a different neuron circuit. [Diagram 53] 1 illustrates one embodiment of a sample-and-hold buffer. [Figure 54] 1 illustrates an embodiment of an ADC circuit. [Figure 55] 1 illustrates an embodiment of an ADC circuit. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0149] <3D VMM system architecture> Figure 35 illustrates a 3D VMM system 3500 comprising multiple dies, such as dies 3501, 3502, 3503, 3504, 3505, and 3506, stacked vertically within package 3522 to form a packaged integrated circuit. 3D VMM system 3500 includes certain functional blocks that are functionally similar to blocks included in VMM system 3400 of Figure 34, although these blocks may be located on different dies, where components included in dies 3501 and 3502 share components included in dies 3503, 3504, 3505, and 3506.

[0150] In this example, the die 3501 includes a respective VMM array 3507 (functionally similar to the VMM 3401 of FIG. 34), a respective input multiplexer 3509, a respective row buffer 3523 (which may, for example, provide a sampled and held buffered voltage to the array input), a respective high voltage decoder 3508 (functionally similar to the high voltage decoder 3403 of FIG. 34), and a respective neuron circuit 3510 (which may, for example, but not limited to, perform array output current scaling functions, min / max limiting functions, differential output conversion, buffering). The input multiplexer 3509 receives and applies analog input signals to the VMM array 3507, and the neuron circuit 3510 receives analog output signals representing neuron outputs from the VMM array 3507. The output signals may be sent to other blocks in the 3D VMM system 3500.

[0151] Die 3502 also includes a respective VMM array 3507 , a respective input multiplexer 3509 , a respective row buffer 3522 , a respective high voltage decoder 3508 , and a respective neuron circuit 3510 .

[0152] In this example, two dies (dies 3501 and 3502) include respective VMM arrays 3507, although it should be understood that additional dies including respective VMM arrays may be included.

[0153] Die 3503 includes high voltage generator 3511 (functionally similar to high voltage generation block 3410 of FIG. 34), analog circuit 3512 (functionally similar to analog circuit 3415 of FIG. 34), and temperature compensation circuit 3513. 3D VMM system 3500 may have thermal challenges that 2D VMM system 3400 does not have. Because dies 3501, 3502, 3503, 3504, 3505, and 3506 are stacked in a vertical configuration and include different types of circuits, each die may experience different thermal operating conditions during operation. For example, some dies may get hotter than others, and the rate of temperature rise may vary between the dies. This introduces the potential for inaccuracies due to thermal changes. Temperature compensation circuit 3513 compensates for changes in temperature experienced between the various dies. Optionally, one or more thermal sensors are located on each die to provide temperature data to temperature compensation circuit 3513. Temperature compensation circuit 3513 then changes trim or configuration settings to compensate for any changes in temperature. Temperature compensation circuit 3513 also compensates for temperature changes on each die to compensate for cell current changes over temperature, such as making the resulting bitline (neuron) currents approximately the same over temperature. Temperature compensation circuit 3513 is also used for neuron circuit 3510, DAC, and ADC circuits to make the operating dynamic ranges (e.g., DAC output range, ADC input range, neuron circuit 3510 output range) approximately the same over temperature.

[0154] Die 3504 includes input circuitry 3514 (similar in function to input circuitry 3406) that includes address decode circuitry 3524, row registers 3525 (which hold activation input values ​​for array rows), and digital-to-analog converters (DACs) 3515. DACs 3515 receive digital signals from row registers 3525 and convert them to analog signals.

[0155] Die 3505 includes an analog-to-digital converter (ADC) 3516. ADC 3516 receives analog signals and converts them to digital signals.

[0156] The die 3506 includes digital circuitry 3517, static random access memory (SRAM) 3518, registers 3519, physical I / O (Input / Output) connections 3520, digital accelerator 3531, and network-on-chip (NOC) 3715. The die 3506 provides control functions to the other dies. The digital circuitry 3517 may include digital logic, a microcontroller, a single instruction multiple data (SIMD) processor, and a processor. The SRAM 3518 and registers 3519 may be used to store system and configuration information used by the digital circuitry 3517 or other circuits or blocks within the 3D VMM system 3500. The physical I / O connections 3520 provide an IO interface to devices external to the VMM system 3500 (such as an external processing unit) or to another package 3522. Digital accelerator 3531 is used for certain neural networks or certain layers within neural networks where additional processing may be required, such as, but not limited to, when small activation sizes exist, when weights stored in cells are dynamic and not fixed, when MAC operations need to be performed, etc. _.NOC 3715 provides network routing functionality within 3D VMM system 3500, for example, by generating control signals that route signals from one block to another.

[0157] Each of the multiple dies is connected to one or more other dies in the multiple dies through vertical interfaces 3521, each connecting two or more dies together. In one embodiment, the vertical interfaces 3521 are implemented as Through-Silicon Vias (TSVs).

[0158] During a read operation of the 3D VMM system 3500, digital inputs are received by the input circuitry 3514. The digital inputs enable row registers 3525, which store activation inputs and apply selected activation inputs in response to the digital inputs to DACs 3515, which convert the digital outputs from the row registers 3525 into respective analog signals. The analog signals generated by the DACs 3515 are provided by the input circuitry 3514 through one or more vertical interfaces 3521 to input multiplexers 3509 and row buffers 3523 on one or more of the dies 3501, 3502, which then apply these signals to one or more rows in a respective VMM array 3507, resulting in outputs being generated by the VMM array 3507. The output from each VMM array 3507 is received by a respective neuron circuit 3510, which provides a buffer function to drive the parasitic capacitance of one or more vertical interfaces 3521 to which the neuron circuit 3510 connects. The neuron circuit 3510 provides an analog signal via one or more vertical interfaces 3521 to an ADC 3516 of the die 3505, which converts the analog signal to a digital signal. Alternatively, the analog signal may bypass the ADC 3516 and remain in analog form. The output of the ADC 3516 is provided via the physical I / O 3520 to a device external to the 3D VMM system 3500 (such as a processing unit or a graphics processing unit) or is applied as an input to the respective VMM array 3507 (representing another layer in the artificial neural network). Alternatively, the analog signals from the neuron circuits 3510 can bypass the ADC 3516, remain in analog form, and be applied as inputs to the respective VMM arrays 3507.

[0159] Figure 36 shows a 3D VMM system 3600 that includes package 3622. 3D VMM system 3600 is similar to 3D VMM system 3500 and includes many of the same components, except for some differences in the arrangement of components and certain additional components. Items that are the same as in Figure 35 include the same item numbers as in Figure 36, where the components included in dies 3601 and 3602 share components included in dies 3603, 3604, 3605, and 3606.

[0160] 3D VMM system 3600 comprises multiple dies, such as dies 3601, 3602, 3603, 3604, 3605, and 3606, stacked vertically within a common package 3522 to form a packaged integrated circuit.

[0161] In this embodiment, die 3601 includes respective VMM arrays 3507, respective input multiplexers 3509, respective registers 3524 (which hold the activation input values ​​for the array rows), respective row buffers 3523, respective high voltage multiplexers 3608, and respective column multiplexers 3610.

[0162] Die 3602 also includes a respective VMM array 3507 , a respective input multiplexer 3509 , a respective register 3524 , a respective row buffer 3523 , a respective high voltage multiplexer 3608 , and a column multiplexer 3610 .

[0163] In this example, two dies (dies 3601 and 3602) include respective VMM arrays 3507, although it should be understood that additional dies including VMM arrays may be included.

[0164] Die 3603 includes a high voltage generator 3511 , an analog circuit 3512 , a temperature compensation circuit 3513 , and a high voltage multiplexer 3608 .

[0165] Die 3604 includes input circuitry 3614 that includes address decode 3524 and DAC 3515 .

[0166] The die 3605 includes a neuron circuit 3510 and an ADC 3516 .

[0167] Die 3606 includes digital circuitry 3517, SRAM 3518, registers 3519, and physical I / O connections 3520.

[0168] Each of the multiple dies is connected via a respective vertical interface 3521 to one or more other dies in the multiple dies.

[0169] During a read operation of the 3D VMM system 3600, an input is received, an address is received by the input circuitry 3614. The address decoder 3524 decodes the address and provides an input to a respective row register 3525 corresponding to the decoded address via one or more vertical interfaces 3521. The output of each row register 3525 is coupled to a DAC 3515 via one or more vertical interfaces 3521, which converts the digital output bits received from the row register 3525 into an analog signal and provides the analog signal via one or more vertical interfaces 3521 to a respective input multiplexer 3509 and row buffer 3523. The row buffer 3523 applies the buffered analog signal to a row input of the respective VMM array 3507 for a selected row (e.g., array control gate or word line). Output from each VMM array 3507 (such as from an array bitline) is received by a respective column multiplexer 3610, the output of which provides a signal via one or more vertical interfaces 3521 to an ADC 3516 of die 3605, which converts the analog signal to a digital signal. Alternatively, the signal may bypass ADC 3516 and remain in analog form. The output of ADC 3516 is provided to digital circuitry 3517 (which may perform activation functions, pooling functions, or other network functions) of die 3606 via one or more vertical interfaces 3521, and the output of digital circuitry 3517 may be provided as input to another respective VMM array 3507 (representing another layer in the artificial neural network) via respective vertical interfaces 3521, through physical I / O 3520 of die 3606, to a device external to 3D VMM system 3600 (such as a processing unit or graphics processing unit), or to input circuitry 3614 of die 3604.

[0170] 37A illustrates a 3D VMM system 3700 comprising multiple dies in two or more vertical stacks. In this example, a first vertical stack comprises dies 3701, 3702, 3703, 3704, 3705, and 3706, and a second vertical stack comprises dies 3707, 3708, 3709, 3710, 3711, and 3712, all of which are included in a common package 3522 to form a single packaged integrated circuit. In one example, the dies in the first vertical stack are physically separate dies from the dies in the second vertical stack. In another example, the dies in the first vertical stack and the dies in the second vertical stack are the same physical die (meaning, for example, that dies 3701 and 3707 are the same die). Here, the components included in dies 3701 and 3702 share the components included in dies 3703, 3704, 3705, and 3706, and the components included in dies 3707 and 3708 share the components included in dies 3709, 3710, 3711, and 3712.

[0171] In the illustrated embodiment, die 3701 and die 3707 each include a respective VMM array 3507, a respective array input 3729 (including a respective input multiplexer 3509, a respective address decoder 3713, a respective row register 3525, and a respective row buffer 3523), a respective high voltage multiplexer 3608, and a respective neuron circuit 3510.

[0172] Die 3702 and die 3708 also each include a respective VMM array 3507, a respective array input 3729 (including a respective input multiplexer 3509, a respective address decoder 3713, a respective row register 3525, and a respective row buffer 3523), a respective high voltage multiplexer 3608, and a respective neuron circuit 3510.

[0173] In this example, four dies (dies 3701, 3702, 3707, and 3708) include respective VMM arrays 3507, although it should be understood that additional dies including VMM arrays may be included.

[0174] Die 3703 and die 3709 each include a respective high voltage decoder 3714 , a respective high voltage generator 3511 , a respective analog circuit 3512 , and a respective temperature compensation circuit 3513 .

[0175] Die 3704 and die 3710 each include a respective input circuit 3514 that includes a respective DAC 3515 .

[0176] Die 3705 and die 3711 each include a respective ADC 3516.

[0177] Die 3706 and die 3712 each include a respective digital circuit 3517 , a respective SRAM 3518 , a respective register 3519 , a respective physical I / O connection 3520 , and a respective NOC (network on chip) connection 3715 .

[0178] Each of the multiple dies is connected to one or more other dies in the multiple dies via one or more vertical interfaces 3521 or horizontal interfaces 3716, which respectively connect two or more dies together. In one embodiment, each of the vertical interfaces 3521 is a through silicon via (TSV). In one embodiment, each of the horizontal interfaces 3716 is a redistribution layer (RDL) connection.

[0179] During a read operation of the 3D VMM system 3700, a digital input is received by a respective input circuit 3514, which converts the digital input to an analog signal via its respective DAC 3515, and the output of each DAC 3515 is coupled to a row register 3525 of a respective array input 3729 via a respective vertical interface 3521 and / or horizontal interface 3716. The output of the row register 3525 of each array input 3729 is provided to an input multiplexer 3509 and a row buffer 3523, which then apply a signal to one or more rows in the VMM array 3507. The output from each VMM array 3507 is received by a respective neuron circuit 3510, which provides a buffer function to drive the parasitic capacitance of one or more vertical interfaces 3521 or horizontal interfaces 3716 to which the neuron circuit 3510 connects. The neuron circuits 3510 provide the buffered analog signals via one or more vertical interfaces 3521 or horizontal interfaces 3716 to respective ADCs 3516 that convert the analog signals to digital signals. The outputs of the ADCs 3516 are provided to respective digital circuits 3517 (performing activation, pooling or network functions), the outputs of which may be provided via respective physical I / Os 3520 to devices external to the 3D VMM system 3700 (such as a processing unit or a graphics processing unit) or to respective input circuits 3514 that are converted by respective DACs 3515, the outputs of which are coupled via respective physical I / Os 3520 to another VMM array 3507 or another package 3522 (representing another layer in the artificial neural network).

[0180] FIG. 37B shows a 3D VMM system 3750 that is similar to 3D VMM system 3750, except that it has another type of VMM array, shown on die 3758 as VMM array 3557, which includes static RAM cells or dynamic RAM cells.

[0181] FIG. 38 illustrates a 3D VMM system 3800 with multiple dies in two vertical stacks. It is possible to have more than two stacks. In this example, a first vertical stack includes dies 3801, 3802, 3803, 3804, 3805, and 3806, and a second vertical stack includes dies 3807, 3808, 3809, 3810, 3811, and 3812, all of which are included in a common package 3522 to form a single packaged integrated circuit. In one example, the dies in the first vertical stack are physically separate dies from the dies in the second vertical stack. In another example, the dies in the first vertical stack and the dies in the second vertical stack are the same physical die (meaning, for example, that dies 3801 and 3807 are the same die). Here, the components included in dies 3801 and 3802 share the components included in dies 3803, 3804, 3805, and 3806, and the components included in dies 3807 and 3808 share the components included in dies 3809, 3810, 3811, and 3812.

[0182] In the illustrated embodiment, die 3801, 3802, 3807, and die 3808 each include a respective VMM array 3507, a respective array input 3729, a respective high voltage multiplexer 3608, a respective column multiplexer 3610, and a respective array input circuit 3729. Each array input 3729 includes an input multiplexer 3509, an address decoder 3713, a row register 3524, and a row buffer 3523.

[0183] In this example, four dies (dies 3801, 3802, 3807, and 3808) contain VMM arrays, although it should be understood that additional dies may be included that contain VMM arrays.

[0184] Die 3803 and die 3809 each include a high voltage decoder 3714, a high voltage generator 3511, an analog circuit 3512, and a temperature compensation circuit 3513.

[0185] Die 3804 and die 3810 each include an input circuit 3514 that includes a DAC 3515 .

[0186] Die 3805 and die 3811 each include an ADC 3516 and a neuron circuit 3510.

[0187] Die 3806 and die 3812 each include digital circuitry 3517, SRAM 3518, registers 3519, physical I / O connections 3520, and NOC connections 3715.

[0188] Each of the multiple dies is respectively connected to one or more other dies in the multiple dies via one or more vertical interfaces 3521 or horizontal interfaces 3716 that connect two or more dies together. In one embodiment, each vertical interface 3521 is implemented as a through silicon via (TSV). In one embodiment, each horizontal interface 3716 is implemented as a redistribution layer (RDL) connection.

[0189] During a read operation of the 3D VMM system 3800, a digital input is received by the input circuit 3514, which converts the digital input to analog form using a DAC 3515 and provides the analog signal to a respective array input circuit 3729 via a respective vertical interface 3521 and / or horizontal interface 3716. The array input circuit 3729 receives the analog signal and an address from the input circuit and then applies the analog signal to a selected row in the VMM array 3507 in response to the address. The output from the VMM array 3507 is received by the column multiplexer 3610, which provides the analog signal to the ADC 3516 via one or more vertical interfaces 3521 or horizontal interfaces 3716, which converts the analog signal to a digital signal. Alternatively, the analog signal may bypass the ADC 3516 and remain in analog form. The output of the ADC 3516 is provided via one or more vertical interfaces 3521 or horizontal interfaces 3716 to digital circuitry 3517 (which may perform activation functions, pooling functions, or other network functions), and the output of the digital circuitry 3517 may be provided via physical I / O 3520 to a device external to the 3D VMM system 3800 (such as a processing unit or graphics processing unit), or to input circuitry 3514 of other VMM arrays 3507 (representing another layer in the artificial neural network), or to another package 3522 via physical I / O 3520.

[0190] FIG. 39A illustrates a 3D VMM system 3900 that includes multiple dies in two vertical stacks. In this example, the first vertical stack includes dies, 3901, 3902, 3903, and 3904, and the second vertical stack includes dies, 3905, 3906, 3907, and 3908, all of which are included in a common package 3522 to form a single packaged integrated circuit. In one example, the dies in the first vertical stack are physically separate dies from the dies in the second vertical stack. In another example, the dies in the first vertical stack and the dies in the second vertical stack are the same physical die (meaning, for example, that dies 3901 and 3905 are the same die). Here, components included in dies 3901 and 3902 share components included in dies 3903 and 3904, and components included in dies 3905 and 3906 share components included in dies 3907 and 3908.

[0191] In the illustrated embodiment, die 3901, 3902, 3905, and die 3906 each include a VMM array 3507, an array input 3729, a high voltage multiplexer 3608, and a column multiplexer 3610.

[0192] In this example, four dies (dies 3901, 3902, 3903, and 3904) contain VMM arrays, although it should be understood that additional dies may be included that contain VMM arrays.

[0193] Die 3903 and die 3907 each include a high voltage decoder 3714, a high voltage generator 3511, an analog circuit 3512, a temperature compensation circuit 3513, an input circuit 3514 including a DAC 3515, a neuron circuit 3510, and an ADC 3516.

[0194] Die 3904 and die 3908 each include digital circuitry 3517, SRAM 3518, registers 3519, physical I / O connections 3520, and NOC connections 3715.

[0195] Each of the multiple dies is connected to one or more other dies in the multiple dies via one or more vertical interfaces 3521 or horizontal interfaces 3716 that connect two or more dies together.

[0196] During a read operation of the 3D VMM system 3900, a digital input is received by the input circuit 3514, which converts the digital input to analog form using a DAC 3515 and provides the analog signal to a respective array input circuit 3729 via a respective vertical interface 3521 and / or horizontal interface 3716. The array input circuit 3729 receives the analog signal and an address from the input circuit and applies the analog signal to a selected row in the VMM array 3507 in response to the received address. The output from the VMM array 3507 is received by the column multiplexer 3610, which provides the analog signal via one or more vertical interfaces 3521 or horizontal interfaces 3716 to the ADC 3516, which converts the analog signal to a digital signal. Alternatively, the analog signal may bypass the ADC 3516 and remain in analog form. The output of the ADC 3516 is provided via respective vertical interface 3521 and / or horizontal interface 3716 to digital circuit 3517, which is then provided via physical I / O 3520 to a device external to the 3D VMM system 3900 (such as a processing unit or a graphics processing unit), or to input circuit 3514, which is applied as an input to the VMM array 3507 (representing another layer in the artificial neural network) or to another package 3522 via physical I / O 3520.

[0197] Figure 39B shows a 3D VMM system 3950 that is similar to that of Figure 39A, except that dies 3951, 3952, 3955, and 3956 now have both input circuitry 3514 and array input 3729. 3D VMM system 3950 includes package 3522 and dies 3951, 3952, 3953, 3954, 3955, 3956, 3957, and 3958. Each of the multiple dies is connected to one or more other dies in the multiple dies via one or more vertical interfaces 3521 or horizontal interfaces 3716 that connect two or more dies together.

[0198] 39C illustrates a 3D VMM system 3980 comprising multiple dies in two vertical stacks. In this example, a first vertical stack comprises dies 3981, 3982, and 3983, and a second vertical stack comprises dies 3984, 3985, and 3986, all of which are included in a common package 3522 to form a single packaged integrated circuit. In one example, the dies in the first vertical stack are physically separate dies from the dies in the second vertical stack. In another example, the dies in the first vertical stack and the dies in the second vertical stack are the same physical die (meaning, for example, that dies 3981 and 3984 are the same die). Here, components included in dies 3981 and 3982 share components included in die 3983, and components included in dies 3984 and 3985 share components included in die 3986.

[0199] In the illustrated embodiment, die 3981, 3982, 3984, and die 3985 each include a VMM array 3507, a high voltage block 3991, an input block 3990, an output block 3992, and an analog block 3993. The input block 3990 may include an input circuit 3514, a DAC 3515, and an array input circuit 3729. The high voltage block 3991 may include a high voltage multiplexer 3608 and a high voltage decoder 3714. The output block 3992 may include a column multiplexer 3610, a neuron circuit 3510, and an ADC 3516. The analog block 3993 may include a high voltage generator 3511, an analog circuit 3512, and a temperature compensation circuit 3513.

[0200] Die 3983 and die 3986 each include digital circuitry 3517, SRAM 3518, registers 3519, physical I / O connections 3520, digital accelerator 3521 (used digitally for multiple and accumulate (MAC) functions), and NOC connections 3715.

[0201] The multiple dies are each connected to one or more other dies in the multiple dies via one or more vertical interfaces 3521 or horizontal interfaces 3716, each of which respectively connects two or more dies together.

[0202] FIG. 40 illustrates a 3D VMM system 4000 that includes multiple dies in two vertical stacks. In this example, a first vertical stack includes dies 4001, 4002, 4003, and 4004, and a second vertical stack includes dies 4005, 4006, 4007, and 4008, all of which are included in a common package 3522 to form a single packaged integrated circuit. In one example, the dies in the first vertical stack are physically separate dies from the dies in the second vertical stack. In another example, the dies in the first vertical stack and the dies in the second vertical stack are the same physical die (meaning, for example, that dies 4001 and 4005 are the same die). Here, components included in dies 4001 and 4002 share components included in dies 4003 and 4004, and components included in dies 4005 and 4006 share components included in dies 4007 and 4008.

[0203] In the illustrated embodiment, dies 4001, 4002, 4005, and 4006 each include a VMM array 3507, an array input 4029 (including an input multiplexer 3509 and / or a decoder 3713 (not shown)), a high voltage multiplexer 3608, and a column multiplexer 3610.

[0204]

[0205] In this example, four dies (dies 4001, 4002, 4003, and 4004) contain VMM arrays, although it should be understood that additional dies may be included that contain VMM arrays.

[0206] The die 4003 and die 4007 each include a high voltage decoder 3714, a high voltage generator 3511, an analog circuit 3512, a temperature compensation circuit 3513, and a neuron circuit 3510.

[0207] Die 4004 and die 4008 each include digital circuitry 3517, SRAM 3518, registers 3519, physical I / O connections 3520, and NOC connections 3715.

[0208] Unlike VMM system 3900, VMM system 4000 does not include DAC 3515 and ADC 3516 since the inputs and outputs are kept in analog form and are not converted between analog and digital forms.

[0209] Each of the multiple dies is connected to one or more other dies in the multiple dies via one or more vertical interfaces 3521 or horizontal interfaces 3716 that connect two or more dies together.

[0210] During a read operation of the 3D VMM system 4000, an analog input (such as a voltage, current, or time-based entity, such as a sequence of pulses) is received by a respective array input 4029, which then applies a signal to one or more rows in the respective VMM array 3507. The output from each VMM array 3507 is received by a column multiplexer 3610, which provides the analog signal (such as a voltage, current, or time-based entity) via one or more vertical interfaces 3521 or horizontal interfaces 3716 to a respective neuron circuit 3510, which provides a buffered signal via physical I / O 3520 to a device external to the 3D VMM system 4000 (such as a processing unit or a graphics processing unit), or to another package 3522, or to the array input 4029 of another VMM array 3507 (representing another layer in an artificial neural network).

[0211] 41-44 show further details regarding an example configuration of the VMM array 3507. FIGS. 41-44 show 3D VMM systems 4100, 4200, 4300, and 4400, respectively. Each of the 3DD VMM systems 4100, 4200, 4300, and 4400 comprises a first die VMM array 3507-1 and a second die VMM array 3507-2 in a common package (not shown). The VMM arrays 3507-1 and 3507-2 each include an array of non-volatile memory cells arranged in m+1 rows and n+1 columns. Each row is coupled to one of the control gate lines labeled CG0,...,CGm, and each column is coupled to one of the bit lines labeled BL0,...,BLn. The cells are located at the intersections of the bit lines and the control gate lines. For example, cell 4101mn is located at row m and column n and is coupled to CGm and BLn in VMM array 3507-1, and cell 4102mn is located at row m and column n and is coupled to CGm and BLn in VMM array 3507-2. Because VMM arrays 3507-1 and 3507-2 are located on different dies, VMM arrays 3507-1 and 3507-2 may be manufactured using different semiconductor processes, if desired. Regardless of whether the same or different semiconductor processes are used to manufacture VMM arrays 3507-1 and 3507-2, cells in VMM array 3507-1 may store a different number of bits than cells in VMM array 3507-2. For example, a cell in VMM array 3507-1, such as cell 4102mn, may store i bits, while a cell in VMM array 3507-2, such as cell 4101mn, may store j bits, where i and j are different valued integers. For example, i may be 3 (meaning that the cells in VMM array 3507-1 each store a 3-bit value) and j may be 5 (meaning that the cells in VMM array 3507-2 each store a 5-bit value).

[0212] In Figure 41, the inputs are provided separately to VMM arrays 3507-1 and 3507-2 via control gate lines, and the outputs are available separately on bit lines. Figure 41 shows exemplary cells 4101mn and 4102mn. If desired, the outputs can be combined elsewhere, such as by converting them to digital form using ADC 3516 (shown in the previous figure) and adding them together using digital circuitry 3517 (shown in the previous figure), if necessary.

[0213] In Figure 42, inputs are provided separately to VMM arrays 3507-1 and 3507-2 via control gate lines. The output of VMM array 3507-1 is available on a first set of bit lines and the output of VMM array 3507-2 is available on a second set of bit lines which are coupled to the second set of bit lines by respective vertical interfaces 3521 which effectively sum the output signals in analog form. If desired, the combined analog outputs can be digitized elsewhere, such as by converting them to digital form using ADC 3516 (shown in the previous figure) if necessary.

[0214] 43, the inputs are provided to VMM array 3507-1 via a first set of control gate lines and to VMM array 3507-2 via a second set of control gate lines which are coupled to the second set of control gate lines by respective vertical interfaces 3521, and the outputs are separately available on bit lines. If desired, the outputs can be combined elsewhere, such as by converting them to digital form using ADC 3516 (shown in the previous figure) and adding them together using digital circuitry 3517 (shown in the previous figure), if necessary.

[0215] In Figure 44, inputs are generally provided via control gate lines to VMM arrays 3507-1 and 3507-2 via vertical interface 3521. Outputs are obtained on bit lines which are combined together via vertical interface 3521, which effectively adds the signals together in analog form. If desired, the combined analog outputs can be digitized elsewhere, such as by converting them to digital form using ADC 3516 (shown in the previous figure), if necessary.

[0216] 45-48 show examples of structural layout options for the die, vertical interfaces, and horizontal interfaces.

[0217] 45 shows a 3D VMM system 4500 that includes a first set of dies (dies 4501, 4502, 4503, and 4504) arranged in a vertical configuration and connected by vertical interfaces 3521, and a second set of dies (dies 4505, 4506, 4507, and 4508) arranged in a vertical configuration and connected by vertical interfaces 3521, where one die may connect to two other dies by each of the vertical interfaces 3521.

[0218] 46 shows a 3D VMM system 4600 including four levels of die arranged in a vertical staggered configuration, where a first level includes die 4601 and 4602, a second level includes die 4603, 4604, and 4605, a third level includes die 4606 and 4607, and a fourth level includes die 4608, 4609, and 4610. The die on different levels are connected to the die on the level above and the die on the level below by respective vertical interfaces 3521, where one die may be connected to four other die by each of the vertical interfaces 3521.

[0219] 47 shows a 3D VMM system 4700 including four levels of die arranged in a vertical staggered configuration, where a first level includes die 4701 and 4702, a second level includes die 4703, 4704, and 4705, a third level includes die 4706 and 4707, and a fourth level includes die 4708, 4709, and 4710. Dies on different levels are connected by respective vertical interfaces 3521, and dies on the same level are connected by respective horizontal interfaces 3716. Here, one die may connect to six other dies by each of the vertical interfaces 3521 and by each of the horizontal interfaces 3716.

[0220] 48 illustrates an example of a physical layout of a 3D VMM system 4800 where only the connector 4801 is shown. The connector 4801 is located within the die and connects to one or more vertical interfaces 3521 and horizontal interfaces 3716. As can be seen, the die and interfaces may be arranged such that the connector 4801 is located in a vertical staggered configuration. <Circuits for use in 3D VMM systems>

[0221] 49-55 show circuits for use in a 3D VMM system as described above.

[0222] FIG. 49 illustrates a neuron circuit 4900 that may be used in the neuron circuit 3510 described above. Specifically, the neuron circuit 3510 may include an instance of the neuron circuit 4900 for each bit line in the VMM array 3507. The neuron circuit 4900 includes a P-channel Metal-Oxide-Semiconductor (PMOS) transistor 4901 and an operational amplifier 4902 arranged as shown. One terminal of the PMOS transistor 4901 is connected to a voltage source. Another terminal of the PMOS transistor 4901 is connected to a gate 4910 of the PMOS transistor, to a bit line in the VMM array 3507, and to a non-inverting terminal of the operational amplifier 4902. The output of the operational amplifier 4902 is connected to an inverting input of the operational amplifier 4902. The current drawn by the bit line I-BL results in a voltage V_IBL being output from the operational amplifier 4902. The operational amplifier 4902 acts as a buffer and V_IBL maintains its level regardless of the load to which it may be connected. For example, when V_IBL is fed to the vertical interface 3521, the vertical interface 3521 may have parasitic capacitance or parasitic current. The neuron circuit 4900 maintains the output voltage V_IBL despite load variations. It is therefore a neuron current buffer circuit. In another embodiment, the neuron current (bit line current) may be scaled by a current mirror before entering this circuit.

[0223] FIG. 50 illustrates a neuron circuit 5000 that may be used in the neuron circuit 3510 described above. Specifically, the neuron circuit 3510 may include an instance of the neuron circuit 5000 for each bit line in the VMM array 3507. The neuron circuit 5000 includes a controlled switch 5001, a reference memory cell 5002, and an operational amplifier 5003 arranged as shown. The current drawn by the bit line I-BL results in a particular voltage at the non-inverting terminal of the operational amplifier 5003. When the switch 5001 is closed, feedback VNEUOUT from the output of the operational amplifier 5003 is provided to the control gate terminal of the reference memory cell 5002. Due to the inherent properties of the operational amplifier, the operational amplifier 5003 modifies the output voltage VNEUOUT until the voltage at its non-inverting terminal is equal to the voltage VREF at its inverting terminal. This is done via feedback to the control gate terminal of the reference memory cell 5002. Neuron circuit 5000 maintains the output voltage VNEUOUT despite parasitic capacitive loading that may receive that voltage, such as vertical interface 3521. Neuron circuit 5000 converts neuron current I-BL to a voltage VNEUOUT using memory cell 5002. Thus, neuron circuit 5000 is a memory cell-based current-to-voltage converter that uses an operational amplifier with feedback.

[0224] FIG. 51 shows a Force and Sense (F / S) driver circuit 5100 (where forcing is caused by an operational amplifier 5103 with an input V at its positive terminal, sensing is done at a target node, node 5101 or 5102, and the voltage at the target node is fed back to the negative terminal of the amplifier, which is made equal to the input voltage V by the action of the amplifier) ​​that may be used in a neuron circuit 3510 or elsewhere to accurately deliver a neuron voltage output across a varying load. The F / S driver circuit 5100 comprises an operational amplifier 5103 and controlled switches 5104, 5105, 5106, and 5107 arranged as shown. Vin is received at the non-inverting input of the operational amplifier 5103, and the output of the operational amplifier 5103 is called VOUT. Switches 5104 and 5106, when closed, respectively connect the output of operational amplifier 5103 and the inverting input of operational amplifier 5103 as VOUTFB to node 5101 in the first die through vertical interface 3521. Switches 5105 and 5107, when closed, respectively connect the output of operational amplifier 5103 and the inverting input of operational amplifier 5103 as VOUTFB to node 5102 in the second die through vertical interface 3521. Drive circuit 5100 provides accurate voltage VOUT to nodes 5101 and 5102, regardless of any loading effects caused by vertical interface 3521 or by the first and second dies. This is because the sense node (driven node) 5101 or 5102 is fed back to the inverting input of operational amplifier 5103.

[0225] 52 illustrates a differential neuron circuit 5200 that may be used in the neuron circuit 3510 described above. Specifically, the neuron circuit 3510 may include an instance of the differential neuron circuit 5200 for each pair of differential bit lines in the VMM array 3507, such as where one column stores a W+ value and another column stores a W− value, with each pair of W+ and W− representing a stored weight.

[0226] As shown, the differential neuron circuit 5200 includes an operational amplifier 5201, variable integration resistors 5202 and 5203, controlled switches 5204, 5205, 5206, and 5207, and sample-and-hold (and / or integration) capacitors 5208 and 5209. The differential neuron circuit 5200 receives differential currents BLw+ from the W+ bit line and BLw- from the W- bit line, and outputs voltages Vout+ and Vout-, respectively. Output voltage Vout+=(BLw+) * R and Vout-=(BLw-) * R, and variable integration resistors 5202 and 5203 each have a value equal to R. Capacitors 5208 and 5209 function as respective sample and hold (S / H) capacitors to hold the output voltage when resistors 5202 and 5203 are removed from the circuit by opening controlled switches 5206, 5207, and the input current is interrupted by opening controlled switches 5204, 5205. A control circuit (not shown) controls the opening and closing of switches 5204, 5205, 5206, and 5207 to provide the integration time. If desired, the differential output voltages Vout+ and Vout- can be input to the ADC 3516, which converts the differential output voltages Vout+ and Vout- into a series of digital output bits Doutx. Optionally, the circuit can integrate the neuron current using capacitors 5208 and 5209 as integrating capacitors to convert the current to a voltage, Vout=Time * I neuron / capacitance. Neuron scaling is provided by variable resistors 5292 and 5203, or by variable integration time and / or variable capacitance if an integrating capacitor approach is used.

[0227] FIG. 53 illustrates a sample and hold buffer 5300 that may be used in the input circuit 3514 described above. Specifically, the input circuit 3514 may include an instance of the sample and hold buffer 5300 for each row in the VMM array 3507. The sample and hold buffer 5300 comprises a controlled switch 5301, a capacitor 5302, and a buffer 5303. The buffer 5303 may be a unity buffer formed from an operational amplifier. In operation, the switch 5301 is closed, thereby allowing an analog value (e.g., from the DAC 3515) to be stored in the capacitor 5302. That value may then be output from the buffer 5303 which drives the row input of the VMM array 3507. The capacitor 5302 may be an actual capacitor or may be an intrinsic capacitor found in a wire, for example.

[0228] FIG. 54 shows a differential successive address register (SAR) analog-to-digital converter (ADC) 5400 that may be used in the ADC 3516 described above.

[0229] The differential serial address register analog-to-digital converter 5400 converts an analog input or a differential analog input to a digital output using a binary search through all possible quantization levels to identify the appropriate digital output.

[0230] The differential serial address register analog-to-digital converter 5400 includes a binary capacitive digital-to-analog converter (CDAC) 5401, a binary CDAC 5402 (complementary to CDAC 5401), a comparator 5403, and a SAR logic and register 5404.

[0231] The differential sequential address register analog-to-digital converter 5400 receives differential current inputs Vinp and Vinn. The SAR logic and register 5404 cycles through all possible digital bit combinations and then controls the switches in the CDACs 5401 and 5402 to couple the voltage sources to the capacitors. When the output of the comparator 5403 flips, the combination of the digital bits in the SAR logic and register 5404 is output as the digital output. Optionally, the SAR logic and register 5404 generates an additional 1-bit digital output DMAJ in the digital output that is "1" if the majority of the bits in the digital value are "1" and is "0" if the majority of the bits in the corresponding digital value are not "1".

[0232] FIG. 55 shows a reference current based SAR ADC circuit 5500 that may be used in the ADC 3516 described above. Specifically, the ADC 3516 may include one or more instances of the ADC circuit 5500, which converts the current I-BL into a digital value Digital Output. The ADC circuit 5500 includes a binary current block 5501, a switch 5502, a comparator 5503, and a SAR logic and register 5504. The binary current block 5501 provides a binary reference current that is compared with the bit line input current in a binary search manner, i.e., searching from the most significant bit (MSB) to the least significant bit (LSB).

[0233] It should be noted that, as used herein, the terms "over" and "on" both inclusively encompass "directly on" (with no intermediate materials, elements, or gaps disposed therebetween) and "indirectly on" (with an intermediate material, element, or gap disposed therebetween). Similarly, the term "adjacent" includes "directly adjacent" (with no intermediate material, element, or gap disposed between them) and "indirectly adjacent" (with an intermediate material, element, or gap disposed between them), "mounted to" includes "directly mounted to" (with no intermediate material, element, or gap disposed between them) and "indirectly mounted to" (with an intermediate material, element, or gap disposed between them), and "electrically coupled" includes "directly electrically coupled to" (without an intermediate material or element disposed between them that electrically connects the elements together) and "indirectly electrically coupled to" (with an intermediate material or element disposed between them that electrically connects the elements together). For example, forming an element "over a substrate" can include forming the element directly on the substrate, with no intermediate materials / elements between them, and forming the element indirectly on the substrate, with one or more intermediate materials / elements between them.

Claims

1. 1. A three dimensional integrated circuit for use in an artificial neural network, comprising: a first die including a first vector matrix multiplication array and a first input multiplexer, the first die being located in a first vertical layer; a second die including input circuitry, the second die being located in a second vertical layer different from the first vertical layer; one or more vertical interfaces coupling the first die and the second die; during a read operation, the input circuit provides an input signal via at least one of the one or more vertical interfaces to the first input multiplexer, the first input multiplexer applies the input signal to one or more rows in the first vector matrix multiplication array, and the first vector matrix multiplication array generates an output.

2. 10. The three dimensional integrated circuit of claim 1, wherein the second die includes a digital-to-analog converter for converting a digital input to an analog input that is provided as the input signal to the input circuitry.

3. The three dimensional integrated circuit of claim 1 , wherein the first die includes a neuron circuit for buffering the output.

4. 10. The three dimensional integrated circuit of claim 1 , wherein the first die includes a column multiplexer for sending the output to the second die or a third die via at least one of the one or more vertical interfaces.

5. 13. The three dimensional integrated circuit of claim 1, further comprising a third die including an analog to digital converter for converting the output from the first die to a digital output, the third die being located in a third vertical layer distinct from the first vertical layer and the second vertical layer.

6. 6. The three dimensional integrated circuit of claim 5, comprising a fourth die including a high voltage generator, an analog circuit, and a temperature compensation circuit, the fourth die located in a fourth vertical layer distinct from the first vertical layer, the second vertical layer, and the third vertical layer.

7. 7. The three dimensional integrated circuit of claim 6, further comprising a fifth die including a second vector matrix multiplication array, a second input multiplexer, a high voltage decoder, and a neuron circuit, the fifth die being located in a fifth vertical layer distinct from the first vertical layer, the second vertical layer, the third vertical layer, and the fourth vertical layer.

8. 10. The three dimensional integrated circuit of claim 1 further comprising a third die including a second vector matrix multiplication array and a second input multiplexer, said third die located in said first vertical layer.

9. The three dimensional integrated circuit of claim 1 , wherein the vector matrix multiplication array comprises a plurality of non-volatile memory cells.

10. The three dimensional integrated circuit of claim 9 , wherein the plurality of non-volatile memory cells comprises stacked gate flash memory cells.

11. The three dimensional integrated circuit of claim 9 , wherein the plurality of non-volatile memory cells comprises split-gate flash memory cells.

12. 1. A method comprising: providing input signals by input circuitry located on the first die via one or more vertical interfaces to input multiplexers located on the second die; applying the input signal to one or more rows in a neural network array by the input multiplexer; generating an output by said neural network array; The method, wherein the first die and the second die are located on different vertical layers.

13. The method of claim 12 , wherein the neural network array comprises a plurality of non-volatile memory cells.

14. The method of claim 13 , wherein the plurality of non-volatile memory cells comprises stacked gate flash memory cells.

15. The method of claim 13 , wherein the plurality of non-volatile memory cells comprises split-gate flash memory cells.

16. 1. An apparatus comprising: a first die including a first vector matrix multiplication array including a plurality of non-volatile memory cells arranged in rows and columns, the first die being located in a first vertical layer; a second die including a second vector matrix multiplication array including a plurality of non-volatile memory cells arranged in rows and columns, the second die being located in a second vertical layer different from the first vertical layer; one or more vertical interfaces coupling the first die and the second die; During a program operation, one or more non-volatile memory cells in the first array are capable of storing i bits and one or more non-volatile memory cells in the second array are capable of storing j bits, where i≠j.

17. 17. The apparatus of claim 16, wherein the first die is manufactured according to a first semiconductor process and the second die is manufactured according to a second semiconductor process that is different from the first semiconductor process.

18. a first set of bit lines coupled to the first vector matrix multiplication array; a second set of bit lines, different from the first set of bit lines, coupled to the second vector matrix multiplication array.

19. a first set of control gate lines coupled to the first vector matrix multiplication array; a second set of control gate lines coupled to the second vector matrix multiplication array, the second set of control gate lines being different from the first set of control gate lines.

20. 20. The device of claim 19, wherein the first set of control gate lines are coupled to the second set of control gate lines by respective vertical interfaces.

21. 20. The apparatus of claim 18, wherein the first set of bit lines are coupled to the second set of bit lines by respective vertical interfaces.

22. a first set of control gate lines coupled to the first vector matrix multiplication array; a second set of control gate lines coupled to the second vector matrix multiplication array.

23. 23. The device of claim 22, wherein the first set of control gate lines are coupled to the second set of control gate lines by respective vertical interfaces.

24. 20. The apparatus of claim 16, wherein the plurality of non-volatile memory cells in the first die and the plurality of non-volatile memory cells in the second die each comprise stacked gate flash memory cells.

25. 20. The apparatus of claim 16, wherein the plurality of non-volatile memory cells in the first die and the plurality of non-volatile memory cells in the second die each comprise split-gate flash memory cells.

26. 1. A method comprising: storing a value comprising i bits in a non-volatile memory cell in a first neural network array in a first die; storing a value comprising j bits in a non-volatile memory cell in a second neural network array in a second die, where i≠j; The method, wherein the first die and the second die are located on different vertical layers.

27. 1. An apparatus comprising: a first die including a first vector matrix multiplication array including a respective plurality of non-volatile memory cells arranged in rows and columns, the first die being located in a first vertical layer; a second die including a second vector matrix multiplication array including a respective plurality of non-volatile memory cells arranged in rows and columns, the second die being located in a second vertical layer different from the first layer; a third die including one or more of a digital-to-analog converter and an analog-to-digital converter.

28. 28. The apparatus of claim 27, wherein the analog-to-digital converter is a capacitor-based successive approximation register type analog-to-digital converter.

29. 28. The apparatus of claim 27, wherein the analog-to-digital converter is a reference current-based successive approximation register type analog-to-digital converter.

30. 30. The apparatus of claim 27, wherein the respective plurality of non-volatile memory cells in the first vector matrix multiplication array and the second vector matrix multiplication array comprise stacked gate flash memory cells.

31. 30. The apparatus of claim 27, wherein the respective plurality of non-volatile memory cells in the first vector matrix multiplication array and the second vector matrix multiplication array comprise split-gate flash memory cells.

32. 1. An apparatus comprising: a first die including a first vector matrix multiplication array including a plurality of non-volatile memory cells arranged in rows and columns, the first die being located in a first vertical layer; a second die including a second vector matrix multiplication array including a plurality of non-volatile memory cells arranged in rows and columns, the second die being located in a second vertical layer different from the first vertical layer; a third die containing a neuron circuit.

33. The neuron circuit comprises: a p-channel metal oxide semiconductor transistor having a first terminal coupled to a voltage source, a gate, and a second terminal coupled to the gate and to a neuron; 33. The apparatus of claim 32, comprising: an operational amplifier comprising a non-inverting input coupled to the second terminal and to the gate of the p-channel metal-oxide-semiconductor transistor, an inverting input, and an output coupled to the inverting input to generate a voltage output in response to current from the neuron.

34. The neuron circuit comprises: Switch, a reference memory cell having a bit line terminal coupled to the neuron, a source line terminal, and a control gate terminal; 33. The apparatus of claim 32, comprising: an operational amplifier comprising an inverting input coupled to a reference voltage, a non-inverting input coupled to the bit line terminal of the reference memory cell, and an output switchably coupled to the control gate terminal of the reference memory cell via the switch.

35. The neuron circuit comprises: an operational amplifier comprising an inverting input, a non-inverting input, a first output coupled to a first output node, and a second output coupled to a second output node; a first variable integrating resistor switchably coupled between the first output node and the inverting input of the operational amplifier via a first switch; a second variable integrating resistor switchably coupled between the second output node and the non-inverting input of the operational amplifier via a second switch; a first capacitor switchably coupled between a first input current from a bit line through a third switch and the first output node; 33. The apparatus of claim 32 comprising: a second capacitor switchably coupled between a second input current from a bit line through a fourth switch and the second output node.

36. 36. The apparatus of claim 35, wherein the first input current and the second input current are differential current signals and the first output node and the second output node comprise differential voltage signals.

37. 37. The apparatus of claim 36, wherein the first input current is received from a W+ bit line and the second input current is received from a W- bit line.

38. 37. The apparatus of claim 36, comprising an analog-to-digital converter for converting the differential voltage signal into a set of digital output bits.

39. 33. The apparatus of claim 32, wherein the plurality of non-volatile memory cells in the first vector matrix multiplication array and the second vector matrix multiplication array comprise stacked gate flash memory cells.

40. 33. The apparatus of claim 32, wherein the plurality of non-volatile memory cells in the first vector matrix multiplication array and the second vector matrix multiplication array comprise split-gate flash memory cells.

41. 1. An apparatus comprising: a first die including a first vector matrix multiplication array including a respective plurality of non-volatile memory cells arranged in rows and columns, the first die being located in a first vertical layer; a second die including a second vector matrix multiplication array including a respective plurality of non-volatile memory cells arranged in rows and columns, the second die being located in a second vertical layer different from the first vertical layer; a third die containing digital circuitry comprising one or more of a microcontroller, digital logic, or a single instruction multiple data processor.

42. 42. The apparatus of claim 41, wherein the third die includes a digital accelerator.

43. 42. The apparatus of claim 41, wherein the third die includes a static random access memory.

44. 42. The apparatus of claim 41, wherein the third die includes physical input / output connections.

45. 42. The apparatus of claim 41, wherein the third die includes a resistor.

46. 42. The apparatus of claim 41, wherein the respective plurality of non-volatile memory cells in the first vector matrix multiplication array and the respective plurality of non-volatile memory cells in the second vector matrix multiplication array comprise stacked gate flash memory cells.

47. 42. The apparatus of claim 41, wherein the respective plurality of non-volatile memory cells in the first vector matrix multiplication array and the respective plurality of non-volatile memory cells in the second vector matrix multiplication array comprise split-gate flash memory cells.

48. 1. An apparatus comprising: a first vertical layer including a first vector matrix multiplication array including a respective plurality of non-volatile memory cells arranged in rows and columns, and a second vector matrix multiplication array including a respective plurality of non-volatile memory cells arranged in rows and columns; one or more respective horizontal interfaces coupling the first vector matrix multiplication array and the second vector matrix multiplication array; a second vertical layer including a third vector matrix multiplication array including a respective plurality of non-volatile memory cells arranged in rows and columns, and a fourth vector matrix multiplication array including a respective plurality of non-volatile memory cells arranged in rows and columns; one or more respective horizontal interfaces coupling the third vector matrix multiplication array and the fourth vector matrix multiplication array.

49. 49. The apparatus of claim 48, wherein the first vector matrix multiplication array is located on a first die, the second vector matrix multiplication array is located on a second die, the third vector matrix multiplication array is located on a third die, and the fourth vector matrix multiplication array is located on a fourth die.

50. 50. The apparatus of claim 49, wherein the first die and the third die are vertically aligned, and the second die and the fourth die are vertically aligned.

51. 50. The apparatus of claim 49, wherein the first die and the second die are vertically staggered with the third die and the fourth die.

52. 49. The apparatus of claim 48, wherein the respective plurality of non-volatile memory cells in the first vector matrix multiplication array, the second vector matrix multiplication array, the third vector matrix multiplication array, and the fourth vector matrix multiplication array comprise stacked gate flash memory cells.

53. 49. The apparatus of claim 48, wherein the respective plurality of non-volatile memory cells in the first vector matrix multiplication array, the second vector matrix multiplication array, the third vector matrix multiplication array, and the fourth matrix multiplication array comprise split-gate flash memory cells.

Citation Information

Patent Citations

  • Integrated Neuro-Processor Comprising Three-Dimensional Memory Array

    US20170270403A1

  • Three-dimensional neural network array

    US20180165573A1

  • 3D neural inference processing unit architectures

    US20210125040A1

  • 3D stacked integrated circuits having functional blocks configured to accelerate artificial neural network (ANN) computation

    US20210151431A1

  • Analog neural memory array storing synapsis weights in differential cell pairs in artificial neural network

    WO2021178003A1