Analog in-memory computation engine and digital in-memory computation engine for performing operations in neural networks
By using non-volatile memory cell arrays as synapses in neural networks, combined with computation engines within analog and digital memories, the problems of low energy efficiency and high computational complexity in existing technologies are solved, achieving efficient neural network computation and fine-tuning, suitable for applications such as facial recognition.
Patent Information
- Application Number
- CN202380096489.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-07-05
- Filing Date
- 2023-07-14
- Publication Date
- 2025-11-11
AI Technical Summary
Existing artificial neural network hardware technologies suffer from low energy efficiency and high computational complexity. In particular, the large number of synapses required makes CMOS-implemented synapses too large, making it difficult to achieve efficient neural network computation.
Using a non-volatile memory cell array as synapses, the synaptic weights of the neural network are finely tuned by independently programming, erasing, and reading each memory cell, combined with analog programming. Vector-matrix multiplication is performed using an in-memory computing engine, eliminating the need for separate multiplication and addition logic circuits.
It achieves efficient neural network computation, reduces energy consumption and improves computational parallelism, and is suitable for facial recognition and other applications. By combining in-memory and in-digital memory computing engines, it optimizes the hardware performance of neural networks.
Smart Images

Figure CN120937080A_ABST
Abstract
Description
[0001] Priority Statement
[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 458,439, filed April 10, 2023, entitled "Neural Network Comprising Analog Computation-in-Memory Engine and Digital Computation-in-Memory Engine", and U.S. Patent Application No. 18 / 218,368, filed July 5, 2023, entitled "Analog Computation-In-Memory Engine and Digital Computation-In-Memory Engine to Perform Operations in a Neural Network". Technical Field
[0003] Numerous examples of systems comprising one or more analog in-memory computing engines and one or more digital in-memory computing engines to perform neural network operations are disclosed. Background Technology
[0004] Artificial neural networks mimic biological neural networks (the central nervous system of animals, especially the brain) and are used to estimate or approximate functions that can depend on a large number of inputs and are often unknown. Artificial neural networks typically consist of interconnected layers of "neurons" that exchange messages with each other.
[0005] Figure 1 An artificial neural network is illustrated, where circles represent the inputs or layers of neurons. Connections (called synapses) are represented by arrows and have numerical weights that can be tuned empirically. This allows the neural network to adapt to its inputs and learn. Typically, a neural network consists of layers with multiple inputs. There are usually one or more intermediate layers of neurons, and an output layer of neurons that provide the output of the neural network. Neurons at each level make decisions individually or collectively based on data received from the synapses.
[0006] One of the major challenges in developing artificial neural networks for high-performance information processing is the lack of sufficient hardware technology. Real-world neural networks rely on a large number of synapses to achieve high connectivity between neurons, i.e., very high computational parallelism. In principle, this complexity can be achieved using digital supercomputers or clusters of graphics processing units. However, compared to biological networks, these methods are generally energy inefficient, in addition to being costly, as biological networks consume less energy primarily due to their ability to perform low-precision analog calculations. CMOS analog circuits have been used in artificial neural networks, but given the large number of neurons and synapses, most CMOS-implemented synapses are excessively large.
[0007] The applicant previously disclosed an artificial (simulated) neural network utilizing one or more non-volatile memory arrays as synapses in U.S. Patent Application Publication 2017 / 0337466A1, which is incorporated herein by reference. The non-volatile memory array operates as a simulated neural memory and includes non-volatile memory cells arranged in rows and columns. The neural network includes a plurality of first synapses configured to receive a first plurality of inputs and generate a first plurality of outputs therefrom, and a plurality of first neurons configured to receive the first plurality of outputs. The first plurality of synapses includes a plurality of memory cells, wherein each memory cell includes: spaced-apart source and drain regions formed in a semiconductor substrate, wherein a channel region extends between the source and drain regions; a floating gate disposed over and insulated from a first portion of the channel region; and a non-floating gate disposed over and insulated from a second portion of the channel region. Each memory cell stores weight values corresponding to a plurality of electrons on the floating gate. The plurality of memory cells multiply the first plurality of inputs by the stored weight values to generate the first plurality of outputs.
[0008] Non-volatile memory cells
[0009] Non-volatile memory is well known. For example, U.S. Patent 5,029,130 (“the 130 patent”), which is incorporated herein by reference, discloses an array of split-gate non-volatile memory cells, which is a type of flash memory cell. Such memory cells 210 in Figure 2As shown in the figure. Each memory cell 210 includes a source region 14 and a drain region 16 formed in a semiconductor substrate 12, with a channel region 18 between the source and drain regions. A floating gate 20 is formed over and insulated from (and controls the conductivity of) a first portion of the channel region 18, and is formed over a portion of the source region 14. A word line terminal 22 (which is typically coupled to a word line) has a first portion disposed over and insulated from (and controlling the conductivity of) a second portion of the channel region 18, and a second portion extending upward and located over the floating gate 20. The floating gate 20 and the word line terminal 22 are insulated from the substrate 12 by a gate oxide. A bit line 24 is coupled to the drain region 16.
[0010] The memory cell 210 is erased by applying a high positive voltage to the word line terminal 22 (where electrons are removed from the floating gate), which causes electrons on the floating gate 20 to tunnel from the floating gate 20 to the word line terminal 22 through the intermediate insulator via the Fowler-Nordheim (FN) tunnel.
[0011] The memory cell 210 is programmed by source-side injection (SSI) with hot electrons (where electrons are placed on the floating gate) by applying a positive voltage to both the word line terminal 22 and the source region 14. Electron flow occurs from the drain region 16 to the source region 14. As electrons reach the gap between the word line terminal 22 and the floating gate 20, they accelerate and become hot. Due to electrostatic attraction from the floating gate 20, some of the heated electrons are injected onto the floating gate 20 through the gate oxide.
[0012] Memory cell 210 is read by applying a positive read voltage to the drain region 16 and word line terminal 22 (which connects the portion of channel region 18 below the word line terminal). If the floating gate 20 is positively charged (i.e., electrons are erased), the portion of channel region 18 below the floating gate 20 is also turned on, and current flows through channel region 18, which is sensed as an erased state or a "1" state. If the floating gate 20 is negatively charged (i.e., programmed electronically), the portion of channel region below the floating gate 20 is mostly or completely turned off, and current does not flow (or very little current) through channel region 18, which is sensed as a programmed state or a "0" state.
[0013] Table 1 depicts the typical voltage and current ranges that can be applied to the terminals of memory cell 210 to perform read, erase, and program operations:
[0014] Table 1: Figure 2 Operation of flash memory cell 210
[0015] WL BL SL Read 2V-3V 0.6V-2V 0V erase Approximately 11V-13V 0V 0V programming 1V-2V 10.5μA-3μA 9V-10V
[0016] Other split-gate memory cell configurations, used as other types of flash memory cells, are known. For example, Figure 3 A four-gate memory cell 310 is depicted, comprising a source region 14, a drain region 16, a floating gate 20 over a first portion of a channel region 18, a select gate 22 (typically coupled to a word line WL) over a second portion of the channel region 18, a control gate 28 over the floating gate 20, and an erase gate 30 over the source region 14. This configuration is described in U.S. Patent 6,747,310, which is incorporated herein by reference for all purposes. Here, all gates except the floating gate 20 are non-floating gates, meaning they are electrically connected to or can be electrically connected to a voltage source. Programming is performed by heated electrons from the channel region 18 that inject themselves into the floating gate 20. Erasing is performed by electrons tunneling from the floating gate 20 to the erase gate 30.
[0017] Table 2 depicts the typical voltage and current ranges that can be applied to the terminals of memory cell 310 to perform read, erase, and program operations:
[0018] Table 2: Figure 3 Operation of flash memory cell 310
[0019] WL / SG BL CG EG SL Read 1.0V-2V 0.6V-2V 0V-2.6V 0V-2.6V 0V erase -0.5V / 0V 0V 0V / -8V 8V-12V 0V programming 1V 0.1μA-1μA 8V-11V 4.5V-9V 4.5V-5V
[0020] Figure 4 A tri-gate memory cell 410 is depicted, which is another type of flash memory cell. Memory cell 410 and... Figure 3 The memory cell 310 is the same as the memory cell 410, except that the memory cell 410 does not have a separate control gate. Except that no control gate bias is applied, the erase operation (thus erasing is performed using an erase gate) and read operation are the same as... Figure 3 The operation is similar. Programming is also performed without a control gate bias, and therefore, a higher voltage is applied to the source line during programming to compensate for the lack of a control gate bias.
[0021] Table 3 depicts the typical voltage and current ranges that can be applied to the terminals of memory cell 410 to perform read, erase, and program operations:
[0022] Table 3: Figure 4 Operation of flash memory cell 410
[0023] WL / SG BL EG SL Read 0.7V-2.2V 0.6V-2V 0V-2.6V 0V erase -0.5V / 0V 0V 11.5V 0V programming 1V 0.2μA-3μA 4.5V 7V-9V
[0024] Figure 5 A stacked-gate memory cell 510 is depicted, which is another type of flash memory cell. Memory cell 510 and... Figure 2The memory cell 210 is similar, except that the floating gate 20 extends over the entire channel region 18, and the control gate 22 (which will be coupled to the word line here) extends over the floating gate 20, separated by an insulating layer (not shown). Erasing is performed by FN tunneling of electrons from FG to the substrate, and programming is performed by channel hot electron (CHE) injection in the region between the channel 18 and the drain region 16, by the flow of electrons from the source region 14 to the drain region 16, and by read operations similar to those of the memory cell 210 with a higher control gate voltage.
[0025] Table 4 depicts the typical voltage range that can be applied to the terminals of memory cell 510 and substrate 12 to perform read, erase, and program operations:
[0026] Table 4: Figure 5 Operation of flash memory cell 510
[0027] CG BL SL substrate Read 2V-5V 0.6V-2V 0V 0V erase -8V to -10V / 0V FLT FLT 8V-10V / 15V-20V programming 8V-12V 3V-5V 0V 0V
[0028] The methods and components described herein can be applied to other non-volatile memory technologies, such as, but not limited to, FINFET split-gate flash or stacked-gate flash memory, NAND flash memory, SONOS (silicon-oxide-nitride-oxide-silicon with charge trapped in nitride), MONOS (metal-oxide-nitride-oxide-silicon with metal charge trapped in nitride), ReRAM (resistive RAM), PCM (phase-change memory), MRAM (magnetic RAM), FeRAM (ferroelectric RAM), CT (charge-trapping) memory, CN (carbon nanotube) memory, OTP (two-level or multi-level one-time programmable) and CeRAM (associated electron RAM), etc.
[0029] To utilize memory arrays comprising one of the aforementioned types of non-volatile memory cells in artificial neural networks, two modifications were made. First, the circuitry was configured such that each memory cell could be individually programmed, erased, and read without adversely affecting the memory state of other memory cells in the array, as explained further below. Second, continuous (simulated) programming of the memory cells was provided.
[0030] Specifically, the memory state (i.e., the charge on the floating gate) of each memory cell in the array can be continuously changed from a fully erased state to a fully programmed state and vice versa, independently and with minimal interference to other memory cells. This means that the cell storage device is essentially analog, or at least can store one discrete value from many discrete values (such as 16 or 64 different values). This allows for very precise and individual tuning of all memory cells in the memory array, making the memory array ideal for both storage and fine-tuning of synaptic weights in neural networks.
[0031] Neural networks using non-volatile memory cell arrays
[0032] Figure 6 This example conceptually illustrates a non-limiting example of a neural network utilizing a non-volatile memory array. This example uses a non-volatile memory array neural network for a facial recognition application, but any other suitable application can also be implemented using a neural network based on a non-volatile memory array. The non-volatile memory array and associated circuitry used in the neural network are a type of in-memory computation (CIM) engine or vector-matrix (VMM) multiplication system.
[0033] In this example, S0 is the input layer, which is a 32×32 pixel RGB image with 5-bit precision (i.e., three 32×32 pixel arrays, one for each color R, G, and B, with 5-bit precision per pixel). The synapse CB1 from the input layer S0 to layer C1 applies different sets of weights in some cases and shared weights in others, and scans the input image with a 3×3 pixel overlapping filter (kernel), shifting the filter by one pixel (or more than one pixel as indicated by the model). Specifically, the values of nine pixels in a 3×3 portion of the image (i.e., called the filter or kernel) are provided to synapse CB1, where these nine input values are multiplied by appropriate weights, and after summing the output of this multiplication, a single output value is determined by the first synapse of CB1 and provided for generating one of the pixels in the feature map of layer C1. The 3×3 filter is then shifted one pixel to the right within the input layer S0 (i.e., adding a column of three pixels to the right and releasing a column of three pixels to the left), thereby providing the nine pixel values from this newly positioned filter to synapse CB1, where they are multiplied by the same weights and a second single output value is determined by the associated synapse. This process continues until the 3×3 filter scans all three colors and all bits (precision values) across the entire 32×32 pixel image of the input layer S0. This process is then repeated using different sets of weights to generate different feature maps for layer C1 until all feature maps for layer C1 are computed.
[0034] At layer C1, in this example, there are 16 feature maps, each with 30×30 pixels. Each pixel is a new feature pixel extracted from the product of the input and the kernel, so each feature map is a two-dimensional array. Therefore, in this example, layer C1 consists of a 16-layer two-dimensional array (remember that the layers and arrays referred to in this article are logical relationships, not necessarily physical relationships; that is, the array does not have to be oriented as a physical two-dimensional array). Each of the 16 feature maps in layer C1 is generated by one set of sixteen different groups of synaptic weights applied to the filter scan. The C1 feature maps may all relate to different aspects of the same image feature, such as boundary identification. For example, the first map (generated using the first weight reassembly, shared for all scans used to generate the first map) may identify circular edges, the second map (generated using the second weight reassembly, different from the first weight reassembly) may identify rectangular edges, or the aspect ratio of some feature, and so on.
[0035] Before transitioning from layer C1 to layer S1, activation function P1 (pooling) is applied, which pools the values from consecutive non-overlapping 2×2 regions in each feature map. The purpose of pooling function P1 is to average the neighboring locations (or a max function could also be used) to, for example, reduce the dependence on edge locations and reduce the data size before moving to the next stage. At layer S1, there are 16 15×15 feature maps (i.e., sixteen different arrays, each with 15×15 pixels). The synapse CB2 from layer S1 to layer C2 scans the map in layer S1 using a 4×4 filter, where the filter is shifted by 1 pixel. At layer C2, there are 22 12×12 feature maps. Before transitioning from layer C2 to layer S2, activation function P2 (pooling) is applied, which pools the values from consecutive non-overlapping 2×2 regions in each feature map. At layer S2, there are 22 6x6 feature maps. An activation function (pooling) is applied to the synapse CB3 from layer S2 to layer C3, where each neuron in layer C3 is connected to each mapping in layer S2 via a corresponding synapse in CB3. There are 64 neurons in layer C3. The synapse CB4 from layer C3 to the output layer S3 completely connects C3 to S3, meaning each neuron in layer C3 is connected to every neuron in layer S3. The output at S3 comprises 10 neurons, with the highest-output neuron determining the class. For example, this output could indicate an identification or classification of the content of the original image.
[0036] Synapses for each layer are implemented using an array or a portion of an array of non-volatile memory cells.
[0037] Figure 7 This is a block diagram of an array that could be used for this purpose. The vector-matrix multiplication (VMM) array 32 includes non-volatile memory cells and serves as synapses between layers (such as...). Figure 6(CB1, CB2, CB3, and CB4 in the original text). Specifically, the VMM array 32 includes an array 33 of non-volatile memory cells, an erase gate and word line gate decoder 34, a control gate decoder 35, a bit line decoder 36, and a source line decoder 37, which decode the corresponding inputs of the non-volatile memory cell array 33. The inputs to the VMM array 32 may come from the erase gate and word line gate decoder 34 or from the control gate decoder 35. In this example, the source line decoder 37 also decodes the outputs of the non-volatile memory cell array 33. Alternatively, the bit line decoder 36 may decode the outputs of the non-volatile memory cell array 33.
[0038] The non-volatile memory cell array 33 serves two purposes. First, it stores weights that will be used by the VMM array 32. Second, the non-volatile memory cell array 33 efficiently multiplies the inputs with the weights stored in the non-volatile memory cell array 33 and adds them together at each output line (source line or bit line) to produce an output that will be used as the input to the next layer or the final layer. By performing multiplication and addition functions, the non-volatile memory cell array 33 eliminates the need for separate multiplication and addition logic circuits and is also highly efficient due to its in-situ memory computation.
[0039] The output of the non-volatile memory cell array 33 is provided to a differential summer (such as a summing operational amplifier or a summing current mirror) 38, which sums the output of the non-volatile memory cell array 33 to create a single value for the convolution. The differential summer 38 is arranged to perform the summation of positive and negative weights.
[0040] The summed output of the difference summer 38 is then provided to the activation function block 39, which modifies the output. Activation function block 39 can provide a sigmoid, tanh, or ReLU function. The modified output of activation function block 39 becomes the next layer's output (e.g., ...). Figure 6 The elements of the feature map of layer C1 are then applied to the next synapse to produce the next feature map layer or the final layer. Thus, in this example, the non-volatile memory cell array 33 constitutes multiple synapses (which receive their input from existing neuron layers or from input layers such as an image database), and the summing operational amplifier 38 and the activation function block 39 constitute multiple neurons.
[0041] Figure 7The inputs to the VMM array 32 (WLx, EGx, CGx and optional BLx and SLx) can be analog, binary or digital (in which case a DAC is provided to convert the digital bits to the appropriate input analog level), and the output can be analog, binary or digital (in which case an output ADC is provided to convert the output analog level to digital bits).
[0042] Figure 8 A block diagram illustrating the use of the multilayer VMM array 32, labeled VMM arrays 32a, 32b, 32c, 32d, and 32e. (See diagram for reference.) Figure 8 As shown, the input (denoted as Inputx) is converted from digital to analog by a digital-to-analog converter 31 and provided to the input VMM array 32a. The converted analog input can be voltage or current. The first-level input D / A conversion can be accomplished by using a function or LUT (lookup table) of appropriate analog level to map Inputx to the input VMM array 32a. Input conversion can also be accomplished by an analog-to-analog (A / A) converter to convert the external analog input into a mapped analog input to the input VMM array 32a.
[0043] The output generated by the input VMM array 32a is provided as input to the next VMM array (hidden level 1) 32b, which in turn generates the output provided as input to the next VMM array (hidden level 2) 32c, and so on. The various layers of the VMM array 32 serve as different layers of synapses and neurons in a convolutional neural network (CNN). Each VMM array 32a, 32b, 32c, 32d, and 32e can be an independent physical non-volatile memory array, or multiple VMM arrays can utilize different portions of the same physical non-volatile memory array, or multiple VMM arrays can utilize overlapping portions of the same physical non-volatile memory array. Figure 8 The example shown contains five layers (32a, 32b, 32c, 32d, 32e): one input layer (32a), two hidden layers (32b, 32c), and two fully connected layers (32d, 32e). Those skilled in the art will understand that this is merely an example, and conversely, a system may include more than two hidden layers and more than two fully connected layers.
[0044] Vector-Matrix Multiplication (VMM) Array
[0045] Figure 9 A neuronal VMM array 900 is depicted, which is particularly suitable for... Figure 3The memory cell 310 shown serves as a synapse and component for neurons between the input layer and the next layer. The VMM array 900 includes a memory array 901 of non-volatile memory cells and a reference array 902 of non-volatile reference memory cells (at the top of the array). Alternatively, another reference array may be placed at the bottom.
[0046] In the VMM array 900, control gate lines (such as control gate line 903) extend vertically (therefore, reference array 902 is orthogonal to control gate line 903 in the row direction), and erase gate lines (such as erase gate line 904) extend horizontally. Here, the inputs to the VMM array 900 are set on control gate lines (CG0, CG1, CG2, CG3), and the outputs of the VMM array 900 appear on source lines (SL0, SL1). In one example, only even rows are used, and in another example, only odd rows are used. The currents placed on each source line (SL0, SL1, respectively) perform a summation function of all currents from the memory cells connected to that particular source line.
[0047] As described herein with respect to neural networks, the non-volatile memory cells of the VMM array 900 (i.e., memory cells 310 of the VMM array 900) can be configured to operate in the subthreshold region.
[0048] Biasing the non-volatile reference memory cell and non-volatile memory cell described in this paper within the weak inversion (subthreshold region):
[0049] Ids = Io * e (Vg-Vth) / nVt =w*Io*e (Vg) / nVt ,
[0050] Where w = e (-Vth) / nVt
[0051] Where Ids is the drain-to-source current; Vg is the gate voltage on the memory cell; Vth is the threshold voltage of the memory cell; Vt is the thermal voltage = k*T / q, where k is the Boltzmann constant, T is the temperature in Kelvin, and q is the electron charge; n is the slope factor = 1 + (Cdep / Cox), where Cdep = the capacitance of the depletion layer, and Cox is the capacitance of the gate oxide layer; Io is the memory cell current at the gate voltage equal to the threshold voltage, and Io is related to (Wt / L)*u*Cox*(n-1)*Vt 2 Proportional, where u is the carrier mobility, and Wt and L are the width and length of the memory cell, respectively.
[0052] For I-to-V logarithmic converters that use memory cells (such as reference memory cells or peripheral memory cells) or transistors to convert input current into input voltage:
[0053] Vg = n * Vt * log[Ids / wp * Io]
[0054] Where wp represents the w in the reference memory cell or the peripheral memory cell.
[0055] For a memory array used as a vector matrix multiplier (VMM) array with current input, the output current is:
[0056] Iout = wa * Io * e (Vg) / nVt ,Right now
[0057] Iout = (wa / wp) * Iin = W * Iin
[0058] W = e (Vthp-Vtha) / nVt
[0059] Here, wa = w for each memory cell in the memory array.
[0060] Vthp is the effective threshold voltage of the peripheral memory cell, and Vtha is the effective threshold voltage of the primary (data) memory cell. Note that the threshold voltage of the transistor is a function of the substrate bulk bias voltage, and the substrate bulk bias voltage, denoted as Vsb, can be modulated to compensate for various conditions at this temperature. The threshold voltage Vth can be expressed as:
[0061] Vth=Vth0+γ(SQRT|Vsb-2*φF)-SQRT|2*φF|)
[0062] Where Vth0 is the threshold voltage with zero substrate bias, φF is the surface potential, and γ is the host effect parameter.
[0063] Word lines or control gates can be used as inputs to memory cells that accept input voltages.
[0064] Alternatively, the flash memory cells of the VMM array described herein can be configured to operate in a linear region:
[0065] Ids=β*(Vgs-Vth)*Vds;β=u*Cox*Wt / L
[0066] W = α(Vgs - Vth)
[0067] This means that the weight W in the linear region is proportional to (Vgs-Vth).
[0068] Word lines, control gates, bit lines, or source lines can be used as inputs to memory cells operating in a linear region. Bit lines or source lines can be used as outputs to memory cells.
[0069] For an IV linear converter, memory cells (such as reference memory cells or peripheral memory cells) or transistors operating in the linear region can be used to linearly convert input / output current into input / output voltage.
[0070] Alternatively, the memory cells of the VMM array described herein can be configured to operate in a saturation region:
[0071] Ids = 1 / 2 * β * (Vgs - Vth) 2 β=u*Cox*Wt / L
[0072] Wα(Vgs-Vth) 2 This means that the weight W is related to (Vgs-Vth). 2 proportional
[0073] Word lines, control gates, or erase gates can be used as inputs to memory cells operating in saturation regions. Bit lines or source lines can be used as outputs of output neurons.
[0074] Alternatively, the memory cells of the VMM array described herein can be used in all regions or combinations thereof (subthreshold, linear, or saturated regions) of each or more layers of a neural network.
[0075] It is described in U.S. Patent No. 10,748,630 Figure 7 Other examples of the VMM array 32 are described herein by reference. As described in this application, source lines or bit lines can be used as neuron outputs (current summation outputs).
[0076] Figure 10 A neuronal VMM array 1000 is depicted, which is particularly suitable for... Figure 2 The memory cell 210 shown serves as a synapse between the input layer and the next layer. The VMM array 1000 includes a memory array 1003 of non-volatile memory cells, a reference array 1001 of first non-volatile reference memory cells, and a reference array 1002 of second non-volatile reference memory cells. The reference arrays 1001 and 1002, arranged in the column direction of the array, are used to convert current inputs flowing into terminals BLR0, BLR1, BLR2, and BLR3 into voltage inputs WL0, WL1, WL2, and WL3. In practice, the first and second non-volatile reference memory cells are diode-connected via a multiplexer 1014 (partially depicted) through which current inputs flow. The reference cells are tuned (e.g., programmed) to a target reference level. The target reference level is provided by a reference microarray matrix (not shown).
[0077] Memory array 1003 serves two purposes. First, it stores the weights used by VMM array 1000 on their respective memory cells. Second, memory array 1003 efficiently multiplies the inputs (i.e., the current inputs provided in terminals BLR0, BLR1, BLR2, and BLR3, which reference arrays 1001 and 1002 convert into input voltages to supply word lines WL0, WL1, WL2, and WL3) by the weights stored in memory array 1003, and then adds all the results (memory cell currents) to produce an output on the corresponding bit lines (BL0 to BLN), which will be the input to the next layer or to the final layer. By performing multiplication and addition functions, memory array 1003 eliminates the need for separate multiplication and addition logic circuits and is also highly efficient. Here, voltage inputs are provided on word lines WL0, WL1, WL2, and WL3, and the outputs appear on the corresponding bit lines BL0 to BLN during read (inference) operations. The current placed on each of the bit lines BL0 to BLN performs a summation function of the currents from all non-volatile memory cells connected to that particular bit line.
[0078] Table 5 depicts the operating voltages and currents used for the VMM array 1000. The columns in the table indicate the voltages applied to the word lines for selected cells, word lines for unselected cells, bit lines for selected cells, bit lines for unselected cells, source lines for selected cells, and source lines for unselected cells. The rows indicate read, erase, and program operations.
[0079] surface_: Figure 10 Operation of VMM array 1000 :
[0080]
[0081] Figure 11 A neuronal VMM array 1100 is depicted, which is particularly suitable for... Figure 2The memory cell 210 shown serves as a synapse and component for neurons between the input layer and the next layer. The VMM array 1100 includes a memory array 1103 of non-volatile memory cells, a reference array 1101 of first non-volatile reference memory cells, and a reference array 1102 of second non-volatile reference memory cells. Reference arrays 1101 and 1102 extend in the row direction of the VMM array 1100. The VMM array is similar to VMM 1000, except that in VMM array 1100, word lines extend in the vertical direction. Here, inputs are set on word lines (WLA0, WLB0, WLA1, WLB2, WLA2, WLB2, WLA3, WLB3), and outputs appear on source lines (SL0, SL1) during read operations. The currents placed on each source line perform a summation function of all currents from the memory cells connected to that particular source line.
[0082] Table 6 depicts the operating voltages and currents used for the VMM array 1100. The columns in the table indicate the voltages applied to the word lines for selected cells, word lines for unselected cells, bit lines for selected cells, bit lines for unselected cells, source lines for selected cells, and source lines for unselected cells. The rows indicate read, erase, and program operations.
[0083] Table 6: Figure 11 Operation of VMM array 1100
[0084]
[0085] Figure 12 A neuronal VMM array 1200 was described, which is particularly suitable for... Figure 3 The memory cell 310 shown serves as a synapse and component for neurons between the input layer and the next layer. The VMM array 1200 includes a memory array 1203 of non-volatile memory cells, a reference array 1201 of first non-volatile reference memory cells, and a reference array 1202 of second non-volatile reference memory cells. Reference arrays 1201 and 1202 are used to convert current inputs flowing into terminals BLR0, BLR1, BLR2, and BLR3 into voltage inputs CG0, CG1, CG2, and CG3. In practice, the first and second non-volatile reference memory cells are diode-connected via a multiplexer 1212 (partially shown), through which current inputs flow via BLR0, BLR1, BLR2, and BLR3. Each multiplexer 1212 includes a corresponding multiplexer 1205 and a common-source cascode transistor 1204 to ensure that the voltage on the bit lines (such as BLRO) of each of the first and second non-volatile reference memory cells remains constant during read operations. The reference cells are tuned to a target reference level.
[0086] Memory array 1203 serves two purposes. First, it stores weights that will be used by VMM array 1200. Second, memory array 1203 efficiently multiplies the inputs (current inputs provided to terminals BLR0, BLR1, BLR2, and BLR3, which reference arrays 1201 and 1202 convert into input voltages to be provided to the control gates (CG0, CG1, CG2, and CG3)) by the weights stored in the memory array, and then sums all the results (cell currents) to produce an output that appears on BL0-BLN and will be the input to the next layer or the final layer. By performing multiplication and addition functions, the memory array eliminates the need for separate multiplication and addition logic circuits and is also highly efficient. Here, the inputs are provided on the control gate lines (CG0, CG1, CG2, and CG3), and the outputs appear on the bit lines (BL0-BLN) during read operations. The currents placed on each bit line perform a summation function of all the currents from the memory cells connected to that particular bit line.
[0087] The VMM array 1200 performs unidirectional tuning of the non-volatile memory cells in the memory array 1203. That is, each non-volatile memory cell is erased and then partially programmed until the desired charge is reached on the floating gate. If too much charge is placed on the floating gate (causing an incorrect value to be stored in the cell), the cell is erased, and the sequence of partial programming operations restarts. As shown, two rows sharing the same erase gate (such as EG0 or EG1) are erased together (this may be referred to as page erasure), and thereafter, each cell is partially programmed until the desired charge is reached on the floating gate.
[0088] Table 7 depicts the operating voltages and currents used for the VMM array 1200. The columns in the table indicate the voltages applied to the word lines for selected cells, word lines for unselected cells, bit lines for selected cells, bit lines for unselected cells, control gates for selected cells, control gates for unselected cells in the same sector as the selected cell, control gates for unselected cells in different sectors from the selected cell, erase gates for selected cells, erase gates for unselected cells, source lines for selected cells, and source lines for unselected cells. Rows indicate read, erase, and program operations.
[0089] Table 7: Figure 12 Operation of VMM array 1200
[0090]
[0091] Figure 13 A neuronal VMM array 1300 was described, which is particularly suitable for... Figure 3The memory cell 310 shown serves as a synapse and component for neurons between the input layer and the next layer. The VMM array 1300 includes a memory array 1303 of non-volatile memory cells, a reference array 1301 of first non-volatile reference memory cells, and a reference array 1302 of second non-volatile reference memory cells. EG lines EGR0, EG0, EG1, and EGR1 extend vertically, while CG lines CG0, CG1, CG2, and CG3 and SL lines WL0, WL1, WL2, and WL3 extend horizontally. The VMM array 1300 is similar to the VMM array 1400, except that the VMM array 1300 implements bidirectional tuning, where each individual cell can be fully erased, partially programmed, and partially erased as needed to achieve a desired amount of charge on the floating gate due to the use of separate EG lines. As shown in the figure, reference arrays 1301 and 1302 convert the input currents in terminals BLR0, BLR1, BLR2, and BLR3 into control gate voltages CG0, CG1, CG2, and CG3 to be applied to memory cells in the row direction (through the operation of reference cells connected via diodes via multiplexer 1314). The current outputs (neurons) are in bit lines BL0-BLN, where each bit line sums all currents from non-volatile memory cells connected to that particular bit line.
[0092] Table 8 depicts the operating voltages and currents used for the VMM array 1300. The columns in the table indicate the voltages applied to the word lines for the selected cell, the word lines for the unselected cell, the bit lines for the selected cell, the bit lines for the unselected cell, the control gate for the selected cell, the control gate for the unselected cell in the same sector as the selected cell, the control gate for the unselected cell in a different sector from the selected cell, the erase gate for the selected cell, the erase gate for the unselected cell, the source line for the selected cell, and the source line for the unselected cell. The rows indicate read, erase, and program operations.
[0093] Table 8: Figure 13 Operation of VMM array 1300
[0094]
[0095] Figure 22 A neuronal VMM array 2200 was depicted, which is particularly suitable for... Figure 2 The memory unit 210 shown serves as a synapse and component for neurons between the input layer and the next layer. In the VMM array 2200, inputs INPUT0, ..., INPUT... N On bit lines BL0, ..., BL respectively NThe data is received and outputs OUTPUT1, OUTPUT2, OUTPUT3 and OUTPUT4 are generated on source lines SL0, SL1, SL2 and SL3 respectively.
[0096] Figure 23 A neuronal VMM array 2300 was depicted, which is particularly suitable for... Figure 2 The memory unit 210 shown serves as a synapse and component for neurons between the input layer and the next layer. In this example, inputs INPUT0, INPUT1, INPUT2, and INPUT3 are received on source lines SL0, SL1, SL2, and SL3, respectively, and outputs OUTPUT0, ..., OUTPUT... N In the bit lines BL0, ..., BL N Generate above.
[0097] Figure 24 The neuronal VMM array 2400 is depicted, which is particularly suitable for... Figure 2 The memory unit 210 shown serves as a synapse and component for neurons between the input layer and the next layer. In this example, the inputs are INPUT0, ..., INPUT0. M On the word lines WL0, ..., WL respectively M The data is received and outputs OUTPUT0, ..., OUTPUT. N In the bit lines BL0, ..., BL N Generate above.
[0098] Figure 25 A neuronal VMM array 2500 was depicted, which is particularly suitable for... Figure 3 The memory unit 310 shown serves as a synapse and component for neurons between the input layer and the next layer. In this example, the inputs are INPUT0, ..., INPUT0. M On the word lines WL0, ..., WL respectively M The data is received and outputs OUTPUT0, ..., OUTPUT. N In the bit lines BL0, ..., BL N Generate above.
[0099] Figure 26 The neuronal VMM array 2600 is depicted, which is particularly suitable for... Figure 4 The memory unit 410 shown serves as a synapse and component for neurons between the input layer and the next layer. In this example, the inputs are INPUT0, ..., INPUT0. nOn the vertical control grid lines CG0, ..., CG respectively N The data is received, and outputs OUTPUT1 and OUTPUT2 are generated on source lines SL0 and SL1.
[0100] Figure 27 The neuronal VMM array 2700 is depicted, which is particularly suitable for... Figure 4 The memory unit 410 shown serves as a synapse and component for neurons between the input layer and the next layer. In this example, the inputs are INPUT0, ..., INPUT0. N They are received on the gates of bit line control gates 2701-1, 2701-2, ..., 2701-(N-1) and 2701-N, respectively, and these gates are coupled to bit lines BL0, ..., BL0, respectively. N Example outputs OUTPUT1 and OUTPUT2 are generated on source lines SL0 and SL1.
[0101] Figure 28 A neuronal VMM array 2800 was depicted, which is particularly suitable for... Figure 3 The memory unit 310 shown Figure 5 The memory cell 510 shown and Figure 7 The memory unit 710 shown serves as a synapse and component for neurons between the input layer and the next layer. In this example, the inputs are INPUT0, ..., INPUT0. M In the word lines WL0, ..., WL M The data is received and outputs OUTPUT0, ..., OUTPUT. N On bit lines BL0, ..., BL respectively N Generate above.
[0102] Figure 29 The neuronal VMM array 2900 is depicted, which is particularly suitable for... Figure 3 The memory unit 310 shown Figure 5 The memory cell 510 shown and Figure 7 The memory unit 710 shown serves as a synapse and component for neurons between the input layer and the next layer. In this example, the inputs are INPUT0, ..., INPUT0. M In the control grid lines CG0, ..., CG M The data is received. Outputs OUTPUT0, ..., OUTPUT N On the vertical source lines SL0, ..., SL respectively N The above is generated, where each source line SLi The source line is coupled to all memory cells in column i.
[0103] Figure 30 A neuronal VMM array 3000 is depicted, which is particularly suitable for... Figure 3 The memory unit 310 shown Figure 5 The memory cell 510 shown and Figure 7 The memory unit 710 shown serves as a synapse and component for neurons between the input layer and the next layer. In this example, the inputs are INPUT0, ..., INPUT0. M In the control grid lines CG0, ..., CG M The data is received. Outputs OUTPUT0, ..., OUTPUT N On the vertical lines BL0, ..., BL respectively N Generate on, where each bit line BL i Bit lines coupled to all memory cells in column i.
[0104] Long Short-Term Memory
[0105] Existing technologies include a concept known as Long Short-Term Memory (LSTM). LSTM cells are commonly used in neural networks. LSTM allows neural networks to remember information for predetermined arbitrary time intervals and use that information in subsequent operations. A typical LSTM cell includes a cell, an input gate, an output gate, and a forget gate. The three gates regulate the flow of information into and out of the cell and the time interval at which information is remembered in the LSTM. Virtual Memory Models (VMMs) are particularly useful in LSTM cells.
[0106] Figure 14 An example LSTM 1400 is depicted. This example LSTM 1400 includes cells 1401, 1402, 1403, and 1404. Cell 1401 receives the input vector x0 and generates the output vector h0 and the cell state vector c0. Cell 1402 receives the input vector x1, the output vector (hidden state) h0 from cell 1401, and the cell state c0 from cell 1401, and generates the output vector h1 and the cell state vector c1. Cell 1403 receives the input vector x2, the output vector (hidden state) h1 from cell 1402, and the cell state c1 from cell 1402, and generates the output vector h2 and the cell state vector c2. Cell 1404 receives the input vector x3, the output vector (hidden state) h2 from cell 1403, and the cell state c2 from cell 1403, and generates the output vector h3. Additional cells can be used, and this four-cell LSTM is merely an example.
[0107] Figure 15 Depicting what can be used Figure 14 The following is a specific example implementation of LSTM cell 1500, comprising cells 1401, 1402, 1403, and 1404. LSTM cell 1500 receives an input vector x(t), a cell state vector c(t-1) from the previous cell, and an output vector h(t-1) from the previous cell, and generates the cell state vector c(t) and the output vector h(t).
[0108] LSTM unit 1500 includes sigmoid function devices 1501, 1502, and 1503, each applying a number between 0 and 1 to control how much of each component in the input vector is allowed to pass through to the output vector. LSTM unit 1500 also includes tanh devices 1504 and 1505 for applying a hyperbolic tangent function to the input vector, multiplier devices 1506, 1507, and 1508 for multiplying two vectors together, and adder device 1509 for adding two vectors together. The output vector h(t) can be provided to the next LSTM unit in the system, or it can be accessed for other purposes.
[0109] Figure 16 LSTM unit 1600 is depicted, which is an example of a specific implementation of LSTM unit 1500. For the reader's convenience, LSTM unit 1600 uses the same numbering as LSTM unit 1500. Sigmoid function devices 1501, 1502, and 1503, and tanh device 1504 each include multiple VMM arrays 1601 and activation function blocks 1602. Thus, it can be seen that VMM arrays are particularly useful in LSTM units used in certain neural network systems. Multiplier devices 1506, 1507, and 1508, and adder device 1509 are implemented digitally or analogically. Activation function blocks 1602 can be implemented digitally or analogously.
[0110] Alternative forms of the LSTM unit 1600 (and another example of a specific implementation of the LSTM unit 1500) are in Figure 17 As shown in [the image]. Figure 17In this context, sigmoid function devices 1501, 1502, and 1503, as well as tanh device 1504, share the same physical hardware (VMM array 1701 and activation function block 1702) in a time-division multiplexing manner. The LSTM unit 1700 also includes a multiplier device 1703 for multiplying two vectors together, an adder device 1708 for adding two vectors together, a tanh device 1505 (which includes an activation function block 1702), a register 1707 for storing the value i(t) when the value i(t) is output from the sigmoid function block 1702, a register 1704 for storing the value f(t)*c(t-1) when it is output from the multiplier device 1703 via multiplexer 1710, a register 1705 for storing the value i(t)*u(t) when it is output from the multiplier device 1703 via multiplexer 1710, a register 1706 for storing the value o(t)*c~(t) when it is output from the multiplier device 1703 via multiplexer 1710, and a multiplexer 1709.
[0111] LSTM unit 1600 contains multiple VMM arrays 1601 and corresponding activation function blocks 1602, while LSTM unit 1700 contains a set of VMM arrays 1701 and activation function blocks 1702, which are used to represent multiple layers in the example of LSTM unit 1700. LSTM unit 1700 will require less space than LSTM 1600 because LSTM unit 1700 only needs 1 / 4 of its space for VMMs and activation function blocks compared to LSTM unit 1600.
[0112] It is also understood that an LSTM cell will typically comprise multiple VMM arrays, each using functionality provided by certain circuit blocks outside the VMM array itself (such as summer and activation function blocks, and high-voltage generation blocks). Providing a separate circuit block for each VMM array would require a significant amount of space within the semiconductor device and would be inefficient to some extent. Therefore, the examples described below reduce the amount of circuitry provided outside the VMM arrays themselves.
[0113] Gate control recursive unit
[0114] A simulated VMM implementation can be used in gated recurrent unit (GRU) systems. A GRU is a gated mechanism in a recurrent neural network. A GRU is similar to an LSTM, but a GRU unit typically contains fewer components than an LSTM unit.
[0115] Figure 18An example GRU 1800 is depicted. This example GRU 1800 includes units 1801, 1802, 1803, and 1804. Unit 1801 receives an input vector x0 and generates an output vector h0. Unit 1802 receives an input vector x1, an output vector h0 from unit 1801, and generates an output vector h1. Unit 1803 receives an input vector x2 and an output vector (hidden state) h1 from unit 1802 and generates an output vector h2. Unit 1804 receives an input vector x3 and an output vector (hidden state) h2 from unit 1803 and generates an output vector h3. Additional units can be used, and this four-unit GRU is merely an example.
[0116] Figure 19 Depicting what can be used Figure 18 Examples of specific implementations of GRU unit 1900 are provided for units 1801, 1802, 1803, and 1804. GRU unit 1900 receives an input vector x(t) and an output vector h(t-1) from a previous GRU unit, and generates an output vector h(t). GRU unit 1900 includes sigmoid function devices 1901 and 1902, each applying a number between 0 and 1 to the components from the output vector h(t-1) and the input vector x(t). GRU unit 1900 also includes a tanh device 1903 for applying a hyperbolic tangent function to the input vector, multiple multiplier devices 1904, 1905, and 1906 for multiplying two vectors together, an adder device 1907 for adding two vectors together, and a complementary device 1908 for subtracting the input from 1 to generate the output.
[0117] Figure 20 GRU unit 2000 is depicted, which is an example of a specific implementation of GRU unit 1900. For the reader's convenience, GRU unit 2000 uses the same numbering as GRU unit 1900. Figure 20 As can be seen, sigmoid function devices 1901 and 1902, and tanh device 1903 each include multiple VMM arrays 2001 and activation function blocks 2002. Therefore, it can be seen that VMM arrays are particularly useful in GRU units used in some neural network systems. Multiplier devices 1904, 1905, and 1906, adder device 1907, and complement device 1908 are implemented digitally or analogically. Activation function blocks 2002 can be implemented digitally or analogously.
[0118] Alternative forms of the GRU unit 2000 (and another example of a specific implementation of the GRU unit 1900) in Figure 21 As shown in [the image]. Figure 21In this configuration, the GRU unit 2100 utilizes a VMM array 2101 and an activation function block 2102, which, when configured as a sigmoid function, applies numbers between 0 and 1 to control how much of each component in the input vector is allowed to pass through to the output vector. Figure 21 In this context, sigmoid function devices 1901 and 1902 and tanh device 1903 share the same physical hardware (VMM array 2101 and activation function block 2102) in a time-division multiplexing manner. GRU unit 2100 also includes a multiplier device 2103 for multiplying two vectors together, an adder device 2105 for adding two vectors together, a complement device 2109 for subtracting an input from 1 to generate an output, a multiplexer 2104, a register 2106 for holding the value h(t-1)*r(t) when it is output from the multiplier device 2103 via the multiplexer 2104, a register 2107 for holding the value h(t-1)*z(t) when it is output from the multiplier device 2103 via the multiplexer 2104, and a register 2108 for holding the value h^(t)*(1-z(t)) when it is output from the multiplier device 2103 via the multiplexer 2104.
[0119] GRU unit 2000 contains multiple sets of VMM arrays 2001 and activation function blocks 2002, while GRU unit 2100 contains a set of VMM arrays 2101 and activation function blocks 2102, which are used to represent multiple layers in the example of GRU unit 2100. GRU unit 2100 will require less space than GRU unit 2000 because GRU unit 2100 only needs 1 / 3 of its space for VMMs and activation function blocks compared to GRU unit 2000.
[0120] It is also understandable that a GRU system would typically include multiple VMM arrays, each using functionality provided by certain circuit blocks outside the VMM array itself (such as summer and activation function blocks, and high-voltage generation blocks). Providing a separate circuit block for each VMM array would require a significant amount of space within the semiconductor device and would be inefficient to some extent. Therefore, the examples described below reduce the amount of circuitry provided outside the VMM array itself.
[0121] Input and output
[0122] The inputs to a VMM array can be analog levels, binary levels, pulses, time-modulated pulses, or digital bits (in which case a DAC is used to convert the digital bits to the appropriate input analog level), and the outputs can be analog levels, binary levels, timed pulses, pulses, or digital bits (in which case an output ADC is used to convert the output analog level to digital bits).
[0123] Differential unit
[0124] Generally, for each memory cell in a VMM array, each weight W can be implemented by a single memory cell, a differential cell, or a hybrid memory cell (the average of two cells). In the case of differential cells, two memory cells are used to implement the weight W as a differential weight (W = W+ - W-). In the case of two hybrid memory cells, two memory cells are used to implement the weight W as the average of the two cells.
[0125] Figure 31 A VMM system 3100 is depicted. In some examples, the weights W stored in the VMM array are stored as differential pairs W+ (positive weights) and W- (negative weights), where W = (W+) - (W-). In the VMM system 3100, half of the bit lines are designated as W+ lines, i.e., bit lines connected to memory cells that will store positive weights W+, and the other half of the bit lines are designated as W- lines, i.e., bit lines connected to memory cells that implement negative weights W-. The W- lines are distributed alternately among the W+ lines. The subtraction operation is performed by summing circuits (such as summing circuits 3101 and 3102), which receive current from the W+ and W- lines. The outputs of the W+ and W- lines are combined to effectively give W = W+ - W- for each (W+, W-) cell pair of all (W+, W-) line pairs. Although the W- lines, which are alternately distributed between the W+ lines, have been described above, in other examples, the W+ and W- lines can be located anywhere in the array.
[0126] Figure 32 Another example is depicted. In the VMM system 3210, positive weights W+ are implemented in a first array 3211 and negative weights W- are implemented in a second array 3212, which is separate from the first array, and the resulting weights are appropriately combined by a summing circuit 3213.
[0127] Figure 33A VMM system 3300 is depicted. Weights W stored in the VMM array are stored as difference pairs W+ (positive weights) and W- (negative weights), where W = (W+) - (W-). The VMM system 3300 includes arrays 3301 and 3302. Half of the bit lines in each of arrays 3301 and 3302 are designated as W+ lines, i.e., bit lines connected to memory cells that will store positive weights W+, and the other half of the bit lines in each of arrays 3301 and 3302 are designated as W- lines, i.e., bit lines connected to memory cells that implement negative weights W-. The W- lines are distributed alternately among the W+ lines. Subtraction operations are performed by summing circuits (such as summing circuits 3303, 3304, 3305, and 3306) that receive current from the W+ and W- lines. The outputs of the W+ line and the W- line from each array 3301, 3302 are combined to effectively give W = W+ - W- for each (W+, W-) cell pair of all (W+, W-) line pairs. Furthermore, the W values from each array 3301 and array 3302 can be further combined by summing circuits 3307 and 3308 such that each W value is the result of subtracting the W value from array 3302 from the W value from array 3301. This means that the final result from summing circuits 3307 and 3308 is one of two differences.
[0128] Number of weights per layer
[0129] Each non-volatile memory cell used in an analog neural memory system can be erasable and programmable to maintain a very specific and precise amount of charge (i.e., the number of electrons) in a floating gate. Each floating gate can hold one of N distinct values, where N is the number of different weights that can be indicated by each cell. Examples of N include 16, 32, 64, 128, and 256.
[0130] Refer again Figure 6 and Figure 8 Neural networks contain multiple layers. Some layers do not require analog neural memory systems. For example, some layers (such as the first layer of a residual neural network (RESNET)) only require units capable of storing two weights ("0" or "1"), i.e., N=2, which are binary weights. Such layers do not need to use analog neural memory systems because a simpler digital system suffices. Using analog neural memory systems for such layers is inefficient because the precision provided by analog neural memory systems (e.g., N=256) is not utilized by layers that require N=2. Summary of the Invention
[0131] Numerous examples of neural networks comprising one or more analog in-memory computation engines and one or more digital in-memory computation engines are disclosed. This allows the use of digital CIM engines when layers contain binary weights instead of analog weights. Attached Figure Description
[0132] Figure 1 The following is a diagram illustrating an artificial neural network.
[0133] Figure 2 The prior art split-gate flash memory cell is described.
[0134] Figure 3 Another prior art split-gate flash memory cell is described.
[0135] Figure 4 Another prior art split-gate flash memory cell is described.
[0136] Figure 5 Another prior art split-gate flash memory cell is described.
[0137] Figure 6 This is a diagram illustrating different levels of an example artificial neural network utilizing one or more non-volatile memory arrays.
[0138] Figure 7 Here is a block diagram of an example VMM system.
[0139] Figure 8 Here is a block diagram illustrating an example artificial neural network utilizing one or more VMM systems.
[0140] Figure 9 Another example of a VMM system is described.
[0141] Figure 10 Another example of a VMM system is described.
[0142] Figure 11 Another example of a VMM system is described.
[0143] Figure 12 Another example of a VMM system is described.
[0144] Figure 13 Another example of a VMM system is described.
[0145] Figure 14 The existing long short-term memory system is described.
[0146] Figure 15 An example cell used in a long short-term memory system is depicted.
[0147] Figure 16 Depicting Figure 15 The specific implementation of the unit is an example.
[0148] Figure 17 Depicting Figure 15 Another example of a specific implementation of the unit.
[0149] Figure 18 A prior art gate-controlled recursive cell system is described.
[0150] Figure 19 An example cell used in a gate-controlled recursive cell system is depicted.
[0151] Figure 20 Depicting Figure 19 The specific implementation of the unit is an example.
[0152] Figure 21 Depicting Figure 19 Another example of a specific implementation of the unit.
[0153] Figure 22 Another example of a VMM system is described.
[0154] Figure 23 Another example of a VMM system is described.
[0155] Figure 24 Another example of a VMM system is described.
[0156] Figure 25 Another example of a VMM system is described.
[0157] Figure 26 Another example of a VMM system is described.
[0158] Figure 27 Another example of a VMM system is described.
[0159] Figure 28 Another example of a VMM system is described.
[0160] Figure 29 Another example of a VMM system is described.
[0161] Figure 30 Another example of a VMM system is described.
[0162] Figure 31 Another example of a VMM system is described.
[0163] Figure 32 Another example of a VMM system is described.
[0164] Figure 33 Another example of a VMM system is described.
[0165] Figure 34 It describes a computing system in analog memory.
[0166] Figure 35 A computing system within a digital memory is described.
[0167] Figure 36 A hybrid system comprising analog in-memory computing systems and digital in-memory computing systems is described.
[0168] Figure 37 It describes a computing engine within simulated memory.
[0169] Figure 38 An example computational engine within a digital memory is described.
[0170] Figure 39 Another example of a computing engine within digital memory is described.
[0171] Figure 40 Another example of a computing engine within digital memory is described.
[0172] Figure 41A The multiplication logic unit is described.
[0173] Figure 41B A 2-bit adder logic unit is described.
[0174] Figure 41C The digital CIM unit is described.
[0175] Figure 42 Another example of a computing engine within digital memory is described.
[0176] Figure 43 The tree of shifters and adders is depicted.
[0177] Figure 44 Another shifter and adder tree is depicted.
[0178] Figure 45 An example dynamic weight engine is described.
[0179] Figure 46 Another example of a dynamic weight engine is described.
[0180] Figure 47 An integral analog-to-digital converter is described.
[0181] Figure 48 An integral analog-to-digital converter is described.
[0182] Figure 49A and Figure 49B The output circuit, including a current-to-voltage converter and an analog-to-digital converter, is described.
[0183] Figure 50 Depicting by Figure 36 The method of execution of a hybrid system. Detailed Implementation
[0184] Simulated in-memory computing engine
[0185] Figure 34 A block diagram of an analog in-memory computing (CIM) engine 3400 is depicted. The analog CIM engine 3400 includes a VMM array 3401 (which may also be referred to as a neural network array), a row decoder 3402, a high-voltage decoder 3403, a column decoder 3404, bitline drivers 3405 (such as bitline control circuitry for programming), input circuitry 3406, output circuitry 3407, control logic unit 3408, and a bias generator 3409. The analog CIM engine 3400 also includes a high-voltage generation block 3410, which includes a charge pump 3411, a charge pump regulator 3412, and a high-voltage level generator 3413. The analog CIM engine 3400 also includes an algorithm controller 3414, analog circuitry 3415, a control engine 3416 (which may include functions such as arithmetic functions, activation functions, embedded microcontroller logic, but is not limited thereto), test control logic components 3417, and static random access memory (SRAM) blocks 3418 (for which they are programmed / erased or weighted) to store intermediate data such as for input circuitry (e.g., activation data) or for output circuitry (neuron output data, partial and output neuron data), or for programming data inputs (e.g., for a whole row or for multiple rows). The VMM array 3401 includes an array of non-volatile memory cells arranged in rows and columns, wherein the non-volatile memory cells are... Figure 2 , Figure 3 , Figure 4 or Figure 5 The types shown are memory cells 210, 310, 410, or 510, or other types known to those skilled in the art. In one example, a non-volatile memory cell is as follows: Figure 2 , Figure 3 or Figure 4 The split-gate flash memory cell in the example. In another example, the non-volatile memory cell is as follows: Figure 5 Stacked gate flash memory cells in the memory.
[0186] Input circuitry 3406 may include circuitry such as a DAC (digital-to-analog converter), a DPC (digital-to-pulse converter or digital-to-time modulated pulse converter), an AAC (analog-to-analog converter, such as a current-to-voltage converter or a logarithmic converter), a PAC (pulse-to-analog level converter), or any other type of converter. Input circuitry 3406 may implement one or more of a normalization, linear, or nonlinear up / down scaling function or arithmetic function. Input circuitry 3406 may implement a temperature compensation function for the input level. Input circuitry 3406 may implement activation functions such as a rectified linear activation function (ReLU) or a sigmoid function. Input circuitry 3406 may store digital activation data that will be applied as an input signal or combined with an input signal during programming or read operations. The digital activation data may be stored in a register. Input circuitry 3406 may include circuitry for driving array terminals, such as CG, WL, EG, and SL lines, which may include sample-and-hold circuitry and buffers. The DAC may be used to convert the digital activation data into an analog input voltage to be applied to the array.
[0187] Output circuitry 3407 may include circuitry such as ITV (current-to-voltage circuitry), ADC (analog-to-digital converter for converting analog neuron outputs into digital bits), AAC (analog-to-analog converter, such as a current-to-voltage converter or a logarithmic converter), APC (analog-to-pulse converter or analog-to-time-modulated pulse converter), or any other type of converter. Output circuitry 3407 may convert array outputs into activation data. Output circuitry 3407 may implement activation functions such as ReLU or sigmoid. Output circuitry 3407 may implement one or more of the following for neuron outputs: statistical normalization, regularization, up / down scaling / gain functions, statistical rounding, or arithmetic functions (e.g., addition, subtraction, division, multiplication, shifting, logarithmic). Output circuitry 3407 may implement temperature compensation functions for neuron outputs or array outputs (such as bitline outputs) to keep the array's power consumption approximately constant or to improve the accuracy of the array (neuron) outputs, such as by keeping the IV slope approximately the same over a temperature range. Output circuitry 3407 may include registers for storing output data.
[0188] Computation in digital memory
[0189] Figure 35 A block diagram of a digital in-memory computing (CIM) engine 3500 is depicted. The digital CIM engine 3500 contains many, but not all, of the components included in the analog CIM engine 3400.
[0190] The digital CIM engine 3500 includes an array 3501, a row decoder 3502, a column decoder 3504, a bitline driver 3505 (such as bitline control circuitry for programming), input circuitry 3506, output circuitry 3507, control logic unit 3508, and bias generator 3509. The digital CIM engine 3500 also includes an algorithm controller 3514, analog circuitry 3515, a control engine 3516, and test control logic unit 3517. The array 3501 includes an array of non-volatile memory cells arranged in rows and columns, wherein the non-volatile memory cells are... Figure 2 , Figure 3 , Figure 4 or Figure 5 The types shown are memory cells 210, 310, 410, or 510, or other types known to those skilled in the art. In one example, a non-volatile memory cell is as follows: Figure 2 , Figure 3 or Figure 4 The split-gate flash memory cell in the example. In another example, the non-volatile memory cell is as follows: Figure 5 The array 3501 operates in the digital domain, where the values stored in each non-volatile memory cell are binary values, rather than analog values as in the analog CIM engine 3400.
[0191] Input circuitry 3506 may include circuitry such as a DAC, DPC, AAC, PAC, or any other type of converter. Input circuitry 3506 may implement one or more of a normalization, linear, or nonlinear up / down scaling function or arithmetic function. Input circuitry 3506 may implement a temperature compensation function for the input level. Input circuitry 3506 may implement activation functions such as ReLU or sigmoid. Input circuitry 3506 may store digital activation data that will be applied as an input signal or combined with an input signal during programming or read operations. The digital activation data may be stored in a register. Input circuitry 3506 may include circuitry for driving array terminals, such as CG, WL, EG, and SL lines, which may include sample-and-hold circuitry and buffers. A DAC may be used to convert the digital activation data into an analog input voltage to be applied to the array. Output circuitry 3507 may include circuitry such as an ITV, ADC, AAC, APC, or any other type of converter. Output circuitry 3507 may convert the array output into activation data. Output circuitry 3507 may implement activation functions such as ReLU or sigmoid. Output circuitry 3507 may implement one or more of the following for neuron output: statistical normalization, regularization, up / down scaling / gain functions, statistical rounding, or arithmetic functions (e.g., addition, subtraction, division, multiplication, shifting, logarithmic). Output circuitry 3507 may implement a temperature compensation function for neuron output or array output (such as bitline output) to keep the array's power consumption approximately constant or to improve the accuracy of the array (neuron) output, such as by keeping the IV slope approximately the same over a temperature range. Output circuitry 3507 may include registers for storing output data.
[0192] In an alternative example, the digital CIM engine 3500 does not utilize any analog circuitry, but operates purely in the digital domain. For instance, the input circuitry 3506 and the output circuitry 3507 can be formed by purely digital logic circuitry operating in the digital domain, and the analog circuitry 3515 can be removed.
[0193] A hybrid system including an analog in-memory computing engine, a digital in-memory computing engine, and a dynamic weighting engine.
[0194] Figure 36 Hybrid system 3600 is described. The hybrid system includes: analog CIM engines 3601, 3602 and 3604; digital CIM engines 3605 and 3607; digital computing engines 3606 and 3608; dynamic weighting engine 3603; and system bus 3609.
[0195] The simulated CIM engines 3601, 3602 and 3604 store the simulated weights (e.g., N=256) in the non-volatile memory cells of the respective VMM array (e.g., VMM array 3401) and perform VMM operations on these simulated weights. Figures 9 to 13 and Figures 22 to 33 The VMM system described herein is an example of a simulated CIM engine.
[0196] Computation in Digital Memory (CIM) engines 3605 and 3607 store digital weights (i.e., N=2) in non-volatile memory cells of the corresponding VMM array (e.g., VMM array 3501) and perform VMM operations on these digital weights. Alternatively, digital CIM engines 3605 and 3607 may store digital weights in non-volatile macros or chips. Unlike analog CIM engines, digital CIM engines, in addition to having non-volatile memory cells with floating gates (such as...), Figures 2 to 5 In addition to the non-volatile memory cells shown, arrays of SRAM or DRAM cells can also be used.
[0197] The digital computing engines 3606 and 3608 are microprocessors, digital signal processors, or other digital computing engines, such as GPUs (Graphics Processing Units), TPUs (Tensor Processing Units), dedicated MAC (Multiply-Add-Accumulate) units, dedicated SIMD (Single Instruction Majority) processors, vector extensions, etc. The digital computing engines 3606 and 3608 can perform integer or floating-point calculations. The digital computing engines 3606 and 3608 can use VMM arrays from analog CIM engines or weights from digital non-volatile memory macros or chips.
[0198] The Dynamic Weight Engine 3603 is a device whose stored weights can be modified by changing the bias voltage or bias current without performing separate erase or programming operations. Because the weights can be modified without performing separate erase or programming operations, they are considered dynamic weights. Weights can be transferred from the VMM array in the analog CIM engine or from digital non-volatile memory macros or chips. Figure 45 and Figure 46 An example of a dynamic weighting engine 3603 is shown, in which a bias voltage is transferred to the gate of a transistor in a cell CIM array. Dynamic weighting control and biasing circuitry are not shown.
[0199] The dynamic weighting engine 3603, the analog CIM engines 3601, 3602 and 3604, the digital CIM engines 3605 and 3607, and the digital computing engines 3606 and 3608 are coupled to the system bus 3609, which enables all coupled devices to communicate with each other.
[0200] The microcontroller 3610, SRAM 3611, vector register 3612, and peripheral control and interface logic 3613 assist certain functions and operations in the hybrid system 3600 (such as, but not limited to, controlling access to SRAM 3611, controlling data flow between various computing engines, controlling communication between the internal system bus 3609 and the external bus, performing activation functions, performing read, erase, and program operations on non-volatile memory (NVM) or volatile memory (VM), or performing weight transfer between computing engines).
[0201] Figure 50 An example method that can be executed by the hybrid system 3600 is described. First, the system performs vector-matrix multiplication in a first layer of the neural network using a digital in-memory computing engine 3605 or 3607 (5001). Second, the system performs vector-matrix multiplication in a second layer of the neural network, different from the first layer, using an analog in-memory computing engine 3601, 3602, or 3604 (5002). Optionally, the system transfers dynamic weights from a dynamic weight engine 3603 to one or more of the analog in-memory computing engines 3601, 3602, or 3604 and / or the digital in-memory computing engines 3605 or 3607 (5003). Optionally, the system stores the dynamic weights in the analog in-memory computing engine 3601, 3602, or 3604.
[0202] Figure 37 A portion of an example of an analog CIM engine 3700 is depicted. Here, two rows and four columns of non-volatile memory cells are shown. The memory cells operate in the subthreshold region and can store analog weights. Activation inputs are provided to each selected row on a word line or control gate, and the memory cells in that row multiply these activation inputs by the stored analog weights, outputting a current representing the product in a bit line coupled to the column of each cell. The currents are then summed along the bit lines (output lines). The array output currents are then converted by an ADC circuit (not shown) to produce digital output bits. The activation inputs can be provided as the output of a pulse-width (PW) converter, a fixed-bias 1-bit digital-to-analog converter, or an n-bit DAC. Figures 9 to 13 and Figures 22 to 33 Other examples show a portion of the simulated CIM engine.
[0203] A portion of the analog CIM engine 3700 can also be used as an analog memory weight storage device. In this case, a word of memory cell is selected (e.g., 64 or 128 cells from the selected row), and the output digital bit from the ADC represents the weight value stored by that portion of the analog CIM engine 3700.
[0204] Figure 38A portion of an example digital CIM engine 3800 is depicted. This portion of the digital CIM engine 3800 includes an array of current-based CIM SRAM memory cells 3801, 3802, 3803, and 3804 arranged in rows and columns. CIM SRAM memory cell 3801 will be described in detail, and it should be understood that the other CIM SRAM memory cells 3801 have the same design. In CIM SRAM memory cell 3801, if the weight stored in the SRAM is "1", NMOS transistor 3816 is turned on, and current NC flows from transistor 3815 towards bit line BL0 to transistor 3817. If the weight stored in the SRAM is "0", NMOS transistor 3816 is turned off, and current NC is prevented from flowing from transistor 3815 towards bit line BL0 to transistor 3817. Current NC is the effective value of the weight stored in the CIM SRAM cell.
[0205] During the operation of the digital CIM engine, an input is applied to the input line IN0. The SRAM memory cell 3801 performs a multiplication of the input on IN0 with the stored weight. If IN0 is "1", the NMOS transistor 3817 is turned on to transfer the current NC (representing the stored weight value) to the bit line BL0, and if IN0 is "0", the NMOS transistor 3817 is turned off and no current NC is transferred to the bit line BL0.
[0206] If the stored weight is "1" and the input on the input line is "1", bit line BL0 will receive current NC because NMOS transistors 3815, 3816, and 3817 are all turned on. If the stored weight is "0" or the input on the input line is "0", bit line BL0 will be disconnected from current NC and will not contain any current attributable to SRAM memory cell 3801.
[0207] The CIM SRAM memory cells 3802, 3803, and 3804, as well as other CIM SRAM memory cells in the digital CIM engine, operate in the same manner as CIM SRAM memory cell 3801.
[0208] The bit lines of the digital CIM engine (such as BL0 and BL1) sum the currents of all CIM SRAM memory cells connected to them. This means that BL0 will contain the sum of all currents (all NC currents with weight = "1") of all CIM SRAM memory cells in column 0, BL1 will contain the sum of all currents of all CIM SRAM memory cells in column 1, and so on. This current can be converted into digital output bits by an ADC circuit or a multi-bit sense amplifier circuit.
[0209] Figure 39 A portion of another example digital CIM engine 3900 is depicted. A portion of the digital CIM engine 3900 includes an array of charge-based CIMSRAM memory cells 3901, 3902, 3903, and 3904 arranged in rows and columns.
[0210] CIM SRAM memory cell 3901 represents CIM SRAM memory cells 3902, 3903, and 3904, as well as other CIM SRAM memory cells in the digital CIM engine, and its operation will now be described. SRAM memory cell 3901 includes: inverters 3911 and 3912, which form a latch; NMOS transistors 3913, 3914, 3916, 3917, and 3918; and capacitor 3915. A charge (representing a weight value) Q (equal to C (the capacitance of the capacitor) * Vdds (which is the supply voltage of the SRAM memory cell)) is stored on capacitor 3915 in the SRAM cell. The latch formed by inverters 3911 and 3912 can store the value "0" or "1" (= the output of inverter 3912), where "0" represents zero charge and "1" represents Q charge, which is the effective weight of SRAM memory cell 3901. When the output of inverter 3912 is "0", meaning = ground (0V), capacitor 3915 will charge to ground through transistor 3916, therefore Q = 0 (representing "0 weight"). When the output of inverter 3912 is "1" (e.g., = Vdds), capacitor 3915 will charge to Vdds through transistor 3916, therefore Q = C * Vdds (representing "1" weight). When transistor 3916 is off (ENSB = "0"), transistor 3917 is on (input IN0 = "1"), and transistor 3918 is on (ENS = "1"), the stored charge is transferred to bit line BL0. For a weight of "1", if input IN0 = "1" (meaning transistor 3917 is on), the charge Q (= C * Vddss) will be transferred to bit line BL0.
[0211] The weights are stored in SRAM memory cell 3901 via a write operation. During the write operation, word line WL0 is asserted, which turns on NMOS transistors 3913 and 3914. A "1" is stored in the latch by driving BLS0B high and BLS0 low, and a "0" is stored in the latch by driving BLS0 high and BLS0B low.
[0212] During the inference or read operation of the digital CIM engine, an input is applied to input line IN0, and ENSB is asserted. SRAM memory cell 3901 performs a multiplication of the input on IN0 with the stored weights. If IN0 is "1", NMOS transistor 3917 is turned on, and if WLN0 is "0", NMOS transistor 3917 is turned off.
[0213] If the stored weight is "1" and if the input INx on the input line is "1", then bit line BL0 will receive the charge of capacitor 3915 because NMOS transistor 3916 is off and transistors 3917 and 3918 are on. If the stored weight is "0" or if the input on input line INx0 is "0", then bit line BL0 will receive 0 charge and will not contain any charge component attributable to SRAM memory cell 3801.
[0214] SRAM memory cells 3902, 3903, and 3904, as well as other SRAM memory cells in the digital CIM engine, operate in the same manner as SRAM memory cell 3901.
[0215] The bit lines of the digital CIM engine (such as BL0 and BL1) sum the charges of all SRAM memory cells connected to them. This means that BL0 will contain the sum of all charges output by all SRAM memory cells in column 0, BL1 will contain the sum of all charges output by all SRAM memory cells in column 1, and so on. This summed charge can be generated by an ADC circuit (such as those described below). Figure 44 and Figure 46 The circuit shown converts the digital output bits into digital output bits.
[0216] Figure 40 A portion of another example digital CIM engine 4000 is depicted. This portion of the digital CIM engine 4000 includes subarrays of CIM digital cells, where each CIM digital cell stores a 1-bit weight w. In this example, subarrays 4010-11 and 4010-12 are shown in detail. Horizontally, the digital CIM engine 4000 includes m subarrays (such as subarrays 4010-11, 4010-12, ..., 4010-1m in row 1). Vertically, the digital CIM engine 4000 includes n subarrays (such as subarrays 4010-11, ..., 4010-n1 in column 1).
[0217] Subarray 4010-11 will now be described as an example. The discussion of subarray 4010-11 also applies to other subarrays. Subarray 4010-11 comprises an array of CIM digital cells (such as CIM digital cell 4001) arranged in rows and columns. In this example, subarray 4010-11 comprises four rows and four columns of CIM digital cells, but it should be understood that, according to the same principles discussed herein, subarray 4010-11 and other subarrays may alternatively comprise any number of rows and columns, such as 8 rows and 8 columns, to accommodate 8-bit inputs and 8-bit weights (meaning that 1 bit is stored in each of the 8 CIM digital cells, collectively storing 8-bit weights), resulting in a 16-bit output (because 8-bit inputs multiplied by 8-bit weights yield a 16-bit output). In this example, subarray 4010-11 implements a 4-bit input multiplied by a 4-bit weight, outputting DOUT[7:0] = IN[3:0] * W[3:0], where each row of 4 CIM digital cells stores 4 weights W[3:0], and all four rows store the same weight W[3:0]. The MUX block is used to enable the output from each subarray 4010x to the main output bus DOUTx.
[0218] The digital CIM engine 4000 receives input for multiple rows. In this example, subarray 4010-11 receives inputs IN[0], IN[1], IN[2], and IN[3] across its four rows. The 4-bit input is multiplied by a 4-bit weight W[3:0] stored in each row. Each digital unit (such as digital unit 4001) receives input at its port I1, multiplies it by a 1-bit weight value (W) stored in the digital unit, and adds the resulting product to any data received at its port I2 from the digital unit in the row above it and any carry bits received at its port CI from the adjacent digital unit to its right in the same row. In the example shown, the leftmost column (Col3) represents the most significant bit of the weight, the column to its right (Col2) represents the next most significant bit of the weight, the next column to its right (Col1) represents the next most significant bit of the weight, and the last column (Col0) represents the least significant bit of the weight.
[0219] Each CIM digital cell (such as CIM digital cell 4001) includes a memory device (such as an SRAM cell, latch, or register) for storing weights W. Each digital cell (such as CIM digital cell 4001) also includes digital multiplication and addition logic. Inputs are received digitally by the digital device in each row of the I1 port of each CIM digital cell. The inputs received on port I1 are multiplied by the weights stored in the SRAM of the CIM digital cell, and the results are then added column by column and output. The output (O) of a particular CIM digital cell is provided as input (I2) to the CIM digital cell located below it in the same column, and the carry output (CO) from the particular CIM digital cell is provided as carry input (CI) to the CIM digital cell to its left in the same row, since that cell represents a more significant bit than the particular CIM digital cell.
[0220] Multiplexer block 4020-11 is used to enable the output from subarray 4010-11 to the main output bus DOUTx in response to the enable signal EN0. Other subarrays have similar multiplexer blocks.
[0221] Figure 41A , Figure 41B and Figure 41C Depicting what can be Figure 40 Example components used in each CIM digital unit (such as CIM digital unit 4001).
[0222] refer to Figure 41A The multiplication logic unit 4101 receives two inputs I1 and I2, and outputs their product O according to the truth table shown in Table 9:
[0223] Table 9: Truth Table of Multiplication Logic Unit 4101
[0224] I1 I2 O 0 0 0 0 1 0 1 0 0 1 1 1
[0225] refer to Figure 41B The 2-bit adder logic unit 4102 receives inputs I1 and I2 and a carry input CI, and generates output O and carry output CO according to the truth table shown in Table 10:
[0226] Table 10: Truth Table of 2-bit Adder Logic Unit 4102
[0227] I1 I2 C1 O CO 0 0 0 0 0 0 1 0 1 0 1 0 0 1 0 1 1 0 0 1 0 0 1 1 0 0 1 1 0 1 1 0 1 0 1 1 1 1 1 1
[0228] refer to Figure 41CThe digital CIM unit 4103 stores the weight W in the SRAM unit 4104. This digital CIM unit uses the multiplication logic unit 4101 to multiply W with the received input I2, and adds the resulting product P to the received inputs I1 and CI, thereby generating outputs O and CO using the 2-bit adder logic unit 4102. In this way, each digital CIM unit 4103 performs a one-bit multiplication operation (W*I2) and adds the result to the received inputs (I1 and CI) to perform vector-matrix multiplication using the stored digital value W.
[0229] Figure 42 A portion of another example digital CIM engine 4200 is depicted. This portion of the digital CIM engine 4200 includes a block array, where each block includes an array of multipliers and a tree of shifters and adders. In this example, blocks 4210-11, 4210-12, 4210-21, and 4210-22 are shown in detail. Horizontally, the digital CIM engine 4200 includes m blocks (such as blocks 4210-11, 4210-12, ..., 4212-1m in row 1). Vertically, the digital CIM engine 4200 includes n blocks, such as blocks 4210-11, 4210-21, ..., 4210-n1 in column 1.
[0230] Block 4210-11 will now be described as an example. The discussion of block 4210-11 also applies to other blocks. Block 4210-11 comprises an array of multipliers arranged in rows and columns, such as multiplier 4201. In this example, block 4210-11 comprises a four-row, four-column multiplier, but it should be understood that, based on the same principles discussed here, block 4210-11 and all other blocks may alternatively comprise any number of rows and columns, such as 8 rows and 8 columns, to accommodate 8-bit inputs and 8-bit weights, resulting in a 16-bit output.
[0231] The digital CIM engine 4200 receives input for multiple rows. In this example, block 4210-11 receives inputs IN[0], IN[1], IN[2], and IN[3] on its four rows. Each multiplier (such as multiplier 4201) receives the input and multiplies it by a 1-bit weight (W) stored in the multiplier, and outputs the product. The product is fed to a shift and adder tree 4202, which adds the products and performs shift operations (as appropriate) to reflect the fact that each row (IN[0], IN[1], IN[2], and IN[3]) represents a different number in the binary input (IN[3:0]). That is, IN[0] represents 2. 0 IN[1] represents 2 1 IN[2] represents 2 2 And IN[3] means 23 Multiplexer block 4203 is used to enable the output from each block (such as blocks 4210-11) to the main output bus DOUTx in response to the enable signal EN0. Other blocks have similar multiplexer blocks.
[0232] Figure 43 A shifter and adder tree 4300 is depicted, which is Figure 42 An example of a specific implementation of the shift and adder tree 4202 used is provided. Each multiplier (such as multiplier 4201) in a block (such as blocks 4210-11) receives a 1-bit input, multiplies that input by a 1-bit weight, and outputs a 1-bit value provided to the shift and adder tree 4202. For each row, the shift and adder tree 4300 receives four 1-bit values, which the shift and adder tree treats as a single 4-bit value arranged in the order corresponding to columns D3, D2, D1, and D0, respectively. The 4-bit value for row 0 (which receives input IN[0]) is X4-0, the 4-bit value for row 1 (which receives input IN[1]) is X4-1, the 4-bit value for row 2 (which receives input IN[2]) is X4-2, and the 4-bit value for row 3 (which receives input IN[3]) is X4-3.
[0233] Shift and adder 4301 first adds X4-0 to X4-1 (whereby the shift and adder shifts the number of X4-1 one digit to the left because row 1 represents the value of the input number to the left of row 0; for example, if X4-1 is 1010, the shifted version will be 10100), resulting in a 6-bit value. Then, shift and adder 4302 adds this 6-bit value to the shifted version of X4-2 (where X4-2 is shifted by two digits), resulting in a 7-bit value. Then, shift and adder 4303 adds this 7-bit value to X4-3 (where X4-3 is shifted by three digits), resulting in an 8-bit value that represents the sum of the products received from the block's multipliers.
[0234] Figure 44 A shifter and adder tree 4400 is depicted, which is Figure 42An example of a specific implementation of the shift and adder tree 4202 used is provided. Each multiplier (such as multiplier 4201) in a block (such as blocks 4210-11) receives a 1-bit input, multiplies that input by a 1-bit weight, and outputs a 1-bit value provided to the shift and adder tree 4202. For each row, the shift and adder tree 4400 receives four 1-bit values, which the shift and adder tree treats as a single 4-bit value arranged in the order corresponding to columns D3, D2, D1, and D0, respectively. The 4-bit value for row 0 (which receives input IN[0]) is X4-0, the 4-bit value for row 1 (which receives input IN[1]) is X4-1, the 4-bit value for row 2 (which receives input IN[2]) is X4-2, and the 4-bit value for row 3 (which receives input IN[3]) is X4-3.
[0235] Shift and adder 4401 first adds X4-0 to X4-1 (whereby X4-1 is shifted one digit to the left because row 1 represents the value of the input number to the left of row 0; for example, if X4-1 is 1010, the shifted version will be 10100), resulting in a 6-bit value X6. Then, shift and adder 4302 adds the shifted version of X4-2 (where X4-2 is shifted two digits) to the shifted version of X4-3 (where X4-3 is shifted three digits), resulting in a 7-bit value X7. Finally, shift and adder 4403 adds X6 and X7 (without shifting), resulting in an 8-bit value X8, which represents the sum of the products received from the block's multipliers.
[0236] Another use of the digital CIM array is for floating-point numbers, where the first array is used for the exponent part and the second array is used for the mantissa part (fractional part), where proper realignment and reorganization of the array outputs are performed to reconstruct the final floating-point number.
[0237] Figure 45 A portion of the dynamic weighting engine 4500 is depicted. The dynamic weighting engine 4500 includes CIM dynamic weighting units 4501, 4502, 4503, and 4504 arranged in rows and columns.
[0238] CIM dynamic weighting unit 4501 represents CIM dynamic weighting units 4502, 4503, and 4504 in the dynamic weighting engine, as well as other CIM memory units, and its operation will now be described. CIM dynamic weighting unit 4501 includes NMOS transistors 4511, 4512, 4514, and 4515, and a capacitor 4513. CIM dynamic weighting unit 4501 can store weights that can be dynamically changed by modifying the input. CIM dynamic weighting unit 4501 is selected when WL0 and WBL0 are asserted. NMOS transistors 4511 and 4514 are turned on when WL0 is asserted. NMOS transistor 4512 is turned on when WBL0 is asserted. A voltage W (representing the weight applied by the CIM dynamic weighting unit 4501) is applied to WWL0 and transferred to the gate of NMOS transistor 4515, and stored by capacitor 4513 coupled between the gate of NMOS transistor 4515 and SL0, which is the source line of row 0. NMOS transistor 4515 draws current from bit line BL0, where the drawn current is a function of voltage W.
[0239] The CIM dynamic weight units 4502, 4503, 4504 and other CIM dynamic weight units in the dynamic weight engine operate in the same way as CIM dynamic weight unit 4501.
[0240] The bit lines of the CIM dynamic weighting engine (such as BL0 and BL1) sum the current drawn by all CIM dynamic weighting cells connected to them. This means that BL0 will contain the sum of all current drawn by all CIM dynamic weighting cells in column 0, BL1 will contain the sum of all current drawn by all eight CIM dynamic weighting cells in column 1, and so on. This current can be controlled by an ADC circuit (such as those described below). Figure 47 (and the circuit shown in Figure 49) is converted into digital output bits.
[0241] Figure 46 A portion of the dynamic weighting engine 4600 is depicted. The dynamic weighting engine 4600 includes CIM dynamic weighting units 4601, 4602, 4603 and 4604 arranged in rows and columns.
[0242] CIM dynamic weighting unit 4601 represents CIM dynamic weighting units 4602, 4603, and 4604, as well as other CIM dynamic weighting units in the dynamic weighting engine, and its operation will now be described. CIM dynamic weighting unit 4601 includes NMOS transistors 4611, 4612, 4613, 4615, and 4616, and a capacitor 4614. CIM dynamic weighting unit 4601 can store weights that can be dynamically changed by modifying the input. CIM dynamic weighting unit 4601 is selected when WL0 and WBL are asserted. When WL0 is asserted, NMOS transistors 4611 and 4615 are turned on. When WBL is asserted, NMOS transistor 4613 is turned on. A voltage W (representing the weight applied by the CIM dynamic weighting unit 4601) is applied to WWL0 and transferred to the gate of NMOS transistor 4616, and stored by capacitor 4614 coupled between the gate of NMOS transistor 4616 and SL0, which is the source line of row 0. NMOS transistor 4616 draws current from bit line BL0, where the drawn current is a function of voltage W.
[0243] The CIM dynamic weight units 4602, 4603, 4604 and other CIM dynamic weight units in the dynamic weight engine operate in the same way as CIM dynamic weight unit 4601.
[0244] The bit lines of the CIM dynamic weighting engine (such as BL0 and BL1) sum the current drawn by all CIM dynamic weighting cells connected to them. This means that BL0 will contain the sum of all current drawn by all CIM dynamic weighting cells in column 0, BL1 will contain the sum of all current drawn by all eight CIM dynamic weighting cells in column 1, and so on. This current can be controlled by an ADC circuit (such as those described below). Figure 47 (and the circuit shown in Figure 49) is converted into digital output bits.
[0245] Figure 47 Describes the use of storing Figure 39 The example of an integrating ADC 4700 is an integrator that converts the charge in the cells of this part of the digital CIM engine 3900 into digital pulses or digital output bits. The integrating ADC 4700 converts the analog output charge in the neuron output block (the charge stored in capacitor 4702CINT) into digital pulses whose width varies proportionally to the magnitude of the analog output current in the neuron output block. Capacitor 4702 represents... Figure 39The ADC 4700 sums the charges of all capacitors from the selected CIM volatile SRAM memory cell (connected in parallel via channel transistors). The ADC 4700 includes a reference current 4707, an operational amplifier 4701, a comparator 4704, an AND gate 4740, and a counter 4720. The charges of the individual capacitors in the CIM volatile SRAM memory cell are summed and then reconfigured across the integrating capacitor 4702 of the operational amplifier 4701. The reference current 4707 discharges all stored charges in the integrating capacitor 4702, and the counter 4720 generates pulses until the discharge is complete. The pulses are represented as follows: Figure 48 The digital output bits are shown.
[0246] Figure 48 Describing the use of from Figure 37 Simulated CIM engine in Figure 38 Digital CIM Engine 3800 Figure 45 Dynamic Weight Engine 4500 and Figure 46 The dynamic weighting engine 4600 in the ADC 4800 converts the current received by one or more columns into digital pulses or digital output bits.
[0247] ADC 4800 will transfer current I NEU This is converted into a digital pulse EC, the width of which varies proportionally to the magnitude of the current. The integrator, comprising an integrating operational amplifier 4801 and an integrating capacitor 4802, is used to input I. NEU 4806 integrates relative to the reference current IREF provided by the current source 4807.
[0248] Optionally, the current source 4807 may include a temperature coefficient of 0 or a temperature coefficient that tracks the neuron current I. NEU The bandgap filter. Tracking I NEU The temperature coefficient can be obtained from a lookup table (not shown) containing values determined during the testing phase.
[0249] During the initialization phase, switch 4808 is closed. Then, the inputs to Vout 4803 and the negative terminal of operational amplifier 4801 become equal to the VREF value. Afterward, switch 4808 is opened, and switch S2 is closed while switch S1 remains open, and the constant reference current IREF 4807 is integrated upward during the fixed time period tref. During the fixed time period tref, Vout rises, and its slope reflects the value of the constant reference current IREF 4807. Afterward, switch S2 is opened, and switch S1 is closed, and during the time period tmeas, the neuron current I... NEU4806 is integrated down over the time interval tmeas (during which Vout decreases), where tmeas is the time to integrate Vout down to VREF, as indicated by the change in the output of comparator 4804.
[0250] When Vout > VREFV, the output of EC 4805 will be high, and vice versa. EC 4805 therefore generates a pulse whose width reflects the time interval tmeas, which in turn is related to the current I. NEU It is proportional to 4806. Therefore, the output neuron current I NEU The 4806 is converted into a digital pulse EC 4805, where the width of the digital pulse EC 4805 is related to the output neuron current I. NEU The value of 4806 changes proportionally.
[0251] Current I NEU 4806 = tmeas / tref * IREF. For example, for a required 10-bit output resolution, tref is equivalent to a time period of 1024 clock cycles. According to I... NEU The values of 4806 and Iref, and the time period tmeas, vary over a period of 0 to 1024 clock cycles. Neuron current I NEU 4806 will affect the charging rate and slope.
[0252] Optionally, the output pulse EC 4805 can be converted into a series of pulses with uniform time intervals to be transmitted to the next block of the circuit, such as the input block of the CIM engine. At the beginning of the time interval tmeas, the output EC 4805, along with the reference clock 4841, is input into AND gate 4840. During the time interval Vout > VREF, the output will be a pulse sequence 4842 (where the frequency of the pulses in the pulse sequence 4842 is the same as the frequency of the clock 4841). The number of pulses is proportional to the time interval tmeas, which is related to the current I. NEU It is directly proportional to 4806.
[0253] Optionally, the pulse sequence 4843 can be input to a counter 4820, which counts the number of pulses in the pulse sequence 4842 and generates a count value 4821, which is a digital count of the number of pulses in the pulse sequence 4842, and this digital count is correlated with the neuronal current I. NEU 4806 is proportional. The count value 4821 includes a set of digital bits. In another example, the integrating ADC 4800 can measure the neuron current I. NEU 4806 is converted into a pulse, where the width of the pulse is related to the neuron current I. NEUThe value of 4806 is inversely proportional. This inversion can be done digitally or analogically and is converted into a series of pulses or digital bits for output to a follower circuit.
[0254] Figure 49A An example output circuit 4900 is depicted that can convert array current into digital output bits. Output circuit 4900 includes a current-to-voltage converter (ITV) 4901 and an analog-to-digital converter (ITV) 4902. ITV 4901 converts the array current into a voltage, which is then digitized by ADC 4902. ADC 4902 can be a successive approximation (SAR) ADC.
[0255] Figure 49B An example output circuit 4950 is depicted, comprising an ITV 4951 that receives a differential input and generates a differential output, and an ADC 4952 that receives a differential input from the differential output of the ITV 4951 and generates a single-ended digital output. Examples of specific implementations of current-to-voltage converters and analog-to-digital converters are described in U.S. Patent Application No. 17 / 521,772, which is incorporated herein by reference.
[0256] Other examples of analog and digital CIM engines may utilize NAND memory, resistive RAM (ReRAM), magnetic RAM (MRAM), or dynamic RAM (DRAM), but are not limited to these.
[0257] It should be noted that, as used herein, the terms "above" and "on" both encompass "directly on" (without intermediate material, elements, or space between) and "indirectly on" (with intermediate material, elements, or space between). Similarly, the term "adjacent" includes "directly adjacent" (without intermediate material, elements, or space between) and "indirectly adjacent" (with intermediate material, elements, or space between), "mounted to" includes "directly mounted to" (without intermediate material, elements, or space between) and "indirectly mounted to" (with intermediate material, elements, or space between), and "electrically coupled to" includes "directly electrically coupled to" (without intermediate material or elements electrically connecting the elements together) and "indirectly electrically coupled to" (with intermediate material or elements electrically connecting the elements together). For example, forming an element "above the substrate" can include forming an element directly on the substrate without intermediate material / elements between them, and forming an element indirectly on the substrate with one or more intermediate materials / elements between them.
Claims
1. A system comprising: An in-memory computing engine is used to perform operations in the first layer of a neural network; and A digital memory-in-memory computing engine for performing operations in a second layer of the neural network that is different from the first layer.
2. The system according to claim 1, wherein the system comprises: A system bus, which is coupled to the computing engine in the analog memory and the computing engine in the digital memory.
3. The system of claim 1, wherein the in-memory computing engine comprises a plurality of non-volatile memory cells arranged in rows and columns.
4. The system of claim 3, wherein the non-volatile memory cell is a stacked gate flash memory cell.
5. The system according to claim 3, wherein the non-volatile memory cell is a split-gate flash memory cell.
6. The system of claim 1, wherein the in-memory computing engine comprises a plurality of static random access memory (SRAM) cells arranged in rows and columns.
7. The system of claim 1, wherein the in-memory computing engine comprises a plurality of CIM digital units arranged in rows and columns.
8. The system according to claim 7, wherein the plurality of CIM digital units respectively include a multiplication logic unit and a 2-bit adder logic unit.
9. The system of claim 7, wherein the in-digital memory computing engine includes a shifter and adder tree coupled to the plurality of CIM digital units.
10. A system comprising: An in-memory computing engine is used to perform operations in the first layer of a neural network; A computing engine within a digital memory, the computing engine within a digital memory being used to perform operations in a second layer of the neural network; and A dynamic weight engine, which performs operations in the third layer of the neural network; The first layer, the second layer, and the third layer are different layers in the neural network.
11. The system of claim 10, wherein the system comprises: A system bus, which is coupled to the in-analog memory computing engine, the in-digital memory computing engine, and the dynamic weighting engine.
12. The system of claim 10, wherein the in-memory computing engine comprises a plurality of non-volatile memory cells arranged in rows and columns.
13. The system of claim 12, wherein the non-volatile memory cell is a stacked gate flash memory cell.
14. The system of claim 12, wherein the non-volatile memory cell is a split-gate flash memory cell.
15. The system of claim 10, wherein the in-memory computing engine comprises a plurality of static random access memory (SRAM) cells arranged in rows and columns.
16. The system of claim 10, wherein the in-memory computing engine comprises a plurality of CIM digital units.
17. The system according to claim 16, wherein the plurality of CIM digital units respectively include a multiplication logic unit and a 2-bit adder logic unit.
18. The system of claim 16, wherein the in-digital memory computing engine includes a shifter and adder tree coupled to the plurality of CIM digital units.
19. A method comprising: Vector-matrix multiplication is performed in the first layer of the neural network using a computation engine within digital memory. as well as Vector-matrix multiplication is performed in a second layer of the neural network, which is different from the first layer, using an in-memory computing engine.
20. The method of claim 19, wherein the in-memory computing engine comprises a plurality of non-volatile memory cells arranged in rows and columns.
21. A method, the method comprising: Vector-matrix multiplication is performed in the first layer of the neural network using a computation engine within digital memory. Vector-matrix multiplication is performed in the second layer of the neural network using an in-memory computing engine; and The dynamic weight engine is used to perform vector-matrix multiplication in the third layer of the neural network; The first layer, the second layer, and the third layer are different layers in the neural network.
22. The method of claim 21, wherein the in-memory computing engine comprises a plurality of non-volatile memory cells arranged in rows and columns.
23. A method, the method comprising: Use one of the analog memory-in-memory computation engine, the digital memory-in-memory computation engine, and the dynamic weight engine to perform vector-matrix multiplication in the first layer of the neural network; as well as Vector-matrix multiplication is performed in the second layer of the neural network using another of the analog in-memory computing engine, the digital in-memory computing engine, and the dynamic weight engine.
24. A method, the method comprising: The weights received from the analog memory computing engine are stored in one or more of the digital memory computing engine, the digital computing engine, and the dynamic weight engine.
Citation Information
Patent Citations
High precision and highly efficient tuning mechanisms and algorithms for analog neuromorphic memory in artificial neural networks
US10748630B2
Deep Learning Neural Network Classifier Using Non-volatile Memory Array
US20170337466A1
Output circuitry for analog neural memory in a deep learning artificial neural network
US20230049032A1
Single transistor non-valatile electrically alterable semiconductor memory device
US5029130A
Flash memory cells with separated self-aligned select and erase gates, and process of fabrication
US6747310B2