Decoding system and physical layout for analog neural memory in deep learning artificial neural network

The improved decoding system and physical layout for analog neural memory systems using non-volatile memory cells address the challenge of high space and energy consumption in neural networks by enabling efficient vector matrix multiplication and reducing the need for separate logic circuits, enhancing energy efficiency and space utilization.

JP2025102761AActive Publication Date: 2025-07-08SILICON STORAGE TECHNOLOGY INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2025031044
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-07-03
Filing Date
2025-02-28
Publication Date
2025-07-08
Estimated Expiration
2039-11-17

AI Technical Summary

Technical Problem

The development of high-performance artificial neural networks is hindered by the lack of appropriate hardware technology that efficiently supports a large number of synapses with low energy consumption and minimal physical space requirements.

Method used

An improved decoding system and physical layout for an analog neural memory system utilizing non-volatile memory cells, specifically CMOS technology and non-volatile memory arrays, which allow for individual programming and reading of memory cells without interference, enabling efficient vector matrix multiplication and reducing the need for separate multiplication and addition logic circuits.

Benefits of technology

This approach enhances the energy efficiency and reduces the physical space required for neural networks by utilizing non-volatile memory cells that can be precisely adjusted and programmed independently, facilitating efficient in-memory computing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025102761000001_ABST
    Figure 2025102761000001_ABST
Patent Text Reader

Abstract

To provide a word line decoder, a control gate decoder, a bit line decoder, a low voltage row decoder, and a high voltage row decoder and various types of physical layout designs for non-volatile flash memory arrays in an analog neural system.SOLUTION: Combined word line and control gate decoder 4100 comprises: a PMOS transistor 4102; an NMOS transistor 4103; row address signals 4104; vertical input word lines 4105; a horizontal word output line 4106 which is coupled to word lines of VMM arrays; an inverter 4107; switches 4108 and 4112; and an isolation transistor 4109, and receives a control gate input 4110 CGIN0 and outputs a control gate line 4111 CG0. The word line output 4106 WL0 and control gate output 4111 CG0 are selected or de-selected at the same times by decoding logic that controls a NAND gate 4101.SELECTED DRAWING: Figure 41
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] (Claiming Priority) This application claims priority to U.S. Provisional Patent Application No. 62 / 840,318, filed Apr. 29, 2019, entitled "DECODING SYSTEM AND PHYSICAL LAYOUT FOR ANALOG NEURAL MEMORY IN DEEP LEARNING ARTIFICIAL NEURAL NETWORK", and U.S. Patent Application No. 16 / 503,355, filed Jul. 3, 2019, entitled "DECODING SYSTEM AND PHYSICAL LAYOUT FOR ANALOG NEURAL MEMORY IN DEEP LEARNING ARTIFICIAL NEURAL NETWORK".

[0002] (Field of the Invention) An improved decoding system and physical layout are disclosed for an analog neural memory system that utilizes non-volatile memory cells.

Background Art

[0003] An artificial neural network mimics a biological neural network (the central nervous system of an animal, particularly the brain), can rely on a large number of inputs, and is used to estimate or approximate a generally unknown function. An artificial neural network generally includes layers of interconnected "neurons" that exchange messages.

[0004] FIG. 1 shows an artificial neural network, in which the circles represent input or layers of neurons. The connections (referred to as synapses) are represented by arrows and have numerical weights that can be adjusted based on experience. Thereby, the neural network adapts to the input and becomes learnable. Typically, a neural network includes a plurality of input layers. Typically, there is one or more intermediate layers of neurons and an output layer of neurons that provides the output of the neural network. At each level, the neurons make decisions individually or jointly based on the data received from the synapses.

[0005] One of the main challenges in the development of artificial neural networks for high-performance information processing is the lack of appropriate hardware technology. In practice, practical neural networks rely on a very large number of synapses, which enables high connectivity between neurons, i.e., a very high degree of parallelization of computational processing. In principle, such complexity can be realized by a digital supercomputer or a dedicated GPU (Graphics Processing Unit) cluster. However, in addition to high costs, these approaches also suffer from poor energy efficiency compared to biological networks, which mainly perform low-precision analog calculations and consume much less energy. CMOS analog circuits have been used for artificial neural networks, but most CMOS-implemented synapses have been too bulky assuming the required large number of neurons and synapses.

[0006] The applicant previously disclosed in U.S. Patent Application No. 15 / 594,439, published as U.S. Patent Publication No. 2017 / 0337466, which is incorporated by reference, an artificial (analog) neural network that utilizes one or more non-volatile memory arrays as synapses. The non-volatile memory arrays operate as analog neural memories. The neural network device includes a first plurality of synapses configured to receive a first plurality of inputs and then generate a first plurality of outputs, and a first plurality of neurons configured to receive the first plurality of outputs. The first plurality of synapses includes a plurality of memory cells, each memory cell including a spaced source region and drain region formed in a semiconductor substrate with a channel region extending therebetween, a floating gate disposed above a first portion of the channel region and insulated from the first portion of the channel region, and a non-floating gate disposed above a second portion of the channel region and insulated from the second portion of the channel region. Each of the plurality of memory cells is configured to store a weight value corresponding to a plurality of electrons on the floating gate. The plurality of memory cells are configured to multiply the stored weight values by the first plurality of inputs to generate the first plurality of outputs.

[0007] Each non-volatile memory cell used in an analog neural memory system must be erased and programmed to hold a very specific and accurate amount of charge, i.e., number of electrons, on the floating gate. For example, each floating gate must hold one of N different values, where N is the number of different weights that can be represented by each cell. Examples of N include 16, 32, 64, 128, and 256.

[0008] One challenge in a vector matrix multiplication (VMM) system is the ability to select a particular cell or group of cells, or in some cases the entire array of cells, for erase, programming, and read operations. A related challenge is to improve the use of physical space within the semiconductor die without losing functionality.

[0009] What is needed is an improved decoding system and physical layout for an analog neural memory system that utilizes non-volatile memory cells.

Summary of the Invention

[0010] An improved decoding system and physical layout are disclosed for an analog neural memory system that utilizes non-volatile memory cells.

[0011]

[0012]

[0013]

[0014]

[0015]

[0016]

[0017]

[0018]

[0019]

[0020]

[0021]

[0022]

[0023]

[0024]

[0025]

[0026]

[0027]

[0028]

[0029]

[0030]

[0031]

[0032]

[0033]

[0034]

[0035]

[0036]

[0037]

[0038]

[0039]

[0040]

[0041]

[0042]

[0043]

[0044]

[0045]

[0046]

[0047]

[0048]

[0049]

[0050]

[0051]

[0052]

[0053]

[0054]

[0055]

[0056]

[0057]

[0058]

[0059]

[0060]

Brief Description of the Drawings

[0061]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27

Figure 28

Figure 29

Figure 30

Figure 31

Figure 32

Figure 33

Figure 34

Figure 35

Figure 36

Figure 37

Figure 38

Figure 39

Figure 40

Figure 41

Figure 42

Figure 43

Figure 44

Figure 45A

Figure 45B

Figure 46

Figure 47

Figure 48

Figure 49

Figure 50

Figure 51

[0062] The artificial neural network of the present invention utilizes a combination of CMOS technology and a non-volatile memory array. Non-volatile memory cell

[0063] Digital non-volatile memory is well known. For example, U.S. Patent No. 5,029,130 (the " '130 patent"), which is incorporated herein by reference, discloses an array of split-gate non-volatile memory cells, which are a type of flash memory cell. Such a memory cell 210 is shown in FIG. 2. Each memory cell 210 includes a source region 14 and a drain region 16 formed in a semiconductor substrate 12, and there is a channel region 18 between the source region 14 and the drain region 16. A floating gate 20 is formed above a first portion of the channel region 18, insulated from the first portion of the channel region 18 (and controlling the conductivity of the first portion of the channel region 18), and formed over a portion of the source region 14. A word line terminal 22 (typically coupled to a word line) is disposed above a second portion of the channel region 18, insulated from the second portion of the channel region 18, and having a first portion that controls the conductivity of the second portion of the channel region 18 and a second portion that extends upwardly over the floating gate 20. The floating gate 20 and the word line terminal 22 are insulated from the substrate 12 by a gate oxide. A bit line terminal 24 is coupled to the drain region 16.

[0064] By applying a high positive voltage to the word line terminal 22, the memory cell 210 is erased (electrons are removed from the floating gate), whereby the electrons in the floating gate 20 pass through the insulator therebetween from the floating gate 20 to the word line terminal 22 via Fowler-Nordheim tunneling.

[0065] The memory cell 210 is programmed by applying a positive voltage to the word line terminal 22 and a positive voltage to the source region 14 (electrons are applied to the floating gate). The electron current flows from the source region 14 (source line terminal) toward the drain region 16. The electrons accelerate and heat up when they reach the gap between the word line terminal 22 and the floating gate 20. A part of the heated electrons is injected into the floating gate 20 through the gate oxide due to the electrostatic attraction from the floating gate 20.

[0066] The memory cell 210 is read by applying a positive read voltage to the drain region 16 and the word line terminal 22 (turning on the portion of the channel region 18 below the word line terminal). When the floating gate 20 is positively charged (i.e., electrons are erased), the portion of the channel region 18 below the floating gate 20 is also turned on, and current flows through the channel region 18, which is detected as the erased state, i.e., the "1" state. When the floating gate 20 is negatively charged (i.e., programmed with electrons), the portion of the channel region below the floating gate 20 is almost or completely turned off, and current does not flow (or hardly flows) through the channel region 18, which is detected as the programmed state, i.e., the "0" state.

[0067] Table 1 shows the typical voltage ranges that can be applied to the terminals of the memory cell 110 to perform read, erase, and program operations. Table 1: Operation of the flash memory cell 210 in FIG. 2

Table 1

[0068] FIG. 3 shows a memory cell 310 similar to the memory cell 210 of FIG. 2 with an additional control gate (CG) terminal 28. The control gate terminal 28 is biased at a high voltage (e.g., 10V) during programming, a low or negative voltage (e.g., 0V / -8V) during erasure, and a low or medium voltage (e.g., 0V / 2.5V) during readout. The other terminals are biased in the same manner as the terminals of FIG. 2.

[0069] FIG. 4 shows a four-gate memory cell 410 including a source region 14, a drain region 16, a floating gate 20 above a first portion of the channel region 18, a select gate 22 (typically coupled to a word line, WL) above a second portion of the channel region 18, a control gate 28 above the floating gate 20, and an erase gate 30 above the source region 14. This configuration is described in U.S. Patent No. 6,747,310, which is hereby incorporated by reference for all purposes. Here, all gates are non-floating gates except for the floating gate 20, i.e., they are electrically connected or connectable to a voltage source. Programming is performed by hot electrons injecting themselves from the channel region 18 into the floating gate 20. Erasure is performed by electrons tunneling from the floating gate 20 to the erase gate 30.

[0070] Table 2 shows typical voltage ranges that can be applied to the terminals of the memory cell 310 to perform read, erase, and program operations. Table 2: Operation of the Flash Memory Cell 410 of FIG. 4

Table 2

[0071] FIG. 5 shows a memory cell 510 similar to the memory cell 410 of FIG. 4, except that the memory cell 510 does not include an erase gate (EG) terminal. Erasure is performed by biasing the substrate 18 to a high voltage and biasing the control gate CG terminal 28 to a low voltage or a negative voltage. Alternatively, erasure is performed by biasing the word line terminal 22 to a positive voltage and biasing the control gate terminal 28 to a negative voltage. Programming and reading are the same as those of FIG. 4.

[0072] FIG. 6 shows a three-gate memory cell 610, which is another type of flash memory cell. The memory cell 610 is identical to the memory cell 410 of FIG. 4, except that the memory cell 610 does not have a separate control gate terminal. (Erasure occurs through the use of an erase gate terminal) The erase operation and the read operation are the same as those of FIG. 4, except that no control gate bias is applied. The programming operation is also performed without a control gate bias. As a result, during the program operation, a higher voltage must be applied to the source line terminal to compensate for the lack of control gate bias.

[0073] Table 3 shows the typical voltage ranges that can be applied to the terminals of the memory cell 610 to perform read, erase, and program operations. Table 3: Operations of the Flash Memory Cell 610 of FIG. 6 [Table 3] "Read 1" is a read mode in which the cell current is the output of the bit line. "Read 2" is a read mode in which the cell current is the output of the source line terminal.

[0074] FIG. 7 shows a stacked gate memory cell 710, which is another type of flash memory cell. The memory cell 710 is similar to the memory cell 210 of FIG. 2, except that the floating gate 20 extends across the entire channel region 18, and the control gate terminal 22 (coupled to the word line) is separated by an insulating layer (not shown) and extends above the floating gate 20. The erase, programming, and read operations operate in a manner similar to those described above for the memory cell 210.

[0075] Table 4 shows typical voltage ranges that can be applied to the terminals of the memory cell 710 and the substrate 12 to perform read, erase, and program operations. Table 4: Operation of the flash memory cell 710 of FIG. 7

Table 4

[0076] "Read 1" is a read mode where the cell current is the output of the bit line. "Read 2" is a read mode where the cell current is the output of the source line terminal. Optionally, in an array including rows and columns of the memory cells 210, 310, 410, 510, 610, or 710, the source line can be coupled to one row of memory cells or two adjacent rows of memory cells. That is, the source line terminal can be shared by adjacent rows of memory cells.

[0077] To utilize a memory array including one of the types of non-volatile memory cells in the above artificial neural network, two modifications are made. First, the lines are configured so that each memory cell can be individually programmed, erased, and read without adversely affecting the memory states of other memory cells in the array, as further described below. Second, continuous (analog) programming of the memory cells is provided.

[0078] Specifically, the memory state of each memory cell in the array (i.e., the charge on the floating gate) can be changed independently, with minimal interference from other memory cells and continuously, from a completely erased state to a completely programmed state. In another embodiment, the memory state of each memory cell in the array (i.e., the charge on the floating gate) can be changed independently, with minimal interference from other memory cells and continuously, from a completely programmed state to a completely erased state and vice versa. This means that cell storage is either analog or can store at least one of a number of discrete values (such as 16 or 64 different values), which allows all cells in the memory array to be adjusted very precisely and individually, making the memory array ideal for storage and allowing for fine-tuning of the synaptic weights of a neural network.

[0079] The methods and means described herein can be applied, without limitation, to other non-volatile memory technologies such as SONOS (silicon-oxide-nitride-oxide-silicon, charge trapping in nitride), MONOS (metal-oxide-nitride-oxide-silicon, metal charge trapping in nitride), ReRAM (resistive random access memory), PCM (phase change memory), MRAM (magnetoresistive random access memory), FeRAM (ferroelectric random access memory), OTP (one-time programmable, bi-level or multi-level), and CeRAM (strongly correlated electron memory). The methods and means described herein can be applied, without limitation, to volatile memory technologies used in neural networks such as SRAM, DRAM, and volatile synaptic cells. Neural Networks Using Non-Volatile Memory Cell Arrays

[0080] FIG. 8 conceptually shows a non-limiting example of a neural network that utilizes the non-volatile memory array of this embodiment. This example uses a non-volatile memory array neural network for a face recognition application, but it is also possible to implement other suitable applications using a non-volatile memory array-based neural network.

[0081] S0 is the input layer, which in this example is a 32×32 pixel RGB image with 5-bit precision (i.e., three 32×32 pixel arrays, one for each of the colors R, G, and B, and each pixel has 5-bit precision). The synapses CB1 going from the input layer S0 to the layer C1 apply a different set of weights to some instances and shared weights to other instances, scanning the input image with an overlapping filter of 3×3 pixels (the kernel) and shifting the filter by 1 pixel (or more than 2 pixels in some models) at a time. Specifically, the 9 pixel values in the 3×3 portion of the image (i.e., what is called the filter or kernel) are provided to the synapses CB1, where these 9 input values are multiplied by appropriate weights, and after summing the outputs of that multiplication, a single output value is determined and given by the first synapse of CB1 to generate one pixel of the layer of the feature map C1. The 3×3 filter is then shifted 1 pixel to the right within the input layer S0 (i.e., a column of 3 pixels is added on the right and a column of 3 pixels is dropped on the left), and thus the 9 pixel values of this newly positioned filter are provided to the synapses CB1, where they are multiplied by the same weights as above, and a second single output value is determined by the associated synapses. This process is continued until the 3×3 filter has scanned over the entire 32×32 pixel image of the input layer S0 for all three colors and all bits (precision values). The process is then repeated using different sets of weights until all the feature maps of layer C1 are calculated, generating different feature maps of C1.

[0082] In this example, in layer C1, there are 16 feature maps each having 30×30 pixels. Each pixel is a new feature pixel extracted from the multiplication of the input and the kernel. Thus, each feature map is a two-dimensional array. Therefore, in this example, layer C1 consists of 16 layers of two-dimensional arrays (note that the layers and arrays referred to in this specification are logical relationships rather than necessarily physical relationships, that is, the arrays are not necessarily oriented in a physical two-dimensional array). Each of the 16 feature maps within layer C1 is generated by one of 16 different sets of synaptic weights applied to the filter scan. All of the C1 feature maps can target different aspects of the same image feature, such as edge identification. For example, the first map (generated using the first set of weights shared by all scans used to generate this first map) can identify circular edges, and the second map (generated using a second set of weights different from the first set of weights) can identify square edges or the aspect ratio of a specific feature, etc.

[0083] Before going from layer C1 to layer S1, an activation function P1 (pooling) that pools values from non-overlapping and consecutive 2×2 regions within each feature map is applied. The purpose of the pooling function is to average neighboring positions (or it is also possible to use the max function), for example, to reduce the dependence on edge positions, and to reduce the data size before going to the next stage. In layer S1, there are 16 15×15 feature maps (i.e., 16 different arrays of 15×15 pixels each). The synapses CB2 going from layer S1 to layer C2 scan the maps in S1 with a 4×4 filter with a 1-pixel filter shift. In layer C2, there are 22 12×12 feature maps. Before going from layer C2 to layer S2, an activation function P2 (pooling) that pools values from non-overlapping and consecutive 2×2 regions within each feature map is applied. In layer S2, there are 22 6×6 feature maps. In the synapses CB3 going from layer S2 to layer C3, an activation function (pooling) is applied, where all neurons in layer C3 are connected to all maps in layer S2 via each synapse of CB3. In layer C3, there are 64 neurons. The synapses CB4 going from layer C3 to the output layer S3 fully connect C3 to S3, that is, all neurons in layer C3 are connected to all neurons in layer S3. The output in S3 contains 10 neurons, and the neuron with the highest output determines the class. This output can, for example, indicate the identification or classification of the content of the original image.

[0084] Each layer of synapses is implemented using an array or a part of an array of non-volatile memory cells.

[0085] Figure 9 is a block diagram of a system that can be used for that purpose. The vector matrix multiplication (VMM) system 32 includes non-volatile memory cells and is used as synapses (such as CB1, CB2, CB3, and CB4 in FIG. 6) between one layer and the next layer. Specifically, the VMM system 32 includes a VMM array 33 including non-volatile memory cells arranged in rows and columns, an erase gate and word line gate decoder 34, a control gate decoder 35, a bit line decoder 36, and a source line decoder 37, and these decoders decode respective inputs to the non-volatile memory cell array 33. Inputs to the VMM array 33 can be made from the erase gate and word line gate decoder 34 or from the control gate decoder 35. The source line decoder 37 in this example also decodes the output of the VMM array 33. Alternatively, the bit line decoder 36 can decode the output of the VMM array 33.

[0086] The VMM array 33 serves two purposes. First, it stores the weights used by the VMM system 32. Second, the VMM array 33 effectively multiplies the inputs by the weights stored in the VMM array 33 and sums them for each output line (source line or bit line) to generate an output, which becomes an input to the next layer or an input to the last layer. By performing the functions of multiplication and addition, the VMM array 33 eliminates the need for separate multiplication and addition logic circuits and is also power-efficient due to in-memory computing.

[0087] The output of the VMM array 33 is supplied to a differential adder (such as an addition op-amp or an addition current mirror) 38 that sums the outputs of the VMM array 33 to create a single value for convolution. The differential adder 38 is arranged to perform the sum of both the positive weight input and the negative weight input and output a single value.

[0088] The total output value of the differential adder 38 is then supplied to an activation function circuit 39 that rectifies the output. The activation function circuit 39 can provide a sigmoid function, a tanh function, a ReLU function, or any other non-linear function. The rectified output value of the activation function circuit 39 becomes an element of the feature map of the next layer (e.g., C1 in FIG. 8) and is then applied to the next synapse to generate the next feature map layer or the last layer. Thus, in this example, the VMM array 33 constitutes a plurality of synapses (receiving inputs from the previous layer of neurons or from an input layer such as an image database), and the adder 38 and the activation function circuit 39 constitute a plurality of neurons.

[0089] The inputs (WLx, EGx, CGx, and optionally BLx and SLx) to the VMM system 32 of FIG. 9 can be at an analog level, a binary level, a digital pulse (in which case a pulse - analog converter PAC may be required to convert the pulse to an appropriate input analog level), or a digital bit (in which case a DAC is provided to convert the digital bit to an appropriate input analog level), and the output can be at an analog level, a binary level, a digital pulse, or a digital bit (in which case an output ADC is provided to convert the output analog level to a digital bit).

[0090] FIG. 10 is a block diagram showing the use of multiple layers of the VMM system 32, labeled as VMM systems 32a, 32b, 32c, 32d, and 32e in the figure. As shown in FIG. 10, an input (denoted as Inputx) is converted from digital to analog by a digital-to-analog converter 31 and provided to the input VMM system 32a. The converted analog input can be a voltage or a current. The input D / A conversion of the first layer can be performed by using a function or a LUT (look-up table) that maps the input Inputx to an appropriate analog level of the matrix multiplier of the input VMM system 32a. The input conversion can also be performed by an analog-to-analog (A / A) converter to convert an external analog input to the mapped analog input to the input VMM system 32a. The input conversion can also be performed by a digital-to-digital pulse (D / P) converter to convert an external digital input to the mapped digital pulse(s) to the input VMM system 32a.

[0091] The output generated by the input VMM system 32a is then provided as input to the next VMM system (hidden level 1) 32b, which then generates an output that is provided as input to the next VMM system (hidden level 2) 32c, and so on. The various layers of the VMM system 32 function as the layers of synapses and neurons of a convolutional neural network (CNN). Each of the VMM systems 32a, 32b, 32c, 32d, and 32e can be a stand-alone physical non-volatile memory array, or multiple VMM systems can utilize different portions of the same physical non-volatile memory array, or multiple VMM systems can utilize overlapping portions of the same physical non-volatile memory system. Each of the VMM systems 32a, 32b, 32c, 32d, and 32e can also be time multiplexed with respect to the various portions of its array or neurons. The example shown in FIG. 10 includes five layers (32a, 32b, 32c, 32d, 32e), namely, one input layer (32a), two hidden layers (32b, 32c), and two fully connected layers (32d, 32e). Those skilled in the art will understand that this is merely illustrative, and alternatively, the system can include more than two hidden layers and more than two fully connected layers. VMM array

[0092] FIG. 11 shows a neuron VMM array 1100 that is particularly suitable for the memory cell 310 shown in FIG. 3 and is utilized as part of the synapses and neurons between the input layer and the next layer. The VMM array 1100 includes a memory array 1101 of non-volatile memory cells and a reference array 1102 of non-volatile reference memory cells (located at the top of the array). Alternatively, another reference array can be located at the bottom.

[0093] In the VMM array 1100, control gate lines such as control gate line 1103 extend in the vertical direction (thus, the reference array 1102 in the row direction is orthogonal to the control gate line 1103), and erase gate lines such as erase gate line 1104 extend in the horizontal direction. Here, the input to the VMM array 1100 is provided to the control gate lines (CG0, CG1, CG2, CG3), and the output of the VMM array 1100 appears on the source lines (SL0, SL1). In one embodiment, only even rows are used, and in another embodiment, only odd rows are used. The current applied to each source line (SL0 and SL1 respectively) performs a summation function of all the currents from the memory cells connected to that particular source line.

[0094] As described herein for the neural network, the non-volatile memory cells of the VMM array 1100, i.e., the flash memory of the VMM array 1100, are preferably configured to operate in the subthreshold region.

[0095] The non-volatile reference memory cells and non-volatile memory cells described herein are biased with weak inversion as follows: Ids = Io * e (Vg-Vth) / nVt = w * Io * e (Vg) / nVt where w = e (-Vth) / nVt is. where Ids is the drain-source current, Vg is the gate voltage of the memory cell, Vth is the threshold voltage of the memory cell, Vt is the thermal voltage = k * T / q, where k is the Boltzmann constant, T is the Kelvin temperature, q is the electronic charge, n is the slope factor = 1+(Cdep / Cox), Cdep = the capacitance of the depletion layer, and Cox is the capacitance of the gate oxide layer, Io is the memory cell current at a gate voltage equal to the threshold voltage, and Io is (Wt / L) * u * Cox * (n - 1) * Vt 2proportional to, where u is the carrier mobility, and Wt and L are the width and length of the memory cell, respectively.

[0096] When using an I-V log converter that converts the input current Ids to an input voltage Vg using a memory cell (such as a reference memory cell or a peripheral memory cell) or a transistor: Vg = n * Vt * log[Ids / wp * Io] where wp is the w of the reference or peripheral memory cell.

[0097] For the memory array used as the vector matrix multiplier VMM array, the output current is as follows: Iout = wa * Io * e (Vg) / nVt that is Iout = (wa / wp) * Iin = W * Iin W = e (Vthp-Vtha) / nVt Iin = wp * Io * e (Vg) / nVt where wa = w for each memory cell of the memory array.

[0098] The word line or control gate can be used as the input of the memory cell for the input voltage.

[0099] Alternatively, the non-volatile memory cells of the VMM array described herein can be configured to operate in the linear region. Ids = β * (Vgs - Vth) * Vds; β = u * Cox * Wt / L W α (Vgs - Vth) That is, the weight W in the linear region is proportional to (Vgs - Vth)

[0100] A word line, a control gate, a bit line, or a source line can be used as an input to a memory cell operating in the linear region. A bit line or a source line can be used as an output of the memory cell.

[0101] For an I-V linear converter, a memory cell (such as a reference memory cell or a peripheral memory cell) or a transistor operating in the linear region, or a resistor can be used to linearly convert an input / output current into an input / output voltage.

[0102] Alternatively, the flash memory cells of the VMM array described herein can be configured to operate in the saturation region. Ids=1 / 2 * β * (Vgs-Vth) 2 ;β=u * Cox * Wt / L Wα(Vgs-Vth) 2 That is, the weight W is proportional to (Vgs-Vth) 2 and proportional to

[0103] A word line, a control gate, or an erase gate can be used as an input to a memory cell operating in the saturation region. A bit line or a source line can be used as an output of the output neuron.

[0104] Alternatively, the memory cells of the VMM array described herein can be used in all regions or combinations thereof (subthreshold, linear, or saturation).

[0105] Another embodiment for the VMM array 32 of FIG. 9 is described in U.S. Patent Application No. 15 / 826,345, which is incorporated herein by reference. As described in the above application, a source line or a bit line can be used as a neuron output (current sum output).

[0106] FIG. 12 shows a neuron VMM array 1200 that is particularly suitable for the memory cell 210 shown in FIG. 2 and is used as a synapse between the input layer and the next layer. The VMM array 1200 includes a memory array 1203 of non-volatile memory cells, a reference array 1201 of first non-volatile reference memory cells, and a reference array 1202 of second non-volatile reference memory cells. The reference arrays 1201 and 1202 arranged in the column direction of the array function to convert the current inputs flowing into the terminals BLR0, BLR1, BLR2, and BLR3 into voltage inputs WL0, WL1, WL2, and WL3. In practice, the first and second non-volatile reference memory cells are diode-connected through a multiplexer 1214 (only a part is shown) in a state where the current input flows in. The reference cells are adjusted (e.g., programmed) to a target reference level. The target reference level is provided by a reference min-array matrix (not shown).

[0107] The memory array 1203 serves two purposes. First, it stores the weights used by the VMM array 1200 in each memory cell. Second, the memory array 1203 effectively multiplies the input (i.e., the current inputs provided to the terminals BLR0, BLR1, BLR2, and BLR3, which are converted by the reference arrays 1201 and 1202 into input voltages and supplied to the word lines WL0, WL1, WL2, and WL3) by the weights stored in the memory cell array 1203, and then adds up all the results (memory cell currents) to generate the output of each bit line (BL0 to BLN), and this output becomes the input to the next layer or the input to the last layer. By the memory array 1203 performing the functions of multiplication and addition, the need for separate multiplication and addition logic circuits is eliminated, and the power efficiency is also good. Here, the voltage inputs are provided to the word lines WL0, WL1, WL2, and WL3, and the outputs appear on each of the bit lines BL0 to BLN during the read (inference) operation. The currents arranged on each of the bit lines BL0 to BLN perform the sum function of the currents from all the non-volatile memory cells connected to that specific bit line.

[0108] Table 5 shows the operating voltages of the VMM array 1200. The columns in the table indicate the voltages applied to the word line of the selected cell, the word line of the unselected cell, the bit line of the selected cell, the bit line of the unselected cell, the source line of the selected cell, and the source line of the unselected cell. FLT indicates floating, i.e., no voltage is applied. The rows indicate the operations of read, erase, and program. Table 5: Operations of the VMM Array 1200 in FIG. 12 [Table 5]

[0109] FIG. 13 shows a neuron VMM array 1300 that is particularly suitable for the memory cell 210 shown in FIG. 2 and is used as part of synapses and neurons between the input layer and the next layer. The VMM array 1300 includes a memory array 1303 of non-volatile memory cells, a reference array 1301 of first non-volatile reference memory cells, and a reference array 1302 of second non-volatile reference memory cells. The reference arrays 1301 and 1302 extend in the row direction of the VMM array 1300. The VMM array is similar to the VMM1000 except that the word lines extend vertically in the VMM array 1300. Here, the input is provided to the word lines (WLA0, WLB0, WLA1, WLB2, WLA2, WLB2, WLA3, WLB3), and the output appears on the source lines (SL0, SL1) during the read operation. The current applied to each source line performs a sum function of all the currents from the memory cells connected to that particular source line.

[0110] Table 6 shows the operating voltages of the VMM array 1300. The columns in the table indicate the voltages applied to the word line of the selected cell, the word line of the unselected cell, the bit line of the selected cell, the bit line of the unselected cell, the source line of the selected cell, and the source line of the unselected cell. The rows indicate the operations of read, erase, and program. Table 6: Operations of the VMM Array 1300 in FIG. 13 [Table 6]

[0111] FIG. 14 shows a neuron VMM array 1400 that is particularly suitable for the memory cell 310 shown in FIG. 3 and is used as part of the synapses and neurons between the input layer and the next layer. The VMM array 1400 includes a memory array 1403 of non-volatile memory cells, a reference array 1401 of first non-volatile reference memory cells, and a reference array 1402 of second non-volatile reference memory cells. The reference arrays 1401 and 1402 function to convert the current inputs flowing into the terminals BLR0, BLR1, BLR2, and BLR3 into voltage inputs CG0, CG1, CG2, and CG3. In practice, the first and second non-volatile reference memory cells are diode-connected through a multiplexer 1412 (only part shown) in a state where the current inputs flow through BLR0, BLR1, BLR2, and BLR3. The multiplexer 1412 includes respective multiplexers 1405 and cascode transistors 1404 to ensure a constant voltage for each bit line (such as BLR0) of the first and second non-volatile reference memory cells during the read operation. The reference cells are adjusted to a target reference level.

[0112] The memory array 1403 serves two purposes. First, it stores the weights used by the VMM array 1400. Second, the memory array 1403 effectively multiplies the input (the current inputs provided to the terminals BLR0, BLR1, BLR2, and BLR3, which the reference arrays 1401 and 1402 convert these current inputs into input voltages and supply to the control gates CG0, CG1, CG2, and CG3) by the weights stored in the memory cell array, and then adds up all the results (cell currents) to generate an output, which appears on BL0 to BLN and becomes the input to the next layer or the input to the last layer. By the memory array performing the functions of multiplication and addition, the need for separate multiplication and addition logic circuits is eliminated, and the power efficiency is also good. Here, the input is provided to the control gate lines (CG0, CG1, CG2, and CG3), and the output appears on the bit lines (BL0 to BLN) during the read operation. The current applied to each bit line performs the total function of all the currents from the memory cells connected to that particular bit line.

[0113] The VMM array 1400 performs unidirectional conditioning of the non-volatile memory cells within the memory array 1403. That is, each non-volatile memory cell is erased and then partially programmed until the desired charge on the floating gate is reached. This can be performed, for example, using the precision programming techniques described below. If too much charge is applied to the floating gate (such as when an incorrect value is stored in the cell), the cell must be erased and the series of partial programming operations must be repeated. As shown, two rows that share the same erase gate (such as EG0 or EG1) need to be erased together (known as page erase), after which each cell is partially programmed until the desired charge on the floating gate is reached.

[0114] Table 7 shows the operating voltages of the VMM array 1400. The columns in the table show the voltages applied to the word line of the selected cell, the word line of the non-selected cell, the bit line of the selected cell, the bit line of the non-selected cell, the control gate of the selected cell, the control gate of the non-selected cell within the same sector as the selected cell, the control gate of the non-selected cell in a different sector from the selected cell, the erase gate of the selected cell, the erase gate of the non-selected cell, the source line of the selected cell, and the source line of the non-selected cell. The rows show the operations of read, erase, and program. Table 7: Operation of the VMM Array 1400 of FIG. 14 [Table 7]

[0115] FIG. 15 shows a neuron VMM array 1500 that is particularly suitable for the memory cell 310 shown in FIG. 3 and is used as part of the synapses and neurons between the input layer and the next layer. The VMM array 1500 includes a memory array 1503 of non-volatile memory cells, a reference array 1501 or a first non-volatile reference memory cell, and a reference array 1502 of second non-volatile reference memory cells. The EG lines EGR0, EG0, EG1, and EGR1 extend vertically, and the CG lines CG0, CG1, CG2, and CG3 and the SL lines WL0, WL1, WL2, and WL3 extend horizontally. The VMM array 1500 is similar to the VMM array 1400 except that the VMM array 1500 performs bidirectional adjustment, and each individual cell can be completely erased, partially programmed, and partially erased as needed to reach the desired charge amount of the floating gate by using an individual EG line. As shown, the reference arrays 1501 and 1502 convert the input current in the terminals BLR0, BLR1, BLR2, and BLR3 into control gate voltages CG0, CG1, CG2, and CG3 (through the action of the diode-connected reference cells via the multiplexer 1514), and these voltages are applied to the memory cells in the row direction. The current outputs (neurons) are in the bit lines BL0 to BLN, and each bit line sums all the currents from the non-volatile memory cells connected to that particular bit line.

[0116] Table 8 shows the operating voltages of the VMM array 1500. The columns in the table show the word lines of the selected cells, the word lines of the non-selected cells, the bit lines of the selected cells, the bit lines of the non-selected cells, the control gates of the selected cells, the control gates of the non-selected cells in the same sector as the selected cells, the control gates of the non-selected cells in a different sector from the selected cells, the erase gates of the selected cells, the erase gates of the non-selected cells, the source lines of the selected cells, and the source lines of the non-selected cells to which the voltages are applied. The rows show the read, erase, and program operations. Table 8: Operation of the VMM Array 1500 in FIG. 15

Table 8

[0117] FIG. 24 shows a neuron VMM array 2400 that is particularly suitable for the memory cell 210 shown in FIG. 2 and is used as part of synapses and neurons between the input layer and the next layer. In the VMM array 2400, the inputs INPUT0, ..., INPUT N are received by bit lines BL0, ..., BL N respectively, and the outputs OUTPUT1, OUTPUT2, OUTPUT3, and OUTPUT4 are generated by source lines SL0, SL1, SL2, and SL3 respectively.

[0118] FIG. 25 shows a neuron VMM array 2500 that is particularly suitable for the memory cell 210 shown in FIG. 2 and is used as part of synapses and neurons between the input layer and the next layer. In this example, the inputs INPUT0, INPUT1, INPUT2, and INPUT3 are received by source lines SL0, SL1, SL2, and SL3 respectively, and the outputs OUTPUT0, ..., OUTPUT N are generated by bit lines BL0, ..., BL N respectively.

[0119] FIG. 26 shows a neuron VMM array 2600 that is particularly suitable for the memory cell 210 shown in FIG. 2 and is used as part of synapses and neurons between the input layer and the next layer. In this example, the inputs INPUT0, ..., INPUT M are received by word lines WL0, ..., WL M respectively, and the outputs OUTPUT0, ..., OUTPUT N are generated by bit lines BL0, ..., BL N respectively.

[0120] FIG. 27 shows a neuron VMM array 2700 that is particularly suitable for the memory cell 310 shown in FIG. 3 and is used as part of synapses and neurons between the input layer and the next layer. In this example, the inputs INPUT0, ..., INPUT M are received by word lines WL0, ..., WL M respectively, and the outputs OUTPUT0, ..., OUTPUT N are generated by bit lines BL0, ..., BL NIt is generated.

[0121] FIG. 28 shows a neuron VMM array 2800 that is particularly suitable for the memory cell 410 shown in FIG. 4 and is used as part of synapses and neurons between an input layer and the next layer. In this example, inputs INPUT0, ..., INPUT n are respectively received by vertical control gate lines CG0, ..., CG N and outputs OUTPUT1 and OUTPUT2 are generated on source lines SL0 and SL1.

[0122] FIG. 29 shows a neuron VMM array 2900 that is particularly suitable for the memory cell 410 shown in FIG. 4 and is used as part of synapses and neurons between an input layer and the next layer. In this example, inputs INPUT0, ..., INPUT N are respectively received by the gates of bit line control gates 2901-1, 2901-2, ..., 2901-(N-1) and 2901-N that are respectively coupled to bit lines BL0, ..., BL N Exemplary outputs OUTPUT1 and OUTPUT2 are generated on source lines SL0 and SL1.

[0123] FIG. 30 shows a neuron VMM array 3000 that is particularly suitable for the memory cell 310 shown in FIG. 3, the memory cell 510 shown in FIG. 5, and the memory cell 710 shown in FIG. 7 and is used as part of synapses and neurons between an input layer and the next layer. In this example, inputs INPUT0, ..., INPUT M are received by word lines WL0, ..., WL M and outputs OUTPUT0, ..., OUTPUT N are respectively generated on bit lines BL0, ..., BL N respectively.

[0124] FIG. 31 shows a neuron VMM array 3100 that is particularly suitable for the memory cell 310 shown in FIG. 3, the memory cell 510 shown in FIG. 5, and the memory cell 710 shown in FIG. 7, and is used as part of synapses and neurons between the input layer and the next layer. In this example, the inputs INPUT0, ..., INPUT M are received on the control gate lines CG0, ..., CG M Outputs OUTPUT0, ..., OUTPUT N are respectively generated on the vertical source lines SL0, ..., SL N and each source line SL i is coupled to the source lines of all the memory cells in column i.

[0125] FIG. 32 shows a neuron VMM array 3200 that is particularly suitable for the memory cell 310 shown in FIG. 3, the memory cell 510 shown in FIG. 5, and the memory cell 710 shown in FIG. 7, and is used as part of synapses and neurons between the input layer and the next layer. In this example, the inputs INPUT0, ..., INPUT M are received on the control gate lines CG0, ..., CG M Outputs OUTPUT0, ..., OUTPUT N are respectively generated on the vertical bit lines BL0, ..., BL N and each bit line BL i is coupled to the bit lines of all the memory cells in column i. Long Short-Term Memory

[0126] The prior art includes the concept known as long short-term memory (LSTM). LSTM units are often used within neural networks. With LSTM, a neural network can store information over an arbitrary period of time and use that information in subsequent operations. Conventional LSTM units include a cell, an input gate, an output gate, and a forget gate. The three gates regulate the flow of information into and out of the cell and the period during which information is stored within the LSTM. VMM is particularly useful in LSTM units.

[0127] FIG. 16 shows an exemplary LSTM 1600. The LSTM 1600 in this example includes cells 1601, 1602, 1603, and 1604. Cell 1601 receives an input vector x0 and generates an output vector h0 and a cell state vector c0. Cell 1602 receives an input vector x1, the output vector (hidden state) h0 from cell 1601, and the cell state c0 from cell 1601, and generates an output vector h1 and a cell state vector c1. Cell 1603 receives an input vector x2, the output vector (hidden state) h1 from cell 1602, and the cell state c1 from cell 1602, and generates an output vector h2 and a cell state vector c2. Cell 1604 receives an input vector x3, the output vector (hidden state) h2 from cell 1603, and the cell state c2 from cell 1603, and generates an output vector h3. Additional cells can also be used, and the LSTM with four cells is merely an example.

[0128] FIG. 17 shows an exemplary implementation of an LSTM cell 1700 that can be used for cells 1601, 1602, 1603, and 1604 in FIG. 16. The LSTM cell 1700 receives an input vector x(t), a cell state vector c(t−1) from a preceding cell, and an output vector h(t−1) from a preceding cell, and generates a cell state vector c(t) and an output vector h(t).

[0129] The LSTM cell 1700 includes sigmoid function devices 1701, 1702, and 1703, each of which controls the degree to which each component of the input vector contributes to the output vector by applying a number between 0 and 1. The LSTM cell 1700 also includes tanh devices 1704 and 1705 for applying a hyperbolic tangent function to the input vector, multiplier devices 1706, 1707, and 1708 for multiplying two vectors, and an adder device 1709 for adding two vectors. The output vector h(t) can be provided to the next LSTM cell in the system or accessed for other purposes.

[0130] FIG. 18 shows an LSTM cell 1800, which is an implementation example of the LSTM cell 1700. For the convenience of the reader, the same numbering method from the LSTM cell 1700 is used in the LSTM cell 1800. The sigmoid function devices 1701, 1702, and 1703, and the tanh device 1704 each include a plurality of VMM arrays 1801 and activation circuit blocks 1802. Therefore, it can be understood that the VMM array is particularly useful in the LSTM cell used in a specific neural network system. The multiplier devices 1706, 1707, and 1708, and the adder device 1709 are implemented in a digital or analog manner. The activation function block 1802 can be implemented in a digital or analog manner.

[0131] An alternative example of the LSTM cell 1800 (and another example of the implementation of the LSTM cell 1700) is shown in FIG. 19. In FIG. 19, the sigmoid function devices 1701, 1702, and 1703, and the tanh device 1704 share the same physical hardware (VMM array 1901 and activation function block 1902) in a time-division multiplexed manner. The LSTM cell 1900 also includes a multiplier device 1903 for multiplying two vectors, an adder device 1908 for adding two vectors, a tanh device 1705 (including the activation circuit block 1902), a register 1907 for storing the value i(t) output from the sigmoid function block 1902, the value f(t) output from the multiplier device 1903 via the multiplexer 1910 * a register 1904 for storing c(t - 1), and the value i(t) output from the multiplier device 1903 via the multiplexer 1910 * a register 1905 for storing u(t), and the value o(t) output from the multiplier device 1903 via the multiplexer 1910 * a register 1906 for storing c~(t), and a multiplexer 1909.

[0132] While the LSTM cell 1800 includes a plurality of sets of VMM arrays 1801 and respective activation function blocks 1802, the LSTM cell 1900 includes only one set of a VMM array 1901 and an activation function block 1902, which is used to represent a plurality of layers in an embodiment of the LSTM cell 1900. The LSTM cell 1900 requires less space than the LSTM 1800 because it only requires 1 / 4 of the space needed for the VMM and the activation function block as compared to the LSTM cell 1800.

[0133] It can be further understood that LSTM units typically include a plurality of VMM arrays, each of which requires the functions provided by specific circuit blocks outside the VMM array, such as adders, activation circuit blocks, and high-voltage generation blocks. Providing separate circuit blocks for each VMM array would require a significant amount of space within the semiconductor device and would be somewhat inefficient. Gated Recurrent Unit

[0134] The analog VMM implementation can be utilized in a gated recurrent unit (GRU) system. A GRU is a gate mechanism within a recurrent neural network. A GRU is similar to an LSTM, except that a GRU cell generally includes fewer components than an LSTM cell.

[0135] FIG. 20 shows an exemplary GRU2000. The GRU2000 in this example includes cells 2001, 2002, 2003, and 2004. Cell 2001 receives an input vector x0 and generates an output vector h0. Cell 2002 receives an input vector x1 and the output vector h0 from cell 2001, and generates an output vector h1. Cell 2003 receives an input vector x2 and the output vector (hidden state) h1 from cell 2002, and generates an output vector h2. Cell 2004 receives an input vector x3 and the output vector (hidden state) h2 from cell 2003, and generates an output vector h3. Additional cells can also be used, and the GRU with four cells is merely an example.

[0136] FIG. 21 shows an exemplary implementation of a GRU cell 2100 that can be used for cells 2001, 2002, 2003, and 2004 of FIG. 20. The GRU cell 2100 receives an input vector x(t) and an output vector h(t - 1) from a preceding GRU cell, and generates an output vector h(t). The GRU cell 2100 includes sigmoid function devices 2101 and 2102, each of which applies a number from 0 to 1 to components from the output vector h(t - 1) and the input vector x(t). The GRU cell 2100 also includes a tanh device 2103 for applying a hyperbolic tangent function to the input vector, a plurality of multiplier devices 2104, 2105, and 2106 for multiplying two vectors, an adder device 2107 for adding two vectors, and a complementary device 2108 for subtracting the input from 1 to generate an output.

[0137] FIG. 22 shows a GRU cell 2200 which is an implementation example of the GRU cell 2100. For the convenience of the reader, the same numbering method from the GRU cell 2100 is used in the GRU cell 2200. As can be seen from FIG. 22, the sigmoid function devices 2101 and 2102, and the tanh device 2103 each include a plurality of VMM arrays 2201 and activation function blocks 2202. Therefore, it can be understood that the VMM array is particularly used in the GRU cell used in a specific neural network system. The multiplier devices 2104, 2105, 2106, the adder device 2107, and the complementary device 2108 are implemented in a digital or analog manner. The activation function block 2202 can be implemented in a digital or analog manner.

[0138] An alternative example of the GRU cell 2200 (and another example of the implementation of the GRU cell 2300) is shown in FIG. 23. In FIG. 23, the GRU cell 2300 uses a VMM array 2301 and an activation function block 2302, and when configured as a sigmoid function, by applying a number from 0 to 1, it controls the degree to which each component of the input vector contributes to the output vector. In FIG. 23, the sigmoid function devices 2101 and 2102, and the tanh device 2103 share the same physical hardware (VMM array 2301 and activation function block 2302) in a time-division multiplexed manner. The GRU cell 2300 also includes a multiplier device 2303 for multiplying two vectors, an adder device 2305 for adding two vectors, a complementary device 2309 for subtracting the input from 1 to generate an output, a multiplexer 2304, and a value h(t - 1) output from the multiplier device 2303 via the multiplexer 2304 * a register 2306 for holding r(t), and a value h(t - 1) output from the multiplier device 2303 via the multiplexer 2304 * a register 2307 for holding z(t), and a value h^(t) output from the multiplier device 2303 via the multiplexer 2304 * (1 - z((t)) a register 2308 for holding, and includes.

[0139] The GRU cell 2200 includes a plurality of sets of a VMM array 2201 and an activation function block 2202, while the GRU cell 2300 includes only one set of a VMM array 2301 and an activation function block 2302 that are used to represent a plurality of layers in an embodiment of the GRU cell 2300. The GRU cell 2300 requires less space than the GRU cell 2200 because it requires only 1 / 3 of the space required for the VMM and the activation function block compared to the GRU cell 2200.

[0140] It can be further understood that a GRU system typically includes a plurality of VMM arrays, each of which requires the functions provided by specific circuit blocks outside the VMM array, such as adders, activation circuit blocks, and high-voltage generation blocks. Providing separate circuit blocks for each VMM array would require a significant amount of space within the semiconductor device and would be somewhat inefficient.

[0141] The input to the VMM array can be at an analog level, a binary level, or a digital bit (in which case a DAC is required to convert the digital bit to an appropriate input analog level), and the output can be at an analog level, a binary level, or a digital bit (in which case an output ADC is required to convert the output analog level to a digital bit).

[0142] For each memory cell within the VMM array, each weight W can be implemented by a single memory cell, or by a differential cell, or by two blended memory cells (the average of two cells). In the case of a differential cell, two memory cells are required to implement the weight W as a differential weight (W = W+ - W-). In the case of two blended memory cells, two memory cells are required to implement the weight W as the average of two cells. Embodiments of a decoding system and physical layout for a VMM array

[0143] Figures 33-51 disclose various decode systems and physical layouts for a VMM array that can be used with any of the types of memory cells described with respect to FIGS. 2-7, or with respect to other non-volatile memory cells.

[0144] FIG. 33 shows a VMM system 3300. The VMM system 3300 includes a VMM array 3301 (which can be based on any of the foregoing VMM array designs, such as VMM arrays 1000, 1100, 1200, 1300, 1400, 1500, 2400, 2500, 2600, 2700, 2800, 2900, 3000, 3100, and 3200, or other VMM designs), a low voltage row decoder 3302, a high voltage row decoder 3303, a column decoder 3304, a column driver 3305, control logic 3306, a bias circuit 3307, a neuron output circuit block 3308, an input VMM circuit block 3309, an algorithm controller 3310, a high voltage generation device block 3311, an analog circuit block 3315, and control logic 3316.

[0145] The input circuit block 3309 functions as an interface from an external input to the input terminals of the memory array 3301. The input circuit block 3309 can include, without limitation, a DAC (digital-to-analog converter), a DPC (digital-to-pulse converter), an APC (analog-to-pulse converter), an IVC (current-to-voltage converter), an AAC (analog-to-analog converter such as a voltage-to-voltage scaler), or an FAC (frequency-to-analog converter). The neuron output block 3308 functions as an interface from the memory array output to an external interface (not shown). The neuron output block 3308 can include, without limitation, an ADC (analog-to-digital converter), an APC (analog-to-pulse converter), a DPC (digital-to-pulse converter), an IVC (current-to-voltage converter), or an IFC (current-to-frequency converter). The neuron output block 3308 may include, without limitation, an activation function, a normalization circuit, and / or a rescaling circuit.

[0146] The low-voltage row decoder 3302 provides bias voltages for read and program operations and provides decode signals to the high-voltage row decoder 3303. The high-voltage row decoder 3303 provides high-voltage bias signals for program and erase operations.

[0147] The algorithm controller 3310 provides control functions for bit lines during program, verify, and erase operations.

[0148] The high-voltage generation device block 3311 includes a charge pump 3312, a charge pump regulator 3313, and a high-voltage generation circuit 3314 that provide a plurality of voltages required for various program, erase, program verification, and read operations.

[0149] FIG. 34 shows a VMM system 3400 particularly suitable for use with a memory cell 410 of the type shown in FIG. 4. The VMM system 3400 includes VMM arrays 3401, 3402, 3402, and 3404 (each of which can be based on any of the VMM array designs described above, such as VMM arrays 1000, 1100, 1200, 1300, 1400, 1500, 2400, 2500, 2600, 2700, 2800, 2900, 3000, and 31000, or other VMM array designs), low voltage row decoders 3405, 3406, 3407, and 3408, a shared high voltage row decoder 3409, word lines or word input lines 3411, 3412, 3413, and 3414, bit lines 3421, 3422, 3423, and 3424, a control gate line 3432, a source line 3434, and an erase gate line 3434. The shared high voltage row decoder 3409 provides the control gate line 3432, the source line 3434, and the erase gate line 3434. In this configuration, the word lines 3411, 3412, 3413, and 3414, and the bit lines 3421, 3422, 3423, and 3424 are parallel to each other. In one embodiment, the word lines and bit lines are arranged in a vertical direction. The control gate line 3432, the source line 3434, and the erase gate line 3436 are parallel to each other and arranged in a horizontal direction, and thus are perpendicular to the word lines or word input lines 3411, 3412, 3413, and 3414, and the bit lines 3421, 3422, 3423, and 3424.

[0150] In the VMM system 3400, the VMM arrays 3401, 3402, 3403, and 3404 share the control gate line 3432, the source line 3434, the erase gate line 3436, and the high voltage row decoder 3409. However, each of the arrays has its own low voltage row decoder such that the low voltage row decoder 3405 is used with the VMM array 3401, the low voltage row decoder 3406 is used with the VMM array 3402, the low voltage row decoder 3407 is used with the VMM array 3403, and the low voltage row decoder 3408 is used with the VMM array 3404. Advantageously in this configuration, the word lines 3411, 3412, 3413, and 3414 are arranged vertically such that the word line 3411 can be routed to the VMM array 3401 only, the word line 3412 can be routed to the VMM array 3402 only, the word line 3413 can be routed to the VMM array 3403 only, and the word line 3414 can be routed to the VMM array 3404 only. This would be very inefficient in a conventional layout where word lines are arranged horizontally for multiple VMM arrays sharing the same high voltage decoder and the same high voltage decode lines.

[0151] FIG. 35 shows a VMM system 3500 that is particularly suitable for use with a memory cell of the type shown in FIG. 4 as the memory cell 410. The VMM system 3500 is similar to the VMM system 3300 of FIG. 33 except that the VMM system 3500 includes separate word lines and a low voltage row decoder for read and programming operations.

[0152] The VMM system 3500 includes VMM arrays 3501, 3502, 3503, and 3504 (each of which can be based on any of the aforementioned VMM designs, such as VMM arrays 1000, 1100, 1200, 1300, 1400, 1500, 2400, 2500, 2600, 2700, 2800, 2900, 3000, 3100, and 3200, or other VMM array designs), low-voltage read row decoders 3505, 3506, 3507, and 3508, a shared low-voltage program row decoder 3530, a shared high-voltage row decoder 3509, read word lines or word input lines 3511, 3512, 3513, and 3514, program pre-decode row lines 3515, bit lines 3521, 3522, 3523, and 3524, control gate lines 3532, source lines 3533, and erase gate lines 3535. The shared high-voltage row decoder 3509 provides the control gate lines 3532, source lines 3533, and erase gate lines 3535. In this layout, the read word lines or word input lines 3511, 3512, 3513, and 3514, the program pre-decode row lines 3515, and the bit lines 3521, 3522, 3523, and 3524 are parallel to each other and are arranged in the vertical direction. The control gate lines 3532, source lines 3533, and erase gate lines 3535 are parallel to each other and are arranged in the horizontal direction, and thus are perpendicular to the read word lines or word input lines 3511, 3512, 3513, and 3514, the program pre-decode row lines 3515, and the bit lines 3521, 3522, 3523, and 3524. In this VMM system 3500, the low-voltage program row decoder 3530 is shared across multiple VMM arrays.

[0153] In the VMM system 3500, the VMM arrays 3501, 3502, 3503, and 3504 share the control gate lines 3532, source lines 3533, erase gate lines 3535, and the high voltage row decoder 3509. However, each of the VMM arrays has its own low voltage read row decoder such that the low voltage read row decoder 3505 is used with the VMM array 3501, the low voltage read row decoder 3506 is used with the VMM array 3502, the low voltage read row decoder 3507 is used with the VMM array 3503, and the low voltage read row decoder 3508 is used with the VMM array 3504. Advantageous in this layout is that the read word lines or word input lines 3511, 3512, 3513, and 3514 are arranged vertically such that the word line 3511 can be routed only to the VMM array 3501, the word line 3512 can be routed only to the VMM array 3502, the word line 3513 can be routed only to the VMM array 3503, and the word line 3514 can be routed only to the VMM array 3504. This would be very inefficient in a conventional layout where word lines are arranged horizontally for multiple arrays sharing the same high voltage decoder and the same high voltage decode lines. In particular, the program pre-decode row line 3515 can be connected to any of the VMM arrays 3501, 3502, 3503, and 3504 via the low voltage program row decoder 3530, and as a result, one or more cells of those VMM arrays can be programmed at once.

[0154] FIG. 36 shows further details regarding a particular aspect of the VMM system 3500, and in particular, details regarding the low voltage row decoders 3505, 3506, 3507, and 3508, exemplified as the low voltage row decoder 3600. The low voltage read row decoder 3600 includes a plurality of switches, such as exemplary switches, to selectively couple the rows and word lines of the cells of each of the VMM arrays 3601, 3602, 3603, and 3604. The low voltage program decoder 3630 includes exemplary NAND gates 3631 and 3632, PMOS transistors 3633 and 3635, and NMOS transistors 3636 and 3636 configured as shown. The NAND gates 3631 and 3632 receive the program pre-decode row line XP3615 as an input. During the program operation, the switches Sp (which can be a CMOS multiplexer or another type of switch) in the low voltage read row decoders 3605, 3605, 3606, and 3608 are closed, and thus the program word lines Wlp0-n are coupled to the word lines in the array to apply a voltage for programming. During the read operation, the read word lines or word input lines 3611, 3612, 3613, and 3614 are selectively coupled to apply a voltage to the word line terminals of the rows in one or more of the arrays 3601, 3602, 3603, and 3604 using the Sr switches (closed) (which can be a CMOS multiplexer or another type of switch) in the low voltage read row decoders 3605, 3606, 3607, and 3608.

[0155] Figure 37 shows a VMM system 3700 that is particularly suitable for use with the type of memory cell shown in FIG. 4 as memory cell 410. The VMM system 3700 includes VMM arrays 3701, 3702, 3702, and 3704 (each of which can be based on any of the aforementioned VMM designs, such as VMM arrays 1000, 1100, 1200, 1300, 1400, 1500, 2400, 2500, 2600, 2700, 2800, 2900, 3000, and 3100, or other VMM array designs), low voltage row decoders 3705, 3706, 3707, and 3708, local high voltage row decoders 3709 and 3710, a global high voltage row decoder 3730, word lines 3711, 3712, 3713, and 3714, bit lines 3721, 3722, 3723, and 3724, high voltage and / or low voltage (HV / LV) pre-decode lines 3732, source lines 3733, and erase gate lines 3734. The shared global high voltage row decoder 3730 provides the HV / LV pre-decode lines 3732, source lines 3733, and erase gate lines 3734. In this layout, the word lines 3711, 3712, 3713, and 3714, and the bit lines 3721, 3722, 3723, and 3724 are parallel to each other and are arranged in the vertical direction. The HV / LV pre-decode lines 3732, source lines 3733, and erase gate lines 3734 are parallel to each other and are arranged in the horizontal direction and are thus perpendicular to the word lines 3711, 3712, 3713, and 3714, and the bit lines 3721, 3722, 3723, and 3724. The HV / LV pre-decode lines 3732 are input to the local high voltage decoders 3709 and 3710. The local high voltage decoder 3709 outputs local control gate lines for the VMM arrays 3701 and 3702. The local high voltage decoder 3710 outputs local control gate lines for the VMM arrays 3703 and 3704. In another embodiment, the local high voltage decoders 3709 and 3710 can each provide local source lines for the VMM arrays 3701 / 3702 and VMM arrays 3703 / 3704.In another embodiment, the local high voltage decoders 3709 and 3710 can each provide local erase gate lines for the VMM arrays 3701 / 3702 and 3703 / 3704, respectively.

[0156] Here, the local high voltage row decoder 3709 is shared by the VMM arrays 3701 and 3702, and the local high voltage row decoder 3710 is shared by the VMM arrays 3703 and 3704. The global high voltage decoder 3730 routes high voltage and low voltage pre-decode signals to local high voltage row decoders such as the local high voltage row decoders 3709 and 3710. Thus, the high voltage decoding function is split between the global high voltage row decoder 3730 and local high voltage decoders such as the local high voltage decoders 3709 and 3710.

[0157] In the VMM system 3700, the VMM arrays 3701, 3702, 3703, and 3704 share the HV / LV pre-decode lines 3732, the source lines 3733, the erase gate lines 3734, and the global high voltage row decoder 3730. However, each of the VMM arrays has its own low voltage row decoder such that the low voltage row decoder 3705 is used with the VMM array 3701, the low voltage row decoder 3706 is used with the VMM array 3702, the low voltage row decoder 3707 is used with the VMM array 3703, and the low voltage row decoder 3708 is used with the VMM array 3704. Advantageous about this layout is that the word lines 3711, 3712, 3713, and 3714 are arranged vertically such that the word line 3711 can be routed to only the VMM array 3701, the word line 3712 can be routed to only the VMM array 3702, the word line 3713 can be routed to only the VMM array 3703, and the word line 3714 can be routed to only the VMM array 3704. Such would be highly inefficient in a conventional layout where word lines are arranged horizontally for multiple arrays sharing a single high voltage decoder.

[0158] FIG. 38 shows a VMM system 3800 that is particularly suitable for use with a memory cell 410 of the type shown in FIG. 4. The VMM system 3800 includes VMM arrays 3801, 3802, 3802, and 3804 (each of which can be based on any of the aforementioned VMM designs, such as VMM arrays 1000, 1100, 1200, 1300, 1400, 1500, 2400, 2500, 2600, 2700, 2800, 2900, 3000, 3100, and 3200, or other VMM array designs), low voltage row decoders 3805, 3806, 3807, and 3808, local high voltage row decoders 3809 and 3810, a global high voltage row decoder 3830, bit lines 3821, 3822, 3823, and 3824, control gate lines or control gate input lines 3811 and 3812, HV / LV pre-decode lines 3833, source lines 3834, and erase gate lines 3835. The shared global high voltage row decoder 3830 provides the HV / LV pre-decode lines 3833, source lines 3834, and erase gate lines 3835. The local high voltage decoders 3809 and 3810 couple the control gate inputs CG3811 and 3812 to the local control gates of the VMM arrays 3801, 3802 and 3803, 3804, respectively. The low voltage row decoders 3805, 3806, 3807, and 3808 provide local (horizontal) word lines to each of the arrays 3801, 3802, 3803, 3804. In this layout, the control gate lines 3811 and 3812, and the bit lines 3821, 3822, 3823, and 3824 are parallel to each other and are arranged in the vertical direction. The source lines 3834 and the erase gate lines 3835 are parallel to each other and are arranged in the horizontal direction and are thus perpendicular to the control gate lines 3811 and 3812 and the bit lines 3821, 3822, 3823, and 3824.

[0159] Similar to the VMM system 3700 of FIG. 37, the local high-voltage row decoder 3809 is shared by the VMM arrays 3801 and 3802, and the local high-voltage row decoder 3810 is shared by the VMM arrays 3803 and 3804. The global high-voltage decoder 3830 routes signals to local high-voltage row decoders such as the local high-voltage row decoders 3809 and 3810. Therefore, the high-voltage decoding function is split between the global high-voltage row decoder 3830 and local high-voltage decoders such as the local high-voltage decoders 3809 and 3810 (which can provide local source lines and / or local erase gate lines).

[0160] In the VMM system 3800, the VMM arrays 3801, 3802, 3803, and 3804 share the HV / LV pre-decode lines 3833, the source lines 3834, the erase gate lines 3835, and the global high-voltage row decoder 3830. However, each of the VMM arrays has its own low-voltage row decoder such that the low-voltage row decoder 3805 is used with the VMM array 3801, the low-voltage row decoder 3806 is used with the VMM array 3802, the low-voltage row decoder 3807 is used with the VMM array 3803, and the low-voltage row decoder 3808 is used with the VMM array 3804. Advantageous in this layout is that the control gate lines 3811 and 3812, which can be read lines or input lines, are arranged vertically such that the control gate line 3811 can be routed to only the VMM arrays 3801 and 3802 and the control gate line 3812 can be routed to only the VMM arrays 3803 and 3804. Such a thing would be impossible in a conventional layout where the word lines are arranged horizontally.

[0161] FIG. 39 shows a VMM system 3900 that is particularly suitable for use with a type of memory cell such as that shown in FIG. 3 as memory cell 310, FIG. 4 as memory cell 410, FIG. 5 as memory cell 510, or FIG. 7 as memory cell 710. The VMM system 3900 includes VMM arrays 3901 and 3902 (each of which can be based on any of the foregoing VMM designs, such as VMM arrays 1000, 1100, 1200, 1300, 1400, 1500, 2400, 2500, 2600, 2700, 2800, 2900, 3000, 3100, and 3200, or other VMM array designs), a low voltage row decoder 3903 (used with arrays 3901 and 3902), a local high voltage row decoder 3905, a global high voltage row decoder 3904, control gate lines 3908 and 3909, and bit lines 3906 and 3907. In this layout, control gate line 3908 is used only by VMM array 3901, and control gate line 3909 is used only by VMM array 3902. The low voltage row decode line 3910 is used as a decode input to the global high voltage row decoder 3904. The global high voltage row decode line 3911 is used as a decode input to the local high voltage decoder 3905.

[0162] The local high voltage row decoder 3905 is shared by VMM arrays 3901 and 3902. The global high voltage decoder 3904 routes signals to local high voltage row decoders of multiple VMM systems, such as the local high voltage row decoder 3905 of VMM system 3900. Thus, the high voltage decode function is split between the global high voltage row decoder 3904 and local high voltage decoders such as local high voltage decoder 3905 as described above.

[0163] In the VMM system 3900, the VMM arrays 3901 and 3902 share word lines (not shown), source gate lines (not shown) if present, erase gate lines (not shown) if present, and the global high voltage row decoder 3904. Here, the VMM arrays 3901 and 3902 share the low voltage row decoder 3903. What is advantageous about this layout is that the VMM arrays 3901 and 3902 do not share control gate lines, allowing each array to be independently accessed using the control gate lines 3908 and 3909 respectively.

[0164] FIG. 51 shows a VMM system 5100 that is particularly suitable for use with a memory cell of the type shown in FIG. 4 as the memory cell 410. The VMM system 5100 includes VMM arrays 5101, 5102, 5103, and 5104 (each of which can be based on any of the aforementioned VMM array designs, such as the VMM arrays 1000, 1100, 1200, 1300, 1400, 1510, 2400, 2510, 2600, 2700, 2800, 2900, 3000, 3100, and 3200, or other VMM array designs), a high voltage decoder 5130, routing blocks 5151 and 5152, input word lines 5111 and 5112, bit lines 5121, 5122, 5123, and 5124, control gate lines 5132, source lines 5133, and erase gate lines 5134. The high voltage decoder 5130 provides the control gate lines 5132, source lines 5133, and erase gate lines 5134. The routing blocks 5151, 5152 are where the vertically received input word lines 5111 and 5112 are each routed to the horizontally extending word lines of the VMM arrays 5101 - 5104. Alternatively, the routing blocks 5151, 5152 may route the control gate input lines 5132 vertically received to the horizontally extending control gate lines 5132 of the VMM arrays.

[0165] Figure 40 shows a low-voltage row decoder 4000 including a NAND gate 4001, a PMOS transistor 4002, and an NMOS transistor 4003. The NAND gate 4001 receives a row address signal 4004. The PMOS transistor 4002 is coupled to a vertical word line input 4005. The output is at a horizontal word line 4006 which is one of many word lines coupled to respective VMM arrays. In this example, there are a total of 16 word lines, and thus there are 16 instances of the row decoder 4000 each outputting one of the 16 word lines. Thus, based on the received row address signal, one word line such as word line 4006 outputs respective signals such as voltage, and the other word lines are set to ground.

[0166] Figure 41 shows a combined co-select / deselect word line and control gate decoder 4100 including a low-voltage row decoder as in Figure 40, here with a NAND gate 4101, a PMOS transistor 4102, an NMOS transistor 4103, a row address signal 4104, a vertical input word line 4105, and a horizontal word output line 4106 coupled to the word lines of the VMM array. The combined word line and control gate decoder 4100 further includes an inverter 4107, switches 4108 and 4112, and an isolation transistor 4109, receives a control gate input 4110 CGIN0, and outputs a control gate line 4111 CG0. The word line output 4106 WL0 and the control gate output CG0 4111 are simultaneously selected or deselected by decode logic (not shown) controlling the NAND gate 4101.

[0167] FIG. 42 shows a bit line decoder 4200 operating on VMM arrays 4201 and 4202. The bit line decoder 4200 includes a column multiplexer 4203 for selecting one or more bit lines for programming and verification (where the verification operation is used to confirm that the cell current has reached a specific target during the synchronization operation (programming or erasing operation)), and a sense amplifier 4204 for performing a read operation with one or more bit lines. As shown, local bit line muxes 4201b and 4202b multiplex local array bit lines onto global bit lines 4220x coupled to the column multiplexer 4203. The sense amplifier includes an ADC or other device. Thus, the bit line decoder 4200 is shared across multiple arrays.

[0168] FIG. 43 shows a VMM system 4300, which includes VMM arrays 4301, 4302, 4303, and 4304, low voltage row decoders 4305 and 4307, local high voltage row decoders 4306 and 4308, a global high voltage row decoder 4309, digital bus inputs QIN[7:0] 4311 and 4312 (here they are inputs to the VMM arrays), and bit lines 4321, 4322, 4323, and 4324. Each low voltage row decoder, such as low voltage row decoder 4305, includes a circuit block row decoder 4335 for each word line, such as an exemplary data input block 4331 (which may consist of eight latches or registers) that outputs a signal 4333 on the word line and a block 4332 (which may include a data-voltage conversion circuit or a data-pulse conversion circuit). Thus, the input to this low voltage row decoder is a digital bus QIN[7:0] with appropriate control logic. In each circuit block column decoder 4335, the digital inputs QIN[7:0] 4311 and 4312 are appropriately latched by synchronous timing means and methods (e.g., by a serial to parallel clocking interface, etc.).

[0169] FIG. 44 shows a neural network array input / output bus multiplexer 4400 that receives outputs from a VMM array (such as from an ADC) and provides those outputs in a grouped and multiplexed fashion to input blocks of other VMM arrays (such as a DAC or DPC). In the example shown, the inputs to the input / output bus multiplexer 4400 include 2048 bits (256 sets, each of 8 bits NEU0, ..., NEU255), and the input / output bus multiplexer 4400 provides those bits in 64 different groups of 32 bits each, multiplexing between the different groups (providing 1 group of 32 bits at any given time) by use of time division multiplexing or the like. Control logic 4401 generates control signals 4402 that control the input / output bus multiplexer 4400.

[0170] FIGS. 45A and 45B show an exemplary layout of a VMM array where the word lines are laid out in a horizontal fashion (FIG. 45A) as contrasted with a vertical fashion (FIG. 45B such as FIGS. 34 or 35).

[0171] FIG. 46 shows an exemplary layout of a VMM array where the word lines are laid out in a vertical fashion (such as FIGS. 34 or 35). However, in this layout, two word lines (such as word lines 4601 and 4602) can occupy the same column but access different rows within the array (due to the gap between them).

[0172] FIG. 47 shows a VMM high voltage decode circuit that includes a word line decoder circuit 4701, a source line decoder circuit 4704, and a high voltage level shifter 4708, which are suitable for use with the type of memory cells shown in FIG. 2.

[0173] The word line decoder circuit 4701 includes PMOS select transistors 4702 (controlled by signal HVO_B) and NMOS deselect transistors 4703 (controlled by signal HVO_B) configured as shown.

[0174] The source line decoder circuit 4704 includes an NMOS monitoring transistor 4705 (controlled by signal HVO), a driving transistor 4706 (controlled by signal HVO), and a deselection transistor 4707 (controlled by signal HVO_B), configured as shown in the figure.

[0175] The high voltage level shifter 4708 receives an enable signal EN and outputs a high voltage signal HV and its complementary signal HVO_B.

[0176] FIG. 48 shows a VMM high voltage decode circuit including an erase gate decoder circuit 4801, a control gate decoder circuit 4804, a source line decoder circuit 4807, and a high voltage level shifter 4811, which are suitable for use in the type of memory cell shown in FIG. 3.

[0177] The erase gate decoder circuit 4801 and the control gate decoder circuit 4804 use the same design as the word line decoder circuit 4701 in FIG. 47.

[0178] The source line decoder circuit 4807 uses the same design as the source line decoder circuit 4704 in FIG. 47.

[0179] The high voltage level shifter 4811 uses the same design as the high voltage level shifter 4708 in FIG. 47.

[0180] FIG. 49 shows a word line driver 4900. The word line driver 4900 selects a word line (such as exemplary word lines WL0, WL1, WL2, and WL3 shown in this specification) and provides a bias voltage to the word line. Each word line is attached to a selection isolation transistor such as a selection transistor 4901 controlled by a control line 4902. Selection transistors such as the selection transistor 4901 isolate the high voltage (for example, 8 - 12V) used during the erase operation from the word line decoder transistors, and the selection transistors can be implemented using IO transistors that operate at a low voltage (for example, 1.8V, 3.3V). Here, during any operation, the control line 4902 is activated and all selection transistors similar to the selection transistor 4901 are turned on. The exemplary bias transistor 4903 (part of the word line decoding circuit) selectively couples the word line to a first bias voltage (such as 3V), and the exemplary bias transistor 4904 (part of the word line decoding circuit) selectively couples the word line to a second bias voltage (lower than the first bias voltage, including ground, a bias between ground and the first bias voltage, or a negative voltage bias to reduce leakage from unused memory rows). During the ANN (Analog Neural Network) read operation, all used word lines are selected and coupled to the first bias voltage. All unused word lines are coupled to the second bias voltage. During other operations such as the program operation, only one word line is selected and the other word lines are coupled to the second bias voltage, which can be a negative bias (for example, -0.3 to -0.5V or more) to reduce array leakage.

[0181] The bias transistors 4903 and 4904 are coupled to the output of the stage 4906 of the shift register 4905. The shift register 4905 enables each row to be independently controlled according to an input data pattern (loaded at the start of the ANN operation).

[0182] FIG. 50 shows a word line driver 5000. The word line driver 5000 is similar to the word line driver 4900, except that each selection transistor is further coupled to a capacitor such as capacitor 5001. Capacitor 5001 can provide a precharge or bias to the word line at the start of operation, enabled by transistor 5002 to sample the voltage on line 5003. Capacitor 5001 acts to sample and hold the input voltage for each word line (S / H). Transistors 5004 and 5005 are off during the ANN operation (array current adder and activation function) of the VMM array, meaning that the voltage on the S / H capacitor 5001 functions as the (floating) voltage source for each word line. Alternatively, capacitor 5001 can be provided by the capacitance of the word line from the VMM array (or as the control gate capacitance if the input is at the control gate).

[0183] As used herein, it should be noted that both the terms "over" and "on" include both "directly over" (with no intervening material, element, or gap therebetween) and "indirectly over" (with an intervening material, element, or gap therebetween). Similarly, the term "adjacent" includes "directly adjacent" (with no intervening material, element, or gap therebetween) and "indirectly adjacent" (with an intervening material, element, or gap therebetween), "attached to" includes "directly attached to" (with no intervening material, element, or gap therebetween) and "indirectly attached to" (with an intervening material, element, or gap therebetween), and "electrically coupled" includes "directly electrically coupled" (with no intervening material or element electrically connecting the elements together therebetween) and "indirectly electrically coupled" (with an intervening material or element electrically connecting the elements together therebetween). For example, forming an element "over a substrate" can include forming the element directly on the substrate without an intervening material / element therebetween, and forming the element indirectly over the substrate with one or more intervening materials / elements therebetween.

Claims

1. An analog neural memory system, A vector matrix multiplication array including an array of non-volatile memory cells arranged in rows and columns, each memory cell including a bit line terminal, a source line terminal, and a word line terminal, the vector matrix multiplication array; A plurality of bit lines, each of the plurality of bit lines being coupled to the bit line terminals of a column of memory cells, the plurality of bit lines; A plurality of word lines, each of the plurality of word lines being coupled to the word line terminals of a row of memory cells, the plurality of word lines; A plurality of source lines, each of the plurality of source lines being coupled to the source line terminals of one or more rows of memory cells, the plurality of source lines, comprising: The plurality of word lines are parallel to the plurality of bit lines and perpendicular to the plurality of source lines, the analog neural memory system.

2. The system according to claim 1, wherein the non-volatile memory cell is a split gate flash memory cell.

3. The system according to claim 1, wherein the non-volatile memory cell is a stacked gate flash memory cell.

4. An analog neural memory system, A vector matrix multiplication array including an array of non-volatile memory cells arranged in rows and columns, each memory cell including a bit line terminal, a control gate terminal, and a word line terminal, the vector matrix multiplication array; A plurality of bit lines, each of the plurality of bit lines being coupled to the bit line terminals of a column of memory cells, the plurality of bit lines; A plurality of control gate lines, each of the plurality of control gate lines being coupled to the control gate terminals of a row of memory cells, the plurality of control gate lines; A plurality of word lines, each of the plurality of word lines being coupled to the word line terminals of a row of memory cells, the plurality of word lines; The plurality of control gate lines are parallel to the plurality of bit lines and perpendicular to the plurality of word lines, the analog neural memory system.

5. The system according to claim 4, wherein the non-volatile memory cell is a split gate flash memory cell.

6. The system according to claim 4, wherein the non-volatile memory cell is a stacked gate flash memory cell.

7. An analog neural memory system, A plurality of vector matrix multiplication arrays, each array including non-volatile memory cells organized in rows and columns, a plurality of vector matrix multiplication arrays; A plurality of low-voltage row decoders, each low-voltage row decoder providing a row decoder function for one of the plurality of vector matrix multiplication arrays, a plurality of low-voltage row decoders; A plurality of global high-voltage row decoders, each global high-voltage row decoder being shared by two of the plurality of vector matrix multiplication arrays and providing a high-voltage signal to two of the plurality of low-voltage row decoders, a plurality of global high-voltage row decoders, an analog-to-neural memory system.

8. The system according to claim 7, wherein the non-volatile memory cell is a split-gate flash memory cell.

9. The system according to claim 7, wherein the non-volatile memory cell is a stacked-gate flash memory cell.

10. An analog-to-neural memory system, A vector matrix multiplication array including an array of non-volatile memory cells organized in rows and columns, each memory cell including a control gate terminal and a word line terminal, a vector matrix multiplication array; A plurality of word lines, each of the plurality of word lines being coupled to the word line terminal of a row of memory cells, a plurality of word lines; A plurality of control gate lines, each of the plurality of control gate lines being coupled to the control gate terminal of a row of memory cells, a plurality of control gate lines; A plurality of decoders, each decoder being selectively coupled to one or both of the plurality of word lines for providing a row decoder function and the plurality of control gate lines for providing a control gate decoder function, an analog-to-neural memory system.

11. The system according to claim 10, wherein the decoder is selectively coupled to the plurality of word lines, and the row decoder function can be selected or deselected.

12. The system according to claim 10, wherein the decoder is selectively coupled to the control gate lines, and the control gate decoder function can be selected or deselected.

13. The system according to claim 11, wherein the decoder is further selectively coupled to the control gate lines, and the control gate decoder function can be selected or deselected.

14. The system according to claim 10, wherein the non-volatile memory cell is a split-gate flash memory cell. **Claim 15** The system according to claim 10, wherein the non-volatile memory cell is a stacked-gate flash memory cell. **Claim 16** An analog neural memory system, A vector matrix multiplication array including an array of non-volatile memory cells organized in rows and columns, each memory cell including a bit line terminal, a source line terminal, a control gate terminal, and a word line terminal, the vector matrix multiplication array; A plurality of bit lines, each of the plurality of bit lines being coupled to the bit line terminals of a column of memory cells; A plurality of word lines, each of the plurality of word lines being coupled to the word line terminals of a row of memory cells; A plurality of control gate lines, each of the plurality of control gate lines being coupled to the control gate terminals of a row of memory cells; A plurality of source lines, each of the plurality of source lines being coupled to the source line terminals of two rows of memory cells, comprising: An output block coupled to the plurality of bit lines; An input block coupled to the plurality of word lines, the plurality of control gate lines, or the plurality of source lines; A multiplexer that receives bits from the output block and provides a portion of the bits to the input block or an input block coupled to another vector matrix multiplication array in response to a control signal. An analog neural memory system. **Claim 17** The system according to claim 16, wherein the non-volatile memory cell is a split-gate flash memory cell. **Claim 18** The system according to claim 16, wherein the non-volatile memory cell is a stacked-gate flash memory cell. **Claim 19** The system according to claim 16, wherein the multiplexer is a time-division multiplexer. **Claim 20** An analog neural memory system, A vector matrix multiplication array including an array of non-volatile memory cells organized in rows and columns, each memory cell including a bit line terminal, a source line terminal, and a word line terminal, the vector matrix multiplication array; A plurality of bit lines, each of the plurality of bit lines being coupled to the bit line terminals of a column of memory cells. A plurality of word lines, each of the plurality of word lines being coupled to the word line terminals of a row of memory cells through one or more routing blocks. A plurality of source lines, each of the plurality of source lines being coupled to the source line terminals of two rows of memory cells. The system includes the plurality of source lines. The plurality of word lines are parallel to the plurality of bit lines and perpendicular to the plurality of source lines. The one or more routing blocks couple the plurality of word lines to the word line terminals of a row of memory cells, and the row is arranged in a direction perpendicular to the plurality of word lines. An analog neuromorphic memory system.

21. The system according to claim 20, wherein the non-volatile memory cell is a split-gate flash memory cell.

22. The system according to claim 20, wherein the non-volatile memory cell is a stacked-gate flash memory cell.

23. An analog neuromorphic memory system, A plurality of global bit lines, A plurality of sense amplifiers, each of the plurality of sense amplifiers being coupled to one of the plurality of global bit lines. A column multiplexer coupled to the plurality of global bit lines for selecting one or more of the plurality of global bit lines for program and verification operations. A plurality of vector matrix multiplication arrays, each vector matrix multiplication array including An array of non-volatile memory cells organized in rows and columns, each memory cell including a bit line terminal. A plurality of local bit lines, each of the plurality of local bit lines being coupled to the bit line terminals of a column of respective memory cells within the array. A plurality of multiplexers for coupling the plurality of local bit lines to the plurality of global bit lines. The plurality of vector matrix multiplication arrays include the plurality of multiplexers. The column multiplexer and the plurality of sense amplifiers are shared by the plurality of vector matrix multiplication arrays. An analog neuromorphic memory system.

24. The system according to claim 23, wherein the non-volatile memory cell is a split-gate flash memory cell. **Claim 25** The system according to claim 23, wherein the non-volatile memory cell is a stacked-gate flash memory cell. **Claim 26** The system according to claim 23, wherein the sense amplifier is an analog-to-digital converter. **Claim 27** An analog neural memory system, a first vector-matrix multiplication array including an array of non-volatile memory cells organized in rows and columns, each memory cell including a bit-line terminal and a control-gate terminal, a second vector-matrix multiplication array including an array of non-volatile memory cells organized in rows and columns, each memory cell including a bit-line terminal, a word-line terminal, and a control-gate terminal, a high-voltage row decoder for applying a high voltage to word-line terminals of cells in a selected row in the first vector-matrix multiplication array and in a selected row in the second vector-matrix multiplication array, a first set of control-gate lines, each of the control-gate lines being coupled to a control-gate terminal of a row of cells in the first vector-matrix multiplication array and not being coupled to the second vector-matrix multiplication array, a second set of control-gate lines, each of the control-gate lines being coupled to a control-gate terminal of a row of cells in the second vector-matrix multiplication array and not being coupled to the first vector-matrix multiplication array, the analog neural memory system comprising the first set and the second set of control-gate lines. **Claim 28** The system according to claim 27, wherein the non-volatile memory cell is a split-gate flash memory cell. **Claim 29** The system according to claim 27, wherein the non-volatile memory cell is a stacked-gate flash memory cell. **Claim 30** An analog neural memory system, a plurality of vector-matrix multiplication arrays, each vector-matrix multiplication array including an array of non-volatile memory cells organized in rows and columns, each memory cell including a word-line terminal. A plurality of read row decoders, each read row decoder being coupled to one of the plurality of vector matrix multiplication arrays for applying a voltage to one or more selected rows during a read operation, the plurality of read row decoders; A shared program row decoder coupled to all of the plurality of vector matrix multiplication arrays for applying a voltage to one or more selected rows in one or more of the vector matrix multiplication arrays during a program operation, an analog neural memory system comprising.

31. The system according to claim 30, wherein the non-volatile memory cell is a split gate flash memory cell.

32. The system according to claim 30, wherein the non-volatile memory cell is a stacked gate flash memory cell.

33. An analog neural memory system, A plurality of vector matrix multiplication arrays, each array including non-volatile memory cells organized in rows and columns, the plurality of vector matrix multiplication arrays; A plurality of low voltage row decoders, each low voltage row decoder providing a row decoder function for one of the plurality of vector matrix multiplication arrays, the plurality of low voltage row decoders; A high voltage row decoder shared by the plurality of vector matrix multiplication arrays and providing a high voltage signal to one or more terminals of one or more non-volatile memory cells within the array, an analog neural memory system comprising.

34. The system according to claim 33, wherein the non-volatile memory cell is a split gate flash memory cell.

35. The system according to claim 33, wherein the non-volatile memory cell is a stacked gate flash memory cell.

Citation Information

Patent Citations

  • Deep learning neural network classifier using non-volatile memory array

    WO2017200883A1

  • Method and apparatus for configuring array columns and rows for accessing flash memory cells

    WO2018034825A1