Decoding System and Physical Layout for an Analog Neural Memory in a Deep Learning Artificial Neural Network

Through improved decoding systems and physical layout, the non-volatile memory cell array is used to solve the problem of low erasing, programming and reading operation efficiency in the prior art, improve the computing parallelism and energy efficiency of neural networks, and optimize the space utilization of memory cells.

CN113748432BActive Publication Date: 2025-07-29SILICON STORAGE TECHNOLOGY INC
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN201980095856.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-07-03
Filing Date
2019-11-17
Publication Date
2025-07-29
Estimated Expiration
2039-11-17

AI Technical Summary

Technical Problem

In the prior art, nonvolatile memory units are difficult to efficiently erase, program and read in simulated neural memory systems, and the physical space in the semiconductor die is not used sufficiently, affecting the computing parallelism and energy efficiency of the neural network.

Method used

Using an improved decoding system and physical layout, a non-volatile memory cell array is used to accurately erase, program and read operations through word lines, control gates, bit lines and source line decoders, optimize the space utilization of memory cells and realize efficient calculations of simulated neural memory.

Benefits of technology

The computational parallelism and energy efficiency of neural networks are improved, precise tuning and fine-tuning of memory cells are achieved, and physical space use is optimized in the semiconductor die.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113748432B_ABST
    Figure CN113748432B_ABST
Patent Text Reader

Abstract

The present invention discloses various embodiments of a word line decoder, a control gate decoder, a bit line decoder, a low voltage row decoder, and a high voltage row decoder, as well as various types of physical layout designs for simulating a non-volatile flash memory array in a neural system. The present invention discloses shared and segmented embodiments of a high voltage row decoder.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Priority Claim

[0002] This application claims priority to U.S. Provisional Patent Application No. 62 / 840,318, filed on April 29, 2019, entitled "DECODING SYSTEM AND PHYSICAL LAYOUT FOR ANALOG NEURAL MEMORY IN DEEP LEARNING ARTIFICIAL NEURAL NETWORK" and U.S. Patent Application No. 16 / 503,355, filed on July 3, 2019, entitled "DECODING SYSTEM AND PHYSICAL LAYOUT FOR ANALOG NEURAL MEMORY IN DEEP LEARNING ARTIFICIAL NEURAL NETWORK". Technical Field

[0003] The present invention discloses an improved decoding system and physical layout for an analog neural memory system using non-volatile memory cells. Background Art

[0004] Artificial neural networks mimic biological neural networks (the central nervous system of animals, especially the brain) and are used to estimate or approximate functions that can depend on a large number of inputs and are typically unknown. Artificial neural networks generally include layers of interconnected "neurons" that exchange messages with each other.

[0005] Figure 1 An artificial neural network is shown, where the circles represent the inputs or layers of neurons. The connections (called synapses) are represented by arrows and have numerical weights that can be adjusted according to experience. This enables the neural network to adapt to the inputs and learn. Generally, a neural network includes a layer of multiple inputs. There is usually one or more intermediate layers of neurons, and an output layer of neurons that provides the output of the neural network. The neurons at each level make decisions separately or jointly based on the data received from the synapses.

[0006] One of the main challenges in developing artificial neural networks for high-performance information processing is the lack of sufficient hardware technology. In fact, practical neural networks rely on a large number of synapses to achieve high connectivity between neurons, that is, very high computational parallelism. In principle, such complexity can be achieved by digital supercomputers or clusters of dedicated graphics processing units. However, compared with biological networks, these methods have mediocre energy efficiency in addition to high costs, and biological networks consume less energy mainly because they perform low-precision analog computations. CMOS analog circuits have been used in artificial neural networks, but due to the need for a large number of neurons and synapses, most CMOS-implemented synapses are too large.

[0007] The applicant previously disclosed in U.S. Patent Application 15 / 594,439 (published as U.S. Patent Publication 2017 / 0337466), which is incorporated herein by reference, an artificial (analog) neural network that uses one or more non-volatile memory arrays as synapses. The non-volatile memory array operates as an analog neural memory. The neural network device includes a first plurality of synapses configured to receive a first plurality of inputs and generate a first plurality of outputs therefrom, and a first plurality of neurons configured to receive the first plurality of outputs. The first plurality of synapses includes a plurality of memory cells, and each memory cell of the memory cells includes: spaced-apart source regions and drain regions formed in a semiconductor substrate, wherein a channel region extends between the source region and the drain region; a floating gate disposed above a first portion of the channel region and insulated from the first portion; and a non-floating gate disposed above a second portion of the channel region and insulated from the second portion. Each of the plurality of memory cells is configured to store a weight value corresponding to a plurality of electrons on the floating gate. The plurality of memory cells are configured to multiply the first plurality of inputs by the stored weight values to generate the first plurality of outputs.

[0008] Each non-volatile memory cell used in the analog neural memory system must be erased and programmed to maintain a very specific and precise amount of charge (i.e., the number of electrons) in the floating gate. For example, each floating gate must maintain one of N different values, where N is the number of different weights that can be indicated by each cell. Examples of N include 16, 32, 64, 128, and 256.

[0009] One challenge in a vector-matrix multiplication (VMM) system is the ability to select a specific cell or group of cells for erase, program, and read operations, or in some cases, select the entire array of cells. A related challenge is to improve the use of physical space within a semiconductor die without losing functionality.

[0010] What is needed is an improved decoding system and physical layout for an analog neural memory system that utilizes non-volatile memory cells. Summary of the Invention

[0011] The present invention discloses an improved decoding system and physical layout for an analog neural memory system that utilizes non - volatile memory cells. Brief Description of the Drawings

[0012] Figure 1 A schematic diagram showing an artificial neural network of the prior art.

[0013] Figure 2 Showing a split - gate flash memory cell of the prior art.

[0014] Figure 3 Showing another split - gate flash memory cell of the prior art.

[0015] Figure 4 Showing another split - gate flash memory cell of the prior art.

[0016] Figure 5 Showing another split - gate flash memory cell of the prior art.

[0017] Figure 6 Showing another split - gate flash memory cell of the prior art.

[0018] Figure 7 Showing a stacked - gate flash memory cell of the prior art.

[0019] Figure 8 A schematic diagram showing different levels of an exemplary artificial neural network using one or more non - volatile memory arrays.

[0020] Figure 9 A block diagram showing a vector - matrix multiplication system.

[0021] Figure 10 A block diagram showing an exemplary artificial neural network using one or more vector - matrix multiplication systems.

[0022] Figure 11 Showing another embodiment of a vector - matrix multiplication system.

[0023] Figure 12 Showing another embodiment of a vector - matrix multiplication system.

[0024] Figure 13 Showing another embodiment of a vector - matrix multiplication system.

[0025] Figure 14 Showing another embodiment of a vector - matrix multiplication system.

[0026] Figure 15 Showing another embodiment of a vector - matrix multiplication system.

[0027] Figure 16 Shows a long short - term memory system of the prior art.

[0028] Figure 17 Shows an exemplary cell used in the long short - term memory system.

[0029] Figure 18 Shows Figure 17 an embodiment of the exemplary cell.

[0030] Figure 19 Shows Figure 17 another embodiment of the exemplary cell.

[0031] Figure 20 Shows a gated recurrent unit system of the prior art.

[0032] Figure 21 Shows an exemplary cell used in the gated recurrent unit system.

[0033] Figure 22 Shows Figure 21 an embodiment of the exemplary cell.

[0034] Figure 23 Shows Figure 21 another embodiment of the exemplary cell.

[0035] Figure 24 Shows another embodiment of the vector - matrix multiplication system.

[0036] Figure 25 Shows another embodiment of the vector - matrix multiplication system.

[0037] Figure 26 Shows another embodiment of the vector - matrix multiplication system.

[0038] Figure 27 Shows another embodiment of the vector - matrix multiplication system.

[0039] Figure 28 Shows another embodiment of the vector - matrix multiplication system.

[0040] Figure 29 Shows another embodiment of the vector - matrix multiplication system.

[0041] Figure 30 Shows another embodiment of the vector - matrix multiplication system.

[0042] Figure 31 Shows another embodiment of the vector - matrix multiplication system.

[0043] Figure 32Shows another embodiment of a vector-matrix multiplication system.

[0044] Figure 33 Shows an exemplary block diagram of a vector-matrix multiplication system.

[0045] Figure 34 Shows an exemplary decoding embodiment of a vector-matrix multiplication system.

[0046] Figure 35 Shows another exemplary decoding embodiment of a vector-matrix multiplication system.

[0047] Figure 36 Shows an exemplary row decoder.

[0048] Figure 37 Shows another exemplary decoding embodiment of a vector-matrix multiplication system.

[0049] Figure 38 Shows another exemplary decoding embodiment of a vector-matrix multiplication system.

[0050] Figure 39 Shows another exemplary decoding embodiment of a vector-matrix multiplication system.

[0051] Figure 40 Shows an embodiment of a low voltage row decoder.

[0052] Figure 41 Shows an embodiment of a combined low voltage row decoder and control gate decoder.

[0053] Figure 42 Shows an embodiment of a bit line decoder.

[0054] Figure 43 Shows a vector-matrix multiplication system and an input block.

[0055] Figure 44 Shows a multiplexer for receiving outputs from an array and providing inputs to one or more arrays in a multiplexed manner.

[0056] Figure 45A And Figure 45B Shows an exemplary layout of a vector-matrix multiplication system.

[0057] Figure 46 Shows an exemplary layout of a vector-matrix multiplication system.

[0058] Figure 47 Shows a word line decoder circuit, a source line decoder circuit, and a high voltage level shifter for use with a vector multiplier matrix.

[0059] Figure 48Shows an erase gate decoder circuit, a control gate decoder circuit, a source line decoder circuit, and a high voltage level shifter for use with a vector multiplier matrix.

[0060] Figure 49 Shows another embodiment of a word line driver for use with a vector multiplier matrix.

[0061] Figure 50 Shows another embodiment of a word line driver for use with a vector multiplier matrix.

[0062] Figure 51 Shows another exemplary decoding embodiment of a vector-matrix multiplication system. Detailed Description

[0063] The artificial neural network of the present invention utilizes a combination of CMOS technology and a non-volatile memory array.

[0064] Non-volatile Memory Cell

[0065] Digital non-volatile memories are well known. For example, U.S. Patent 5,029,130 (“the '130 patent”), which is incorporated herein by reference, discloses an array of split-gate non-volatile memory cells, which are a type of flash memory cell. Such memory cells 210 are shown in Figure 2 Each memory cell 210 includes a source region 14 and a drain region 16 formed in a semiconductor substrate 12, with a channel region 18 therebetween. A floating gate 20 is formed above and insulated from (and controls the conductivity of) a first portion of the channel region 18, and is formed above a portion of the source region 14. A word line terminal 22 (which is typically coupled to a word line) has a first portion disposed above a second portion of the channel region 18 and insulated from (and controls the conductivity of) the second portion of the channel region, and a second portion that extends upward and is located above the floating gate 20. The floating gate 20 and the word line terminal 22 are insulated from the substrate 12 by a gate oxide. A bit line terminal 24 is coupled to the drain region 16.

[0066] The memory cell 210 is erased by placing a high positive voltage on the word line terminal 22 (where electrons are removed from the floating gate), which causes electrons on the floating gate 20 to tunnel through the intervening insulator from the floating gate 20 to the word line terminal 22 via Fowler-Nordheim tunneling.

[0067] The memory cell 210 (where electrons are placed on the floating gate) is programmed by placing a positive voltage on the word line terminal 22 and a positive voltage on the source region 14. The electron current will flow from the source region 14 (source line terminal) to the drain region 16. When the electrons reach the gap between the word line terminal 22 and the floating gate 20, the electrons will accelerate and heat up. Due to the electrostatic attraction from the floating gate 20, some of the heated electrons will be injected onto the floating gate 20 through the gate oxide.

[0068] The memory cell 210 is read by placing a positive read voltage on the drain region 16 and the word line terminal 22 (which turns on the portion of the channel region 18 under the word line terminal). If the floating gate 20 is positively charged (i.e., the electrons are erased), then the portion of the channel region 18 under the floating gate 20 is also turned on, and current will flow through the channel region 18, which is sensed as the erased state or "1" state. If the floating gate 20 is negatively charged (i.e., programmed with electrons), then the portion of the channel region 18 under the floating gate 20 is mostly or completely turned off, and current will not (or very little current) flow through the channel region 18, which is sensed as the programmed state or "0" state.

[0069] Table 1 shows the typical voltage ranges that can be applied to the terminals of the memory cell 110 for performing read, erase, and program operations:

[0070] Table 1: Figure 2 Operation of the flash memory cell 210

[0071] WL BL SL Read 1 0.5-3V 0.1-2V 0V Read 2 0.5-3V 0-2V 2-0.1V Erase Approximately 11 - 13V 0V 0V Program 1V - 2V 1 - 3μA 9-10V

[0072] "Read 1" is a read mode in which the cell current is output on the bit line. "Read 2" is a read mode in which the cell current is output on the source line terminal.

[0073] Figure 3 The memory cell 310 is shown, which is similar to the Figure 2 memory cell 210, but with the addition of a control gate (CG) terminal 28. The control gate terminal 28 is biased at a high voltage (e.g., 10V) during programming, at a low voltage or negative voltage (e.g., 0V / -8V) during erase, and at a low voltage or medium voltage (e.g., 0V / 2.5V) during read. The other terminals are biased similar to Figure 2 that.

[0074] Figure 4FIG. 0 shows a four-gate memory cell 410, which includes a source region 14, a drain region 16, a floating gate 20 over a first portion of a channel region 18, a select gate 22 (usually coupled to a word line WL) over a second portion of the channel region 18, a control gate 28 over the floating gate 20, and an erase gate 30 over the source region 14. This configuration is described in U.S. Patent 6,747,310, which is incorporated herein by reference for all purposes. Here, all gates except the floating gate 20 are non-floating gates, which means they are electrically connected or can be electrically connected to a voltage source. Programming is performed by heated electrons from the channel region 18 that inject themselves into the floating gate 20. Erasure is performed by electrons tunneling from the floating gate 20 to the erase gate 30.

[0075] Table 2 shows typical voltage ranges that can be applied to the terminals of the memory cell 310 to perform read, erase, and program operations:

[0076] Table 2: Figure 4 Operation of the flash memory cell 410

[0077]

[0078]

[0079] "Read 1" is a read mode in which the cell current is output on the bit line. "Read 2" is a read mode in which the cell current is output on the source line terminal.

[0080] Figure 5 FIG. 19 shows a memory cell 510, which is similar to the memory cell 410 of FIG. 0 except that it does not include an erase gate EG terminal. Erasure is performed by biasing the substrate 18 to a high voltage and biasing the control gate CG terminal 28 to a low voltage or a negative voltage. Alternatively, erasure is performed by biasing the word line terminal 22 to a positive voltage and biasing the control gate terminal 28 to a negative voltage. Programming and reading are similar to those of FIG. 0. Figure 4 FIG. 0 Figure 4 FIG. 0

[0081] Figure 6 FIG. 27 shows a three-gate memory cell 610, which is another type of flash memory cell. The memory cell 610 is the same as the memory cell 410 of FIG. 0 except that the memory cell 610 does not have a separate control gate terminal. Except for not applying a control gate bias, the erase operation (erasing by using the erase gate terminal) and the read operation are similar to those of FIG. 0. In the absence of a control gate bias, the program operation is also completed, and as a result, a higher voltage must be applied to the source line terminal during the program operation to compensate for the lack of a control gate bias. Figure 4 FIG. 0 Figure 4 FIG. 0

[0082] Table 3 shows the typical voltage ranges that can be applied to the terminals of memory cell 610 for performing read, erase, and program operations:

[0083] Table 3: Figure 6 Operation of the flash memory cell 610

[0084] WL / SG BL EG SL Read 1 0.5-2.2V 0.1-2V 0-2.6V 0V Read 2 0.5-2.2V 0-2V 0-2.6V 2-0.1V Erase - 0.5V / 0V 0V 11.5V 0V Program 1V 2 - 3μA 4.5V 7-9V

[0085] "Read 1" is a read mode in which the cell current is output on the bit line. "Read 2" is a read mode in which the cell current is output on the source line terminal.

[0086] Figure 7 Shows a stacked-gate memory cell 710, which is another type of flash memory cell. Memory cell 710 is similar to Figure 2 memory cell 210, except that the floating gate 20 extends over the entire channel region 18 and the control gate terminal 22 (which will be coupled to the word line here) extends over the floating gate 20, separated by an insulating layer (not shown). The erase, program, and read operations operate in a similar manner as previously described for memory cell 210.

[0087] Table 4 shows the typical voltage ranges that can be applied to the terminals of memory cell 710 and substrate 12 for performing read, erase, and program operations:

[0088] Table 4: Figure 7 Operation of the flash memory cell 710

[0089] CG BL SL Substrate Read 1 0-5V 0.1–2V 0-2V 0V Read 2 0.5-2V 0-2V 2-0.1V 0V Erase - 8 to - 10V / 0V FLT FLT 8 - 10V / 15 - 20V Program 8-12V 3 - 5V / 0V 0V / 3 - 5V 0V

[0090] "Read 1" is a read mode in which the cell current is output on the bit line. "Read 2" is a read mode in which the cell current is output on the source line terminal. Optionally, in an array including rows and columns of memory cells 210, 310, 410, 510, 610, or 710, the source line can be coupled to a row of memory cells or two adjacent rows of memory cells. That is, the source line terminal can be shared by adjacent rows of memory cells.

[0091] To utilize a memory array including one of the above types of non-volatile memory cells in an artificial neural network, two modifications were made. First, the circuitry was configured such that each memory cell can be individually programmed, erased, and read without adversely affecting the memory states of other memory cells in the array, as further explained below. Second, continuous (analog) programming of the memory cells was provided.

[0092] Specifically, the memory state of each memory cell in the array (i.e., the charge on the floating gate) can be continuously changed from a fully erased state to a fully programmed state independently and with minimal interference to other memory cells. In another embodiment, the memory state of each memory cell in the array (i.e., the charge on the floating gate) can be continuously changed from a fully programmed state to a fully erased state independently and with minimal interference to other memory cells, and vice versa. This means that the cell storage device is analog or can at least store one discrete value out of many discrete values (such as 16 or 64 different values), which allows for very precise and individual tuning of all cells in the memory array, and which makes the memory array ideal for storing and fine-tuning the synaptic weights of a neural network.

[0093] The methods and devices described herein can be applied to other non-volatile memory technologies such as, but not limited to, SONOS (silicon-oxide-nitride-oxide-silicon, charge trapped in nitride), MONOS (metal-oxide-nitride-oxide-silicon, metal charge trapped in nitride), ReRAM (resistive ram), PCM (phase change memory), MRAM (magnetic ram), FeRAM (ferroelectric ram), OTP (one-time programmable in bilayer or multilayer) and CeRAM (correlated electron ram), etc. The methods and devices described herein can be applied to volatile memory technologies for neural networks such as, but not limited to, SRAM, DRAM and volatile synaptic units.

[0094] Neural Network with Non-volatile Memory Cell Array

[0095] Figure 8 A non-limiting example of a neural network using a non-volatile memory array in this embodiment is conceptually shown. This example uses the non-volatile memory array neural network for a face recognition application, but any other suitable application can also be implemented using a neural network based on a non-volatile memory array.

[0096] For this example, S0 is the input layer, which is a 32x32 pixel RGB image with 5-bit precision (i.e., three 32x32 pixel arrays, one for each color R, G, and B, and each pixel is 5-bit precision). The synapses CB1 from the input layer S0 to layer C1 apply different sets of weights in some cases and shared weights in other cases, and scan the input image with a 3x3 pixel overlapping filter (kernel), shifting the filter by 1 pixel (or more than 1 pixel as indicated by the model). Specifically, the values of 9 pixels in a 3x3 portion of the image (i.e., what is called the filter or kernel) are provided to the synapses CB1, where these 9 input values are multiplied by appropriate weights, and after summing the output of this multiplication, a single output value is determined and provided by the first synapse of CB1 for a pixel in one of the layers C1 of the resulting feature map. Then the 3x3 filter is shifted one pixel to the right within the input layer S0 (i.e., adding the column of three pixels on the right and discarding the column of three pixels on the left), whereby the 9 pixel values in this newly positioned filter are provided to the synapses CB1, where they are multiplied by the same weights and a second single output value is determined by the associated synapse. This process continues until the 3x3 filter has scanned all three colors and all bits (precision values) over the entire 32x32 pixel image of the input layer S0. Then this process is repeated using different sets of weights to generate different feature maps for C1 until all the feature maps for layer C1 have been computed.

[0097] At layer C1, in this example, there are 16 feature maps, each with 30x30 pixels. Each pixel is a new feature pixel extracted from the product of the input and the kernel, so each feature map is a two-dimensional array, and thus in this example, layer C1 consists of a 16-layer two-dimensional array (remember that the layers and arrays referred to in this text are logical relationships and do not have to be physical relationships, i.e., the arrays do not have to be oriented as physical two-dimensional arrays). Each of the 16 feature maps in layer C1 is generated by one of sixteen different sets of synaptic weights applied to the filter scans. The C1 feature maps can all relate to different aspects of the same image feature, such as edge recognition. For example, the first map (generated using the first set of weights, shared for all scans used to generate this first map) may identify circular edges, the second map (generated using a second set of weights different from the first set) may identify rectangular edges, or the aspect ratio of certain features, and so on.

[0098] Before transitioning from layer C1 to layer S1, an activation function P1 (pooling) is applied that pools the values from consecutive non-overlapping 2x2 regions in each feature map. The purpose of the pooling function is to take the mean (or alternatively the max function) of neighboring locations to, for example, reduce the dependence on edge locations and reduce the data size before entering the next stage. At layer S1, there are 16 feature maps of 15x15 (i.e., sixteen different arrays of 15x15 pixels each). The synapse CB2 from layer S1 to layer C2 scans the maps in S1 using a 4x4 filter with a 1-pixel shift of the filter. At layer C2, there are 22 feature maps of 12x12. Before transitioning from layer C2 to layer S2, an activation function P2 (pooling) is applied that pools the values from consecutive non-overlapping 2x2 regions in each feature map. At layer S2, there are 22 feature maps of 6x6. The activation function (pooling) is applied to the synapse CB3 from layer S2 to layer C3, where each neuron in layer C3 is connected via the corresponding synapse of CB3 to each map in layer S2. At layer C3, there are 64 neurons. The synapse CB4 from layer C3 to the output layer S3 fully connects C3 to S3, i.e., each neuron in layer C3 is connected to each neuron in layer S3. The output at S3 includes 10 neurons, and the neuron with the highest output determines the class. For example, this output can indicate the recognition or classification of the content of the original image.

[0099] Each layer's synapse is implemented using an array or a portion of an array of non-volatile memory cells.

[0100] Figure 9 is a block diagram of a system that can be used for this purpose. The vector-matrix multiplication (VMM) system 32 includes non-volatile memory cells and serves as the synapse between one layer and the next layer (such as Figure 6 CB1, CB2, CB3, and CB4 in). Specifically, the VMM system 32 includes a VMM array 33 having non-volatile memory cells arranged in rows and columns, an erase gate and word line decoder 34, a control gate decoder 35, a bit line decoder 36, and a source line decoder 37 that decode the respective inputs to the non-volatile memory cell array 33. The input to the VMM array 33 can come from the erase gate and word line decoder 34 or from the control gate decoder 35. In this example, the source line decoder 37 also decodes the output of the VMM array 33. Alternatively, the bit line decoder 36 can decode the output of the VMM array 33.

[0101] The VMM array 33 serves two purposes. First, it stores the weights to be used by the VMM system 32. Second, the VMM array 33 effectively multiplies the inputs with the weights stored in the VMM array 33 and each output line (source line or bit line) sums them up to produce an output, which will be used as the input for the next layer or the final layer. By performing the multiplication and addition functions, the VMM array 33 eliminates the need for separate multiplication and addition logic circuits and is also highly efficient due to its in-situ memory computing.

[0102] The output of the VMM array 33 is provided to a differential summer (such as a summing operational amplifier or a summing current mirror) 38, which sums up the output of the VMM array 33 to create a single value for this convolution. The differential summer 38 is arranged to perform the summation of both positive and negative weight inputs to output a single value.

[0103] Then the output value of the differential summer 38 is provided to an activation function circuit 39 after summation, which corrects the output. The activation function circuit 39 can provide sigmoid, tanh, ReLU functions or any other non-linear functions. The corrected output value of the activation function circuit 39 becomes an element of the feature map for the next layer (e.g., Figure 8 layer C1 in) and is then applied to the next synapse to produce the next feature map layer or the final layer. Thus, in this example, the VMM array 33 constitutes multiple synapses (which receive their inputs from an existing neuron layer or from an input layer such as an image database), and the summer 38 and the activation function circuit 39 constitute multiple neurons.

[0104] Figure 9 The inputs (WLx, EGx, CGx and optionally BLx and SLx) to the VMM system 32 in can be analog levels, binary levels, digital pulses (in which case a pulse - analog converter PAC may be required to convert the pulses to a suitable input analog level) or digital bits (in which case a DAC is provided to convert the digital bits to a suitable input analog level); the outputs can be analog levels, binary levels, digital pulses or digital bits (in which case an output ADC is provided to convert the output analog level to digital bits).

[0105] Figure 10 A block diagram showing the use of a multi-layer VMM system 32 (here labeled as VMM systems 32a, 32b, 32c, 32d and 32e). As Figure 10As shown, the input (denoted as Inputx) is converted from digital to analog by the digital-to-analog converter 31 and provided to the input VMM system 32a. The converted analog input can be either voltage or current. The input D / A conversion of the first layer can be accomplished by using a function or a LUT (look-up table) that maps the input Inputx to an appropriate analog level of the matrix multiplier of the input VMM system 32a. The input conversion can also be done by an analog-to-analog (A / A) converter to convert an external analog input into a mapped analog input to the input VMM system 32a. The input conversion can also be done by a digital-to-digital pulse (D / P) converter to convert an external digital input into one or more mapped digital pulses to the input VMM system 32a.

[0106] The output generated by the input VMM system 32a is provided as an input to the next VMM system (hidden level 1) 32b, which in turn generates an output provided as an input to the next VMM system (hidden level 2) 32c, and so on. Each layer of the VMM system 32 serves as different layers of synapses and neurons of a convolutional neural network (CNN). Each of the VMM systems 32a, 32b, 32c, 32d, and 32e can be an independent physical non-volatile memory array, or multiple VMM systems can utilize different portions of the same non-volatile memory array, or multiple VMM systems can utilize overlapping portions of the same physical non-volatile memory system. Each of the VMM systems 32a, 32b, 32c, 32d, and 32e can also be time-division multiplexed for different portions of its array or neurons. Figure 10 The example shown includes five layers (32a, 32b, 32c, 32d, 32e): an input layer (32a), two hidden layers (32b, 32c), and two fully connected layers (32d, 32e). A person of ordinary skill in the art will know that this is merely exemplary, and conversely, the system can include more than two hidden layers and more than two fully connected layers.

[0107] VMM Array

[0108] Figure 11 The neuron VMM array 1100 is shown, which is particularly suitable for Figure 3 the memory cell 310 shown and serves as the synapses and components of the neurons between the input layer and the next layer. The VMM array 1100 includes a memory array 1101 of non-volatile memory cells and a reference array 1102 of non-volatile reference memory cells (at the top of the array). Alternatively, another reference array can be placed at the bottom.

[0109] In the VMM array 1100, control gate lines (such as control gate line 1103) extend in the vertical direction (so the reference array 1102 is orthogonal to the control gate line 1103 in the row direction), and erase gate lines (such as erase gate line 1104) extend in the horizontal direction. Here, the inputs of the VMM array 1100 are set on the control gate lines (CG0, CG1, CG2, CG3), and the outputs of the VMM array 1100 appear on the source lines (SL0, SL1). In one embodiment, only even rows are used, and in another embodiment, only odd rows are used. The current placed on each source line (SL0 and SL1 respectively) performs a summation function of all the currents from the memory cells connected to that particular source line.

[0110] As described herein for neural networks, the non-volatile memory cells of the VMM array 1100 (i.e., the flash memory of the VMM array 1100) are preferably configured to operate in the subthreshold region.

[0111] Bias the non-volatile reference memory cells and non-volatile memory cells described herein in weak inversion:

[0112] Ids = Io*e (Vg-Vth) / nVt = w*Io*e (Vg) / nVt ,

[0113] where w = e (-Vth) / nVt

[0114] where Ids is the drain-to-source current; Vg is the gate voltage on the memory cell; Vth is the threshold voltage of the memory cell; Vt is the thermal voltage = k*T / q, where k is the Boltzmann constant, T is the temperature in Kelvin, and q is the electron charge; n is the slope factor = 1 + (Cdep / Cox), where Cdep = the capacitance of the depletion layer, and Cox is the capacitance of the gate oxide layer; Io is the memory cell current at a gate voltage equal to the threshold voltage, and Io is proportional to (Wt / L)*u*Cox*(n - 1)*Vt 2 is proportional to, where u is the carrier mobility, and Wt and L are the width and length of the memory cell respectively.

[0115] For an I-to-V logarithmic converter that uses a memory cell (such as a reference memory cell or a peripheral memory cell) or a transistor to convert the input current Ids to an input voltage Vg:

[0116] Vg = n*Vt*log[Ids / wp*Io]

[0117] Here, wp is the w of the reference memory cell or the peripheral memory cell.

[0118] For a memory array used as a vector matrix multiplier (VMM) array, the output current is:

[0119] Iout = wa * Io * e (Vg) / nVt , that is

[0120] Iout = (wa / wp) * Iin = W * Iin

[0121] W = e (Vthp-Vtha) / nVt

[0122] Iin = wp * Io * e (Vg) / nVt

[0123] Here, wa = w of each memory cell in the memory array.

[0124] The word line or control gate can be used as the input of the memory cell for the input voltage.

[0125] Alternatively, the non-volatile memory cells of the VMM array described herein can be configured to operate in the linear region:

[0126] Ids = β * (Vgs - Vth) * Vds; β = u * Cox * Wt / L,

[0127] W α (Vgs - Vth),

[0128] which means that the weight W in the linear region is proportional to (Vgs - Vth)

[0129] The word line or control gate or bit line or source line can be used as the input of the memory cell operating in the linear region. The bit line or source line can be used as the output of the memory cell.

[0130] For an I-to-V linear converter, a memory cell (such as a reference memory cell or a peripheral memory cell) or a transistor or a resistor operating in the linear region can be used to linearly convert the input / output current into the input / output voltage.

[0131] Alternatively, the flash memory cells of the VMM array described herein can be configured to operate in the saturation region:

[0132] Ids = 1 / 2 * β * (Vgs - Vth) 2 ; β = u * Cox * Wt / L

[0133] W α (Vgs - Vth) 2 , which means that the weight W is proportional to (Vgs - Vth) 2 proportional

[0134] A word line, a control gate, or an erase gate can be used as an input to a memory cell operating in the saturation region. A bit line or a source line can be used as an output of an output neuron.

[0135] Alternatively, the memory cells of the VMM arrays described herein can be used in all regions or combinations thereof (subthreshold, linear, or saturation regions).

[0136] U.S. Patent Application 15 / 826,345 describes Figure 9 other embodiments of the VMM array 32, which is incorporated herein by reference. As described herein, a source line or a bit line can be used as a neuron output (current summing output).

[0137] Figure 12 FIG. shows a neuron VMM array 1200, which is particularly suitable for Figure 2 the memory cell 210 shown, and serves as a synapse between the input layer and the next layer. The VMM array 1200 includes a memory array 1203 of non-volatile memory cells, a reference array 1201 of first non-volatile reference memory cells, and a reference array 1202 of second non-volatile reference memory cells. The reference arrays 1201 and 1202 arranged along the column direction of the array are used to convert the current inputs flowing into the terminals BLR0, BLR1, BLR2, and BLR3 into voltage inputs WL0, WL1, WL2, and WL3. In fact, the first non-volatile reference memory cells and the second non-volatile reference memory cells are diode-connected through a multiplexer 1214 (only partially shown), and the current inputs flow into them. The reference cells are tuned (e.g., programmed) to a target reference level. The target reference level is provided by a reference microarray matrix (not shown).

[0138] The memory array 1203 serves two purposes. First, it stores the weights to be used by the VMM array 1200 on its respective memory cells. Second, the memory array 1203 effectively multiplies the inputs (i.e., the current inputs provided at terminals BLR0, BLR1, BLR2, and BLR3, which the reference arrays 1201 and 1202 convert to input voltages to be provided to word lines WL0, WL1, WL2, and WL3) by the weights stored in the memory array 1203 and then sums all the results (memory cell currents) to produce an output on the respective bit lines (BL0 - BLN), which will be the input to the next layer or the input to the final layer. By performing the multiplication and addition functions, the memory array 1203 eliminates the need for separate multiplication and addition logic circuits and is also highly efficient. Here, the voltage inputs are provided on the word lines (WL0, WL1, WL2, and WL3), and the outputs appear on the respective bit lines (BL0 - BLN) during a read (inference) operation. The current placed on each of the bit lines in BL0 - BLN performs a summation function of the currents from all the non-volatile memory cells connected to that particular bit line.

[0139] Table 5 shows the operating voltages for the VMM array 1200. The columns in the table indicate the voltages placed on the word lines for the selected cells, the word lines for the unselected cells, the bit lines for the selected cells, the bit lines for the unselected cells, the source lines for the selected cells, and the source lines for the unselected cells, where FLT indicates floating, i.e., no voltage is applied. The rows indicate the read, erase, and program operations.

[0140] Table 5: Figure 12 Operation of the VMM array 1200

[0141] WL WL - Unselected BL BL - Unselected SL SL - Unselected Read 0.5-3.5V - 0.5V / 0V 0.1 - 2V (Ineuron) 0.6V - 2V / FLT 0V 0V Erase Approximately 5 - 13V 0V 0V 0V 0V 0V Program 1V - 2V - 0.5V / 0V 0.1 - 3uA Vinh Approximately 2.5V 4-10V 0 - 1V / FLT

[0142] Figure 13 A neuron VMM array 1300 is shown, which is particularly applicable to Figure 2The memory cell 210 shown is used as a synapse and component of a neuron between the input layer and the next layer. The VMM array 1300 includes a memory array 1303 of non-volatile memory cells, a reference array 1301 of first non-volatile reference memory cells, and a reference array 1302 of second non-volatile reference memory cells. The reference arrays 1301 and 1302 extend in the row direction of the VMM array 1300. The VMM array is similar to the VMM 1000, except that in the VMM array 1300, the word lines extend in the vertical direction. Here, the inputs are set on the word lines (WLA0, WLB0, WLA1, WLB2, WLA2, WLB2, WLA3, WLB3), and the outputs appear on the source lines (SL0, SL1) during a read operation. The current placed on each source line performs a summation function of all the currents from the memory cells connected to that particular source line.

[0143] Table 6 shows the operating voltages for the VMM array 1300. The columns in the table indicate the voltages placed on the word lines for the selected cells, the word lines for the unselected cells, the bit lines for the selected cells, the bit lines for the unselected cells, the source lines for the selected cells, and the source lines for the unselected cells. The rows indicate the read, erase, and program operations.

[0144] Table 6: Figure 13 Operation of the VMM array 1300

[0145]

[0146] Figure 14 The neuron VMM array 1400 is shown, which is particularly applicable to Figure 3 the memory cell 310 shown and is used as a synapse and component of a neuron between the input layer and the next layer. The VMM array 1400 includes a memory array 1403 of non-volatile memory cells, a reference array 1401 of first non-volatile reference memory cells, and a reference array 1402 of second non-volatile reference memory cells. The reference arrays 1401 and 1402 are used to convert the current inputs flowing into the terminals BLR0, BLR1, BLR2, and BLR3 into voltage inputs CG0, CG1, CG2, and CG3. In fact, the first non-volatile reference memory cells and the second non-volatile reference memory cells are diode-connected through a multiplexer 1412 (only partially shown), where the current inputs flow into them through BLR0, BLR1, BLR2, and BLR3. Each multiplexer 1412 includes a corresponding multiplexer 1405 and a cascode transistor 1404 to ensure a constant voltage on the bit line (such as BLR0) of each of the first non-volatile reference memory cells and the second non-volatile reference memory cells during a read operation. The reference cells are tuned to a target reference level.

[0147] The memory array 1403 serves two purposes. First, it stores the weights to be used by the VMM array 1400. Second, the memory array 1403 effectively multiplies the inputs (the current inputs provided to terminals BLR0, BLR1, BLR2, and BLR3, which the reference arrays 1401 and 1402 convert into input voltages to be provided to the control gates CG0, CG1, CG2, and CG3) by the weights stored in the memory array, and then sums all the results (cell currents) to produce an output that appears on BL0 - BLN and will be the input to the next layer or the input to the final layer. By performing the multiplication and addition functions, the memory array eliminates the need for separate multiplication and addition logic circuits and is also highly efficient. Here, the inputs are provided on the control gate lines (CG0, CG1, CG2, and CG3), and the outputs appear on the bit lines (BL0–BLN) during a read operation. The current placed on each bit line performs the summation function of all the currents from the memory cells connected to that particular bit line.

[0148] The VMM array 1400 implements unidirectional tuning for the non-volatile memory cells in the memory array 1403. That is, each non-volatile memory cell is erased and then partially programmed until the desired charge on the floating gate is reached. This can be performed, for example, using the precise programming techniques described below. If too much charge is placed on the floating gate (such that an incorrect value is stored in the cell), the cell must be erased and the sequence of the partial programming operations must be restarted. As shown, two rows sharing the same erase gate (such as EG0 or EG1) need to be erased together (which is called page erase), and thereafter, each cell is partially programmed until the desired charge on the floating gate is reached.

[0149] Table 7 shows the operating voltages for the VMM array 1400. The columns in the table indicate the voltages placed on the word line for the selected cell, the word line for the unselected cell, the bit line for the selected cell, the bit line for the unselected cell, the control gate for the selected cell, the control gate for the unselected cell in the same sector as the selected cell, the control gate for the unselected cell in a different sector from the selected cell, the erase gate for the selected cell, the erase gate for the unselected cell, the source line for the selected cell, and the source line for the unselected cell. The rows indicate the read, erase, and program operations.

[0150] Table 7: Figure 14 Operation of the VMM array 1400

[0151]

[0152] Figure 15 The neuron VMM array 1500 is shown, which is particularly suitable for Figure 3The memory cell 310 shown and serves as a synapse and component of a neuron between the input layer and the next layer. The VMM array 1500 includes a memory array 1503 of non-volatile memory cells, a reference array 1501 of first non-volatile reference memory cells, and a reference array 1502 of second non-volatile reference memory cells. The EG lines EGR0, EG0, EG1, and EGR1 extend vertically, while the CG lines CG0, CG1, CG2, and CG3 and the SL lines WL0, WL1, WL2, and WL3 extend horizontally. The VMM array 1500 is similar to the VMM array 1400, except that the VMM array 1500 implements bidirectional tuning, where each individual cell can be fully erased, partially programmed, and partially erased as needed to achieve a desired charge amount on the floating gate due to the use of individual EG lines. As shown, the reference arrays 1501 and 1502 convert the input current in the terminals BLR0, BLR1, BLR2, and BLR3 into control gate voltages CG0, CG1, CG2, and CG3 to be applied to the memory cells in the row direction (by the action of diode-connected reference cells via the multiplexer 1514). The current outputs (neurons) are in the bit lines BL0 - BLN, where each bit line sums all the currents from the non-volatile memory cells connected to that specific bit line.

[0153] Table 8 shows the operating voltages for the VMM array 1500. The columns in the table indicate the voltages placed on the word line for the selected cell, the word line for the unselected cell, the bit line for the selected cell, the bit line for the unselected cell, the control gate for the selected cell, the control gate for the unselected cell in the same sector as the selected cell, the control gate for the unselected cell in a different sector from the selected cell, the erase gate for the selected cell, the erase gate for the unselected cell, the source line for the selected cell, and the source line for the unselected cell. The rows indicate read, erase, and program operations.

[0154] Table 8: Figure 15 Operation of the VMM array 1500

[0155]

[0156] Figure 24 Shows the neuron VMM array 2400, which is particularly suitable for Figure 2 the memory cell 210 shown and serves as a synapse and component of a neuron between the input layer and the next layer. In the VMM array 2400, the inputs INPUT0…, INPUT N are received on the bit lines BL0,... BL N respectively, and the outputs OUTPUT1, OUTPUT2, OUTPUT3, and OUTPUT4 are generated on the source lines SL0, SL1, SL2, and SL3 respectively.

[0157] Figure 25 Shows the neuron VMM array 2500, which is particularly suitable for Figure 2 the memory cell 210 shown, and serves as the synapse and component of the neuron between the input layer and the next layer. In this example, the inputs INPUT0, INPUT1, INPUT2, and INPUT3 are received on the source lines SL0, SL1, SL2, and SL3 respectively, and the outputs OUTPUT0,... OUTPUT N are generated on the bit lines BL0,…,BL N .

[0158] Figure 26 Shows the neuron VMM array 2600, which is particularly suitable for Figure 2 the memory cell 210 shown, and serves as the synapse and component of the neuron between the input layer and the next layer. In this example, the inputs INPUT0,…,INPUT M are received on the word lines WL0,…,WL M respectively, and the outputs OUTPUT0,... OUTPUT N are generated on the bit lines BL0,…,BL N .

[0159] Figure 27 Shows the neuron VMM array 2700, which is particularly suitable for Figure 3 the memory cell 310 shown, and serves as the synapse and component of the neuron between the input layer and the next layer. In this example, the inputs INPUT0,…,INPUT M are received on the word lines WL0,…,WL M respectively, and the outputs OUTPUT0,... OUTPUT N are generated on the bit lines BL0,…,BL N .

[0160] Figure 28 Shows the neuron VMM array 2800, which is particularly suitable for Figure 4 the memory cell 410 shown, and serves as the synapse and component of the neuron between the input layer and the next layer. In this example, the inputs INPUT0,…,INPUT n are received on the vertical control gate lines CG0,…,CG N respectively, and the outputs OUTPUT1 and OUTPUT2 are generated on the source lines SL0 and SL1.

[0161] Figure 29Shows a neuron VMM array 2900, which is particularly suitable for Figure 4 the memory cell 410 shown, and serves as the synapse and component of the neuron between the input layer and the next layer. In this example, the inputs INPUT0 to INPUT N are received on the gates of bit line control gates 2901-1, 2901-2 to 2901-(N-1) and 2901-N respectively, and these gates are coupled to bit lines BL0 to BL N respectively. Exemplary outputs OUTPUT1 and OUTPUT2 are generated on source lines SL0 and SL1.

[0162] Figure 30 Shows a neuron VMM array 3000, which is particularly suitable for Figure 3 the memory cell 310 shown, Figure 5 the memory cell 510 shown, and Figure 7 the memory cell 710 shown, and serves as the synapse and component of the neuron between the input layer and the next layer. In this example, the inputs INPUT0,…,INPUT M are received on word lines WL0,…,WL M and the outputs OUTPUT0,…,OUTPUT N are generated on bit lines BL0,…,BL N respectively.

[0163] Figure 31 Shows a neuron VMM array 3100, which is particularly suitable for Figure 3 the memory cell 310 shown, Figure 5 the memory cell 510 shown, and Figure 7 the memory cell 710 shown, and serves as the synapse and component of the neuron between the input layer and the next layer. In this example, the inputs INPUT0 to INPUT M are received on control gates CG0 to CG M and the outputs OUTPUT0,…,OUTPUT N are generated on vertical source lines SL0,…,SL N respectively, where each source line SL i is coupled to the source lines of all memory cells in column i.

[0164] Figure 32 Shows a neuron VMM array 3200, which is particularly suitable for Figure 3 the memory cell 310 shown, Figure 5 the memory cell 510 shown, and Figure 7The memory cell 710 shown, and serves as the synapse and component of the neuron between the input layer and the next layer. In this example, the inputs INPUT0 to INPUT M are received on the control gate lines CG0 to CG M Outputs OUTPUT0, …, OUTPUT N are respectively generated on the vertical bit lines BL0, …, BL N where each bit line BL i is coupled to the bit lines of all the memory cells in column i.

[0165] Long Short-Term Memory

[0166] The prior art includes the concept known as Long Short-Term Memory (LSTM). LSTM cells are commonly used in neural networks. LSTM allows a neural network to remember information over an arbitrary predetermined time interval and use that information in subsequent operations. Conventional LSTM cells include a cell, an input gate, an output gate, and a forget gate. The three gates regulate the flow of information into and out of the cell and the time interval during which information is remembered in the LSTM. VMM can be particularly useful in LSTM cells.

[0167] Figure 16 An exemplary LSTM 1600 is shown. The LSTM 1600 in this example includes cells 1601, 1602, 1603, and 1604. Cell 1601 receives the input vector x0 and generates the output vector h0 and the cell state vector c0. Cell 1602 receives the input vector x1, the output vector (hidden state) h0 from cell 1601, and the cell state c0 from cell 1601, and generates the output vector h1 and the cell state vector c1. Cell 1603 receives the input vector x2, the output vector (hidden state) h1 from cell 1602, and the cell state c1 from cell 1602, and generates the output vector h2 and the cell state vector c2. Cell 1604 receives the input vector x3, the output vector (hidden state) h2 from cell 1603, and the cell state c2 from cell 1603, and generates the output vector h3. Additional cells can be used, and the LSTM with four cells is only an example.

[0168] Figure 17 An exemplary embodiment of the LSTM cell 1700 that can be used for Figure 16 the cells 1601, 1602, 1603, and 1604 is shown. The LSTM cell 1700 receives the input vector x(t), the cell state vector c(t - 1) from the previous cell, and the output vector h(t - 1) from the previous cell, and generates the cell state vector c(t) and the output vector h(t).

[0169] The LSTM cell 1700 includes sigmoid function devices 1701, 1702, and 1703, each of which applies a number between 0 and 1 to control the amount of each component in the input vector that is allowed to pass through to the output vector. The LSTM cell 1700 also includes tanh devices 1704 and 1705 for applying the hyperbolic tangent function to the input vector, multiplier devices 1706, 1707, and 1708 for multiplying two vectors together, and adder device 1709 for adding two vectors together. The output vector h(t) can be provided to the next LSTM cell in the system, or it can be accessed for other purposes.

[0170] Figure 18 An LSTM cell 1800 is shown, which is an example of a specific implementation of the LSTM cell 1700. For the convenience of the reader, the same numbers are used in the LSTM cell 1800 as in the LSTM cell 1700. The sigmoid function devices 1701, 1702, and 1703 and the tanh device 1704 each include a plurality of VMM arrays 1801 and activation circuit blocks 1802. Thus, it can be seen that VMM arrays are particularly useful in LSTM cells used in certain neural network systems. The multiplier devices 1706, 1707, and 1708 and the adder device 1709 are implemented either digitally or analogously. The activation function block 1802 can be implemented either digitally or analogously.

[0171] An alternative form of the LSTM cell 1800 (and another example of a specific implementation of the LSTM cell 1700) is shown in Figure 19 In Figure 19 the sigmoid function devices 1701, 1702, and 1703 and the tanh device 1704 share the same physical hardware (VMM array 1901 and activation function block 1902) in a time-division multiplexing manner. The LSTM cell 1900 also includes a multiplier device 1903 for multiplying two vectors together, an adder device 1908 for adding two vectors together, a tanh device 1705 (which includes an activation circuit block 1902), a register 1907 for storing the value i(t) when the value i(t) is output from the sigmoid function block 1902, a register 1904 for storing the value when the value f(t)*c(t-1) is output from the multiplier device 1903 through the multiplexer 1910, a register 1905 for storing the value when the value i(t)*u(t) is output from the multiplier device 1903 through the multiplexer 1910, a register 1906 for storing the value when the value o(t)*c~(t) is output from the multiplier device 1903 through the multiplexer 1910, and a multiplexer 1909.

[0172] The LSTM cell 1800 includes multiple groups of VMM arrays 1801 and corresponding activation function blocks 1802, while the LSTM cell 1900 only includes one group of VMM arrays 1901 and activation function block 1902, which are used to represent multiple layers in the implementation of the LSTM cell 1900. The LSTM cell 1900 will require less space than the LSTM 1800, because compared with the LSTM cell 1800, the LSTM cell 1900 only needs 1 / 4 of its space for the VMM and activation function blocks.

[0173] It can also be understood that an LSTM cell generally will include multiple VMM arrays, and each VMM array requires functions provided by some circuit blocks outside the VMM array, such as summing and activation circuit blocks and high voltage generation circuit blocks. Providing separate circuit blocks for each VMM array will require a large amount of space within the semiconductor device and will be somewhat inefficient.

[0174] Gated Recurrent Unit

[0175] The analog VMM implementation can be used for gated recurrent unit (GRU) systems. A GRU is a gating mechanism in a recurrent neural network. A GRU is similar to an LSTM, except that a GRU cell generally includes fewer components than an LSTM cell.

[0176] Figure 20 An exemplary GRU 2000 is shown. The GRU 2000 in this example includes cells 2001, 2002, 2003, and 2004. The cell 2001 receives the input vector x0 and generates the output vector h0. The cell 2002 receives the input vector x1, the output vector h0 from the cell 2001, and generates the output vector h1. The cell 2003 receives the input vector x2 and the output vector (hidden state) h1 from the cell 2002, and generates the output vector h2. The cell 2004 receives the input vector x3 and the output vector (hidden state) h2 from the cell 2003, and generates the output vector h3. Additional cells can be used, and a GRU with four cells is just an example.

[0177] Figure 21 Shown can be used for Figure 20Exemplary specific implementations of the GRU units 2100 of units 2001, 2002, 2003, and 2004. The GRU unit 2100 receives an input vector x(t) and an output vector h(t-1) from the previous GRU unit and generates an output vector h(t). The GRU unit 2100 includes sigmoid function devices 2101 and 2102, each of which applies a number between 0 and 1 to components from the output vector h(t-1) and the input vector x(t). The GRU unit 2100 also includes a tanh device 2103 for applying the hyperbolic tangent function to the input vector, a plurality of multiplier devices 2104, 2105, and 2106 for multiplying two vectors together, an adder device 2107 for adding two vectors together, and a complementary device 2108 for subtracting the input from 1 to generate an output.

[0178] Figure 22 Shows the GRU unit 2200, which is an example of a specific implementation of the GRU unit 2100. For the convenience of the reader, the same numbers are used in the GRU unit 2200 as in the GRU unit 2100. As Figure 22 shown, the sigmoid function devices 2101 and 2102 and the tanh device 2103 each include a plurality of VMM arrays 2201 and activation function blocks 2202. Thus, it can be seen that the VMM array is particularly useful in GRU units used in certain neural network systems. The multiplier devices 2104, 2105, and 2106, the adder device 2107, and the complementary device 2108 are implemented digitally or analogously. The activation function block 2202 can be implemented digitally or analogously.

[0179] An alternative form of the GRU unit 2200 (and another example of a specific implementation of the GRU unit 2300) is shown in Figure 23 In Figure 23 the GRU unit 2300 utilizes a VMM array 2301 and an activation function block 2302 that, when configured as a sigmoid function, applies a number between 0 and 1 to control the amount of each component in the input vector that is allowed to reach the output vector. In Figure 23In [the figure], the sigmoid function devices 2101 and 2102, and the tanh device 2103 share the same physical hardware (VMM array 2301 and activation function block 2302) in a time-division multiplexing manner. The GRU unit 2300 also includes a multiplier device 2303 that multiplies two vectors together, an adder device 2305 that adds two vectors together, a complementary device 2309 that subtracts the input from 1 to generate an output, a multiplexer 2304, a register 2306 that holds the value when the value h(t - 1)*r(t) is output from the multiplier device 2303 through the multiplexer 2304, a register 2307 that holds the value when the value h(t - 1)*z(t) is output from the multiplier device 2303 through the multiplexer 2304, and a register 2308 that holds the value when the value h^(t)*(1 - z(t)) is output from the multiplier device 2303 through the multiplexer 2304.

[0180] The GRU unit 2200 includes multiple sets of VMM arrays 2201 and activation function blocks 2202, while the GRU unit 2300 includes only one set of VMM arrays 2301 and activation function blocks 2302, which are used to represent multiple layers in the implementation of the GRU unit 2300. The GRU unit 2300 will require less space than the GRU unit 2200 because, compared with the GRU unit 2200, the GRU unit 2300 only needs 1 / 3 of its space for the VMM and activation function blocks.

[0181] It can also be understood that a GRU system generally includes multiple VMM arrays, and each VMM array requires functions provided by certain circuit blocks outside the VMM array (such as summing and activation circuit blocks and high-voltage generation blocks). Providing separate circuit blocks for each VMM array will require a large amount of space within the semiconductor device and will be somewhat inefficient.

[0182] The input of the VMM array can be an analog level, a binary level, or a digital bit (in which case, a DAC is required to convert the digital bit into an appropriate input analog level), and the output can be an analog level, a binary level, or a digital bit (in which case, an output ADC is required to convert the output analog level into a digital bit).

[0183] For each memory cell in the VMM array, each weight W can be implemented by a single memory cell, or by a differential cell, or by two hybrid memory cells (the average of 2 cells). In the case of a differential cell, two memory cells are required to implement the weight W as a differential weight (W = W+ – W-). In the case of two hybrid memory cells, two memory cells are required to implement the weight W as the average of the two cells.

[0184] Decoding System and Physical Layout Implementation for VMM Array

[0185] Figures 33 - 51 Disclosed are various decoding systems and physical layouts for a VMM array, which can be used with any of the memory cell types previously described Figures 2 - 7 or with other non-volatile memory cells.

[0186] Figure 33 A VMM system 3300 is shown. The VMM system 3300 includes a VMM array 3301 (which can be based on any of the previously discussed VMM array designs, such as VMM arrays 1000, 1100, 1200, 1300, 1400, 1500, 2400, 2500, 2600, 2700, 2800, 2900, 3000, 3100, and 3200 or other VMM designs), a low-voltage row decoder 3302, a high-voltage row decoder 3303, a column decoder 3304, a column driver 3305, a control logic component 3306, a bias circuit 3307, a neuron output circuit block 3308, an input VMM circuit block 3309, an algorithm controller 3310, a high-voltage generator block 3311, an analog circuit block 3315, and a control logic component 3316.

[0187] The input circuit block 3309 serves as an interface to the input terminals from an external input to the memory array 3301. The input circuit block 3309 can include, but is not limited to, a DAC (digital-to-analog converter), a DPC (digital-pulse converter), an APC (analog-pulse converter), an IVC (current-voltage converter), an AAC (analog-analog converter, such as a voltage-voltage scaler), or an FAC (frequency-analog converter). The neuron output block 3308 serves as an interface from the memory array to an external interface (not shown). The neuron output block 3308 can include, but is not limited to, an ADC (analog-to-digital converter), an APC (analog-pulse converter), a DPC (digital-analog converter), an IVC (current-voltage converter), or an IFC (current-frequency converter). The neuron output block 3308 can include, but is not limited to, an activation function, a normalization circuit, and / or a rescaling circuit.

[0188] The low-voltage row decoder 3302 provides a bias voltage for read operations and programming operations and provides a decoded signal to the high-voltage row decoder 3303. The high-voltage row decoder 3303 provides a high-voltage bias signal for programming operations and erase operations.

[0189] The algorithm controller 3310 provides a control function for the bit lines during programming, verification, and erase operations.

[0190] The high voltage generator block 3311 includes a charge pump 3312, a charge pump regulator 3313, and a high voltage generation circuit 3314 which provides multiple voltages required for various programming, erasing, program verification, and read operations.

[0191] Figure 34 FIG. VMM system 3400 is shown, which is particularly suitable for use with Figure 4 memory cells of the type shown as memory cell 410. The VMM system 3400 includes VMM arrays 3401, 3402, 3402 and 3404 (each of which can be based on any of the foregoing VMM array designs, such as VMM arrays 1000, 1100, 1200, 1300, 1400, 1500, 2400, 2500, 2600, 2700, 2800, 2900, 3000 and 31000 or other VMM array designs); low voltage row decoders 3405, 3406, 3407 and 3408; a shared high voltage row decoder 3409; word lines or word input lines 3411, 3412, 3413 and 3414; bit lines 3421, 3422, 3423 and 3424; control gate lines 3432, source lines 3434 and erase gate lines 3434. The shared high voltage row decoder 3409 provides the control gate lines 3432, source lines 3434 and erase gate lines 3434. In this arrangement, the word lines 3411, 3412, 3413 and 3414 and the bit lines 3421, 3422, 3423 and 3424 are parallel to each other. In one embodiment, the word lines and bit lines are arranged in the vertical direction. The control gate lines 3432, source lines 3434 and erase gate lines 3436 are parallel to each other and arranged in the horizontal direction, thus perpendicular to the word lines or word input lines 3411, 3412, 3413 and 34, and the bit lines 3421, 3422, 3423 and 3424.

[0192] In the VMM system 3400, the VMM arrays 3401, 3402, 3403, and 3404 share the control gate lines 3432, source lines 3434, erase gate lines 3436, and the high voltage row decoder 3409. However, each of the arrays has its own low voltage row decoder such that the low voltage row decoder 3405 is used with the VMM array 3401; the low voltage row decoder 3406 is used with the VMM array 3402; the low voltage row decoder 3407 is used with the VMM array 3403; and the low voltage row decoder 3408 is used with the VMM array 3404. Advantageously, the word lines 3411, 3412, 3413, and 3414 are arranged in a vertical direction such that the word line 3411 can be routed only to the VMM array 3401, the word line 3412 can be routed only to the VMM array 3402, the word line 3413 can be routed only to the VMM array 3403, and the word line 3414 can be routed only to the VMM array 3404. This would be very inefficient with a conventional layout in which the word lines are arranged in a horizontal direction for multiple VMM arrays sharing the same high voltage decoder and the same high voltage decoding lines.

[0193] Figure 35 FIG. shows a VMM system 3500, which is particularly suitable for use with Figure 4 memory cells of the type shown as memory cells 410 in. The VMM system 3500 is similar to Figure 33 the VMM system 3300, except that the VMM system 3500 includes separate word lines and low voltage row decoders for read operations and programming operations.

[0194] The VMM system 3500 includes VMM arrays 3501, 3502, 3503, and 3504 (each of which can be based on any of the aforementioned VMM designs, such as VMM arrays 1000, 1100, 1200, 1300, 1400, 1500, 2400, 2500, 2600, 2700, 2800, 2900, 3000, 3100, and 3200 or other VMM array designs); low-voltage read row decoders 3505, 3506, 3507, and 3508; a shared low-voltage programming row decoder 3530; a shared high-voltage row decoder 3509; read word lines or word input lines 3511, 3512, 3513, and 3514; a programming pre-decoded row line 3515; bit lines 3521, 3522, 3523, and 3524; control gate lines 3532, source lines 3533, and erase gate lines 3535. The shared high-voltage row decoder 3509 provides the control gate lines 3532, source lines 3533, and erase gate lines 3535. In this layout, the read word lines or word input lines 3511, 3512, 3513, and 3514, the programming pre-decoded row line 3515, and the bit lines 3521, 3522, 3523, and 3524 are parallel to each other and arranged in the vertical direction. The control gate lines 3532, source lines 3533, and erase gate lines 3535 are parallel to each other and arranged in the horizontal direction and are thus perpendicular to the read word lines or word input lines 3511, 3512, 3513, and 3514, the programming pre-decoded row line 3515, and the bit lines 3521, 3522, 3523, and 3524. In this VMM system 3500, the low-voltage programming row decoder 3530 is shared across multiple VMM arrays.

[0195] In the VMM system 3500, the VMM arrays 3501, 3502, 3503, and 3504 share the control gate lines 3532, source lines 3533, erase gate lines 3535, and the high-voltage row decoder 3509. However, each VMM array has its own low-voltage read row decoder, such that the low-voltage read row decoder 3505 is used with the VMM array 3501; the low-voltage read row decoder 3506 is used with the VMM array 3502; the low-voltage read row decoder 3507 is used with the VMM array 3503; and the low-voltage read row decoder 3508 is used with the VMM array 3504. An advantage of this layout is that the read word lines or word input lines 3511, 3512, 3513, and 3514 are arranged in the vertical direction such that the word line 3511 can be routed only to the VMM array 3501, the word line 3512 can be routed only to the VMM array 3502, the word line 3513 can be routed only to the VMM array 3503, and the word line 3514 can be routed only to the VMM array 3504. This would be very inefficient with a conventional layout, in which the word lines are arranged horizontally for multiple arrays sharing the same high-voltage decoder and the same high-voltage decode lines. Notably, the programming pre-decode row line 3515 can be connected to any one of the VMM arrays 3501, 3502, 3503, and 3504 via the low-voltage programming row decoder 3530, such that the cells in one or more of those VMM arrays can be programmed at one time.

[0196] Figure 36Additional details regarding certain aspects of the VMM system 3500 are shown, specifically, details regarding the low voltage row decoders 3505, 3506, 3507, and 3508, which are illustrated as the low voltage row decoder 3600. The low voltage read row decoder 3600 includes a plurality of switches, such as the exemplary switches shown, to selectively couple the word lines to the cell rows in the VMM arrays 3601, 3602, 3603, and 3604, respectively. The low voltage programming decoder 3630 includes the exemplary NAND gates 3631 and 3632, PMOS transistors 3633 and 3635, and NMOS transistors 3636 and 3636 configured as shown. The NAND gates 3631 and 3632 receive the programming pre-decoded row line XP 3615 as an input. During a programming operation, the switches Sp (which can be a CMOS multiplexer or another type of switch) in the low voltage read row decoders 3605, 3605, 3606, and 3608 are closed, and thus the programming word lines Wlp0-n are coupled to the word lines in the array to apply the voltage for programming. During a read operation, the read word lines or word input lines 3611, 3612, 3613, and 3614 are selectively coupled to apply a voltage to the word line terminals of the rows within one or more of the arrays 3601, 3602, 3603, and 3604 using the Sr switches (closed) within the low voltage read row decoders 3605, 3606, 3607, and 3608 (which can be a CMOS multiplexer or another type of switch).

[0197] Figure 37 The VMM system 3700 is shown, which is particularly applicable to be used with Figure 4VMM system 3700 includes VMM arrays 3701, 3702, 3703, and 3704 (each of which may be based on any of the aforementioned VMM designs, such as VMM arrays 1000, 1100, 1200, 1300, 1400, 1500, 2400, 2500, 2600, 2700, 2800, 2900, 3000, and 3100, or other VMM array designs); Voltage row decoders 3705, 3706, 3707, and 3708; local high voltage row decoders 3709 and 3710; global high voltage row decoder 3730; word lines 3711, 3712, 3713, and 3714; bit lines 3721, 3722, 3723, and 3724; high voltage and / or low voltage (HV / LV) pre-decoding line 3732, source line 3733, and erase gate line 3734. Shared global high voltage row decoder 3730 provides HV / LV pre-decoding line 3732, source line 3733, and erase gate line 3734. In this layout, word lines 3711, 3712, 3713, and 3714 and bit lines 3721, 3722, 3723, and 3724 are parallel to each other and arranged in a vertical direction. HV / LV pre-decode lines 3732, source lines 3733, and erase gate lines 3734 are parallel to each other and arranged in a horizontal direction, and are therefore perpendicular to word lines 3711, 3712, 3713, and 3714, and bit lines 3721, 3722, 3723, and 3724. HV / LV pre-decode lines 3732 are input to local high voltage decoders 3709 and 3710. Local high voltage decoder 3709 outputs local control gate lines for VMM arrays 3701 and 3702. Local high voltage decoder 3710 outputs local control gate lines for VMM arrays 3703 and 3704. In another embodiment, local high voltage decoders 3709 and 3710 can provide local source lines for VMM arrays 3701 / 3702 and VMM arrays 3703 / 3704, respectively. In another embodiment, local high voltage decoders 3709 and 3710 may provide local erase gate lines for VMM arrays 3701 / 3702 and VMM arrays 3703 / 3704, respectively.

[0198] Here, local high voltage row decoder 3709 is shared by VMM arrays 3701 and 3702, and local high voltage row decoder 3710 is shared by VMM arrays 3703 and 3704. Global high voltage decoder 3730 routes high voltage and low voltage pre-decode signals to local high voltage row decoders, such as local high voltage row decoders 3709 and 3710. Thus, the high voltage decoding function is divided between global high voltage row decoder 3730 and local high voltage decoders, such as local high voltage decoders 3709 and 3710.

[0199] In the VMM system 3700, the VMM arrays 3701, 3702, 3703, and 3704 share the HV / LV pre-decoding lines 3732, the source lines 3733, the erase gate lines 3734, and the global high voltage row decoder 3730. However, each of the VMM arrays has its own low voltage row decoder such that the low voltage row decoder 3705 is used with the VMM array 3701; the low voltage row decoder 3706 is used with the VMM array 3702; the low voltage row decoder 3707 is used with the VMM array 3703; and the low voltage row decoder 3708 is used with the VMM array 3704. Advantageously for this layout, the word lines 3711, 3712, 3713, and 3714 are arranged in the vertical direction such that the word line 3711 can be routed only to the VMM array 3701, the word line 3712 can be routed only to the VMM array 3702, the word line 3713 can be routed only to the VMM array 3703, and the word line 3714 can be routed only to the VMM array 3704. This would be very inefficient with a conventional layout in which, for multiple arrays sharing a single high voltage decoder, the word lines are arranged in the horizontal direction.

[0200] Figure 38 Figure 3800 shows a VMM system that is particularly suitable for use with Figure 4is used in conjunction with memory cells of the type shown as memory cell 410. The VMM system 3800 includes VMM arrays 3801, 3802, 3802, and 3804 (each of which can be based on any of the foregoing VMM designs, such as VMM arrays 1000, 1100, 1200, 1300, 1400, 1500, 2400, 2500, 2600, 2700, 2800, 2900, 3000, 3100, and 3200 or other VMM array designs); low voltage row decoders 3805, 3806, 3807, and 3808; local high voltage row decoders 3809 and 3810; global high voltage row decoder 3830; bit lines 3821, 3822, 3823, and 3824; control gate lines or control gate input lines 3811 and 3812, HV / LV pre-decoding lines 3833, source lines 3834, and erase gate lines 3835. The shared global high voltage row decoder 3830 provides the HV / LV pre-decoding lines 3833, source lines 3834, and erase gate lines 3835. The local high voltage decoders 3809 and 3810 couple the control gate inputs CG 3811 and 3812 to the local control gates of VMM arrays 3801, 3802, and 3803, 3804, respectively. The low voltage row decoders 3805, 3806, 3807, and 3808 provide local (horizontal) word lines to arrays 3801, 3802, 3803, 3804, respectively. In this layout, the control gate lines 3811 and 3812 and the bit lines 3821, 3822, 3823, and 3824 are parallel to each other and arranged in the vertical direction. The source lines 3834 and the erase gate lines 3835 are parallel to each other and arranged in the horizontal direction, thus perpendicular to the control gate lines 3811 and 3812 and the bit lines 3821, 3822, 3823, and 3824.

[0201] As in Figure 37 the VMM system 3700, the local high voltage row decoder 3809 is shared by VMM arrays 3801 and 3802, and the local high voltage row decoder 3810 is shared by VMM arrays 3803 and 3804. The global high voltage decoder 3830 routes signals to the local high voltage row decoders, such as local high voltage row decoders 3809 and 3810. Thus, the high voltage decoding function is divided between the global high voltage row decoder 3830 and local high voltage decoders such as local high voltage decoders 3809 and 3810 (which can provide local source lines and / or local erase gate lines).

[0202] In VMM system 3800, VMM arrays 3801, 3802, 3803, and 3804 share HV / LV pre-decode lines 3833, source lines 3834, erase gate lines 3835, and a global high-voltage row decoder 3830. However, each VMM array has its own low-voltage row decoder, such that low-voltage row decoder 3805 is used with VMM array 3801; low-voltage row decoder 3806 is used with VMM array 3802; low-voltage row decoder 3807 is used with VMM array 3803; and low-voltage row decoder 3808 is used with VMM array 3804. This layout is advantageous in that control gate lines 3811 and 3812 (which can be read lines or input lines) are arranged in a vertical direction, so that control gate line 3811 can be routed only to VMM arrays 3801 and 3802, and control gate line 3812 can be routed only to VMM arrays 3803 and 3804. This would not be possible using a conventional layout where word lines are arranged in a horizontal direction.

[0203] Figure 39 A VMM system 3900 is shown, which is particularly suitable for use with Figure 3 310, Figure 4 410, Figure 5 is shown as memory cell 510 or Figure 7 7. A VMM system 3900 is used with memory cells of the type shown as memory cells 710 in FIG. 7. VMM system 3900 includes VMM arrays 3901 and 3902 (each of which can be based on any of the aforementioned VMM designs, such as VMM arrays 1000, 1100, 1200, 1300, 1400, 1500, 2400, 2500, 2600, 2700, 2800, 2900, 3000, 3100, and 3200, or other VMM array designs); a low-voltage row decoder 3903 (used with arrays 3901 and 3902); a local high-voltage row decoder 3905, a global high-voltage row decoder 3904; control gate lines 3908 and 3909; and bit lines 3906 and 3907. In this layout, control gate line 3908 is used only by VMM array 3901, and control gate line 3909 is used only by VMM array 3902. Low voltage row decode line 3910 serves as the decode input to global high voltage row decoder 3904. Global high voltage row decode line 3911 serves as the decode input to local high voltage decoder 3905.

[0204] Local high voltage row decoder 3905 is shared by VMM arrays 3901 and 3902. Global high voltage decoder 3904 routes signals to local high voltage row decoders of multiple VMM systems, such as local high voltage row decoder 3905 of VMM system 3900. Thus, the high voltage decoding function is divided between global high voltage row decoder 3904 and local high voltage decoders, such as local high voltage decoder 3905, as described above.

[0205] In VMM system 3900, VMM arrays 3901 and 3902 share word lines (not shown), source gate lines (if present) (not shown), erase gate lines (if present) (not shown), and a global high-voltage row decoder 3904. Here, VMM arrays 3901 and 3902 share a low-voltage row decoder 3903. Advantageously, VMM arrays 3901 and 3902 do not share control gate lines, enabling each array to be independently accessed using control gate lines 3908 and 3909, respectively.

[0206] Figure 51 A VMM system 5100 is shown, which is particularly suitable for use with Figure 4 4. VMM system 5100 includes VMM arrays 5101, 5102, 5103, and 5104 (each of which can be based on any of the aforementioned VMM array designs, such as VMM arrays 1000, 1100, 1200, 1300, 1400, 1510, 2400, 2510, 2600, 2700, 2800, 2900, 3000, 3100, and 3200, or other VMM array designs); a high voltage decoder 5130; routing blocks 5151 and 5152; input word lines 5111 and 5112, bit lines 5121, 5122, 5123, and 5124; a control gate line 5132, a source line 5133, and an erase gate line 5134. The high voltage decoder 5130 provides control gate lines 5132, source lines 5133, and erase gate lines 5134. Routing blocks 5151 and 5152 are where the vertically received input word lines 5111 and 5112, respectively, are routed to the horizontally running word lines of the VMM arrays 5101-5104. Alternatively, routing blocks 5151 and 5152 can route the vertically received control gate input lines 5132 to the horizontally running control gate lines 5132 of the VMM array.

[0207] Figure 40Shows a low-voltage row decoder 4000, which includes a NAND gate 4001, a PMOS transistor 4002, and an NMOS transistor 4003. The NAND gate 4001 receives a row address signal 4004. The PMOS transistor 4002 is coupled to a vertical word line input 4005. The output is on a horizontal word line 4006, which is one of a number of word lines coupled to a corresponding VMM array. In this example, there are a total of 16 word lines, and thus there will be 16 instantiations of the row decoder 4000, each instantiation outputting one of the 16 word lines. Thus, based on the received row address signal, one word line (such as word line 4006) will output a corresponding signal (such as a voltage), and the other word lines will be set to ground.

[0208] Figure 41 Shows a combined common select / deselect word line and control gate decoder 4100, which includes a low-voltage row decoder as Figure 40 described herein, including a NAND gate 4101, a PMOS transistor 4102, an NMOS transistor 4103, a row address signal 4104, a vertical input word line 4105, and a horizontal word output line 4106 coupled to the word lines of the VMM array. The combined word line and control gate decoder 4100 also includes an inverter 4107, switches 4108 and 4112, and an isolation transistor 4109, and receives a control gate input 4110 CGIN0 and outputs a control gate line 4111 CG0. The word line output 4106 WL0 and the control gate output CG0 4111 are simultaneously selected or deselected by decoding logic components (not shown) that control the NAND gate 4101.

[0209] Figure 42 Shows a bit line decoder 4200, which operates on VMM arrays 4201 and 4202. The bit line decoder 4200 includes a column multiplexer 4203 (for selecting one or more bit lines for programming and verification, where the verification operation is used to confirm that the cell current reaches a certain target during a tuning operation (programming or erasing operation)), and a sense amplifier 4204 (for performing a read operation on one or more bit lines). As shown, local bit line muxes 4201b and 4202b multiplex local array bit lines to global bit lines 4220x to be coupled to the column multiplexer 4203. The sense amplifier includes an ADC or other device. Thus, the bit line decoder 4200 is shared across multiple arrays.

[0210] Figure 43Shows a VMM system 4300, which includes VMM arrays 4301, 4302, 4303, and 4304; low-voltage row decoders 4305 and 4307; local high-voltage row decoders 4306 and 4308, a global high-voltage row decoder 4309, digital bus inputs QIN[7:0] 4311 and 4312 (here they are inputs to the VMM arrays), and bit lines 4321, 4322, 4323, and 4324. Each low-voltage row decoder, such as low-voltage row decoder 4305, includes a circuit block row decoder 4335 for each word line, such as an exemplary data input block 4331 (which may consist of 8 latches or registers) and a block 4332 (which may include a data-voltage converter circuit or a data-pulse converter circuit), which outputs a signal 4333 on the word line. Thus, the input to this low-voltage row decoder is the digital bus QIN[7:0] with appropriate control logic components. For each circuit block row decoder 4335, the digital inputs QIN[7:0] 4311 and 4312 are appropriately latched, such as by a synchronous clock device and method (such as through a serial-to-parallel clock interface).

[0211] Figure 44 Shows a neural network array input-output bus multiplexer 4400, which receives outputs from a VMM array (such as from an ADC) and provides those outputs to the input blocks of other VMM arrays (such as a DAC or DPC) in a multiplexed manner. In the example shown, the inputs to the input-output bus multiplexer 4400 include 2048 bits (256 groups, NEU0...NEU255, each group 8 bits), and the input-output bus multiplexer 4400 provides those bits in 64 different groups (each group 32 bits), where it multiplexes between different groups, such as by using time-division multiplexing (where it provides 1 group of 32 bits at any given time). Control logic component 4401 generates a control signal 4402 to control the input-output bus multiplexer 4400.

[0212] Figure 45A and Figure 45B Shows an exemplary layout of a VMM array, where the word lines are arranged horizontally ( Figure 45A ) and vertically ( Figure 45B , such as in Figure 34 or Figure 35 ).

[0213] Figure 46 Shows an exemplary layout of a VMM array, where the word lines are arranged vertically (such as in Figure 34 or Figure 35 ). However, in this layout, two word lines (such as word lines 4601 and 4602) can occupy the same column, but access different rows in the array (due to the gap between them).

[0214] Figure 47 shows a VMM high-voltage decoding circuit, which includes a word line decoder circuit 4701, a source line decoder circuit 4704, and a high-voltage level shifter 4708 suitable for use with Figure 2 memory cells of the type shown.

[0215] The word line decoder circuit 4701 includes a PMOS selection transistor 4702 (controlled by the signal HVO_B) and an NMOS deselection transistor 4703 (controlled by the signal HVO_B) configured as shown.

[0216] The source line decoder circuit 4704 includes an NMOS monitoring transistor 4705 (controlled by the signal HVO), a driving transistor 4706 (controlled by the signal HVO), and a deselection transistor 4707 (controlled by the signal HVO_B) configured as shown.

[0217] The high-voltage level shifter 4708 receives an enable signal EN and outputs a high-voltage signal HV and its complementary signal HVO_B.

[0218] Figure 48 shows a VMM high-voltage decoding circuit, which includes an erase gate decoder circuit 4801, a control gate decoder circuit 4804, a source line decoder circuit 4807, and a high-voltage level shifter 4811 suitable for use with Figure 3 memory cells of the type shown.

[0219] The erase gate decoder circuit 4801 and the control gate decoder circuit 4804 use the same design as the word line decoder circuit 4701 in Figure 47 .

[0220] The source line decoder circuit 4807 uses the same design as the source line decoder circuit 4704 in Figure 47 .

[0221] The high-voltage level shifter 4811 uses the same design as the high-voltage level shifter 4708 in Figure 47 .

[0222] Figure 49The word line driver 4900 is shown. The word line driver 4900 selects word lines (such as the exemplary word lines WL0, WL1, WL2, and WL3 shown here) and provides a bias voltage to the word lines. Each word line is attached to a select isolation transistor controlled by a control line 4902, such as select transistor 4901. The select transistor (such as select transistor 4901) isolates the high voltage (e.g., 8 - 12V) used during the erase operation from the word line decoding transistors, which can be implemented with IO transistors operating at a low voltage (e.g., 1.8V, 3.3V). Here, during any operation, the control line 4902 is activated, and all select transistors similar to select transistor 4901 are turned on. The exemplary bias transistor 4903 (a part of the word line decoding circuit) selectively couples the word line to a first bias voltage (such as 3V), and the exemplary bias transistor 4904 (a part of the word line decoding circuit) selectively couples the word line to a second bias voltage (lower than the first bias voltage, including ground, a bias between ground and the first bias voltage, or a negative voltage bias to reduce leakage from unused memory rows). During an ANN (analog neural network) read operation, all used word lines are selected and tied to the first bias voltage. All unused word lines are tied to the second bias voltage. During other operations such as a programming operation, only one word line is selected and the other word lines are tied to the second bias voltage, which can be a negative bias (e.g., - 0.3V to - 0.5V or greater) to reduce array leakage.

[0223] The bias transistors 4903 and 4904 are coupled to the output of stage 4906 of the shift register 4905. The shift register 4905 enables each row to be independently controlled according to an input data pattern (loaded at the start of the ANN operation).

[0224] Figure 50 The word line driver 5000 is shown. The word line driver 5000 is similar to the word line driver 4900, except that each select transistor is further coupled to a capacitor (such as capacitor 5001). The capacitor 5001 can provide pre - charge or bias to the word line at the start of operation, enabled by transistor 5002 to sample the voltage on line 5003. The capacitor 5001 is used for sample - hold (S / H) of the input voltage of each word line. Transistors 5004 and 5005 are turned off during the ANN operation (array current adder and activation function) of the VMM array, which means the voltage on the S / H capacitor 5001 will be used as the (floating) voltage source for the corresponding word line. Alternatively, the capacitor 5001 can be provided by the word line capacitance from the VMM array (or as the control gate capacitance if the input is on the control gate).

[0225] It should be noted that, as used herein, the terms "above" and "on" both inclusively include "directly on" (with no intervening material, element, or space therebetween) and "indirectly on" (with intervening material, element, or space therebetween). Similarly, the term "adjacent" includes "directly adjacent" (with no intervening material, element, or space therebetween) and "indirectly adjacent" (with intervening material, element, or space therebetween), "mounted to" includes "directly mounted to" (with no intervening material, element, or space therebetween) and "indirectly mounted to" (with intervening material, element, or space therebetween), and "electrically coupled to" includes "directly electrically coupled to" (with no intervening material or element electrically connecting the elements together) and "indirectly electrically coupled to" (with intervening material or element electrically connecting the elements together). For example, forming an element "above a substrate" can include directly forming the element on the substrate with no intervening material / element therebetween, and indirectly forming the element on the substrate with one or more intervening material / elements therebetween.

Claims

1. A simulated neural memory system, comprising: A first vector - matrix multiplication array, the first vector - matrix multiplication array comprising an array of non - volatile memory cells organized in rows and columns, wherein each memory cell comprises a bit - line terminal, a source - line terminal, and a word - line terminal; A second vector - matrix multiplication array, the second vector - matrix multiplication array comprising an array of non - volatile memory cells organized in rows and columns, wherein each memory cell comprises a bit - line terminal, a source - line terminal, and a word - line terminal; A first plurality of bit - lines, the first plurality of bit - lines coupled to the bit - line terminals of the column memory cells in the first vector - matrix multiplication array; A second plurality of bit - lines, the second plurality of bit - lines coupled to the bit - line terminals of the column memory cells in the second vector - matrix multiplication array; A first low - voltage row decoder, the first low - voltage row decoder coupled to the word - line terminals of the row memory cells in the first vector - matrix multiplication array; A second low - voltage row decoder, the second low - voltage row decoder coupled to the word - line terminals of the row memory cells in the second vector - matrix multiplication array; A first plurality of word - lines, the first plurality of word - lines coupled to the word - line terminals of the row memory cells in the first vector - matrix multiplication array through the first low - voltage row decoder; A second plurality of word - lines, the second plurality of word - lines coupled to the word - line terminals of the row memory cells in the second vector - matrix multiplication array through the second low - voltage row decoder; And A plurality of source - lines, wherein each of the plurality of source - lines is coupled to the source - line terminals of the row memory cells in the first vector - matrix multiplication array and coupled to the source - line terminals of the row memory cells in the second vector - matrix multiplication array; Wherein the first plurality of word - lines and the second plurality of word - lines are parallel to the first plurality of bit - lines and the second plurality of bit - lines and perpendicular to the plurality of source - lines, and during a neuron read operation of the first vector - matrix multiplication array or the second vector - matrix multiplication array, an output equal to the sum of the products generated for each cell in the corresponding first vector - matrix multiplication array or second vector - matrix multiplication array is provided on the plurality of source - lines, wherein each product is equal to the input current received at the word - line terminal of the cell multiplied by the weight stored in the cell.

2. The system according to claim 1, wherein the non - volatile memory cell is a split - gate flash memory cell.

3. The system according to claim 1, wherein the non - volatile memory cell is a stacked - gate flash memory cell.

4. The system according to claim 1, further comprising: One or more routing blocks that couple the plurality of word - lines to the word - line terminals of a row of memory cells, wherein the row is arranged in a direction perpendicular to the plurality of word - lines.

5. The system according to claim 4, wherein the non - volatile memory cell is a split - gate flash memory cell.

6. The system according to claim 4, wherein the non-volatile memory cell is a stacked-gate flash memory cell.

7. An analog neural memory system, comprising: A first vector-matrix multiplication array, the first vector-matrix multiplication array including an array of non-volatile memory cells organized in rows and columns, wherein each memory cell includes a bit line terminal, a control gate terminal, and a word line terminal; A second vector-matrix multiplication array, the second vector-matrix multiplication array including an array of non-volatile memory cells organized in rows and columns, wherein each memory cell includes a bit line terminal, a control gate terminal, and a word line terminal; A first plurality of bit lines, the first plurality of bit lines coupled to the bit line terminals of the column memory cells in the first vector-matrix multiplication array; A second plurality of bit lines, the second plurality of bit lines coupled to the bit line terminals of the column memory cells in the second vector-matrix multiplication array; A first low voltage row decoder, the first low voltage row decoder coupled to the word line terminals of the row memory cells in the first vector-matrix multiplication array; A second low voltage row decoder, the second low voltage row decoder coupled to the word line terminals of the row memory cells in the second vector-matrix multiplication array; A plurality of control gate lines, wherein each of the plurality of control gate lines is coupled to the control gate terminals of a row of memory cells; And A first plurality of word lines, the first plurality of word lines coupled to the word line terminals of the row memory cells in the first vector-matrix multiplication array through the first low voltage row decoder; A second plurality of word lines, the second plurality of word lines coupled to the word line terminals of the row memory cells in the second vector-matrix multiplication array through the second low voltage row decoder; Wherein the plurality of control gate lines are parallel to the plurality of bit lines and perpendicular to the plurality of word lines, and during a neuron read operation of the first vector-matrix multiplication array or the second vector-matrix multiplication array, an output equal to the sum of the products generated for each cell in the corresponding first vector-matrix multiplication array or second vector-matrix multiplication array is provided on the plurality of bit lines, wherein each product is equal to the input current received on the control gate terminal of the cell multiplied by the weight stored in the cell.

8. The system according to claim 7, wherein the non-volatile memory cell is a split-gate flash memory cell.

9. The system according to claim 7, wherein the non-volatile memory cell is a stacked-gate flash memory cell.

Citation Information

Patent Citations

  • Deep learning neural network classifier using non-volatile memory array

    US11308383B2

  • Deep Learning Neural Network Classifier Using Non-volatile Memory Array

    US20170337466A1

  • High Precision And Highly Efficient Tuning Mechanisms And Algorithms For Analog Neuromorphic Memory In Artificial Neural Networks

    US20190164617A1

  • Single transistor non-valatile electrically alterable semiconductor memory device

    US5029130A

  • Flash memory cells with separated self-aligned select and erase gates, and process of fabrication

    US6747310B2