Configurable input blocks and output blocks and physical layout for analog neural memory in deep learning artificial neural network
The configurable input and output blocks for analog neural memory systems using non-volatile memory cells address the inefficiencies in existing neural networks by optimizing space and energy usage, enabling flexible and efficient vector matrix multiplication.
Patent Information
- Application Number
- JP2025030034
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-06-21
- Filing Date
- 2025-02-27
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2039-11-18
Smart Images

Figure 2025106236000001_ABST
Abstract
Description
Technical Field
[0001] (Claim of Priority) This application claims priority to U.S. Provisional Patent Application No. 62 / 842,279, filed May 2, 2019, entitled "CONFIGURABLE INPUT BLOCKS AND OUTPUT BLOCKS AND PHYSICAL LAYOUT FOR ANALOG NEURAL MEMORY IN DEEP LEARNING ARTIFICIAL NEURAL NETWORK", and U.S. Patent Application No. 16 / 449,201, filed Jun. 21, 2019, entitled "CONFIGURABLE INPUT BLOCKS AND OUTPUT BLOCKS AND PHYSICAL LAYOUT FOR ANALOG NEURAL MEMORY IN DEEP LEARNING ARTIFICIAL NEURAL NETWORK".
[0002] (Field of the Invention) Disclosed are configurable input blocks and output blocks, and related physical layouts, for an analog neural memory system that utilizes non-volatile memory cells.
Background Art
[0003] An artificial neural network mimics a biological neural network (the central nervous system of an animal, particularly the brain), can rely on a large number of inputs, and is used to estimate or approximate a generally unknown function. An artificial neural network generally includes layers of interconnected "neurons" that exchange messages.
[0004] Figure 1 shows an artificial neural network, in which the circles represent input or neuron layers. Connections (referred to as synapses) are represented by arrows and have numerical weights that can be adjusted based on experience. As a result, the neural network adapts to the input and becomes learnable. Typically, a neural network includes multiple input layers. Typically, there is one or more intermediate layers of neurons and an output layer of neurons that provides the output of the neural network. At each level, the neurons make decisions individually or jointly based on the data received from the synapses.
[0005] One of the major challenges in the development of artificial neural networks for high-performance information processing is the lack of appropriate hardware technologies. In practice, practical neural networks rely on a very large number of synapses, which enables high connectivity between neurons, i.e., a very high degree of parallelization of computational processing. In principle, such complexity can be realized by a digital supercomputer or a dedicated GPU (graphics processing unit) cluster. However, in addition to high costs, these approaches also suffer from poor energy efficiency compared to biological networks, which mainly perform low-precision analog calculations and consume much less energy. CMOS analog circuits have been used in artificial neural networks, but most CMOS implementation synapses have been too bulky assuming the required large number of neurons and synapses.
[0006] The applicant has previously disclosed in U.S. Patent Application No. 15 / 594,439, published as U.S. Patent Publication No. 2017 / 0337466, incorporated by reference, an artificial (analog) neural network that utilizes one or more non-volatile memory arrays as synapses. The non-volatile memory arrays operate as analog neural memories. The neural network device includes a first plurality of synapses configured to receive a first plurality of inputs and generate a first plurality of outputs therefrom, and a first plurality of neurons configured to receive the first plurality of outputs. The first plurality of synapses includes a plurality of memory cells, each memory cell including a spaced source region and drain region formed in a semiconductor substrate with a channel region extending therebetween, a floating gate disposed above a first portion of the channel region and insulated from the first portion of the channel region, and a non-floating gate disposed above a second portion of the channel region and insulated from the second portion of the channel region. Each of the plurality of memory cells is configured to store a weight value corresponding to a number of electrons on the floating gate. The plurality of memory cells is configured to multiply the stored weight values by the first plurality of inputs to generate the first plurality of outputs.
[0007] Each non-volatile memory cell used in an analog neural memory system must be erased and programmed to hold a very specific and accurate amount of charge, i.e., number of electrons, within the floating gate. For example, each floating gate must hold one of N different values, where N is the number of different weights that can be represented by each cell. Examples of N include 16, 32, 64, 128, and 256.
[0008] One challenge in a vector matrix multiplication (VMM) system is the ability to quickly and accurately deliver the output from a VMM as an input to another VMM while efficiently utilizing the physical space within a semiconductor die.
[0009] What is required is a configurable input block and output block for an analog neural memory system that utilizes non-volatile memory cells, as well as a physical layout.
Summary of the Invention
[0010] Disclosed are a configurable input block and output block for an analog neural memory system that utilizes non-volatile memory cells, and a related physical layout.
[0011] One embodiment of an analog neural memory system is a plurality of vector-matrix multiplication arrays, each array including non-volatile memory cells arranged in rows and columns, the plurality of vector-matrix multiplication arrays, and an input block capable of providing an input to a configurable number N of the plurality of vector-matrix multiplication arrays, where N can range between 1 and the total number of arrays in the plurality of vector-matrix multiplication arrays, the input block, and the array that receives the input provides an output in response to the input.
[0012] Another embodiment of an analog neural memory system is a plurality of vector-matrix multiplication arrays, each of the plurality of vector-matrix multiplication arrays including non-volatile memory cells arranged in rows and columns, the plurality of vector-matrix multiplication arrays, and an output block capable of providing an output from a configurable number N of the plurality of vector-matrix multiplication arrays, where N can range between 1 and the total number of arrays in the plurality of vector-matrix multiplication arrays, the output block, and the output is provided in response to the received input.
[0013] Another embodiment of the analog neural memory system is a plurality of vector matrix multiplication arrays, each array including non-volatile memory cells arranged in rows and columns, the plurality of vector matrix multiplication arrays, and an output block for performing a verification operation after a programming operation on a configurable number N of the vector matrix multiplication arrays, where N can range between 1 and the total number of arrays in the plurality of vector matrix multiplication arrays.
[0014] Another embodiment of the analog neural memory system is a plurality of vector matrix multiplication arrays, each array including non-volatile memory cells arranged in rows and columns, the plurality of vector matrix multiplication arrays, an input block capable of providing an input to a first configurable number N of the vector matrix multiplication arrays, where N can range between 1 and the total number of arrays in the plurality of vector matrix multiplication arrays, an output block capable of providing an output from a second configurable number M of the vector matrix multiplication arrays, where M can range between 1 and the total number of arrays in the plurality of vector matrix multiplication arrays, and the output block generates the output in response to the input.
[0015] Another embodiment of the analog neural memory system is a plurality of vector matrix multiplication arrays, each vector matrix multiplication array including non-volatile memory cells arranged in rows and columns, the plurality of vector matrix multiplication arrays, and an output block capable of receiving an output neuron current from one or more of the vector matrix multiplication arrays and generating digital output bits using a ramp-type analog-to-digital converter.
[0016] Another embodiment of the analog neural memory system is a plurality of vector matrix multiplication arrays, each vector matrix multiplication array comprising a plurality of vector matrix multiplication arrays including non-volatile memory cells, and an input block capable of converting a plurality of digital input bits into a binary-indexed time addition signal as a timing input to at least one of the vector matrix multiplication arrays.
[0017] An embodiment of a method for performing output conversion on an analog neural memory including a plurality of vector matrix multiplication arrays, each vector matrix multiplication array including non-volatile memory cells, includes receiving an output neuron current from one or more of the plurality of vector matrix multiplication arrays, and generating digital output bits using the output neuron current and a ramp-type analog-to-digital converter, the converter operating in a coarse comparison mode and a fine comparison mode.
[0018]
[0019]
[0020]
[0021]
[0022]
[0023]
[0024]
[0025]
[0026]
[0027]
[0028]
[0029]
[0030]
[0031]
[0032]
[0033]
[0034]
[0035]
[0036]
[0037]
[0038]
[0039]
[0040]
[0041]
[0042]
[0043]
[0044]
[0045]
[0046]
[0047]
[0048]
[0049]
[0050]
[0051]
[0052]
[0053]
[0054]
[0055]
[0056]
[0057]
[0058]
[0059]
[0060]
[0061]
[0062]
[0063]
[0064]
[0065]
[0066]
[0067]
[0068]
[0069]
[0070]
[0071]
[0072]
[0073]
[0074]
[0075]
[0076]
[0077]
[0078]
Brief Description of the Drawings
[0079]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
Figure 26
Figure 27
Figure 28
Figure 29
Figure 30
Figure 31
Figure 32
Figure 33
Figure 34
Figure 35A
Figure 35B
Figure 36
Figure 37A
Figure 37B
Figure 38
Figure 39
Figure 40
Figure 41
Figure 42
Figure 43A
Figure 43B
Figure 44A
Figure 44B
Figure 44C
Figure 45A
Figure 45B
Figure 45C
Figure 46
Figure 47A
Figure 47B
Figure 48
Figure 49
Figure 50A
Figure 50B
Figure 51A
Figure 51B
Figure 52
DETAILED DESCRIPTION OF THE INVENTION
[0080] The artificial neural network of the present invention utilizes a combination of CMOS technology and a non-volatile memory array. Non-volatile memory cell
[0081] Digital non-volatile memories are well known. For example, U.S. Patent No. 5,029,130 (the “'130 patent”), which is incorporated herein by reference, discloses an array of split-gate non-volatile memory cells, which are a type of flash memory cell. Such a memory cell 210 is shown in FIG. 2. Each memory cell 210 includes a source region 14 and a drain region 16 formed in a semiconductor substrate 12, and a channel region 18 is between the source region 14 and the drain region 16. A floating gate 20 is formed above a first portion of the channel region 18, insulated from the first portion of the channel region 18 (and controlling the conductivity of the first portion of the channel region 18), and is formed over a portion of the source region 14. A word line terminal 22 (typically coupled to a word line) is disposed above a second portion of the channel region 18, insulated from the second portion of the channel region 18, and has a first portion (controlling the conductivity of the second portion of the channel region 18) and a second portion extending upward over the floating gate 20. The floating gate 20 and the word line terminal 22 are insulated from the substrate 12 by a gate oxide. A bit line 24 is coupled to the drain region 16.
[0082] By applying a high positive voltage to the word line terminal 22, the memory cell 210 is erased (electrons are removed from the floating gate). As a result, the electrons in the floating gate 20 pass through the insulator between them from the floating gate 20 to the word line terminal 22 via Fowler-Nordheim tunneling.
[0083] The memory cell 210 is programmed (electrons are applied to the floating gate) by applying a positive voltage to the word line terminal 22 and a positive voltage to the source region 14. The electron current flows from the source region 14 towards the drain region 16. The electrons accelerate and heat up when they reach the gap between the word line terminal 22 and the floating gate 20. A part of the heated electrons is injected into the floating gate 20 through the gate oxide due to the electrostatic attraction from the floating gate 20.
[0084] The memory cell 210 is read by applying a positive read voltage to the drain region 16 and the word line terminal 22 (turning on the part of the channel region 18 below the word line terminal). When the floating gate 20 is positively charged (i.e., electrons are erased), the part of the channel region 18 below the floating gate 20 is also turned on, and the current flows through the channel region 18, which is detected as the erased state, i.e., the "1" state. When the floating gate 20 is negatively charged (i.e., programmed with electrons), the part of the channel region below the floating gate 20 is almost or completely turned off, and the current does not flow (or hardly flows) through the channel region 18, which is detected as the programmed state, i.e., the "0" state.
[0085] Table 1 shows the typical voltage ranges that can be applied to the terminals of the memory cell 110 to perform read, erase, and program operations. Table 1: Operation of the flash memory cell 210 in FIG. 2
Table 1
[0086] FIG. 3 shows a memory cell 310 similar to the memory cell 210 of FIG. 2 with an additional control gate (CG) 28. The control gate 28 is biased at a high voltage (e.g., 10V) during programming, a low or negative voltage (e.g., 0V / -8V) during erasure, and a low or medium voltage (e.g., 0V / 2.5V) during readout. The other terminals are biased in the same manner as the terminals of FIG. 2.
[0087] FIG. 4 shows a four-gate memory cell 410 including a source region 14, a drain region 16, a floating gate 20 above a first portion of the channel region 18, a select gate 22 (typically coupled to a word line, WL) above a second portion of the channel region 18, a control gate 28 above the floating gate 20, and an erase gate 30 above the source region 14. This configuration is described in U.S. Patent No. 6,747,310, which is incorporated herein by reference for all purposes. Here, all gates are non-floating gates except for the floating gate 20, i.e., they are electrically connected or connectable to a voltage source. Programming is performed by injecting hot electrons from the channel region 18 into the floating gate 20 itself. Erasure is performed by tunneling electrons from the floating gate 20 to the erase gate 30.
[0088] Table 2 shows typical voltage ranges that can be applied to the terminals of the memory cell 310 to perform readout, erasure, and program operations. Table 2: Operation of the Flash Memory Cell 410 of FIG. 4
Table 2
[0089] FIG. 5 shows a memory cell 510 similar to the memory cell 410 of FIG. 4, except that the memory cell 510 does not include an erase gate (EG). Erasure is performed by biasing the substrate 18 to a high voltage and biasing the control gate CG28 to a low voltage or a negative voltage. Alternatively, erasure is performed by biasing the word line 22 to a positive voltage and biasing the control gate 28 to a negative voltage. Programming and reading are the same as those in FIG. 4.
[0090] FIG. 6 shows a three-gate memory cell 610, which is another type of flash memory cell. The memory cell 610 is identical to the memory cell 410 of FIG. 4, except that the memory cell 610 does not have a separate control gate. (Erasure occurs through the use of an erase gate) The erasure operation and the read operation are the same as those in FIG. 4, except that no control gate bias is applied. The programming operation is also performed without a control gate bias. As a result, during the program operation, a higher voltage must be applied to the source line to compensate for the lack of control gate bias.
[0091] Table 3 shows the typical voltage ranges that can be applied to the terminals of the memory cell 610 to perform read, erase, and program operations. Table 3: Operations of the Flash Memory Cell 610 in FIG. 6
Table 3
[0092] FIG. 7 shows a stacked gate memory cell 710, which is another type of flash memory cell. Memory cell 710 is similar to memory cell 210 of FIG. 2, except that floating gate 20 extends over the entire channel region 18 and control gate 22 (coupled to the word line) extends over floating gate 20 separated by an insulating layer (not shown). Erase, programming, and read operations operate in a similar manner as described above for memory cell 210.
[0093] Table 4 shows typical voltage ranges that can be applied to the terminals of memory cell 710 and substrate 12 to perform read, erase, and program operations. Table 4 Operation of Flash Memory Cell 710 of FIG. 7
Table 4
[0094] "Read 1" is a read mode in which the cell current is output to the bit line. "Read 2" is a read mode in which the cell current is output to the source line. Optionally, in an array including rows and columns of memory cells 210, 310, 410, 510, 610, or 710, the source line can be coupled to one row of memory cells or two adjacent rows of memory cells. That is, the source line can be shared by adjacent rows of memory cells.
[0095] To utilize a memory array including one of the types of non-volatile memory cells in the artificial neural network described above, two modifications are made. First, the lines are configured so that each memory cell can be individually programmed, erased, and read without adversely affecting the memory states of other memory cells in the array, as further described below. Second, continuous (analog) programming of the memory cells is provided.
[0096] Specifically, the memory state of each memory cell in the array (i.e., the charge on the floating gate) can be changed independently, with minimal interference from other memory cells, continuously from a completely erased state to a completely programmed state. In another embodiment, the memory state of each memory cell in the array (i.e., the charge on the floating gate) can be changed independently, with minimal interference from other memory cells, continuously from a completely programmed state to a completely erased state and vice versa. This means that cell storage is either analog or can store at least one of a number of discrete values (such as 16 or 64 different values), which allows all cells in the memory array to be adjusted very precisely and individually, making the memory array ideal for storage and allowing for fine-tuning of the synaptic weights of a neural network.
[0097] The methods and means described herein can be applied, without limitation, to other non-volatile memory technologies such as SONOS (silicon-oxide-nitride-oxide-silicon, charge trapping in nitride), MONOS (metal-oxide-nitride-oxide-silicon, metal charge trapping in nitride), ReRAM (resistive random access memory), PCM (phase change memory), MRAM (magnetoresistive random access memory), FeRAM (ferroelectric random access memory), OTP (one-time programmable, bi-level or multi-level), and CeRAM (strongly correlated electron memory). The methods and means described herein can be applied, without limitation, to volatile memory technologies used in neural networks such as SRAM, DRAM, and volatile synaptic cells. Neural Network Using Non-Volatile Memory Cell Array
[0098] FIG. 8 conceptually shows a non-limiting example of a neural network that utilizes the non-volatile memory array of this embodiment. This example uses a non-volatile memory array neural network for a face recognition application, but it is also possible to implement other suitable applications using a non-volatile memory array-based neural network.
[0099] S0 is the input layer, which in this example is a 32×32 pixel RGB image with 5-bit precision (i.e., three 32×32 pixel arrays, one for each of the colors R, G, and B, and each pixel has 5-bit precision). The synapses CB1 from the input layer S0 to the layer C1 apply a different set of weights to some instances and shared weights to other instances, scanning the input image with a 3×3 pixel overlapping filter (kernel) and shifting the filter by 1 pixel (or more than 2 pixels in some models) at a time. Specifically, the 9 pixel values in the 3×3 portion of the image (i.e., what is called the filter or kernel) are provided to the synapses CB1, where these 9 input values are multiplied by appropriate weights, and after summing the outputs of that multiplication, a single output value is determined and given by the first synapse of CB1 to generate one pixel of the layer of the feature map C1. The 3×3 filter is then shifted 1 pixel to the right within the input layer S0 (i.e., a 3-pixel column is added on the right and a 3-pixel column is dropped on the left), whereby the 9 pixel values of this newly positioned filter are provided to the synapses CB1, where they are multiplied by the same weights as above, and a second single output value is determined by the relevant synapses. This process is continued until the 3×3 filter has scanned over the entire 32×32 pixel image of the input layer S0 for all three colors and all bits (precision values). The process is then repeated using different sets of weights until all the feature maps of layer C1 are calculated, generating different feature maps of C1.
[0100] In this example, in layer C1, there are 16 feature maps each having 30×30 pixels. Each pixel is a new feature pixel extracted from the multiplication of the input and the kernel. Thus, each feature map is a two-dimensional array. Therefore, in this example, layer C1 consists of 16 layers of two-dimensional arrays (note that the layers and arrays referred to in this specification are logical relationships rather than necessarily physical relationships, that is, the arrays are not necessarily oriented in a physical two-dimensional array). Each of the 16 feature maps in layer C1 is generated by one of 16 different sets of synaptic weights applied to the filter scan. All of the C1 feature maps can target different aspects of the same image feature, such as boundary identification. For example, the first map (generated using a first set of weights shared by all scans used to generate this first map) can identify circular edges, and the second map (generated using a second set of weights different from the first set of weights) can identify square edges or the aspect ratio of a specific feature, etc.
[0101] Before going from layer C1 to layer S1, an activation function P1 (pooling) that pools values from non-overlapping, contiguous 2×2 regions within each feature map is applied. The purpose of the pooling function is to average neighboring positions (or it is also possible to use the max function), for example, to reduce the dependence on edge positions, and to reduce the data size before going to the next stage. In layer S1, there are 16 15×15 feature maps (i.e., 16 different arrays of 15×15 pixels each). The synapses CB2 going from layer S1 to layer C2 scan the maps in S1 with a 4×4 filter with a 1-pixel filter shift. In layer C2, there are 22 12×12 feature maps. Before going from layer C2 to layer S2, an activation function P2 (pooling) that pools values from non-overlapping, contiguous 2×2 regions within each feature map is applied. In layer S2, there are 22 6×6 feature maps. In the synapses CB3 going from layer S2 to layer C3, an activation function (pooling) is applied, where all neurons in layer C3 are connected to all maps in layer S2 via each synapse of CB3. In layer C3, there are 64 neurons. The synapses CB4 going from layer C3 to the output layer S3 fully connect C3 to S3, i.e., all neurons in layer C3 are connected to all neurons in layer S3. The output in S3 contains 10 neurons, and the neuron with the highest output determines the class. This output can indicate, for example, the identification or classification of the content of the original image.
[0102] Each layer of synapses is implemented using an array or a part of an array of non-volatile memory cells.
[0103] FIG. 9 is a block diagram of an array that can be used for that purpose. The vector matrix multiplication (VMM) system 32 includes non-volatile memory cells and is utilized as synapses (such as CB1, CB2, CB3, and CB4 in FIG. 6) between one layer and the next layer. Specifically, the VMM system 32 includes an array 33 of non-volatile memory cells arranged in rows and columns, an erase gate and word line gate decoder 34, a control gate decoder 35, a bit line decoder 36, and a source line decoder 37, and those decoders decode respective inputs to the non-volatile memory cell array 33. Inputs to the VMM array 33 can be made from the erase gate and word line gate decoder 34 or from the control gate decoder 35. The source line decoder 37 in this example also decodes the output of the VMM array 33. Alternatively, the bit line decoder 36 can decode the output of the VMM array 33.
[0104] The VMM array 33 serves two purposes. First, it stores the weights used by the VMM system 32. Second, the VMM array 33 effectively multiplies the inputs by the weights stored in the VMM array 33, adds them for each output line (source line or bit line), and generates an output, which becomes the input to the next layer or the input to the last layer. By performing the functions of multiplication and addition, the VMM array 33 eliminates the need for separate multiplication and addition logic circuits and is also power-efficient due to in-memory computing on the spot.
[0105] The output of the VMM array 33 is supplied to a differential adder (such as an addition op-amp or an addition current mirror) 38 that sums the outputs of the VMM array 33 to create a single value for convolution. The differential adder 38 is arranged to perform the sum of the positive and negative weights.
[0106] The total output value of the differential adder 38 is then supplied to an activation function circuit 39 that rectifies the output. The activation function circuit 39 may provide a sigmoid function, a tanh function, a ReLU function, or any other non-linear function. The rectified output value of the activation function circuit 39 becomes an element of the feature map of the next layer (e.g., C1 in FIG. 8) and is then applied to the next synapse to generate the next feature map layer or the last layer. Thus, in this example, the VMM array 33 constitutes a plurality of synapses (receiving inputs from the previous layer of neurons or from an input layer such as an image database), and the adder 38 and the activation function circuit 39 constitute a plurality of neurons.
[0107] The inputs (WLx, EGx, CGx, and optionally BLx and SLx) to the VMM system 32 of FIG. 9 can be at an analog level, a binary level, a digital pulse (in which case a pulse-analog converter PAC may be required to convert the pulse to an appropriate input analog level), or a digital bit (in which case a DAC is provided to convert the digital bit to an appropriate input analog level), and the output can be at an analog level, a binary level, a digital pulse, or a digital bit (in which case an output ADC is provided to convert the output analog level to digital bits).
[0108] FIG. 10 is a block diagram showing the use of multiple layers of the VMM system 32, labeled as VMM systems 32a, 32b, 32c, 32d, and 32e in the figure. As shown in FIG. 10, an input (denoted as Inputx) is converted from digital to analog by a digital - analog converter 31 and provided to the input VMM system 32a. The converted analog input can be a voltage or a current. The input D / A conversion of the first layer can be performed by using a function or a LUT (look - up table) that maps the input Inputx to an appropriate analog level of the matrix multiplier of the input VMM system 32a. The input conversion can also be performed by an analog - analog (A / A) converter to convert an external analog input to the mapped analog input to the input VMM system 32a. The input conversion can also be performed by a digital - digital pulse (D / P) converter to convert an external digital input to the mapped digital pulse(s) to the input VMM system 32a.
[0109] The output generated by the input VMM system 32a is provided as input to the next VMM system (hidden level 1) 32b, which then generates an output that is provided as input to the next VMM system (hidden level 2) 32c, and so on. The various layers of the VMM system 32 function as the various layers of the synapses and neurons of a convolutional neural network (CNN). The VMM systems 32a, 32b, 32c, 32d, and 32e can each be a stand-alone physical non-volatile memory array, or multiple VMM systems can utilize different portions of the same physical non-volatile memory array, or multiple VMM systems can utilize overlapping portions of the same physical non-volatile memory system. Each VMM system 32a, 32b, 32c, 32d, and 32e can also be time-multiplexed with respect to the various portions of its array or neurons. The example shown in FIG. 10 includes five layers (32a, 32b, 32c, 32d, 32e), namely, one input layer (32a), two hidden layers (32b, 32c), and two fully connected layers (32d, 32e). One of ordinary skill in the art will understand that this is merely exemplary, and that instead the system can include more than two hidden layers and more than two fully connected layers. Further, the different layers can use different combinations of n-bit memory cells that include two levels of memory cells (meaning only two levels of "0" and "1", where different cells support multiple different levels). VMM array
[0110] FIG. 11 shows a neuron VMM array 1100 that is particularly suitable for the memory cell 310 shown in FIG. 3 and that is utilized as part of the synapses and neurons between the input layer and the next layer. The VMM array 1100 includes a memory array 1101 of non-volatile memory cells and a reference array 1102 of non-volatile reference memory cells (located at the top of the array). Alternatively, another reference array can be located at the bottom.
[0111] In the VMM array 1100, control gate lines such as the control gate line 1103 extend in the vertical direction (therefore, the reference array 1102 in the row direction is orthogonal to the control gate line 1103), and erasure gate lines such as the erasure gate line 1104 extend in the horizontal direction. Here, the input to the VMM array 1100 is provided to the control gate lines (CG0, CG1, CG2, CG3), and the output of the VMM array 1100 appears on the source lines (SL0, SL1). In one embodiment, only even rows are used, and in another embodiment, only odd rows are used. The current applied to each source line (SL0 and SL1 respectively) performs a summation function of all the currents from the memory cells connected to that particular source line.
[0112] As described herein for neural networks, the non-volatile memory cells of the VMM array 1100, i.e., the flash memory of the VMM array 1100, are preferably configured to operate in the subthreshold region.
[0113] The non-volatile reference memory cells and non-volatile memory cells described herein are biased with weak inversion as follows: Ids = Io * e (Vg-Vth) / nVt = w * Io * e (Vg) / nVt where w = e (-Vth) / nVt where Vg is the gate voltage to the memory cell, Vth is the threshold voltage of the memory cell, Vt is the thermal voltage = k * T / q (where k is the Boltzmann constant, T is the temperature in Kelvin, and q is the electron charge), n is the slope factor = 1+(Cdep / Cox) (where Cdep = the capacitance of the depletion layer and Cox is the capacitance of the gate oxide layer), and Io is the memory cell current at a gate voltage equal to the threshold voltage. Io is (Wt / L) * u * Cox * (n - 1) * Vt 2proportional to, where u is the carrier mobility, and Wt and L are the width and length of the memory cell, respectively.
[0114] When using an I-V log converter that converts an input current into an input voltage using a memory cell (such as a reference memory cell or a peripheral memory cell) or a transistor: Vg = n * Vt * log[Ids / wp * Io] where wp is the w of the reference or peripheral memory cell.
[0115] For the memory array used as a vector matrix multiplier VMM array, the output current is as follows: Iout = wa * Io * e (Vg) / nVt , that is, Iout = (wa / wp) * Iin = W * Iin W = e (Vthp-Vtha) / nVt where wa = w for each memory cell of the memory array.
[0116] The word line or control gate can be used as the input of the memory cell for the input voltage.
[0117] Alternatively, the flash memory cells of the VMM array described in this specification can be configured to operate in the linear region. Ids = beta * (Vgs - Vth) * Vds, beta = u * Cox * Wt / L, where Wt and L are the respective width and length of the transistor. W = α(Vgs - Vth), that is, the weight W is proportional to (Vgs - Vth).
[0118] A word line, a control gate, a bit line, or a source line can be used as an input to a memory cell operating in the linear region. A bit line or a source line can be used as an output of the memory cell.
[0119] For an I-V linear converter, a memory cell (such as a reference memory cell or a peripheral memory cell) or a transistor or a resistor operating in the linear region can be used to linearly convert an input / output current into an input / output voltage.
[0120] Alternatively, the flash memory cells of the VMM array described herein can be configured to operate in the saturation region. Ids=1 / 2 * beta * (Vgs-Vth) 2 , beta = u * Cox * Wt / L W = α(Vgs-Vth) 2 , that is, the weight W is proportional to (Vgs-Vth) 2 .
[0121] A word line, a control gate, or an erase gate can be used as an input to a memory cell operating in the saturation region. A bit line or a source line can be used as an output of the output neuron.
[0122] Alternatively, the flash memory cells of the VMM array described herein can be used in all regions or combinations thereof (subthreshold, linear, or saturation).
[0123] Another embodiment for the VMM array 32 of FIG. 9 is described in U.S. Patent Application No. 15 / 826,345, which is incorporated herein by reference. As described in the above application, a source line or a bit line can be used as a neuron output (current sum output).
[0124] FIG. 12 shows a neuron VMM array 1200 that is particularly suitable for the memory cell 210 shown in FIG. 2 and is used as a synapse between the input layer and the next layer. The VMM array 1200 includes a memory array 1203 of non-volatile memory cells, a reference array 1201 of first non-volatile reference memory cells, and a reference array 1202 of second non-volatile reference memory cells. The reference arrays 1201 and 1202 arranged in the column direction of the array function to convert the current inputs flowing into the terminals BLR0, BLR1, BLR2, and BLR3 into voltage inputs WL0, WL1, WL2, and WL3. In practice, the first and second non-volatile reference memory cells are diode-connected through a multiplexer 1214 (only part shown) in a state where the current inputs flow in. The reference cells are adjusted (e.g., programmed) to a target reference level. The target reference level is provided by a reference mini-array matrix (not shown).
[0125] The memory array 1203 serves two purposes. First, it stores the weights used by the VMM array 1200 in each memory cell. Second, the memory array 1203 effectively multiplies the input (i.e., the current inputs provided to the terminals BLR0, BLR1, BLR2, and BLR3, which are converted by the reference arrays 1201 and 1202 into input voltages and supplied to the word lines WL0, WL1, WL2, and WL3) by the weights stored in the memory array 1203, and then adds up all the results (memory cell currents) to generate the outputs of the respective bit lines (BL0 to BLN), and this output becomes the input to the next layer or the input to the last layer. By the memory array 1203 performing the multiplication and addition functions, the need for separate multiplication and addition logic circuits is eliminated, and the power efficiency is also good. Here, the voltage inputs are provided to the word lines WL0, WL1, WL2, and WL3, and the outputs appear on the respective bit lines BL0 to BLN during the read (inference) operation. The currents arranged on each bit line BL0 to BLN perform the total function of the currents from all the non-volatile memory cells connected to that particular bit line.
[0126] Table 5 shows the operating voltages of the VMM array 1200. The columns in the table show the voltages applied to the word lines of the selected cells, the word lines of the non-selected cells, the bit lines of the selected cells, the bit lines of the non-selected cells, the source lines of the selected cells, and the source lines of the non-selected cells. FLT indicates floating, i.e., no voltage is applied. The rows indicate the operations of read, erase, and program. Table 5: Operations of the VMM Array 1200 in FIG. 12
Table 5
[0127] FIG. 13 shows a neuron VMM array 1300 that is particularly suitable for the memory cell 210 shown in FIG. 2 and is used as part of synapses and neurons between the input layer and the next layer. The VMM array 1300 includes a memory array 1303 of non-volatile memory cells, a reference array 1301 of first non-volatile reference memory cells, and a reference array 1302 of second non-volatile reference memory cells. The reference arrays 1301 and 1302 extend in the row direction of the VMM array 1300. The VMM array is similar to the VMM1100 except that the word lines extend vertically in the VMM array 1300. Here, the input is provided to the word lines (WLA0, WLB0, WLA1, WLB2, WLA2, WLB2, WLA3, WLB3), and the output appears on the source lines (SL0, SL1) during the read operation. The current applied to each source line performs a summation function of all the currents from the memory cells connected to that particular source line.
[0128] Table 6 shows the operating voltages of the VMM array 1300. The columns in the table show the voltages applied to the word lines of the selected cells, the word lines of the non-selected cells, the bit lines of the selected cells, the bit lines of the non-selected cells, the source lines of the selected cells, and the source lines of the non-selected cells. The rows indicate the operations of read, erase, and program. Table 6: Operations of the VMM Array 1300 in FIG. 13
Table 6
[0129] FIG. 14 shows a neuron VMM array 1400 that is particularly suitable for the memory cell 310 shown in FIG. 3 and is used as part of synapses and neurons between the input layer and the next layer. The VMM array 1400 includes a memory array 1403 of non-volatile memory cells, a reference array 1401 of first non-volatile reference memory cells, and a reference array 1402 of second non-volatile reference memory cells. The reference arrays 1401 and 1402 function to convert the current inputs flowing into the terminals BLR0, BLR1, BLR2, and BLR3 into voltage inputs CG0, CG1, CG2, and CG3. In practice, the first and second non-volatile reference memory cells are diode-connected through a multiplexer 1412 (only part shown) in a state where the current inputs flow through BLR0, BLR1, BLR2, and BLR3. The multiplexer 1412 includes respective multiplexers 1405 and cascode transistors 1404 to ensure a constant voltage for each bit line (such as BLR0) of the first and second non-volatile reference memory cells during a read operation. The reference cells are adjusted to a target reference level.
[0130] The memory array 1403 serves two purposes. First, it stores the weights used by the VMM array 1400. Second, the memory array 1403 effectively multiplies the input (the current inputs provided to the terminals BLR0, BLR1, BLR2, and BLR3, which the reference arrays 1401 and 1402 convert these current inputs into input voltages and supply to the control gates CG0, CG1, CG2, and CG3) by the weights stored in the memory cell array, and then adds up all the results (cell currents) to generate an output, which appears on BL0 to BLN and serves as an input to the next layer or an input to the last layer. By the memory array performing the functions of multiplication and addition, the need for separate multiplication and addition logic circuits is eliminated, and the power efficiency is also good. Here, the input is provided to the control gate lines (CG0, CG1, CG2, and CG3), and the output appears on the bit lines (BL0 to BLN) during a read operation. The current applied to each bit line performs the total function of all the currents from the memory cells connected to that particular bit line.
[0131] The VMM array 1400 performs a one-way adjustment of the non-volatile memory cells in the memory array 1403. That is, each non-volatile memory cell is erased and then partially programmed until the desired charge on the floating gate is reached. This can be carried out, for example, using the precise programming techniques described below. If too much charge is applied to the floating gate (such as an incorrect value being stored in the cell), the cell must be erased and a series of partial programming operations must be repeated. As shown, two rows sharing the same erase gate (such as EG0 or EG1) need to be erased together (known as page erase), and then each cell is partially programmed until the desired charge on the floating gate is reached.
[0132] Table 7 shows the operating voltages of the VMM array 1400. The columns in the table show the voltages applied to the word line of the selected cell, the word line of the non-selected cell, the bit line of the selected cell, the bit line of the non-selected cell, the control gate of the selected cell, the control gate of the non-selected cell in the same sector as the selected cell, the control gate of the non-selected cell in a different sector from the selected cell, the erase gate of the selected cell, the erase gate of the non-selected cell, the source line of the selected cell, and the source line of the non-selected cell. The rows show each operation of read, erase, and program. Table 7: Operations of the VMM Array 1400 in FIG. 14 [Table 7]
[0133] FIG. 15 shows a neuron VMM array 1500 that is particularly suitable for the memory cell 310 shown in FIG. 3 and is used as part of the synapses and neurons between the input layer and the next layer. The VMM array 1500 includes a memory array 1503 of non-volatile memory cells, a reference array 1501 or a first non-volatile reference memory cell, and a reference array 1502 of second non-volatile reference memory cells. The EG lines EGR0, EG0, EG1, and EGR1 extend vertically, and the CG lines CG0, CG1, CG2, and CG3 and the SL lines WL0, WL1, WL2, and WL3 extend horizontally. The VMM array 1500 is similar to the VMM array 1400 except that the VMM array 1500 implements bidirectional adjustment, and each individual cell can be completely erased, partially programmed, and partially erased as needed to reach the desired charge amount of the floating gate by using an individual EG line. As shown, the reference arrays 1501 and 1502 convert the input current in the terminals BLR0, BLR1, BLR2, and BLR3 into control gate voltages CG0, CG1, CG2, and CG3 (through the action of diode-connected reference cells via the multiplexer 1514), and these voltages are applied to the memory cells in the row direction. The current output (neuron) is in the bit lines BL0 to BLN, and each bit line sums all the currents from the non-volatile memory cells connected to that particular bit line.
[0134] Table 8 shows the operating voltages of the VMM array 1500. The columns in the table show the voltages applied to the word line of the selected cell, the word line of the non-selected cell, the bit line of the selected cell, the bit line of the non-selected cell, the control gate of the selected cell, the control gate of the non-selected cell in the same sector as the selected cell, the control gate of the non-selected cell in a different sector from the selected cell, the erase gate of the selected cell, the erase gate of the non-selected cell, the source line of the selected cell, and the source line of the non-selected cell. The rows show each operation of read, erase, and program. Table 8: Operation of the VMM Array 1500 in FIG. 15
Table 8
[0135] FIG. 24 shows a neuron VMM array 2400 that is particularly suitable for the memory cell 210 shown in FIG. 2 and is used as part of the synapses and neurons between the input layer and the next layer. In the VMM array 2400, the inputs INPUT0, ..., INPUT N are received by bit lines BL0, ..., BL N respectively, and the outputs OUTPUT1, OUTPUT2, OUTPUT3, and OUTPUT4 are generated on source lines SL0, SL1, SL2, and SL3 respectively.
[0136] FIG. 25 shows a neuron VMM array 2500 that is particularly suitable for the memory cell 210 shown in FIG. 2 and is used as part of the synapses and neurons between the input layer and the next layer. In this example, the inputs INPUT0, INPUT1, INPUT2, and INPUT3 are received by source lines SL0, SL1, SL2, and SL3 respectively, and the outputs OUTPUT0, ..., OUTPUT N are generated on bit lines BL0, ..., BL N respectively.
[0137] FIG. 26 shows a neuron VMM array 2600 that is particularly suitable for the memory cell 210 shown in FIG. 2 and is used as part of the synapses and neurons between the input layer and the next layer. In this example, the inputs INPUT0, ..., INPUT M are received by word lines WL0, ..., WL M respectively, and the outputs OUTPUT0, ..., OUTPUT N are generated on bit lines BL0, ..., BL N respectively.
[0138] FIG. 27 shows a neuron VMM array 2700 that is particularly suitable for the memory cell 310 shown in FIG. 3 and is used as part of the synapses and neurons between the input layer and the next layer. In this example, the inputs INPUT0, ..., INPUT M are received by word lines WL0, ..., WL M respectively, and the outputs OUTPUT0, ..., OUTPUT N are generated on bit lines BL0, ..., BL Nis generated.
[0139] FIG. 28 shows a neuron VMM array 2800 that is particularly suitable for the memory cell 410 shown in FIG. 4 and is used as part of synapses and neurons between an input layer and the next layer. In this example, inputs INPUT0, ..., INPUT n are respectively received by vertical control gate lines CG0, ..., CG N and outputs OUTPUT1 and OUTPUT2 are generated on source lines SL0 and SL1.
[0140] FIG. 29 shows a neuron VMM array 2900 that is particularly suitable for the memory cell 410 shown in FIG. 4 and is used as part of synapses and neurons between an input layer and the next layer. In this example, inputs INPUT0, ..., INPUT N are respectively received by gates of bit line control gates 2901-1, 2901-2, ..., 2901-(N-1), and 2901-N that are respectively coupled to bit lines BL0, ..., BL N Exemplary outputs OUTPUT1 and OUTPUT2 are generated on source lines SL0 and SL1.
[0141] FIG. 30 shows a neuron VMM array 3000 that is particularly suitable for the memory cell 310 shown in FIG. 3, the memory cell 510 shown in FIG. 5, and the memory cell 710 shown in FIG. 7, and is used as part of synapses and neurons between an input layer and the next layer. In this example, inputs INPUT0, ..., INPUT M are received by word lines WL0, ..., WL M and outputs OUTPUT0, ..., OUTPUT N are respectively generated on bit lines BL0, ..., BL N respectively.
[0142] FIG. 31 shows a neuron VMM array 3100 that is particularly suitable for the memory cell 310 shown in FIG. 3, the memory cell 510 shown in FIG. 5, and the memory cell 710 shown in FIG. 7, and is used as part of synapses and neurons between the input layer and the next layer. In this example, the inputs INPUT0, ..., INPUT M are received on the control gate lines CG0, ..., CG M Outputs OUTPUT0, ..., OUTPUT N are respectively generated on the vertical source lines SL0, ..., SL N and each source line SL i is coupled to the source line terminals of all the memory cells in column i.
[0143] FIG. 32 shows a neuron VMM array 3200 that is particularly suitable for the memory cell 310 shown in FIG. 3, the memory cell 510 shown in FIG. 5, and the memory cell 710 shown in FIG. 7, and is used as part of synapses and neurons between the input layer and the next layer. In this example, the inputs INPUT0, ..., INPUT M are received on the control gate lines CG0, ..., CG M Outputs OUTPUT0, ..., OUTPUT N are respectively generated on the vertical bit lines BL0, ..., BL N and each bit line BL i is coupled to the bit line terminals of all the memory cells in column i. Long short-term memory
[0144] The prior art includes the concept known as long short-term memory (LSTM). LSTM units are often used within neural networks. With LSTM, a neural network can store information over an arbitrary period of time and use that information in subsequent operations. Conventional LSTM units include a cell, an input gate, an output gate, and a forget gate. The three gates regulate the flow of information into and out of the cell and the period during which information is stored within the LSTM. VMM is particularly useful in LSTM units.
[0145] Figure 16 shows an exemplary LSTM 1600. The LSTM 1600 in this example includes cells 1601, 1602, 1603, and 1604. Cell 1601 receives an input vector x0 and generates an output vector h0 and a cell state vector c0. Cell 1602 receives an input vector x1, an output vector (hidden state) h0 from cell 1601, and a cell state c0 from cell 1601, and generates an output vector h1 and a cell state vector c1. Cell 1603 receives an input vector x2, an output vector (hidden state) h1 from cell 1602, and a cell state c1 from cell 1602, and generates an output vector h2 and a cell state vector c2. Cell 1604 receives an input vector x3, an output vector (hidden state) h2 from cell 1603, and a cell state c2 from cell 1603, and generates an output vector h3. Additional cells can also be used, and the LSTM with four cells is merely an example.
[0146] Figure 17 shows an exemplary implementation of an LSTM cell 1700 that can be used for cells 1601, 1602, 1603, and 1604 in Figure 16. The LSTM cell 1700 receives an input vector x(t), a cell state vector c(t - 1) from a preceding cell, and an output vector h(t - 1) from a preceding cell, and generates a cell state vector c(t) and an output vector h(t).
[0147] The LSTM cell 1700 includes sigmoid function devices 1701, 1702, and 1703, each of which controls the degree to which each component of the input vector contributes to the output vector by applying a number between 0 and 1. The LSTM cell 1700 also includes tanh devices 1704 and 1705 for applying a hyperbolic tangent function to the input vector, multiplier devices 1706, 1707, and 1708 for multiplying two vectors, and an adder device 1709 for adding two vectors. The output vector h(t) can be provided to the next LSTM cell in the system or accessed for other purposes.
[0148] FIG. 18 shows an LSTM cell 1800 which is an implementation example of the LSTM cell 1700. For the convenience of the reader, the same numbering method from the LSTM cell 1700 is used in the LSTM cell 1800. The sigmoid function devices 1701, 1702, and 1703, and the tanh device 1704 each include a plurality of VMM arrays 1801 and activation circuit blocks 1802. Therefore, it can be understood that the VMM array is particularly useful in the LSTM cell used in a specific neural network system. The multiplier devices 1706, 1707, and 1708, and the adder device 1709 are implemented in a digital or analog manner. The activation function block 1802 can be implemented in a digital or analog manner.
[0149] An alternative example of the LSTM cell 1800 (and another example of the implementation of the LSTM cell 1700) is shown in FIG. 19. In FIG. 19, the sigmoid function devices 1701, 1702 and 1703, and the tanh device 1704 share the same physical hardware (VMM array 1901 and activation function block 1902) in a time-division multiplexed manner. The LSTM cell 1900 also includes a multiplier device 1903 for multiplying two vectors, an adder device 1908 for adding two vectors, a tanh device 1705 (including the activation circuit block 1902), a register 1907 for storing the value i(t) output from the sigmoid function block 1902, the value f(t) output from the multiplier device 1903 via the multiplexer 1910 * a register 1904 for storing c(t−1), the value i(t) output from the multiplier device 1903 via the multiplexer 1910 * a register 1905 for storing u(t), the value o(t) output from the multiplier device 1903 via the multiplexer 1910 * a register 1906 for storing c~(t), and a multiplexer 1909.
[0150] While the LSTM cell 1800 includes a plurality of sets of VMM arrays 1801 and respective activation function blocks 1802, the LSTM cell 1900 includes only one set of a VMM array 1901 and an activation function block 1902, which is used to represent a plurality of layers in an embodiment of the LSTM cell 1900. The LSTM cell 1900 requires less space than the LSTM 1800 because it requires only 1 / 4 of the space required for the VMM and the activation function block as compared to the LSTM cell 1800.
[0151] It can be further understood that LSTM units typically include a plurality of VMM arrays, each of which requires the functions provided by specific circuit blocks outside the VMM array, such as adders, activation circuit blocks, and high-voltage generation blocks. Providing separate circuit blocks for each VMM array would require a significant amount of space within the semiconductor device and would be somewhat inefficient. Gated Recurrent Unit
[0152] The analog VMM implementation can be utilized in a gated recurrent unit (GRU) system. A GRU is a gate mechanism within a recurrent neural network. A GRU is similar to an LSTM, except that a GRU cell generally includes fewer components than an LSTM cell.
[0153] Figure 20 shows an exemplary GRU2000. The GRU2000 in this example includes cells 2001, 2002, 2003, and 2004. Cell 2001 receives the input vector x0 and generates the output vector h0. Cell 2002 receives the input vector x1 and the output vector h0 from cell 2001 and generates the output vector h1. Cell 2003 receives the input vector x2 and the output vector (hidden state) h1 from cell 2002 and generates the output vector h2. Cell 2004 receives the input vector x3 and the output vector (hidden state) h2 from cell 2003 and generates the output vector h3. Additional cells can also be used, and the GRU with four cells is merely an example.
[0154] Figure 21 shows an exemplary implementation of a GRU cell 2100 that can be used for cells 2001, 2002, 2003, and 2004 in Figure 20. The GRU cell 2100 receives the input vector x(t) and the output vector h(t - 1) from the preceding GRU cell and generates the output vector h(t). The GRU cell 2100 includes sigmoid function devices 2101 and 2102, each of which applies a number between 0 and 1 to components from the output vector h(t - 1) and the input vector x(t). The GRU cell 2100 also includes a tanh device 2103 for applying the hyperbolic tangent function to the input vector, a plurality of multiplier devices 2104, 2105, and 2106 for multiplying two vectors, an adder device 2107 for adding two vectors, and a complementary device 2108 for subtracting the input from 1 to generate the output.
[0155] FIG. 22 shows a GRU cell 2200 which is an implementation example of the GRU cell 2100. For the convenience of the reader, the same numbering method from the GRU cell 2100 is used in the GRU cell 2200. As can be seen from FIG. 22, the sigmoid function devices 2101 and 2102, and the tanh device 2103 each include a plurality of VMM arrays 2201 and activation function blocks 2202. Therefore, it can be understood that the VMM array is particularly used in the GRU cell used in a specific neural network system. The multiplier devices 2104, 2105, 2106, the adder device 2107, and the complementary device 2108 are implemented in a digital or analog manner. The activation function block 2202 can be implemented in a digital or analog manner.
[0156] An alternative example of the GRU cell 2200 (and another example of the implementation of the GRU cell 2300) is shown in FIG. 23. In FIG. 23, the GRU cell 2300 uses a VMM array 2301 and an activation function block 2302, and when configured as a sigmoid function, by applying a number from 0 to 1, it controls the degree to which each component of the input vector contributes to the output vector. In FIG. 23, the sigmoid function devices 2101 and 2102, and the tanh device 2103 share the same physical hardware (VMM array 2301 and activation function block 2302) in a time-division multiplexed manner. The GRU cell 2300 also includes a multiplier device 2303 for multiplying two vectors, an adder device 2305 for adding two vectors, a complementary device 2309 for subtracting the input from 1 to generate an output, a multiplexer 2304, and a value h(t - 1) output from the multiplier device 2303 via the multiplexer 2304 * A register 2306 for holding r(t), and a value h(t - 1) output from the multiplier device 2303 via the multiplexer 2304 * A register 2307 for holding z(t), and a value ĥ(t) output from the multiplier device 2303 via the multiplexer 2304 * (1 - z((t)) and a register 2308 for holding it.
[0157] The GRU cell 2200 includes a plurality of sets of a VMM array 2201 and an activation function block 2202, while the GRU cell 2300 includes only one set of a VMM array 2301 and an activation function block 2302, which is used to represent multiple layers in an embodiment of the GRU cell 2300. The GRU cell 2300 requires less space than the GRU cell 2200 because, compared with the GRU cell 2200, the space required for the VMM and the activation function block is only 1 / 3.
[0158] It can be further understood that a GRU system typically includes a plurality of VMM arrays, each of which requires the functions provided by specific circuit blocks outside the VMM array, such as adders, activation circuit blocks, and high-voltage generation blocks. Providing separate circuit blocks for each VMM array would require a significant amount of space within the semiconductor device and would be somewhat inefficient.
[0159] The input to the VMM array can be at an analog level, a binary level, or a digital bit (in which case a DAC is required to convert the digital bit to an appropriate input analog level), and the output can be at an analog level, a binary level, or a digital bit (in which case an output ADC is required to convert the output analog level to a digital bit).
[0160] For each memory cell in the VMM array, each weight W can be implemented by a single memory cell, or by differential cells, or by two blended memory cells (the average of two cells). In the case of differential cells, two memory cells are required to implement the weight W as a differential weight (W = W+ - W-). In the case of two blended memory cells, two memory cells are required to implement the weight W as the average of two cells. Configurable input / output system for VMM array
[0161] FIG. 33 shows a VMM system 3300. The VMM system 3300 can be based on any of the foregoing VMM array designs, such as VMM array 3301 (VMM arrays 1100, 1200, 1300, 1400, 1500, 2400, 2500, 2600, 2700, 2800, 2900, 3000, 3100, and 3200 or other VMM array designs), a low voltage row decoder 3302, a high voltage row decoder 3303, a column decoder 3304, a column driver 3305, control logic 3306, a bias circuit 3307, a neuron output circuit block 3308, an input VMM circuit block 3309, an algorithm controller 3310, a high voltage generator block 3311, an analog circuit block 3315, and control logic 3316.
[0162] The input circuit block 3309 functions as an interface from an external input to the input terminals of the memory array 3301. The input circuit block 3309 can include, without limitation, a DAC (digital - analog converter), a DPC (digital - pulse converter), an APC (analog - pulse converter), an IVC (current - voltage converter), an AAC (analog - analog converter such as a voltage - voltage scaler), or an FAC (frequency - analog converter). The neuron output block 3308 functions as an interface from the memory array output to an external interface (not shown). The neuron output block 3308 can include, without limitation, an ADC (analog - digital converter), an APC (analog - pulse converter), a DPC (digital - pulse converter), an IVC (current - voltage converter), or an IFC (current - frequency converter). The neuron output block 3308 may include, without limitation, an activation function, a normalization circuit, and / or a rescaling circuit.
[0163] FIG. 34 shows a VMM system 3400, which system includes VMM arrays 3401, 3402, 3403, and 3404, high voltage row decoders 3405 and 3406, low voltage row decoders 3407 and 3408, input blocks 3409 and 3410 (each similar to input block 3309 of FIG. 33), and output blocks 3411 and 3412. In this configuration, VMM arrays 3401 and 3403 share a set of bit lines and output block 3411, and VMM arrays 3402 and 3404 share a set of bit lines and output block 3412. VMM arrays 3401 and 3403 can be read simultaneously (thereby effectively combined into a single larger array), or can be read at different times. Output blocks 3411 and 3412 (similar to output block 3308 of FIG. 33) are configurable to process a read operation from one array at a time (such as a read from only array 3401 or 3403) or a read operation from multiple arrays at a time (such as a read from both arrays 3401 and 3403).
[0164] FIG. 35A shows a VMM system 3500, which system includes VMM arrays 3503, 3504, and 3505, a shared global high voltage row decoder 3506, local high voltage row decoders 3507 and 3508, a shared low voltage row decoder 3509, and an input block 3510. In this configuration, VMM arrays 3503, 3504, and 3505 share input block 3510. VMM arrays 3503, 3504, and 3505 can receive inputs (e.g., voltages or pulses on word lines, control gate lines, erase gate lines, or source lines) simultaneously via input block 3510 (thereby effectively combined into a single larger VMM array), or can receive inputs at different times via input block 3510 (thereby effectively operating as three separate VMM arrays having the same input block). Input block 3510 is configurable to provide inputs to one array at a time, or to multiple arrays at a time.
[0165] FIG. 35B shows a VMM system 3550, which system includes VMM arrays 3511, 3512, 3513, and 3514, a global high voltage decoder 3515, local high voltage row decoders 3516, 3517, and 3518, a shared low voltage row decoder 3519, and an input block 3520. In this configuration, the VMM arrays 3511, 3512, 3513, and 3514 share the input block 3520. The VMM arrays 3511, 3512, 3513, and 3514 can simultaneously receive inputs (e.g., voltages or pulses on word lines, control gate lines, erase gate lines, or source lines) via the input block 3520 (which are thereby effectively combined into a single larger array), or can receive inputs via the input block 3520 at different times (which thereby effectively operates as three separate VMM arrays having the same input block 3520). The input block 3520 is configurable to provide inputs to one array at a time, or to multiple arrays at a time. For example, the input block 3510 of FIG. 35A is configured to provide inputs to three arrays, and the input block 3520 is configured to provide inputs to four arrays.
[0166] FIG. 36 shows a VMM system 3600, which system includes a horizontal set 3601 and a horizontal set 3611. The horizontal set 3601 includes VMM arrays 3602 and 3603, a shared global high voltage row decoder 3604, a local high voltage row decoder 3605, a shared low voltage row decoder 3606, and an input block 3607. The VMM arrays 3602 and 3603 share the input block 3607. The input block 3607 is configurable to provide inputs to one array or to multiple arrays at a time.
[0167] The horizontal set 3611 includes VMM arrays 3612 and 3613, a shared global high-voltage decoder 3614, a local high-voltage row decoder 3615, a shared low-voltage row decoder 3616, and an input block 3617. The VMM arrays 3612 and 3613 share the input block 3617. The input block 3617 is configurable to provide inputs to one array at a time or to multiple arrays at a time.
[0168] In a first configuration, the horizontal set 3601 utilizes output blocks 3608 and 3609, and the horizontal set 3611 utilizes output blocks 3618 and 3619. The output blocks 3608, 3609, 3618, and 3619 can output current, digital pulses, or digital bits as outputs. In one embodiment where digital bits are output, the output blocks 3608, 3609, 3618, and 3619 each output eight digital output bits.
[0169] In a second configuration, the output blocks 3608 and 3609 are disabled, the VMM arrays 3602 and 3612 share the output block 3618, and the VMM arrays 3603 and 3613 share the output block 3619. The VMM arrays 3602 and 3612 can be read simultaneously, thereby effectively combined into a single larger vertical array (i.e., the number of rows per bit line increases), or they can be read at different times. When the VMM arrays 3602 and 3612 are read simultaneously, in one embodiment where each output block outputs a value in the range of 8 bits when coupled to only one array, the output blocks 3608 and 3609 each output a value in the range of 9 bits. This is due to the increased dynamic range of the output neurons multiplied by using two arrays as a single large array. In this case, if the next array only requires an 8-bit dynamic range, the output may need to be rescaled or normalized (e.g., scaled down from 9 bits to 8 bits). In another embodiment, the number of output bits can be maintained the same when increasing the number of vertical arrays.
[0170] Similarly, VMM arrays 3603 and 3613 can be read simultaneously (thereby effectively combined into a single larger array), or can be read at different times. Output blocks 3618 and 3619 are configurable to handle read operations from one array at a time, or from multiple arrays at a time.
[0171] In VMM systems 3400, 3500, 3550, and 3600, if the system is configurable to utilize a different number of arrays with each input block and / or output block, then the input block or output block itself must also be configurable. For example, in VMM system 3600, if output blocks 3608, 3609, 3612, and 3619 each output an 8-bit output when coupled to a single array, then output blocks 3618 and 3619 must each be configured to output a 9-bit output when coupled to two arrays (e.g., arrays 3602 and 3612, and arrays 3603 and 3609, respectively). Then, if those outputs are provided to an input block of another VMM system and the input block expects an 8-bit input rather than a 9-bit input, the output must first be normalized. A number of analog and digital techniques are known for converting an N-bit value to an M-bit value. In the foregoing example, N is 9 and M is 8, but one of ordinary skill in the art will understand that N and M can be any positive integers.
[0172] In VMM systems 3400, 3500, 3550, and 3600, additional arrays can be coupled to the input and output blocks. For example, in VMM system 3400, three or more arrays can be coupled to input block 3409, three or more arrays can be coupled to input block 3410, in VMM system 3500, four or more arrays can be coupled to input block 3510, in VMM system 3550, five or more arrays can be coupled to input block 3520, in VMM system 3600, three or more arrays can be coupled to input block 3607, three or more arrays can be coupled to input block 3617, three or more arrays can be coupled to output block 3618, and three or more arrays can be coupled to output block 3619. In those situations, the relevant input and output blocks need to be further configured to accommodate the additional arrays.
[0173] Output blocks 3411 and 3412 of VMM system 3400, and output blocks 3618 and 3619 need to be configurable for verification operations after programming operations, and the verification operations are affected by the number of arrays connected to the output blocks. Further, in program / erase verification (which is used to adjust and means generating a specific charge in the floating gate of the memory to generate a desired cell current), the accuracy of the output block circuit (e.g., 10 bits) needs to be greater than the accuracy required for inferential reading (e.g., 8 bits). For example, the verification accuracy is at least 1 bit greater than the inferential accuracy, e.g., 1 - 5 bits greater. This is necessary to ensure sufficient margin between one level and the next, for reasons such as, without limitation, verification result distribution, data retention drift, temperature, or variations.
[0174] In addition, the input blocks 3409, 3410, 3510, 3520, 3607, and 3617 and the output blocks 3411, 3412, 3608, 3609, 3618, and 3619 in FIGS. 34, 35A, 35B, and 36 need to be configurable for the calibration process because the number of arrays connected to the output blocks affects the calibration. Examples of calibration processes include processes for compensating for changes due to offset, leakage, manufacturing processes, and temperature changes.
[0175] In the following section, various adjustable components used in the input and output blocks are disclosed to enable the input and output blocks to be configured based on the number of arrays coupled to the input or output block. Components of Input and Output Blocks
[0176] FIG. 37A shows an integrated dual hybrid ramp analog-to-digital converter (ADC) 3700, which can be used in output blocks such as output blocks 3411, 3412, 3608, 3609, 3618, and 3619 in FIGS. 34 and 36, and output neurons, I NEU 3706 is the output current from the VMM array received by the output block. The integrated dual hybrid ramp analog-to-digital converter (ADC) 3700 is I NEU converts 3706 into a series of digital / analog pulses or digital output bits. FIG. 37B shows the operating waveforms of the integrated ADC 3700 of FIG. 37A. Output waveforms 3710, 3711, and 3714 are for one current level. Output waveforms 3712, 3713, and 3715 are for another, higher current level. Waveforms 3710 and 3712 have a pulse width proportional to the value of the output current. Waveforms 3711 and 3713 have a number of pulses proportional to the value of the output current. Waveforms 3714 and 3715 have digital output bits proportional to the value of the output current.
[0177] In one embodiment, the ADC 3700 is I NEU3706 (the analog output current received by the output block from the VMM array) is converted, as shown in the example illustrated in FIG. 38, into a digital pulse whose width varies in proportion to the magnitude of the analog output current in the neuron output block. The ADC 3700 includes an integrator composed of an integrating operational amplifier 3701 and an adjustable integrating capacitor 3702 that integrates I NEU 3706. Optionally, IREF 3707 can include a bandgap filter having a temperature coefficient of 0 or a temperature coefficient that tracks the neuron current I NEU 3706. The latter can be obtained from a reference array that includes values determined during the test phase, if necessary. During the initialization phase, the switch 3708 is closed. Then, the input to Vout 3703 and the negative terminal of the operational amplifier 3701 becomes equal to the VREF value. Thereafter, the switch 3708 is opened, and for a certain time tref, the switch S1 is closed and the neuron current I NEU 3706 is integrated upward. During the certain time tref, Vout rises and its slope changes as the neuron current changes. Thereafter, during the period tmeas, by opening the switch S1 and closing the switch S2, a constant reference current IREF is integrated downward over the time tmeas (during this period, Vout drops), and tmeas is the time required to integrate Vout downward to VREF.
[0178] The output EC 3705 goes high when VOUT > VREFV and low otherwise. Thus, EC 3705 generates a pulse whose width reflects the period tmeas, and as a result, this width is proportional to the current I NEU 3706 (pulses 3710 and 3712 in FIG. 37B).
[0179] Optionally, the output pulse EC 3705 can be converted into a series of pulses of a uniform period for transmission to the next stage of circuitry, such as the input block of another VMM array. At the start of period tmeas, the output EC 3705 is input to an AND gate 3740 together with a reference clock 3741. The output becomes a pulse train 3742 (the frequency of the pulses in the pulse train 3742 is the same as the frequency of the clock 3741) during the period when VOUT > VREF. The number of pulses is proportional to the period tmeas, and the period tmeas is proportional to the current I NEU 3706 (waveforms 3711 and 3713 in FIG. 37B).
[0180] Optionally, the pulse train 3743 can be input to a counter 3720, which counts the number of pulses in the pulse train 3742 and generates a count value 3721 that is a digital count of the number of pulses in the pulse train 3742 that is directly proportional to the neuron current I NEU 3706. The count value 3721 includes a set of digital bits (waveforms 3714 and 3715 in FIG. 37B).
[0181] In another embodiment, the integrating dual-slope ADC 3700 can convert the neuron current I NEU 3706 into pulses, and the width of the pulses is inversely proportional to the magnitude of the neuron current I NEU 3706. This inversion can be performed in a digital or analog manner and can be converted into a series of pulses or digital bits for output to subsequent circuitry.
[0182] The adjustable integrating capacitor 3702 and the adjustable reference current IREF 3707 are adjusted according to the number N of arrays connected to the integrating dual-slope analog-to-digital converter (ADC) 3700. For example, when N arrays are connected to the integrating dual-slope analog-to-digital converter (ADC) 3700, the adjustable integrating capacitor 3702 is adjusted by 1 / N, or the adjustable reference current IREF 3707 is adjusted by N.
[0183] Optionally, calibration steps can be performed while the VMM array and the ADC 3700 are at or above their operating temperature to offset any leakage current present in the VMM array or the control circuit, and then that offset value can be subtracted from Ineu in FIG. 37A. The calibration steps can also be performed to compensate for process variations or voltage supply variations in addition to temperature variations.
[0184] The method of operating the output circuit block includes first performing calibration for offset and voltage supply variation compensation. Next, output conversion is performed (such as converting the neuron current to a pulse or digital bits), and then data normalization is performed to match the output range to the input range of the next VMM array. Data normalization may include data compression or output data quantization (such as reducing the number of bits, for example, from 10 bits to 8 bits). Activation may be performed after output conversion or after data normalization, compression, or quantization. Examples of calibration algorithms are described below with reference to FIGS. 49, 50A, 50B, and 51.
[0185] FIG. 39 shows a current-voltage converter 3900 that can optionally be used to convert the neuron output current to a voltage, which can be applied as an input (e.g., to a WL line or a CG line) of a VMM memory array, for example. Thus, the current-voltage converter 3900 can be used in the input blocks 3409, 3410, 3510, 3520, 3607, and 3617 of FIGS. 34, 35A, 35B, and 36 when those blocks receive an analog current as an input (as opposed to a pulse or digital data).
[0186] The current-voltage converter 3900 includes an operational amplifier 3901, an adjustable capacitor 3902, a switch 3903, a switch 3904, and a current source 3905 that represents the neuron current INEU received by the input block here. During current-voltage operation, the switch 3903 is opened and the switch 3904 is closed. The output Vout increases in amplitude in proportion to the magnitude of the neuron current INEU 3905.
[0187] Figure 40 shows a digital data - voltage converter 4000 that can optionally be used to convert digital data received as signal DIN into a voltage, which can be applied, for example, as an input to a VMM memory array (e.g., of WL lines or CG lines). When switch 4002 is closed, the data input of signal DIN allows the IREF_u reference current 4001 to enter capacitor 4003 and generate a voltage at its terminals. Thus, the digital data - voltage converter 4000 can be used in input blocks 3409, 3410, 3510, 3520, 3607, and 3617 of FIGS. 34, 35A, 35B, and 36 when those blocks receive digital data as an input (as opposed to a pulse or analog current). Additionally, the digital data - voltage converter 4000 can be configured such that the digital data received as an input as signal DIN flows directly to output OUT by opening switches 4002 and 4004 and closing switch 4005. Thus, switches 4002, 4004, and 4005 are configured to allow output OUT to receive either the voltage of capacitor 4003 or the digital data received as signal DIN directly. In the illustrated embodiment, signal DIN is received as a data pulse.
[0188] The digital data-voltage pulse converter 4000 includes an adjustable reference current 4001, a switch 4002, a variable capacitor 4003, a switch 4004, and a switch 4005. The adjustable reference current 4001 and the variable capacitor 4003 can be configured to have different values to adjust for differences in the size of the array to which the digital data-voltage pulse converter 400 is attached. During operation, digital data controls the switch 4002 such that the switch 4002 closes whenever the digital data is high. When the switch closes, the adjustable reference current 4001 charges the variable capacitor 4003. The switch 4004 can be closed whenever it is desired to provide an output to the node OUT, such as when the array is ready to be read. Alternatively, the switch 4004 can be opened and the switch 4005 can be closed to pass the data input as an output.
[0189] FIG. 41 shows a configurable analog-to-digital converter 4100 that can optionally be used to convert an analog neuron current to digital data. The configurable analog-to-digital converter 4100 can be used in output blocks such as output blocks 3411, 3412, 3608, 3609, 3618, and 3619 in FIGS. 34 and 36, and the output neuron, I NEU 4101 is the output current received by the output block.
[0190] The configurable analog-to-digital converter 4100 includes a current source 4101, a variable resistor 4102, and an analog-to-digital converter 4103. The current INEU 4101 drops across the variable resistor 4102 Rneu to produce a voltage Vneu = Ineu * Rneu. The ADC 4103 (without limitation, an integrating ADC, a SAR ADC, a flash ADC, or a sigma-delta type ADC, etc.) converts this voltage to digital bits.
[0191] FIG. 42 shows a configurable current-to-voltage converter 4200 that can optionally be used to convert an analog neuron current to a voltage, which can be applied as an input (e.g., to a WL line or a CG line) to the VMM memory array. Thus, the configurable current-to-voltage converter 4200 can be used in input blocks 3409, 3410, 3510, 3520, 3607, and 3617 of FIGS. 34, 35A, 35B, and 36 when those blocks receive an analog current as an input (as opposed to a pulse or digital data). The configurable current-to-voltage converter 4200 includes an adjustable resistor Rin 4202 that receives an input current Iin 4201 (the received input current described above) and generates Vin 4203 (=Iin * Rin).
[0192] FIGS. 43A and 43B show a digital bit-pulse width converter 4300 used within an input block, a row decoder, or an output block. The pulse width output from the digital bit-pulse width converter 4300 is proportional to the value of the digital bit.
[0193] The digital bit-pulse width converter includes a binary counter 4301. The state Q[N:0] of the binary counter 4301 can be loaded by serial data or parallel data within a load sequence. The row control logic 4310 outputs a voltage pulse WLEN having a pulse width proportional to the value of a digital data input provided from a block such as the integrating ADC of FIG. 37.
[0194] FIG. 43B shows the waveform of the output pulse width, which is proportional to the digital bit value. First, the data within the received digital bit is inverted, and the inverted digital bit is loaded into the counter 4301 either serially or in parallel. Then, the row pulse width is generated by the row control logic 4310 as shown in waveform 4320 by counting in binary fashion until the maximum counter value is reached.
[0195] An example using a 4-bit value in the DIN is shown in Table 9. Table 9: Digital input bits to the output pulse width [Table 9]
[0196] Optionally, a pulse train - pulse converter can be used to convert an output including a pulse train into a single pulse whose width varies in proportion to the number of pulses in the pulse train, and can be used as an input to the VMM array applied to the word lines or control gates in the VMM array. An example of a pulse train - pulse converter is a binary counter having control logic.
[0197] Another embodiment utilizes an up binary counter and digital comparison logic. That is, the output pulse width is generated by counting using an up binary counter until the digital output of the binary counter becomes the same as the digital input bits.
[0198] Another embodiment utilizes a down binary counter. First, the down binary counter is loaded serially or in parallel with a digital data input pattern. Next, the output pulse width is generated by counting down the down binary counter until the digital output of the binary counter reaches the minimum value, i.e., the "0" logic state.
[0199] FIG. 44A shows a digital data - pulse line converter 4400 including a binary - indexed pulse stage 4401 - i, where i ranges from 0 to N (i.e., the least significant bit LSB to the most significant bit MSB). The line converter 4400 is used to provide a line input to the array. Each stage 4401 - i includes a latch 4402 - i, a switch 4403 - i, and a line digital binary - indexed pulse input 4404 - i (RDIN_Ti). For example, the binary - indexed pulse input 4404 - 0 (RDIN_T0) is one time unit, i.e., 1 *It has a pulse width equal to tpls1unit. The binary-indexed pulse input 4404-1 (RDIN_T1) has two time units, i.e., 2 * It has a pulse width equal to tpls1unit. The binary-indexed pulse input 4404-2 (RDIN_T2) has four time units, i.e., 4 * It has a pulse width equal to tpls1unit. The binary-indexed pulse input 4403-3 (RDIN_T3) has eight time units, i.e., 8 * It has a pulse width equal to tpls1unit. The digital data of the pattern DINi (from the neuron output) in each row is stored in the latch 4402-i. When the output Qi of the latch 4402-i is "1", the binary-indexed pulse input 4404-i (RDIN_Ti) is transferred to the time addition converter node 4408 via the switch 4403-i. Each time addition converter node 4408 is connected to each input of the NAND gate 4404, and the output of the NAND gate 4404 generates the output WLIN / CGIN 4409 of the row converter via the level-shifting inverter 4405. The time addition converter node 4408 adds the binary-indexed pulse inputs 4404-i sequentially in time according to the common clocking signal CLK. This is because the binary-indexed pulse input 4404-i (RDIN_Ti) is activated sequentially one digital bit at a time, for example, from LSB to MSB, or from MSB to LSB, or in any random bit pattern.
[0200] FIG. 44B shows an exemplary waveform 4420. Here, exemplary signals for the row digital binary-indexed pulse inputs 4404-i, specifically 4404-0, 4404-1, 4404-2, and 4404-3, and exemplary outputs from the level-shifting inverter 4405 labeled as WL0 and WL3 are shown, where WL0 and WL3 are generated from the circuitry of the row converter 4400. In this example, WL0 is generated by the row digital inputs 4403-0 and 4403-3 of its row decoder being asserted (WL0: Q0 = "1", Q3 = "1"), and WL3 is generated by the row digital inputs 4403-1 and 4403-2 of its row decoder being asserted (WL3: Q1 = "1", Q2 = "1"). If none of the row digital inputs 4403-x are asserted, there is no pulse on WL0 or WL3 (the control logic for this case is not shown in FIG. 44A). The inputs from the other rows of the digital-pulse row converter 4400, i.e., the other inputs to the NAND gates 4404, are assumed to be high during this period.
[0201] FIG. 44C shows a row digital pulse generator 4410 that generates the row digital binary-indexed pulse inputs 4403-i (RDIN_Ti), and the width of the pulse is proportional to the binary value of the digital bit, as described above in connection with FIG. 44A.
[0202] Figure 45A shows a ramp-type analog-to-digital converter 4400, which includes a current source 4401 (representing the received neuron current Ineu), a switch 4402, a variable configurable capacitor 4403, and a comparator 4404. This comparator receives, as the non-inverting input, the voltage generated across the variable configurable capacitor 4403, denoted as Vneu, and, as the inverting input, a configurable reference voltage Vreframp, and generates an output Cout. Vreframp is raised in discrete levels for each comparison clock cycle. The comparator 4404 compares Vneu with Vreframp, and as a result, the output Cout becomes "1" when Vneu > Vreframp and "0" otherwise. Thus, the output Cout becomes a pulse, and its width varies according to Ineu. A larger Ineu keeps Cout at "1" for a longer period, resulting in an expanded width of the pulse of the output Cout. The digital counter 4420 converts each pulse of the output Cout into a digital output bit, as shown in Figure 45B for two different Ineu currents, denoted as OT1A and OT2A respectively. Alternatively, the ramp voltage Vreframp is a continuous ramp voltage 4455 as shown in the graph 4450 of Figure 45B. A multi-ramp embodiment is shown in Figure 45C for shortening the conversion time by using a coarse-fine ramp conversion algorithm. First, a coarse reference ramp reference voltage 4471 is raised rapidly to find the sub-range of each Ineu. Next, fine reference ramp reference voltages 4472, namely Vreframp1 and Vreframp2, are each used for each sub-range to convert the Ineu current within the corresponding sub-range. As shown, there are two sub-ranges for the fine reference ramp voltage. Three or more coarse / fine steps or two sub-ranges are possible.
[0203] FIG. 52 shows a comparator 5200 that can be optionally used in place of comparators 3704 and 4404 of FIGS. 37A and 45A. The comparator 5200 can be a static comparator (not necessarily using a clock signal) or a dynamic comparator (using a comparison clock signal). When the comparator 5200 is a dynamic comparator, it can include a clocked cross-coupled inverter comparator, a StrongARM comparator, or other known dynamic comparators. The comparator 5200 operates as a coarse comparator when the coarse enable 5203 is asserted, and the comparator 5200 operates as a fine comparator when the fine enable 5204 is asserted. The select signal 5206 can be optionally used to indicate the coarse comparator mode or the fine enable mode, or can be optionally used to configure the comparator 5200 to operate as a static comparator or a dynamic comparator. When the comparator 5200 functions as a dynamic comparator, the comparator 5200 receives a clock signal 5205. When operating as a dynamic comparator, if the comparator is a coarse comparator, the comparison clock signal 5205 is a first clock signal of a first frequency, and if the comparator is a fine comparator, the clock signal 5205 is a second clock signal of a second frequency greater than the first frequency. The comparator 5200, when operated as a coarse comparator, has lower accuracy and slower speed, but will use less power compared to the situation where the comparator 5200 operates as a fine comparator. Thus, the dynamic comparator used for coarse comparison can utilize a slow comparison clock, while the dynamic comparator used for fine comparison can utilize a fast comparison clock during the conversion ramping period.
[0204] Comparator 5200 compares the array output 5201 against the reference voltage 5202 and generates an output 5205, similar to comparators 3704 and 4404 in FIGS. 37A and 45A. When comparator 5200 is operating as a coarse comparator, the reference voltage 5202 can be an offset voltage.
[0205]
[0206] During the conversion period to generate digital output bits as shown in FIGS. 37B and 45B / 45C, comparator 5200 can function as a coarse comparator and a fine comparator during the coarse comparison period and the fine comparison period, respectively. At the start of this digital output bit conversion, a fine comparison period or a hybrid coarse-fine comparison period (coarse comparison in parallel with fine comparison) is executed over a certain period. Next, the coarse comparison period is executed, and then finally the fine comparison is executed to complete the conversion.
[0207] FIG. 46 shows an algorithmic analog-to-digital output converter 4600 that includes the gains of switch 4601, switch 4602, sample-and-hold (S / H) circuit 4603, 1-bit analog-to-digital converter (ADC) 4604, 1-bit digital-to-analog converter (DAC) 4605, adder 4606, and 2-residue operational amplifier (2x operational amplifier) 4607. The algorithmic analog-to-digital output converter 4600 generates a converted digital output 4608 in response to an analog input Vin and control signals applied to switches 4602 and 4602. The input received at the analog input Vin (e.g., Vneu of FIG. 45A) is first sampled by the S / H circuit 4603 via switch 4602, and then conversion is performed in N clock cycles for N bits. For each conversion clock cycle, the 1-bit ADC 4604 compares the S / H voltage 4609 with a reference voltage (e.g., VREF / 2, where VREF is the full-scale voltage for N bits) and outputs a digital bit (e.g., “0” if the input <= VREF / 2, “1” if the input > VREF / 2). This digital bit, which is the digital output signal 4608, is then converted to an analog voltage (e.g., either VREF / 2 or 0V) by the 1-bit DAC 4605 and supplied to the adder 4606 to be subtracted from the S / H voltage 4609. The 2x residue operational amplifier 4607 then amplifies the adder differential voltage output to a conversion residue voltage 4610, which is supplied to the S / H circuit 4603 via switch 4601 for the next clock cycle. To reduce the effects of offsets such as those from the ADC 4604 and the residue operational amplifier 4607, a 1.5-bit (i.e., 3-level) algorithmic ADC can be used instead of this 1-bit (i.e., 2-level) algorithmic ADC. A 1.5-bit or 2-bit (i.e., 4-level) DAC is required for the 1.5-bit algorithmic ADC.
[0208] FIG. 47A shows a successive approximation (SAR) analog-to-digital converter 4700 applied to an output neuron to convert a cell current representing the output neuron into digital output bits. The SAR ADC 4700 includes a SAR 4701, a digital-to-analog converter 4702, and a comparator 4703. The cell current can drop across a resistor to generate a voltage VCELL, which is applied to the inverting input of the comparator 4703. Alternatively, the cell current can charge a sample-and-hold capacitor to generate a voltage VCELL (such as Vneu as shown in FIG. 45A). Then, a binary search is used by the SAR 4701 to calculate each bit from the most significant bit (MSB) to the least significant bit (LSB). The DAC 4702 is used to set an appropriate analog reference voltage to the comparator 4703 based on the digital bits (DN~D0) from the SAR 4701. Then, the output of the comparator 4703 is fed back to the SAR 4701 to select the next analog level of the analog reference voltage for the comparator 4703. As shown in FIG. 47B, in the example of 4-bit digital output bits, there are four evaluation periods. The first pulse evaluates DOUT3 by setting the analog level of the analog reference voltage for the comparator 4703 to the midpoint of the range, and then the second pulse evaluates DOUT2 by setting the analog level of the analog reference voltage for the comparator 4703 to the middle between the midpoint of the range and the maximum point of the range, or to the middle between the midpoint of the range and the minimum point of the range. Further steps follow this, with each step refining the analog reference voltage level for the comparator 4703. The successive outputs of the SAR 4701 are the output digital bits. An alternative SAR ADC circuit is a switched capacitor (SC) circuit with only one reference level and a local SC ratio that continuously generates a ratio reference level for successive comparisons.
[0209] FIG. 48 shows a sigma-delta type analog-to-digital converter 4800 applied to an output neuron to convert a cell current 4806 (ICELL or Ineu) into a digital output bit 4807. An integrator including an operational amplifier 4801 and a configurable capacitor 4805 (Cint) integrates the sum of the current from the cell current 4806 and a configurable reference current obtained from a 1-bit current DAC 4804 that converts the digital output 4807 into a current. A comparator 4802 compares the integrated output voltage Vint from the comparator 4801 with a reference voltage VREF2, and the output of the comparator 4802 is supplied to the D input of a clocked DFF 4803. The clocked DFF 4803 provides a digital output stream 4807 in response to the output of the comparator 4802. The digital output stream 4807 may be supplied to a digital filter before being output as the digital output bit 4807. The clock period of the clocked DFF 4803 is configurable for different Ineu ranges.
[0210] Next, calibration methods 4900, 5010, 5020, and 5100 will be described with reference to FIGS. 49, 50A, 50B, and 51, respectively. Methods 4900, 5010, 5020, and 5100 compensate for leakage and / or offset. The leakage can include one or more of array leakage and circuit leakage. The array leakage can include one or more of memory cell leakage and leakage from one or more of a decoding circuit and a column writing circuit. The offset can include one or more of an array offset and a circuit offset. The array offset can include an offset from array variations caused by one or more of a memory cell capacitance and a cell junction capacitance. The circuit offset can include an offset from one or more of a decoding circuit and a column writing circuit.
[0211] Figure 49 shows a calibration method 4900 for compensating for leakage and / or offset. A leakage and / or offset calibration step is performed (step 4901). The leakage and / or offset is measured, and the measured amount is stored as leakage_value and / or offset_value (step 4902). The LSB is determined using the following formula: LSB = leakage_value and / or offset_value + deltaLmin. Optionally, deltaLMin is a current value that compensates for variations between levels due to process, temperature, noise, or wear, and ensures that the separation between levels is sufficient. deltaLmin can optionally be determined from an evaluation of the sample data characteristics. (step 4903). The MSB is determined using the following formula: MSB = LSB + (N - 1) * deltaL, where N is the number of levels, and deltaL is the delta level amount equal to the average or ideal difference between two consecutive levels. (step 4904). In one embodiment, DeltaL is equal to the LSB. In another embodiment, DeltaL is determined from an evaluation of the sample data characteristics. DeltaL may have a uniform or non-uniform value for different pairs of consecutive levels.
[0212] For example, in the case of a 6-bit memory cell, there are 64 levels of current, each level is related to the weights in a neural network application, and N = 64. To create a baseline value, the minimum offset current may be injected at this step during calibration and measurement steps.
[0213] Table 10 shows exemplary values for a 4-bit cell. Table 10: Exemplary levels for a 4-bit cell (16 levels):
Table 10
[0214] FIG. 50A and FIG. 50B show a calibration method 5000 that includes one or more of a real-time calibration method 5010 and a background calibration method 5020.
[0215] In the real-time calibration method 5010, leakage and / or offset calibration is performed, including measuring leakage and / or offset and storing the measured values as leakage_value and / or offset_value (step 5011). The LSB is determined using the following formula: LSB level = leakage_value and / or offset_value + deltaLmin. (Step 5012). The MSB is determined using the following formula: MSB = LSB + (N - 1) * deltaL, where N is the number of levels (step 5013). The explanations of deltaLmin and deltaL regarding FIG. 49 also apply to FIG. 50A in the same way. Numerical examples are as follows: when leakage and offset = 200 pA, deltaLmin = 300 pA, LSB = 500 pA, deltaL = 400 pA, and N = 16, MSB = 500 pA + (16 - 1) * 400 pA = 6500 pA.
[0216] In the background calibration method 5020, offset_value and / or leakage_value + temperature data is stored in a fuse (e.g., a look-up table for offset and / or leakage vs. temperature) (step 5021). This is done once or periodically in the background calibration step. offset_value and / or leakage_value + temperature data is called (step 5022). Temperature adjustment for offset_value and / or leakage_value is performed according to the look-up table or by the device transistor equation (step 5023). Then, the LSB is determined using the following formula: LSB level = offset_value and / or leakage_value + deltaLmin (step 5024). The MSB is determined using the following formula: MSB = LSB + (N - 1) *deltaL (Step 5025). The explanations of deltaLmin and deltaL regarding FIG. 49 are equally applicable in FIG. 50B. The temperature adjustment can be performed by a look-up table or extrapolated from a device equation (e.g., a subthreshold, linear, or saturation equation).
[0217] FIG. 51A shows a calibration and conversion method 5100 with automatic cancellation of leakage and / or offset. Leakage and / or offset calibration is performed (Step 5101). The leakage and / or offset is measured, for example, by ADC conversion, and the measured digital output is stored in a counter (Step 5102). The conversion of the neuron output is enabled, and a countdown of the counter is performed until the counter reaches zero (thereby compensating for the leakage and / or offset initially stored in the counter), and then a count-up is performed for the digital output bits (Step 5103).
[0218] FIG. 51B shows a calibration and conversion method 5110 with automatic cancellation of leakage and / or offset, which is a variation of method 5100. Leakage and / or offset calibration is performed (Step 5111). The leakage and / or offset is measured, for example, by ADC conversion, and the measured digital output is stored in a register (Step 5112). The conversion of the neuron output is enabled, a count-up is performed for the digital output bits, and then the stored digital output is subtracted (Step 5113).
[0219] As used herein, it should be noted that both the terms "over" and "on" include both "directly over" (with no intervening material, element, or gap therebetween) and "indirectly over" (with an intervening material, element, or gap therebetween). Similarly, the term "adjacent" includes "directly adjacent" (with no intervening material, element, or gap therebetween) and "indirectly adjacent" (with an intervening material, element, or gap therebetween), "attached to" includes "directly attached to" (with no intervening material, element, or gap therebetween) and "indirectly attached to" (with an intervening material, element, or gap therebetween), and "electrically coupled" includes "directly electrically coupled" (with no intervening material or element electrically connecting the elements together therebetween) and "indirectly electrically coupled" (with an intervening material or element electrically connecting the elements together therebetween). For example, forming an element "over a substrate" can include forming the element directly on the substrate without an intervening material / element therebetween, and forming the element indirectly over the substrate with one or more intervening materials / elements therebetween.
Claims
1. An analog neural memory system comprising: A plurality of vector matrix multiplication arrays, each array comprising non-volatile memory cells arranged in rows and columns; An input block capable of providing an input to a configurable number N of said plurality of vector matrix multiplication arrays, where N can range between 1 and the total number of arrays in said plurality of vector matrix multiplication arrays; An analog neural memory system, wherein the array receiving the input provides an output in response to the input.
2. The system according to claim 1, wherein the input is generated by the input block in response to an analog current received by the input block.
3. The system according to claim 1, wherein the input is generated by the input block in response to a variable-length pulse received by the input block.
4. The system according to claim 1, wherein the input is generated by the input block in response to a series of uniform pulses received by the input block.
5. The system according to claim 1, wherein the input is generated by the input block in response to a set of bits received by the input block.
6. The system according to claim 1, wherein the non-volatile memory cell is a split-gate flash memory cell.
7. The system according to claim 1, wherein the non-volatile memory cell is a stacked-gate flash memory cell.
8. An analog neural memory system comprising: A plurality of vector matrix multiplication arrays, each of the plurality of vector matrix multiplication arrays comprising non-volatile memory cells arranged in rows and columns; An output block capable of providing an output from a configurable number N of said plurality of vector matrix multiplication arrays, where N can range between 1 and the total number of arrays in said plurality of vector matrix multiplication arrays; An analog neural memory system, wherein the output is provided in response to a received input.
9. The output block is The system according to claim 8, comprising an analog-to-digital converter for converting the analog current received from the N vector matrix multiplication arrays into the output, the output comprising a series of digital pulses.
10. The system according to claim 9, wherein the analog-to-digital converter comprises a comparator.
11. The system according to claim 10, wherein the comparator can be configured to operate in response to a first clock signal or a second clock signal, and the frequency of the second clock signal is greater than the frequency of the first clock signal.
12. The system according to claim 9, wherein the analog-to-digital converter comprises an integrating analog-to-digital converter.
13. The system according to claim 9, wherein the analog-to-digital converter comprises a ramp-type analog-to-digital converter.
14. The system according to claim 9, wherein the analog-to-digital converter comprises an algorithmic analog-to-digital converter.
15. The system according to claim 9, wherein the analog-to-digital converter comprises a sigma-delta type analog-to-digital converter.
16. The system according to claim 9, wherein the analog-to-digital converter comprises a successive approximation type analog-to-digital converter.
17. The system The system according to claim 9, further comprising a digital data-voltage converter for converting the series of digital pulses into a voltage.
18. The system The system according to claim 9, further comprising an integrating analog-to-digital data converter for converting the analog current into a set of digital bits.
19. The system The system according to claim 18, further comprising a digital bit-pulse width converter for converting the set of digital bits into one or more pulses, the width of the one or more pulses being proportional to the value of the set of digital bits.
20. The system The system according to claim 9, further comprising a current-voltage converter for converting the output analog current into a voltage.
21. The system according to claim 8, wherein the output is a variable-length pulse.
22. The system according to claim 8, wherein the output is a series of uniform pulses.
23. The system according to claim 8, wherein the output is a set of bits.
24. The system according to claim 8, wherein the non-volatile memory cell is a split-gate flash memory cell.
25. The system according to claim 8, wherein the non-volatile memory cell is a stacked-gate flash memory cell.
26. The system according to claim 8, wherein the output block performs calibration to compensate for temperature.
27. The system according to claim 8, wherein the output block performs calibration to compensate for process variations or voltage supply variations.
28. An analog neural memory system, comprising: a plurality of vector matrix multiplication arrays, each array including non-volatile memory cells arranged in rows and columns; an output block for performing a verification operation after a programming operation on a configurable number N of the vector matrix multiplication arrays, where N can range between 1 and the total number of arrays in the plurality of vector matrix multiplication arrays.
29. The system according to claim 28, wherein the accuracy of the verification operation exceeds the inference accuracy.
30. The system according to claim 29, wherein the inference is performed by an integrating ADC.
31. An analog neural memory system, comprising: a plurality of vector matrix multiplication arrays, each array including non-volatile memory cells arranged in rows and columns; an input block capable of providing an input to a first configurable number N of the vector matrix multiplication arrays, where N can range between 1 and the total number of arrays in the plurality of vector matrix multiplication arrays; an output block capable of providing an output from a second configurable number M of the vector matrix multiplication arrays, where M can range between 1 and the total number of arrays in the plurality of vector matrix multiplication arrays; wherein the output block generates the output in response to the input.
32. The system according to claim 31, wherein the input is generated by the input block in response to an analog current received by the input block.
33. The input is the system according to claim 31, generated by the input block in response to a variable-length pulse received by the input block.
34. The input is the system according to claim 31, generated by the input block in response to a series of uniform pulses received by the input block.
35. The input is the system according to claim 31, generated by the input block in response to a set of bits received by the input block.
36. The output is an analog current, for the system according to claim 31.
37. The output is a variable-length pulse, for the system according to claim 31.
38. The output is a series of uniform pulses, for the system according to claim 31.
39. The output is a set of bits, for the system according to claim 31.
40. The output block includes an analog-to-digital converter including a comparator, for the system according to claim 31.
41. The comparator can be configured to operate in response to a first clock signal or a second clock signal, and the frequency of the second clock signal is greater than the frequency of the first clock signal, for the system according to claim 40.
42. The comparator can be configured to operate in a coarse comparison period or a fine comparison period during conversion, for the system according to claim 40.
43. The non-volatile memory cell is a split-gate flash memory cell, for the system according to claim 31.
44. The non-volatile memory cell is a stacked-gate flash memory cell, for the system according to claim 31.
45. The output block performs calibration to compensate for temperature, for the system according to claim 31.
46. The output block performs calibration to compensate for process variations, for the system according to claim 31.
47. The output block performs calibration to compensate for voltage supply variations, for the system according to claim 31.
48. An analog neural memory system, A plurality of vector matrix multiplication arrays, each vector matrix multiplication array including non-volatile memory cells arranged in rows and columns, and a plurality of vector matrix multiplication arrays An output block that receives an output neuron current from one or more of the vector matrix multiplication arrays and is capable of generating digital output bits using a ramp-type analog-to-digital converter, and an analog neural memory system comprising the same. **Claim 49** The system according to claim 48, further comprising a discrete or continuous ramping reference voltage. **Claim 50** The system according to claim 48, further comprising a sample-and-hold circuit and a comparator, wherein the ramping reference voltage is applied to an input of the comparator. **Claim 51** The system according to claim 50, wherein the ramping reference voltage includes a coarse voltage ramp followed by a plurality of fine voltage ramps. **Claim 52** The system according to claim 51, wherein the coarse voltage ramp includes a plurality of coarse ramping voltages. **Claim 53** An analog neural memory system, A plurality of vector matrix multiplication arrays, each vector matrix multiplication array including a plurality of vector matrix multiplication arrays including non-volatile memory cells, An input block capable of converting a plurality of digital input bits into a binary-indexed time addition signal as a timing input to at least one of the vector matrix multiplication arrays, and an analog neural memory system comprising the same. **Claim 54** The system according to claim 53, wherein the input block generates a binary-indexed pulse for each digit input bit. **Claim 55** The system according to claim 53, wherein the input block includes a storage latch for each input digital bit. **Claim 56** The system according to claim 53, further comprising a generator for generating a binary-indexed pulse. **Claim 57** The system according to claim 53, wherein the input block includes a row decoder. **Claim 58** The system according to claim 53, wherein the binary-indexed time addition signal is generated according to digital input bits for each row. **Claim 59** The system according to claim 53, wherein the time addition is from the LSB to the MSB or in any random order. **Claim 60** A method for performing output conversion on an analog neural memory including a plurality of vector matrix multiplication arrays, each vector matrix multiplication array including non-volatile memory cells, the method comprising: Receiving an output neuron current from one or more of the plurality of vector matrix multiplication arrays; Generating digital output bits using the output neuron current and a ramp type analog-to-digital converter, the converter operating in a coarse comparison mode and a fine comparison mode; a method comprising: Claim 61 The method of claim 60, wherein the generating step utilizes a dynamic comparator. Claim 62 The method of claim 61, wherein the dynamic comparator is configured to be different with respect to the coarse comparison mode and the fine comparison mode. Claim 63 The method of claim 62, wherein the dynamic comparator receives a first comparison clock for the coarse comparison mode and a second comparison clock for the fine comparison mode, and the frequency of the second comparison clock exceeds the frequency of the first comparison clock.
Citation Information
Patent Citations
Neural network
JP1990181284A
Coupler by neuro-chip
JP1990201586A
Single-slope analog-to-digital converter
JP2010503253A
Analog-digital conversion circuit and imaging apparatus mounted with the same
JP2011066773A
A / d converter circuit
JP2013223112A