Programmable Output Block for Analog Neural Memory in Deep Learning Artificial Neural Networks

The programmable output block in the VMM array system addresses the challenge of accurately measuring and transferring VMM array outputs by configuring the ADC gain and resolution, ensuring efficient and precise data conversion.

JP7700283B2Active Publication Date: 2025-06-30SILICON STORAGE TECHNOLOGY INC
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2023578766
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-07-05
Filing Date
2021-10-05
Publication Date
2025-06-30
Estimated Expiration
2041-10-05

AI Technical Summary

Technical Problem

Existing systems face challenges in accurately measuring and transferring the output of Vector Matrix Multiplication (VMM) arrays due to issues like leakage current, leading to loss of information.

Method used

A programmable output block is designed to configure the gain and resolution of the Analog-to-Digital Converter (ADC) within the output block, allowing for efficient conversion and transfer of the VMM array output.

Benefits of technology

The programmable output block effectively addresses the accuracy and information loss issues by enabling configurable gain and resolution, ensuring precise conversion and transfer of VMM array outputs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007700283000010
    Figure 0007700283000010
  • Figure 0007700283000011
    Figure 0007700283000011
  • Figure 0007700283000012
    Figure 0007700283000012
Patent Text Reader

Abstract

Numerous embodiments are disclosed for a programmable output block for use with a VMM array in an artificial neural network. In one embodiment, the gain of the output block can be configured by a configuration signal. In another embodiment, the resolution of the ADC in the output block can be configured by a configuration signal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] (Claims of Priority) This application is a continuation-in-part of U.S. patent application Ser. No. 16 / 353,830, filed Mar. 14, 2019, titled "System for Converting Neuron Current Into Neuron Current-Based Time Pulses in an Analog Neural Memory in a Deep Learning Artificial Neural Network", which claims priority from U.S. Provisional Patent Application No. 62 / 814,813, filed Mar. 6, 2019, titled "System for Converting Neuron Current Into Neuron Current-Based Time Pulses in an Analog Neural Memory in a Deep Learning Artificial Neural Network" and U.S. Provisional Patent Application No. 62 / 794,492, filed Jan. 18, 2019, titled "System for Converting Neuron Current Into Neuron Current-Based Time Pulses in an Analog Neural Memory in a Deep Learning Artificial Neural Network", all of which are hereby incorporated by reference.

[0002] (Field of the Invention) Disclosed are a number of embodiments for programmable output blocks for use with vector-by-matrix multiplication (VMM) arrays within an artificial neural network.

Background Art

[0003] An artificial neural network mimics a biological neural network (the central nervous system of an animal, especially the brain), can depend on a large number of inputs, and is used to estimate or approximate a generally unknown function. An artificial neural network generally includes layers of interconnected "neurons" that exchange messages with each other.

[0004] Figure 1 shows an artificial neural network, in which the circles represent layers of inputs or neurons. Connections (called synapses) are represented by arrows and have numerical weights that can be tuned based on experience. Thereby, the neural network adapts to the inputs and becomes learnable. Typically, a neural network includes multiple input layers. Typically, there is one or more intermediate layers of neurons and an output layer of neurons that provides the output of the neural network. At each level, the neurons make decisions individually or collectively based on the data received from the synapses.

[0005] One of the main challenges in the development of artificial neural networks for high-performance information processing is the lack of appropriate hardware technology. In practice, practical neural networks rely on a very large number of synapses, which enables high connectivity between neurons, that is, a very high degree of parallelization of computational processing. In principle, such complexity can be realized by a digital supercomputer or a dedicated graphics processing unit cluster. However, in addition to high costs, these approaches also suffer from poor energy efficiency compared to biological networks, which mainly perform low-precision analog calculations and consume much less energy. CMOS analog circuits have been used in artificial neural networks, but most CMOS-implemented synapses have been too bulky assuming the required large number of neurons and synapses.

[0006] The applicant has previously disclosed in U.S. Patent Application No. 15 / 594,439, published as U.S. Patent Application Publication No. 2017 / 0337466, incorporated by reference, an artificial (analog) neural network that utilizes one or more non-volatile memory arrays as synapses. The non-volatile memory arrays operate as analog neuromorphic memories. The neural network device includes a first plurality of synapses configured to receive a first plurality of inputs and then generate a first plurality of outputs, and a first plurality of neurons configured to receive the first plurality of outputs. The first plurality of synapses includes a plurality of memory cells, each of the memory cells including a spaced source region and drain region formed in a semiconductor substrate with a channel region extending therebetween, a floating gate insulated and disposed above a first portion of the channel region, and a non-floating gate insulated and disposed above a second portion of the channel region. Each of the plurality of memory cells is configured to store a weight value corresponding to a plurality of electrons in the floating gate. The plurality of memory cells are configured to multiply the stored weight values by the first plurality of inputs to generate the first plurality of outputs.

[0007] Each non-volatile memory cell used in an analog neuromorphic memory system must hold a charge, i.e., the number of electrons, in the floating gate in a very specific and accurate amount in response to erase and program. For example, each floating gate must hold one of N different values, where N is the number of different weights that can be represented by each cell. Examples of N include 16, 32, 64, 128, and 256.

[0008] One challenge in a system that utilizes a VMM array is the ability to accurately measure the output of the VMM array and transfer that output to another stage, such as an input block of another VMM array. Many approaches are known, but each has certain drawbacks, such as loss of information due to leakage current.

[0009] What is needed is an improved output block for receiving the output current of a VMM array and converting that output current into a form more suitable for transfer to another stage of an electronic device.

Summary of the Invention

[0010] A number of embodiments for a programmable output block for use with a VMM array within an artificial neural network are disclosed. In one embodiment, the gain of the output block can be configured by a configuration signal. In another embodiment, the resolution of the ADC within the output block can be configured by a configuration signal.

[0011]

[0012]

[0013]

[0014]

[0015]

[0016]

[0017]

[0018]

[0019]

[0020]

[0021]

[0022]

[0023]

[0024]

[0025]

[0026]

[0027]

[0028]

[0029]

[0030]

[0031]

[0032]

[0033]

[0034]

[0035]

[0036]

[0037]

[0038]

[0039]

[0040]

[0041]

[0042]

[0043]

[0044]

[0045]

[0046]

[0047]

[0048]

[0049]

[0050]

[0051]

[0052]

[0053]

[0054]

[0055]

[0056]

[0057]

[0058]

[0059]

[0060]

[0061]

[0062]

[0063]

[0064]

[0065]

[0066]

[0067]

[0068]

[0069]

[0070]

[0071]

[0072]

[0073]

[0074]

[0075]

[0076]

[0077]

[0078]

Brief Description of the Drawings

[0079]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27

Figure 28

Figure 29

Figure 30

Figure 31

Figure 32

Figure 33

Figure 34A

Figure 34B

Figure 35A

Figure 35B

Figure 36A

Figure 36B

Figure 36C

Figure 37

Figure 38

Figure 39

Figure 40

Figure 41

Figure 42

Figure 43

Figure 44

Figure 45

Figure 46

Figure 47

Figure 48

Figure 49

Figure 50

Figure 51

Figure 52

Figure 53

Figure 54

Figure 55A

Figure 55B

Figure 56

Figure 57

Figure 58

Figure 59

Figure 60

Figure 61A

Figure 61B

Figure 61C

Figure 62

Figure 63

Figure 64

Figure 65

Figure 66

Best Mode for Carrying Out the Invention

[0080] The artificial neural network of the present invention utilizes a combination of CMOS technology and a non-volatile memory array. Non-volatile memory cell

[0081] Digital non-volatile memories are well known. For example, U.S. Patent No. 5,029,130 (the “’130 Patent”), which is incorporated herein by reference, discloses an array of split-gate non-volatile memory cells, which are a type of flash memory cell. Such a memory cell 210 is shown in FIG. 2. Each memory cell 210 includes a source region 14 and a drain region 16 formed in a semiconductor substrate 12, with a channel region 18 therebetween. A floating gate 20 is formed insulated above a first portion of the channel region 18 (and controls the conductivity of the first portion of the channel region 18), and extends over a portion of the source region 14. A word line terminal 22 (typically coupled to a word line) is disposed insulated above a second portion of the channel region 18, having a first portion (which controls the conductivity of the second portion of the channel region 18) and a second portion that extends upwardly above the floating gate 20. The floating gate 20 and the word line terminal 22 are insulated from the substrate 12 by a gate oxide. A bit line 24 is coupled to the drain region 16.

[0082] By applying a high positive voltage to the word line terminal 22, the memory cell 210 is erased (electrons are removed from the floating gate), whereby electrons in the floating gate 20 pass through the insulator therebetween from the floating gate 20 to the word line terminal 22 via Fowler-Nordheim tunneling.

[0083] The memory cell 210 is programmed (electrons are applied to the floating gate) by applying a positive voltage to the word line terminal 22 and a positive voltage to the source region 14. An electron current will flow from the source region 14 towards the drain region 16. The electrons accelerate and heat up when they reach the gap between the word line terminal 22 and the floating gate 20. A portion of the heated electrons is injected into the floating gate 20 through the gate oxide due to the electrostatic attraction from the floating gate 20.

[0084] The memory cell 210 is read by applying a positive read voltage to the drain region 16 and the word line terminal 22 (turning on the portion of the channel region 18 below the word line terminal). When the floating gate 20 is positively charged (i.e., electrons are erased), the portion of the channel region 18 below the floating gate 20 is also turned on, and current flows through the channel region 18, which is detected as the erased state, i.e., the "1" state. When the floating gate 20 is negatively charged (i.e., programmed with electrons), the portion of the channel region below the floating gate 20 becomes almost or completely off, and current does not flow (or hardly flows) through the channel region 18, which is detected as the programmed state, i.e., the "0" state.

[0085] Table 1 shows the typical voltage ranges that can be applied to the terminals of the memory cell 110 to perform read, erase, and program operations. Table 1: Operation of the flash memory cell 210 of FIG. 2 [Table 1]

[0086] FIG. 3 shows a memory cell 310 similar to the memory cell 210 of FIG. 2 with an added control gate (CG) 28. The control gate 28 is biased at a high voltage (e.g., 10V) during programming, a low or negative voltage (e.g., 0V / -8V) during erase, and a low or medium voltage (e.g., 0V / 2.5V) during read. The other terminals are biased in the same manner as the terminals of FIG. 2.

[0087] FIG. 4 shows a four-gate memory cell 410 including a source region 14, a drain region 16, a floating gate 20 above a first portion of a channel region 18, a select gate 22 (typically coupled to a word line, WL) above a second portion of the channel region 18, a control gate 28 above the floating gate 20, and an erase gate 30 above the source region 14. This configuration is described in U.S. Patent No. 6,747,310, which is incorporated herein by reference for all purposes. Here, all gates are non-floating gates except for the floating gate 20, i.e., they are electrically connected or connectable to a voltage source. Programming is performed by injecting hot electrons themselves from the channel region 18 into the floating gate 20. Erasure is performed by tunneling electrons from the floating gate 20 to the erase gate 30.

[0088] Table 2 shows typical voltage ranges that can be applied to the terminals of the memory cell 310 to perform read, erase, and program operations. Table 2: Operation of the Flash Memory Cell 410 of FIG. 4

Table 2

[0089] FIG. 5 shows a memory cell 510 similar to the memory cell 410 of FIG. 4 except that the memory cell 510 does not include an erase gate (EG). Erasure is performed by biasing the substrate 18 to a high voltage and biasing the control gate CG28 to a low voltage or a negative voltage. Alternatively, erasure is performed by biasing the word line 22 to a positive voltage and biasing the control gate 28 to a negative voltage. Programming and reading are the same as those of FIG. 4.

[0090] FIG. 6 shows a three-gate memory cell 610, which is another type of flash memory cell. Memory cell 610 is identical to memory cell 410 of FIG. 4, except that memory cell 610 does not have a separate control gate. (Erasure occurs through the use of an erase gate) The erase operation and the read operation are the same as those of FIG. 4, except that no control gate bias is applied. The programming operation is also performed without a control gate bias. As a result, during the programming operation, a higher voltage must be applied to the source line to compensate for the lack of control gate bias.

[0091] Table 3 shows the typical voltage ranges that can be applied to the terminals of memory cell 610 to perform read, erase, and program operations. Table 3: Operation of Flash Memory Cell 610 of FIG. 6

Table 3

[0092] FIG. 7 shows a stacked gate memory cell 710, which is another type of flash memory cell. Memory cell 710 is the same as memory cell 210 of FIG. 2, except that the floating gate 20 extends over the entire channel region 18 and the control gate 22 (which is coupled to the word line here) is separated by an insulating layer (not shown) and extends over the floating gate 20. The erase, programming, and read operations operate in the same manner as those described above for memory cell 210.

[0093] Table 4 shows the typical voltage ranges that can be applied to the terminals of memory cell 710 and the substrate 12 to perform read, erase, and program operations. Table 4: Operation of Flash Memory Cell 710 of FIG. 7

Table 4

[0094] To utilize a memory array that includes one of the types of non-volatile memory cells in the above artificial neural network, two modifications are made. First, the wires are configured so that each memory cell can be individually programmed, erased, and read without adversely affecting the memory states of other memory cells in the array, as further described below. Second, continuous (analog) programming of the memory cells is provided.

[0095] Specifically, the memory state (i.e., the charge on the floating gate) of each memory cell in the array can be changed independently, with minimal disruption to other memory cells, continuously from a fully erased state to a fully programmed state. In another embodiment, the memory state (i.e., the charge on the floating gate) of each memory cell in the array can be changed independently, with minimal disruption to other memory cells, continuously from a fully programmed state to a fully erased state and vice versa. This means that the cell memory is either analog or can store at least one of a number of discrete values (such as 16 or 64 different values), which allows all cells in the memory array to be very precisely and individually tunable, and makes the memory array ideal for fine-tuning adjustments to memory and the synaptic weights of neural networks.

[0096] The methods and means described herein can be applied to other non-volatile memory technologies such as SONOS (silicon-oxide-nitride-oxide-silicon, charge trap in nitride), MONOS (metal-oxide-nitride-oxide-silicon, metal charge trap in nitride), ReRAM (resistive ram), PCM (phase change memory), MRAM (magnetic ram), FeRAM (ferroelectric ram), OTP (bi-level or multi-level one time programmable), CeRAM (correlated electron ram). The methods and means described herein can be applied to volatile memory technologies used in neural networks such as SRAM, DRAM, and / or volatile synaptic cells without limitation. Neural network using a non-volatile memory cell array

[0097] FIG. 8 conceptually illustrates a non-limiting example of a neural network that utilizes the non-volatile memory array of the present embodiment. This example uses a non-volatile memory array neural network for a face recognition application, but it is also possible to implement other suitable applications using a non-volatile memory array-based neural network.

[0098] S0 is the input layer, which in this example is a 32×32 pixel RGB image with 5-bit precision (i.e., three 32×32 pixel arrays, one for each of the colors R, G, and B, and each pixel has 5-bit precision). The synapses CB1 going from the input layer S0 to the layer C1 apply a different set of weights to some instances and shared weights to other instances, scanning the input image with overlapping 3×3 pixel filters (kernels) and shifting the filter by 1 pixel (or more than 2 pixels in some models) at a time. Specifically, the 9 pixel values in the 3×3 portion of the image (i.e., what is called the filter or kernel) are provided to the synapses CB1, where these 9 input values are multiplied by appropriate weights, and after summing the output of that multiplication, a single output value is determined and given by the first synapse of CB1 to generate one pixel of the layer of the feature map C1. The 3×3 filter is then shifted 1 pixel to the right within the input layer S0 (i.e., a 3-pixel column is added on the right and a 3-pixel column is dropped on the left), and the 9 pixel values of this newly positioned filter are provided to the synapses CB1, where they are multiplied by the same weights as above and a second single output value is determined by the relevant synapse. This process is continued until the 3×3 filter has scanned the entire 32×32 pixel image of the input layer S0 for all three colors and all bits (precision values). The process is then repeated using different sets of weights until all of the feature maps of layer C1 are calculated, generating different feature maps of C1.

[0099] In this example, in layer C1, there are 16 feature maps each having 30×30 pixels. Each pixel is a new feature pixel extracted from the multiplication of the input and the kernel. Thus, each feature map is a two-dimensional array. Therefore, in this example, layer C1 consists of 16 layers of two-dimensional arrays (note that the layers and arrays referred to in this specification are logical relationships rather than necessarily physical relationships, that is, the arrays are not necessarily oriented as physical two-dimensional arrays). Each of the 16 feature maps in layer C1 is generated by one of 16 different sets of synaptic weights applied to the filter scan. All of the C1 feature maps can target different aspects of the same image feature, such as edge identification. For example, the first map (generated using the first set of weights shared by all scans used to generate this first map) can identify circular edges, and the second map (generated using a second set of weights different from the first set of weights) can identify rectangular edges or the aspect ratio of a specific feature, etc.

[0100] Before going from layer C1 to layer S1, an activation function P1 (pooling) is applied that pools the values from non-overlapping and consecutive 2×2 regions within each feature map. The purpose of the pooling function is to average the neighboring positions (or it is also possible to use the max function), for example, to reduce the dependence on edge positions, and to reduce the data size before going to the next stage. In layer S1, there are 16 15×15 feature maps (i.e., 16 different arrays of 15×15 pixels each). The synapses CB2 going from layer S1 to layer C2 scan the maps in S1 with a 4×4 filter with a 1-pixel filter shift. In layer C2, there are 22 12×12 feature maps. Before going from layer C2 to layer S2, an activation function P2 (pooling) is applied that pools the values from non-overlapping and consecutive 2×2 regions within each feature map. In layer S2, there are 22 6×6 feature maps. In the synapses CB3 going from layer S2 to layer C3, an activation function (pooling) is applied, where all neurons in layer C3 are connected to all maps in layer S2 via each synapse of CB3. In layer C3, there are 64 neurons. The synapses CB4 going from layer C3 to the output layer S3 fully connect C3 to S3, i.e., all neurons in layer C3 are connected to all neurons in layer S3. The output in S3 contains 10 neurons, and the neuron with the highest output determines the class. This output can indicate, for example, the identification or classification (categorization) of the content of the original image.

[0101] Each layer of synapses is implemented using an array or a part of an array of non-volatile memory cells.

[0102] FIG. 9 is a block diagram of an array that can be used for that purpose. The vector matrix multiplication (VMM) array 32 includes non-volatile memory cells and is utilized as synapses (such as CB1, CB2, CB3, and CB4 in FIG. 6) between one layer and the next layer. Specifically, the VMM array 32 includes an array 33 of non-volatile memory cells, an erase gate and word line gate decoder 34, a control gate decoder 35, a bit line decoder 36, and a source line decoder 37, and these decoders decode respective inputs to the non-volatile memory cell array 33. Inputs to the VMM array 32 can be made from the erase gate and word line gate decoder 34 or from the control gate decoder 35. The source line decoder 37 in this example also decodes the output of the non-volatile memory cell array 33. Alternatively, the bit line decoder 36 can decode the output of the non-volatile memory cell array 33.

[0103] The non-volatile memory cell array 33 serves two purposes. First, the non-volatile memory cell array 33 stores the weights used by the VMM array 32. Second, the non-volatile memory cell array 33 effectively multiplies the inputs by the weights stored in the non-volatile memory cell array 33 and adds them for each output line (source line or bit line) to generate an output, which becomes the input to the next layer or the input to the last layer. By the non-volatile memory cell array 33 performing the functions of multiplication and addition, the need for separate multiplication and addition logic circuits is eliminated, and the power efficiency is also good due to the in-memory calculation.

[0104] The output of the non-volatile memory cell array 33 is supplied to a differential summing device (such as a summing operational amplifier or a summing current mirror) 38 that sums the output of the non-volatile memory cell array 33 to create a single value for convolution. The differential summing device 38 is arranged to perform the sum of the positive and negative weights.

[0105] The total output value of the differential adder 38 is then supplied to an activation function circuit 39 that rectifies the output. The activation function circuit 39 can provide a sigmoid, tanh, or ReLU function. The rectified output value of the activation function circuit 39 becomes an element of the feature map of the next layer (e.g., C1 in FIG. 8) and is then applied to the next synapse to generate the next feature map layer or the last layer. Thus, in this example, the nonvolatile memory cell array 33 constitutes a plurality of synapses (receiving inputs from the previous layer of neurons or from an input layer such as an image database), and the adder 38 and the activation function circuit 39 constitute a plurality of neurons.

[0106] The inputs (WLx, EGx, CGx, and optionally BLx and SLx) to the VMM array 32 of FIG. 9 can be at an analog level, binary level, digital pulse (in which case a pulse - analog converter PAC may be required to convert the pulse to an appropriate input analog level), or digital bits (in which case a DAC is provided to convert the digital bits to an appropriate input analog level), and the outputs can be at an analog level, binary level, digital pulse, or digital bits (in which case an output ADC is provided to convert the output analog level to digital bits).

[0107] FIG. 10 is a block diagram showing the use of multiple layers of the VMM array 32, labeled as VMM arrays 32a, 32b, 32c, 32d, and 32e in the figure. As shown in FIG. 10, an input (denoted as Inputx) is converted from digital to analog by a digital-to-analog converter 31 and provided to the input VMM array 32a. The converted analog input can be a voltage or a current. The input D / A conversion of the first layer can be performed by using a function or a LUT (look-up table) that maps the input Inputx to an appropriate analog level of the matrix multiplier of the input VMM array 32a. The input conversion can also be performed by an analog-to-analog (A / A) converter to convert an external analog input to the mapped analog input to the input VMM array 32a. The input conversion can also be performed by a digital-to-digital pulse (D / P) converter to convert an external digital input to the mapped digital pulse(s) to the input VMM array 32a.

[0108] The output generated by the input VMM array 32a is then provided as input to the next VMM array (hidden level 1) 32b, which in turn generates an output that is provided as input to the input VMM array (hidden level 2) 32c, and so on. The various layers of the VMM array 32 function as the respective layers of synapses and neurons of a convolutional neural network (CNN). Each of the VMM arrays 32a, 32b, 32c, 32d, and 32e can be a stand-alone physical non-volatile memory array, or multiple VMM arrays can utilize different portions of the same physical non-volatile memory array, or multiple VMM arrays can utilize overlapping portions of the same physical non-volatile memory array. Each of the VMM arrays 32a, 32b, 32c, 32d, and 32e can also be time-multiplexed with respect to the various portions of that array or neuron. The example shown in FIG. 10 includes five layers (32a, 32b, 32c, 32d, 32e), namely, one input layer (32a), two hidden layers (32b, 32c), and two fully connected layers (32d, 32e). One of ordinary skill in the art will understand that this is merely exemplary, and alternatively, the system can include more than two hidden layers and more than two fully connected layers. Vector Matrix Multiplication (VMM) Array

[0109] FIG. 11 shows a neuron VMM array 1100 that is particularly suitable for the memory cell 310 shown in FIG. 3 and is utilized as a portion of the synapses and neurons between the input layer and the next layer. The VMM array 1100 includes a memory array 1101 of non-volatile memory cells and a reference array 1102 of non-volatile reference memory cells (located at the top of the array). Alternatively, another reference array can be located at the bottom.

[0110] In the VMM array 1100, control gate lines such as control gate line 1103 extend vertically (thus, the reference array 1102 in the row direction is orthogonal to the control gate line 1103), and erase gate lines such as erase gate line 1104 extend horizontally. Here, the input to the VMM array 1100 is provided to the control gate lines (CG0, CG1, CG2, CG3), and the output of the VMM array 1100 appears on the source lines (SL0, SL1). In one embodiment, only even rows are used, and in another embodiment, only odd rows are used. The current of each source line (SL0 and SL1 respectively) performs a summation function of all the currents from the memory cells connected to that particular source line.

[0111] As described herein for the neural network, the non-volatile memory cells of the VMM array 1100, i.e., the flash memory of the VMM array 1100, are preferably configured to operate in the subthreshold region.

[0112] The non-volatile reference memory cells and non-volatile memory cells described herein are biased with weak inversion as follows: Ids = Io * e (Vg-Vth) / kVt = w * Io * e (Vg) / kVt where w = e (-Vth) / kVt is.

[0113] When using an I-V logarithmic converter that uses a memory cell (such as a reference memory cell or a peripheral memory cell) or a transistor to convert an input current into an input voltage: Vg = k * Vt * log[Ids / wp * Io] where wp is the w of the reference or peripheral memory cell.

[0114] Regarding the memory array used as the vector matrix multiplier VMM array, the output current is as follows: Iout = wa * Io* e (Vg) / kVt That is Iout = (wa / wp) * Iin = W * Iin W = e (Vthp-Vtha) / kVt Wherein, wa is w of each memory cell of the memory array.

[0115] The word line or control gate can be used as the input of the memory cell for the input voltage.

[0116] Alternatively, the flash memory cells of the VMM array described herein can be configured to operate in the linear region. Ids = β * (Vgs - Vth) * Vds; β = u * Cox * W / L W = α(Vgs - Vth)

[0117] The word line or control gate or bit line or source line can be used as the input of the memory cell operating in the linear region. The bit line or source line can be used as the output of the memory cell.

[0118] For an I - V linear converter, a memory cell (such as a reference memory cell or a peripheral memory cell) or a transistor or a resistor operating in the linear region can be used to linearly convert the input - output current into an input - output voltage.

[0119] Another embodiment for the VMM array 32 of FIG. 9 is described in U.S. Patent Application No. 15 / 826,345, which is incorporated herein by reference. As described in the above application, the source line or bit line can be used as the neuron output (sum - of - currents output). Alternatively, the flash memory cells of the VMM array described herein can be configured to operate in the saturation region. Ids = α1 / 2 * β * (Vgs - Vth) 2; β = u * Cox * W / L W = α(Vgs - Vth) 2

[0120] The word line, control gate, or erase gate can be used as the input of a memory cell operating in the saturation region. The bit line or source line can be used as the output of the output neuron.

[0121] Alternatively, the flash memory cells of the VMM array described herein can be used in all regions or combinations thereof (subthreshold, linear, or saturation).

[0122] FIG. 12 shows a neuron VMM array 1200 particularly suitable for the memory cell 210 shown in FIG. 2 and is utilized as a synapse between the input layer and the next layer. The VMM array 1200 includes a memory array 1203 of non-volatile memory cells, a reference array 1201 of first non-volatile reference memory cells, and a reference array 1202 of second non-volatile reference memory cells. The reference arrays 1201 and 1202 arranged in the column direction of the array function to convert the current inputs flowing into the terminals BLR0, BLR1, BLR2, and BLR3 into voltage inputs WL0, WL1, WL2, and WL3. In practice, the first and second non-volatile reference memory cells are diode-connected through a multiplexer 1214 (only a part is shown) in a state where the current inputs flow in. The reference cells are tuned (e.g., programmed) to a target reference level. The target reference level is provided by a reference min-array matrix (not shown).

[0123] The memory array 1203 serves two purposes. First, the memory array 1203 stores the weights used by the VMM array 1200 in respective memory cells. Second, the memory array 1203 effectively multiplies the weights stored in the memory array 1203 by the input (i.e., the current inputs provided to terminals BLR0, BLR1, BLR2, and BLR3, which the reference arrays 1201 and 1202 convert into input voltages and supply to word lines WL0, WL1, WL2, and WL3), and then adds up all the results (memory cell currents) to generate the output of each bit line (BL0~BLN), and this output becomes the input to the next layer or the input to the last layer. By performing the functions of multiplication and addition, the memory array 1203 eliminates the need for separate multiplication and addition logic circuits and also has good power efficiency. Here, the voltage input is provided to word lines WL0, WL1, WL2, and WL3, and the output appears on respective bit lines BL0~BLN during the read (inference) operation. The current of each of bit lines BL0~BLN performs the total function of the currents from all the non-volatile memory cells connected to that specific bit line.

[0124] Table 5 shows the operating voltages of the VMM array 1200. The columns in the table indicate the voltages applied to the word line of the selected cell, the word line of the non-selected cell, the bit line of the selected cell, the bit line of the non-selected cell, the source line of the selected cell, and the source line of the non-selected cell. The rows indicate the operations of read, erase, and program. Table 5: Operations of the VMM array 1200 in FIG. 12 [Table 5]

[0125] FIG. 13 shows a neuron VMM array 1300 that is particularly suitable for the memory cell 210 shown in FIG. 2 and is used as part of synapses and neurons between the input layer and the next layer. The VMM array 1300 includes a memory array 1303 of non-volatile memory cells, a reference array 1301 of first non-volatile reference memory cells, and a reference array 1302 of second non-volatile reference memory cells. The reference arrays 1301 and 1302 extend in the row direction of the VMM array 1300. The VMM array is similar to the VMM 1000, except that the word lines extend in the vertical direction in the VMM array 1300. Here, the input is provided to the word lines (WLA0, WLB0, WLA1, WLB 1 , WLA2, WLB2, WLA3, WLB3), and the output appears on the source lines (SL0, SL1) during a read operation. The current of each source line performs a sum function of all the currents from the memory cells connected to that particular source line.

[0126] Table 6 shows the operating voltages of the VMM array 1300. The columns in the table show the voltages applied to the word line of the selected cell, the word line of the non-selected cell, the bit line of the selected cell, the bit line of the non-selected cell, the source line of the selected cell, and the source line of the non-selected cell. The rows show the operations of read, erase, and program. Table 6: Operations of the VMM Array 1300 in FIG. 13 [Table 6]

[0127] FIG. 14 shows a neuron VMM array 1400 that is particularly suitable for the memory cell 310 shown in FIG. 3 and is used as part of synapses and neurons between the input layer and the next layer. The VMM array 1400 includes a memory array 1403 of non-volatile memory cells, a reference array 1401 of first non-volatile reference memory cells, and a reference array 1402 of second non-volatile reference memory cells. The reference arrays 1401 and 1402 function to convert the current inputs flowing into the terminals BLR0, BLR1, BLR2, and BLR3 into voltage inputs CG0, CG1, CG2, and CG3. In practice, the first and second non-volatile reference memory cells are diode-connected through a multiplexer 1412 (only part shown) in a state where the current inputs flow through BLR0, BLR1, BLR2, and BLR3. The multiplexer 1412 includes respective multiplexers 1405 and cascode transistors 1404 to ensure a constant voltage for each bit line (such as BLR0) of the first and second non-volatile reference memory cells during the read operation. The reference cells are tuned to a target reference level.

[0128] The memory array 1403 serves two purposes. First, the memory array 1403 stores the weights used by the VMM array 1400. Second, the memory array 1403 effectively multiplies the weights stored in the memory array by the inputs (the current inputs provided to the terminals BLR0, BLR1, BLR2, and BLR3, and the reference arrays 1401 and 1402 convert these current inputs into input voltages and supply them to the control gates (CG0, CG1, CG2, and CG3)), and then adds up all the results (cell currents) to generate an output that appears on BL0 to BLN and serves as an input to the next layer or the last layer. By the memory array performing the multiplication and addition functions, the need for separate multiplication and addition logic circuits is eliminated, and the power efficiency is also good. Here, the inputs are provided to the control gate lines (CG0, CG1, CG2, and CG3), and the outputs appear on the bit lines (BL0 to BLN) during the read operation. The current of each bit line performs the summation function of all the currents from the memory cells connected to that specific bit line.

[0129] The VMM array 1400 implements unidirectional tuning of non-volatile memory cells within the memory array 1403. That is, each non-volatile memory cell is erased and then partially programmed until the desired charge on the floating gate is reached. This can be performed, for example, using the precision programming techniques described below. If too much charge is applied to the floating gate (such as when an incorrect value is stored in the cell), the cell must be erased and a series of partial programming operations must be repeated. As shown, two rows that share the same erase gate (such as EG0 or EG1) must be erased together (known as page erase), after which each cell is partially programmed until the desired charge on the floating gate is reached.

[0130] Table 7 shows the operating voltages of the VMM array 1400. The columns in the table show the voltages applied to the word line of the selected cell, the word line of the non-selected cell, the bit line of the selected cell, the bit line of the non-selected cell, the control gate of the selected cell, the control gate of the non-selected cell within the same sector as the selected cell, the control gate of the non-selected cell within a different sector from the selected cell, the erase gate of the selected cell, the erase gate of the non-selected cell, the source line of the selected cell, and the source line of the non-selected cell. The rows show the read, erase, and program operations. Table 7: Operation of the VMM Array 1400 of FIG. 14

Table 7

[0131] FIG. 15 shows a neuron VMM array 1500 that is particularly suitable for the memory cell 310 shown in FIG. 3 and is utilized as part of a synapse and neuron between an input layer and the next layer. The VMM array 1500 includes a memory array 1503 of non-volatile memory cells and of the first non-volatile reference memory cell a reference array 150 1 and, a reference array 1502 of second non-volatile reference memory cells. The EG lines EGR0, EG0, EG1, and EGR1 extend vertically, and the CG lines CG0, CG1, CG2, and CG3 and the SL lines WL0, WL1, WL2, and WL3 extend horizontally. The VMM array 1500 is similar to the VMM array 1400 except that the VMM array 1500 implements bidirectional tuning. Each individual cell can be completely erased, partially programmed, and partially erased as needed to reach the desired charge level on the floating gate by using an individual EG line. As shown, the reference arrays 1501 and 1502 convert the input current at the terminals BLR0, BLR1, BLR2, and BLR3 into control gate voltages CG0, CG1, CG2, and CG3 (through the action of diode-connected reference cells via the multiplexer 1514), and these voltages are applied to the memory cells in the row direction. The current output (neuron) is in the bit lines BL0~BLN, and each bit line sums all the currents from the non-volatile memory cells connected to that particular bit line.

[0132] Table 8 shows the operating voltages of the VMM array 1500. The columns in the table show the voltages applied to the word line of the selected cell, the word line of the non-selected cell, the bit line of the selected cell, the bit line of the non-selected cell, the control gate of the selected cell, the control gate of the non-selected cell in the same sector as the selected cell, the control gate of the non-selected cell in a different sector from the selected cell, the erase gate of the selected cell, the erase gate of the non-selected cell, the source line of the selected cell, and the source line of the non-selected cell. The rows show the read, erase, and program operations. Table 8: Operation of the VMM Array 1500 in FIG. 15

Table 8

[0133] FIG. 24 shows a neuron VMM array 2400 that is particularly suitable for the memory cell 210 shown in FIG. 2 and is used as part of a synapse and a neuron between the input layer and the next layer. In the VMM array 2400, the inputs INPUT0...., INPUTn are respectively received by bit lines BL0, ... BL n and outputs OUTPUT1, OUTPUT2, OUTPUT3, and OUTPUT4 are respectively generated by source lines SL0, SL1, SL2, and SL3.

[0134] FIG. 25 shows a neuron VMM array 2500 that is particularly suitable for the memory cell 210 shown in FIG. 2 and is used as part of synapses and neurons between the input layer and the next layer. In this example, inputs INPUT0, INPUT1, INPUT2, and INPUT3 are respectively received by source lines SL0, SL1, SL2, and SL3, and outputs OUTPUT0, ... OUTPUT N are bit lines BL0, ..., BL N generated by.

[0135] FIG. 26 shows a neuron VMM array 2600 that is particularly suitable for the memory cell 210 shown in FIG. 2 and is used as part of synapses and neurons between the input layer and the next layer. In this example, inputs INPUT0, ..., INPUT M are word lines WL0, ..., WL M respectively received by, and outputs OUTPUT0, ... OUTPUT N are bit lines BL0, ..., BL N generated by.

[0136] FIG. 27 shows a neuron VMM array 2700 that is particularly suitable for the memory cell 310 shown in FIG. 3 and is used as part of synapses and neurons between the input layer and the next layer. In this example, inputs INPUT0, ..., INPUT M are word lines WL0, ..., WL M respectively received by, and outputs OUTPUT0, ... OUTPUT N are bit lines BL0, ..., BL N generated by.

[0137] FIG. 28 shows a neuron VMM array 2800 that is particularly suitable for the memory cell 410 shown in FIG. 4 and is used as part of synapses and neurons between the input layer and the next layer. In this example, the inputs INPUT0, ..., INPUT n are received by bit lines BL0, ..., BL N respectively, and the outputs OUTPUT1 and OUTPUT2 are generated by the erase gate lines EG0 and EG1.

[0138] FIG. 29 shows a neuron VMM array 2900 that is particularly suitable for the memory cell 410 shown in FIG. 4 and is used as part of synapses and neurons between the input layer and the next layer. In this example, the inputs INPUT0, ..., INPUT N are received by the gates of bit line control gates 2901-1, 2901-2, ..., 2901-(N-1), and 2901-N respectively, which are coupled to bit lines BL0, ..., BL N respectively. Exemplary outputs OUTPUT1 and OUTPUT2 are generated by the erase gate lines SL0 and SL1.

[0139] FIG. 30 shows a neuron VMM array 3000 that is particularly suitable for the memory cell 310 shown in FIG. 3, the memory cell 510 shown in FIG. 5, and the memory cell 710 shown in FIG. 7, and is used as part of synapses and neurons between the input layer and the next layer. In this example, the inputs INPUT0, ..., INPUT M are received by word lines WL0, ..., WL M and the outputs OUTPUT0, ..., OUTPUT N are generated by bit lines BL0, ..., BL N respectively.

[0140] FIG. 31 shows a neuron VMM array 3100 that is particularly suitable for the memory cell 310 shown in FIG. 3, the memory cell 510 shown in FIG. 5, and the memory cell 710 shown in FIG. 7, and is used as part of synapses and neurons between the input layer and the next layer. In this example, the inputs INPUT0, ..., INPUT Mis the control gate lines CG0, ..., CG M is received. The outputs OUTPUT0, ..., OUTPUT N are generated by the source lines SL0, ..., SL N and each source line SL i is coupled to the source line terminals of all memory cells within column i.

[0141] FIG. 32 shows a VMM system 3200. The VMM system 3200 includes a VMM array 3201 (which can be based on any of the VMM designs described above, such as VMM900, 1000, 1100, 1200, and 1320, or other VMM designs), a low voltage row decoder 3202, a high voltage row decoder 3203, a reference cell low voltage column decoder 3204 (shown in the column direction, which means it provides an input to output a transformation in the row direction), a bit line multiplexer 3205, control logic 3206, an analog circuit 3207, a neuron output block 3208, an input VMM circuit block 3209, a pre - decoder 3210, a test circuit 3211, an erase program control logic EPCTL3212, an analog and high voltage generation circuit 3213, a bit line PE driver 3214, redundant arrays 3215 and 3216, an NVR sector 3217, and a reference sector 3218. The input circuit block 3209 functions as an interface from an external input to the input terminals of the memory array. The neuron output block 3208 functions as an interface from the memory array output to an external interface.

[0142] The low-voltage row decoder 3202 provides bias voltages for read and program operations and provides decoded signals to the high-voltage row decoder 3203. The high-voltage row decoder 3203 provides high-voltage bias signals for program and erase operations. The reference cell low-voltage column decoder 3204 provides a decoding function for reference cells. The bit line PE driver 3214 provides a control function for bit lines during program, verify, and erase operations. The analog and high-voltage generation circuit 3213 is a shared bias block that provides multiple voltages required for various program, erase, program verify, and read operations. The redundant arrays 3215 and 3216 provide array redundancy to replace defective array portions. The NVR (Non-Volatile Register alias Information Sector) sector 3217 is an array sector used to store user information, device ID, password, security key, trim bits, configuration bits, and manufacturing information without limitation.

[0143] Figure 33 shows an analog neuromemory system 3300. The analog neuromemory system 3300 includes macroblocks 3301a, 3301b, 3301c, 3301d, 3301e, 330If, 3301g, and 3301h, neuron output (such as summing circuit and sample and hold S / H circuit) blocks 3302a, 3302b, 3302c, 3302d, 3302e, 3302f, 3302g, and 3302h, and input circuit blocks 3303a, 3303b, 3303c, 3303d, 3303e, 3303f, 3303g, and 3304h. Each of the macroblocks 3301a, 3301b, 3301c, 3301d, 3301e, and 3301f is a VMM subsystem including a VMM array containing rows and columns of non-volatile memory cells such as flash memory cells. The neuromemory subsystem 3333 includes the macroblock 3301, the input block 3303, and the neuron output block 3302. The neuromemory subsystem 3333 may have its own digital control block.

[0144] The analog neuromemory system 3300 further includes a system control block 3304, an analog low voltage block 3305, a high voltage block 3306, and a timing control circuit 3670, which will be discussed in more detail below with respect to FIG. 36.

[0145] The system control block 3304 may include one or more microcontroller cores, such as an ARM / MIPS / RISC_V core, to handle general control functions and arithmetic operations. The system control block 3304 may also include a SIMD (Single Instruction Multiple Data) unit that operates on multiple data using a single instruction. The system control block 3304 may include a DSP core. The system control block 3304 may include, without limitation, functions such as pool, average, minimum, maximum, softmax, addition, subtraction, multiplication, division, log, antilog, ReLu, sigmoid, tanh, etc., and hardware or software for performing data compression. The system control block 3304 may include hardware or software for performing functions such as an activation approximator / quantizer / normalizer. The system control block 3304 may include the ability to perform functions such as an approximator / quantizer / normalizer for input data. 。Ni The control block of the neuromemory subsystem 3333 may include similar elements of the system control block 3304, such as a microcontroller core, a SIMD core, a DSP core, and other functional units.

[0146] In one embodiment, each of the neuron output blocks 3302a, 3302b, 3302c, 3302d, 3302e, 3302f, 3302g, and 3302h includes a low-impedance output buffer (e.g., an operational amplifier) circuit capable of driving configurable long interconnects. In one embodiment, each of the input circuit blocks 3303a, 3303b, 3303c, 3303d, 3303e, 3303f, 3303g, and 3303h provides a high-impedance current output for addition. In another embodiment, each of the neuron output blocks 3302a, 3302b, 3302c, 3302d, 3302e, 3302f, 3302g, and 3302h includes an activation circuit, in which case an additional low-impedance buffer is required to drive the output.

[0147] In another embodiment, each of the neuron output blocks 3302a, 3302b, 3302c, 3302d, 3302e, 3302f, 3302g, and 3302h includes an analog-to-digital conversion block that outputs digital bits instead of analog signals. In this embodiment, each of the input circuit blocks 3303a, 3303b, 3303c, 3303d, 3303e, 3303f, 3303g, and 3303h includes a digital-to-analog conversion block that receives digital bits from the respective neuron output blocks and converts the digital bits into an analog signal.

[0148] Accordingly, the neuron output blocks 3302a, 3302b, 3302c, 3302d, 3302e, 3302f, 3302g, and 3302h receive the output current from the macro blocks 3301a, 3301b, 3301c, 3301d, 3301e, and 3301f and optionally convert the output current into an analog voltage, digital bits, or one or more digital pulses whose width or number of pulses varies according to the value of the output current. Similarly, the input circuit blocks 3303a, 3303b, 3303c, 3303d, 3303e, 3303f, 3303g, and 3303h optionally receive an analog current, analog voltage, digital bits, or a digital pulse whose width or number of pulses varies in response to the value of the output current and provide an analog current to the macro blocks 3301a, 3301b, 3301c, 3301d, 3301e, and 3301f. The input circuit blocks 3303a, 3303b, 3303c, 3303d, 3303e, 3303f, 3303g, and 3303h optionally include a voltage-current converter, an analog or digital counter for counting the number of digital pulses in the input signal or the length of the width of the digital pulses in the input signal, or a digital-to-analog converter.

[0149] Optionally, when converting the output current into an analog voltage, digital bits, or one or more digital pulses, the neuron output blocks 3302a, 3302b, 3302c, 3302d, 3302e, 3302f, 3302g, and 3302h can apply a programmable gain. This can be called a programmable neuron.

[0150] FIG. 59 shows an example of a programmable neuron output block 5900, which receives an output neuron current Ineu from a VMM array and a gain configuration 5901, and generates an output 5902 representing the output neuron current Ineu with a gain G, where the value of G is set in response to the gain configuration 5901. The gain configuration 5901 can be an analog signal or digital bits. In one embodiment, the programmable neuron output block 5900 includes a gain control circuit 5903, and the gain control circuit 5903 includes a variable resistor 5904 or a variable capacitor 5905 controlled by the gain configuration 5901 to generate the gain G. The output 5902 can be an analog voltage, an analog current, digital bits, or one or more digital pulses. In some embodiments, the gain configuration 5901 is used to trim the gain G of the programmable neuron output block 5900 to compensate for undesirable phenomena such as leakage current.

[0151] Optionally, each programmable neuron output block 5900 can have a different gain configuration 5901. This allows, for example, different gains (such as for scaling the array output) to be implemented in different layers of the neural network.

[0152] In another embodiment, the gain configuration 5901 depends in part on, for example, the input size, which means how many rows are enabled to generate the output neuron current Ineu.

[0153] In another embodiment, the gain configuration 5901 depends in part on the values of all the rows input to the VMM array. For example, for an 8-bit row input to the VMM array, the maximum value is 256 (2^8) for the input of one row, 1024 for the input of four rows, etc. For example, when 256 rows are enabled, a determination is made regarding the sum value of these rows, and the gain configuration 5901 is modified in response to this value.

[0154] In another embodiment, the gain configuration 5901 depends on the output neuron range. For example, when the output neuron current Ineu is within a first range, a first gain G1 is applied by the gain configuration 5901. When the output neuron current Ineu is within a second range, a second gain G2 is applied via the gain configuration 5901. Although this has been described in terms of ranges, one of ordinary skill in the art will recognize that more ranges can be implemented without limitation. Long short-term memory

[0155] The prior art includes a concept known as long short-term memory (LSTM). LSTM units are often used within neural networks. With LSTM, a neural network can store information over an arbitrary period of time and use that information in subsequent operations. Conventional LSTM units include a cell, an input gate, an output gate, and a forget gate. The three gates regulate the flow of information into and out of the cell and the period during which information is stored within the LSTM, and the VMM is particularly useful in LSTM units.

[0156] FIG. 16 shows an exemplary LSTM 1600. The LSTM 1600 in this example includes cells 1601, 1602, 1603, and 1604. Cell 1601 receives an input vector x0 and generates an output vector h0 and a cell state vector c0. Cell 1602 receives an input vector x1, an output vector (hidden state) h0 from cell 1601, and a cell state c0 from cell 1601, and generates an output vector h1 and a cell state vector c1. Cell 1603 receives an input vector x2, an output vector (hidden state) h1 from cell 1602, and a cell state c1 from cell 1602, and generates an output vector h2 and a cell state vector c2. Cell 1604 receives an input vector x3, an output vector (hidden state) h2 from cell 1603, and a cell state c2 from cell 1603, and generates an output vector h3. Additional cells can also be used, and the LSTM having four cells is merely an example.

[0157] FIG. 17 shows an exemplary implementation of an LSTM cell 1700 that can be used for cells 1601, 1602, 1603, and 1604 of FIG. 16. The LSTM cell 1700 receives an input vector x(t), a cell state vector c(t−1) from a preceding cell, and an output vector h(t−1) from a preceding cell, and generates a cell state vector c(t) and an output vector h(t).

[0158] The LSTM cell 1700 includes sigmoid function devices 1701, 1702, and 1703, each of which applies a number from 0 to 1 to control the degree to which each component of the input vector contributes to the output vector. The LSTM cell 1700 also includes tanh devices 1704 and 1705 for applying a hyperbolic tangent function to the input vector, multiplier devices 1706, 1707, and 1708 for multiplying two vectors, and an adder device 1709 for adding two vectors. The output vector h(t) can be provided to the next LSTM cell in the system or accessed for other purposes.

[0159] FIG. 18 shows an LSTM cell 1800 that is an example of an implementation of the LSTM cell 1700. For the convenience of the reader, the same numbering scheme from the LSTM cell 1700 is used in the LSTM cell 1800. The sigmoid function devices 1701, 1702, and 1703, and the tanh device 1704 each include a plurality of VMM arrays 1801 and activation circuit blocks 1802. Thus, it can be seen that the VMM array is particularly useful in LSTM cells used in a particular neural network system. The multiplier devices 1706, 1707, and 1708, and the adder device 1709 are implemented in a digital or analog manner. The activation function block 1802 can be implemented in a digital or analog manner.

[0160] An alternative to the LSTM cell 1800 (and another example of an implementation of the LSTM cell 1700) is shown in FIG. 19. In FIG. 19, the sigmoid function devices 1701, 1702, and 1703, and the tanh device 1704 share the same physical hardware (the VMM array 1901 and the activation function block 1902) in a time-division multiplexed manner. The LSTM cell 1900 also includes a multiplier device 1903 for multiplying two vectors, an adder device 1908 for adding two vectors, a tanh device 1705 (including the activation circuit block 1902), a register 1907 for storing the value i(t) when i(t) is output from the sigmoid function block 1902, and the value f(t) * c(t - 1), a register 1904 for storing it when its value is output from the multiplier device 1903 via the multiplexer 1910, and the value i(t) * u(t), a register 1905 for storing it when its value is output from the multiplier device 1903 via the multiplexer 1910, and the value o(t) * c~(t), a register 1906 for storing it when its value is output from the multiplier device 1903 via the multiplexer 1910, and the multiplexer 1909.

[0161] While the LSTM cell 1800 includes multiple sets of the VMM array 1801 and the respective activation function blocks 1802, the LSTM cell 1900 includes only one set of the VMM array 1901 and the activation function block 1902, which is used to represent multiple layers in an embodiment of the LSTM cell 1900. The LSTM cell 1900 requires less space than the LSTM 1800 because it requires only 1 / 4 of the space for the VMM and the activation function block compared to the LSTM cell 1800.

[0162] An LSTM unit typically includes a plurality of VMM arrays, each of which further requires functions provided by specific circuit blocks outside the VMM array, such as an adder, an activation circuit block, and a high voltage generation block. Providing separate circuit blocks for each VMM array would require a significant amount of space within the semiconductor device and would be somewhat inefficient. Gated Recurrent Unit

[0163] The analog VMM implementation can be utilized in a gated recurrent unit (GRU) system. A GRU is a gate mechanism within a recurrent neural network. A GRU is similar to an LSTM, except that a GRU cell generally includes fewer components than an LSTM cell.

[0164] FIG. 20 shows an exemplary GRU 2000. The GRU 2000 in this example includes cells 2001, 2002, 2003, and 2004. Cell 2001 receives an input vector x0 and generates an output vector h0. Cell 2002 receives an input vector x1 and the output vector h0 from cell 2001 and generates an output vector h1. Cell 2003 receives an input vector x2 and the output vector (hidden state) h1 from cell 2002 and generates an output vector h2. Cell 2004 receives an input vector x3 and the output vector (hidden state) h2 from cell 2003 and generates an output vector h3. Additional cells can also be used, and a GRU with four cells is merely an example.

[0165] FIG. 21 shows an exemplary implementation of a GRU cell 2100 that can be used in the cells 2001, 2002, 2003, and 2004 of FIG. 20. The GRU cell 2100 receives an input vector x(t) and an output vector h(t−1) from a preceding GRU cell, and generates an output vector h(t). The GRU cell 2100 includes sigmoid function devices 2101 and 2102, each of which applies a number between 0 and 1 to components from the output vector h(t−1) and the input vector x(t). The GRU cell 2100 also includes a tanh device 2103 for applying a hyperbolic tangent function to the input vector, a plurality of multiplier devices 2104, 2105, and 2106 for multiplying two vectors, an adder device 2107 for adding two vectors, and a complement device 2108 for subtracting the input from 1 to generate an output.

[0166] FIG. 22 shows a GRU cell 2200 that is an example of an implementation of the GRU cell 2100. For the convenience of the reader, the same numbering method from the GRU cell 2100 is used in the GRU cell 2200. As can be seen from FIG. 22, the sigmoid function devices 2101 and 2102, and the tanh device 2103 each include a plurality of VMM arrays 2201 and activation function blocks 2202. Thus, it can be seen that the VMM array is particularly used in the GRU cell used in a specific neural network system. The multiplier devices 2104, 2105, 2106, the adder device 2107, and the complement device 2108 are implemented in a digital or analog manner. The activation function block 2202 can be implemented in a digital or analog manner.

[0167] An alternative example of GRU cell 2200 (and another example of an implementation of GRU cell 2300) is shown in FIG. 23. In FIG. 23, GRU cell 2300 utilizes VMM array 2301 and activation function block 2302, and when configured as a sigmoid function, it applies a number between 0 and 1 to control the degree to which each component of the input vector contributes to the output vector. In FIG. 23, sigmoid function devices 2101 and 2102, and tanh device 2103 share the same physical hardware (VMM array 2301 and activation function block 2302) in a time-division multiplexed manner. GRU cell 2300 also includes a multiplier device 2303 for multiplying two vectors, an adder device 2305 for adding two vectors, a complement device 2309 for subtracting the input from 1 to generate an output, a multiplexer 2304, and the value h(t - 1) * r(t), a register 2306 for holding the value when it is output from multiplier device 2303 via multiplexer 2304, and the value h(t - 1) * z(t), a register 2307 for holding the value when it is output from multiplier device 2303 via multiplexer 2304, and the value h^(t) * (1 - z(t)), a register 2308 for holding the value when it is output from multiplier device 2303 via multiplexer 2304, are included.

[0168] While GRU cell 2200 includes multiple sets of VMM array 2201 and activation function block 2202, GRU cell 2300 includes only one set of VMM array 2301 and activation function block 2302, which is used to represent multiple layers in an embodiment of GRU cell 2300. GRU cell 2300 requires less space than GRU cell 2200 because it requires only 1 / 3 of the space for the VMM and activation function blocks compared to GRU cell 2200.

[0169] The GRU system typically includes a plurality of VMM arrays, each of which further requires functions provided by specific circuit blocks outside the VMM array, such as a summing device, an activation circuit block, and a high-voltage generation block. Providing separate circuit blocks for each VMM array would require a significant amount of space within the semiconductor device and would be somewhat inefficient.

[0170] The input to the VMM array can be at an analog level, a binary level, or a digital bit (in which case a DAC is required to convert the digital bit to an appropriate input analog level), and the output can be at an analog level, a binary level, or a digital bit (in which case an output ADC is required to convert the output analog level to a digital bit).

[0171] For each memory cell within the VMM array, each weight w can be implemented by a single memory cell, or by differential cells, or by two blended memory cells (the average of two cells). In the case of differential cells, two memory cells are required to implement the weight w as a differential weight (w = w+ - w-). In the case of two blended memory cells, two memory cells are required to implement the weight w as the average of two cells. Output circuit

[0172] FIG. 34A shows an integrating dual-mixed slope analog-to-digital converter (ADC) 3400 that is applied to output neuron I NEU 3406 to convert the output neuron current to a digital pulse or a digital output bit.

[0173] In one embodiment, the ADC3400 converts the analog output current within the neuron output blocks (such as neuron output blocks 3302a, 3302b, 3302c, 3302d, 3302e, 3302f, 3302g, and 3302h in FIG. 32) into digital pulses, the width of which varies in proportion to the magnitude of the analog output current within the neuron output blocks. An integrator including an integrating operational amplifier 3401 and an integrating capacitor 3402 integrates the memory array current I NEU 3406 (which is the output neuron current) with respect to the reference current IREF3407.

[0174] Optionally, IREF3407 can include a bandgap filter having a temperature coefficient of 0 or a temperature coefficient that tracks the neuron current I NEU 3406. The latter can optionally be obtained from a reference array including values determined during a test phase.

[0175] Optionally, calibration steps can be performed while the circuit is above the operating temperature to offset any leakage current present in the array or control circuit, and then the offset value can be subtracted from Ineu in FIG. 34B or FIG. 35B.

[0176] During the initialization phase, switch 3408 is closed. Then, the inputs to Vout3403 and the negative terminal of operational amplifier 3401 become VREF. Thereafter, as shown in FIG. 34B, switch 3 408 is opened, and for a fixed time tref, the neuron current I NEU 3406 is up-integrated. During the fixed time tref, Vout rises and its slope changes as the neuron current changes. Thereafter, during the period tmeas, a fixed reference current IREF is down integrated over the time tmeas (the period during which Vout drops), and tmeas is the time required to down integrate Vout to VREF.

[0177] Output EC3405 goes high when VOUT > VREFV and low otherwise. Thus, EC3405 generates a pulse whose width reflects period tmeas, and period tmeas in turn is proportional to current I NEU 3406. In Figure 34B, EC3405 is shown as waveform 3410 for the example where tmeas = Ineu1 and waveform 3412 for the example where tmeas = Ineu2. Thus, output neuron current I NEU 3406 is converted to digital pulse EC3405, and the width of digital pulse EC3405 varies in proportion to the magnitude of output neuron current I NEU 3406.

[0178] Current I NEU 3406 is = tmeas / tref * IREF. For example, for a desired output bit resolution of 10 bits, tref is a time equal to 1024 clock cycles. Period tmeas varies from a period equal to 0 clock cycles to 1024 clock cycles depending on the value of INEU3406 and the value of Iref. Figure 34B shows examples of two different values of I NEU 3406 where I NEU 3406 = Ineu1 and I NEU 3406 = Ineu2. Thus, neuron current I NEU 3406 affects the rate and slope of charging.

[0179] Optionally, output pulse EC3405 can be converted to a series of pulses of uniform period for transmission to the next circuit stage, such as an input block of another VMM array. At the start of period tmeas, output EC3405 is input to AND gate 3440 using reference clock 3441. The output becomes pulse train 3442 (the frequency of the pulses in pulse train 3442 is the same as the frequency of clock 3441) during the period VOUT > VREF. The number of pulses is proportional to period tmeas, and period tmeas is proportional to current I NEU 3406.

[0180] Optionally, the pulse train 3443 can be input to the counter 3420, which counts the number of pulses in the pulse train 3442 and generates a count value 3421 that is a digital count of the number of pulses in the pulse train 3442 that is proportional to the neuron current I NEU 3406. The count value 3421 includes a set of digital bits. In another embodiment, the integrating dual slope ADC 3400 can convert the neuron current I NEU 3407 into pulses, and the width of the pulses is inversely proportional to the magnitude of the neuron current I NEU 3407. This inversion can be performed in a digital or analog manner and can be converted into a series of pulses or digital bits for output to a subsequent circuit.

[0181] FIG. 35A shows an integrating dual hybrid slope ADC 3500 that is applied to the output neuron, I NEU 3504 to convert it into a digital pulse or a series of digital output bits with a varying width of the cell current. For example, the ADC 3500 can be used to convert the analog output current within a neuron output block (such as the neuron output blocks 3302a, 3302b, 3302c, 3302d, 3302e, 3302f, 3302g, and 3302h of FIG. 32) into a set of digital output bits. An integrator including an integrating operational amplifier 3501 and an integrating capacitor 3502 integrates the neuron current I NEU 3504 with respect to the reference current IREF3503. The switch 3505 can be closed to reset VOUT.

[0182] During the initialization phase, the switch 3505 is closed and VOUT is charged to the voltage V BIAS . Then, as shown in FIG. 35B, the switch 3505 is opened and the cell current I NEU 3504 is integrally up-integrated for a certain time tref. Thereafter, the reference current IREF3503 is integrally down-integrated for a time tmeas until Vout drops to ground. The current I NEU 3504 = tmeas Ineu / treU *IREF. For example, in the case of a desired output bit resolution of 10 bits, tref is a time equal to 1024 clock cycles. The period tmeas varies from a period equal to 0 clock cycles to 1024 clock cycles depending on the value of I NEU 3504 and Iref. FIG. 35B shows examples of two different Ineu values, such as one having a current Ineu1 and one having a current Ineu2. Thus, the neuron current I NEU 3504 affects the charging and discharging speed and slope.

[0183] Output 3506 goes high when VOUT > VREF and low otherwise. Thus, output 3506 generates a pulse whose width reflects the period tmeas, and tmeas in turn is proportional to the current I NEU 3404. In FIG. 35B, output 3506 is shown as waveform 3512 for the example where tmeas = Ineu1 and waveform 3515 for the example where tmeas = Ineu2. Thus, the output neuron current I NEU 3504 is converted to a pulse, output 3506, and the width of the pulse varies in proportion to the magnitude of the output neuron current I NEU 3504.

[0184] Optionally, output 3506 can be converted to a series of pulses of a uniform period for transmission to the next circuit stage, such as an input block of another VMM array. At the start of the period tmeas, output 3506 is input to AND gate 3508 using reference clock 3507. The output becomes pulse train 3509 (the frequency of the pulses in pulse train 3509 is the same as the frequency of reference clock 3507) during the period VOUT > VREF. The number of pulses is proportional to the period tmeas, and the period tmeas is proportional to the current I NEU 3504.

[0185] Optionally, pulse train 3509 can be input to counter 3510, and counter 3510 counts the number of pulses in pulse train 3509 and, as shown by waveforms 3514, 3517, the neuron current INEU Generate a count value 3511 that is a digital count of the number of pulses in the pulse train 3509 that is directly proportional to 3504. The count value 3511 includes a set of digital bits.

[0186] In another embodiment, the integrating dual-slope ADC 3500 is a neuron current I NEU 3504 can be converted into pulses, and the width of the pulses is inversely proportional to the magnitude of the neuron current I NEU 3504. This inversion can be performed in a digital or analog manner and converted into one or more pulses or digital bits for output to subsequent circuitry.

[0187] FIG. 35B shows the count values 3511 (digital bits) for two neuron current values Ineu1 and Ineu2, respectively, for I NEU 3504.

[0188] FIGS. 36A and 36B show waveforms associated with exemplary methods 3600 and 3650 executed by the VMM during operation. In each method 3600 and 3650, the word lines WL0, WL1, and WL2 receive various different inputs, which can optionally be converted into analog voltage waveforms and applied to the word lines. In these examples, the voltage VC represents the voltage across the integrating capacitor 3402 or 3502 in FIGS. 34A and 35A, respectively, within the output block of the first VMM or within the ADC 3400 or 3500, and the OT pulse (= "1") represents the period during which the output of the neuron (proportional to the value of the neuron) is captured using the integrating dual-slope ADC 3400 or 3500. As described with reference to FIGS. 34 and 35, the output of the output block may be pulses of a width that varies in proportion to the output neuron current of the first VMM, or a series of pulses of a uniform width where the number of pulses varies in proportion to the neuron current of the first VMM. These pulses can then be applied as inputs to the second VMM.

[0189] During method 3600, a series of pulses (such as pulse train 3442 or pulse train 3509), or an analog voltage derived from a series of pulses, is applied to the word lines of the second VMM array. Alternatively, a series of pulses, or an analog voltage derived from a series of pulses, can be applied to the control gates of the cells within the second VMM array. The number of pulses (or clock cycles) directly corresponds to the magnitude of the input. In this particular example, the magnitude of the input for WL1 is four times that of WL0 (4 pulses versus 1 pulse).

[0190] During method 3650, single pulses of various widths (such as EC3405 or output 3506), or an analog voltage derived (derive d ) from a single pulse, is applied to the word lines of the second VMM array, where the pulses have variable pulse widths. Alternatively, the pulse or the analog voltage derived from the pulse can be applied to the control gate. The width of the single pulse directly corresponds to the magnitude of the input. For example, the magnitude of the input for WL1 is four times that of WL0 (the WL1 pulse width is four times that of the WL0 pulse width).

[0191] Further, referring to FIG. 36C, the timing control circuit 3670 can be used to manage the output and input interfaces of the VMM array and manage the power of the VMM system by sequentially partitioning the conversion of various outputs or various inputs. FIG. 56 depicts a power management method 5600. The first step is to receive a plurality of inputs of the vector matrix multiplication array (step 5601). The second step is to organize the plurality of inputs into sets of the plurality of inputs (step 5602). The third step is to sequentially provide each of the sets of the plurality of inputs to the array (step 5603).

[0192] One embodiment of the power management method 5600 is as follows. Inputs to the VMM system (such as word lines or control gates of the VMM array) can be applied sequentially over time. For example, in the case of a VMM array having 512 word line inputs, the word line inputs can be divided into four groups, WL0-127, WL128-255, WL256-383, and WL383-511. Each group can be enabled at different times, and the output read operation can be performed (converting the neuron current into digital bits) on the group corresponding to one of the four word line groups by an output integration circuit such as those in FIGS. 34-36. Then, after each of the four groups has been read in order, the output digital bit results are combined together. This operation can be controlled by the timing control circuit 3670.

[0193] In another embodiment, the timing control circuit 3670 performs power management in a vector matrix multiplication system such as the analog neuromemory system 3300 of FIG. 33. The timing control circuit 3670 can apply the inputs to the VMM subsystem 3333 sequentially over time, such as by enabling the input circuit blocks 3303a, 3303b, 3303c, 3303d, 3303e, 3303f, 3303g, and 3303h at different times. Similarly, the timing control circuit 3670 can read out the outputs from the VMM subsystem 333 sequentially over time, such as by enabling the neuron output blocks 3302a, 3302b, 3302c, 3302d, 3302e, 3302f, 3302g, and 3302h at different times.

[0194] FIG. 57 shows a power management method 5700. The first step is to receive a plurality of outputs from the vector matrix multiplication array (step 5701). The next step is to organize the plurality of outputs from the array into a plurality of sets of outputs (step 5702). The next step is to sequentially provide each of the plurality of sets of outputs to a converter circuit (step 5703).

[0195] One embodiment of the power management method 5700 is as follows. Power management can be implemented by the timing control circuit 3670 by sequentially reading out the neuron output groups at different times, that is, by multiplexing the output circuits (such as output ADC circuits) across a plurality of neuron outputs (bit lines). The bit lines can be arranged in different groups, and the output circuits operate sequentially in one group at a time under the control of the timing control circuit 3670.

[0196] FIG. 58 shows a power management method 5800. In a vector matrix multiplication system including a plurality of arrays, the first step is to receive a plurality of inputs. The next step is to sequentially enable one or more of the plurality of arrays to receive some or all of the plurality of inputs (step 5802).

[0197] One embodiment of the power management method 5800 is as follows. The timing control circuit 3670 can operate in one neural network layer at a time. For example, when one neural network layer is represented in the first VMM array and the second neural network layer is represented in the second VMM array, the output read operation (such as when the neuron output is converted to a digital bit) is sequentially executed in one VMM array at a time, thereby managing the power of the VMM system.

[0198] In another embodiment, the timing control circuit 3670 can operate by sequentially enabling a plurality of neuromemory subsystems 3333 or a plurality of macros 3301 as shown in FIG. 33.

[0199] In another embodiment, the timing control circuit 3670 can operate by sequentially enabling a plurality of neuromemory subsystems 3333 or a plurality of macros 3301 as shown in FIG. 33 without discharging the array bias during the inactive period (which means the pause period between sequentially enabled on and off). (For example, the word line WL and / or the bit line BL are input to the control gate CG, and the bit line BL is biased as the output, or the control gate CG and / or the bit line BL are input to the word line WL, and the bit line BL is biased as the output). This is to save power from unnecessary discharge and recharge of the array bias that is used multiple times during one or more read operations (for example, during inference or classification operations).

[0200] FIGS. 37 to 44 show various circuits that can be used in the VMM input blocks such as the input circuit blocks 3303a, 3303b, 3303c, 3303d, 3303e, 3303f, 3303g, and 3303h of FIG. 33, or the neuron output blocks such as the neuron output blocks 3302a, 3302b, 3302c, 3302d, 3302e, 3302f, 3302g, and 3302h of FIG. 33.

[0201] FIG. 37 shows a pulse-voltage converter 3700 that can optionally be used to convert a digital pulse generated by an integrating dual-slope ADC 3400 or 3500 into a voltage, which can be applied, for example, as an input to a VMM memory array (such as the WL line or the CG line). The pulse-voltage converter 3700 includes a reference current generator 3701 that generates a reference current IREF, a capacitor 3702, and a switch 3703. The input is used to control the switch 3703. When a pulse is received at the input, the switch is closed and charge is accumulated in the capacitor 3702, and as a result, the voltage of the capacitor 3702 after the input signal is complete will indicate the number of pulses received. The capacitor can optionally be a word line or a control gate capacitance.

[0202] FIG. 38 shows a current-voltage converter 3800 that can optionally be used to convert a neuron output current to a voltage, which can be applied, for example, as an input to a VMM memory array (e.g., a WL line or a CG line). The current-voltage converter 3800 here includes a current generator 3801 representing the received neuron current, ineu (or Iin), and a variable resistor 3802. The output Vout increases in size as the neuron current increases. The variable resistor 3802 can be adjusted to increase or decrease the maximum range of Vout as desired.

[0203] FIG. 39 shows a current-voltage converter 3900 that can optionally be used to convert a neuron output current to a voltage, which can be applied, for example, as an input to a VMM memory array (e.g., a WL line or a CG line). The current-voltage converter 3900 includes an operational amplifier 3901, a capacitor 3902, a switch 3903, a switch 3904, and a current source 3905 representing the neuron current ICELL here. During operation, switch 3903 is opened and switch 3904 is closed. The output Vout increases in amplitude in proportion to the magnitude of the neuron current ICELL 3905.

[0204] FIG. 40 shows a current-logarithmic voltage converter 4000 that can optionally be used to convert a neuron output current to a logarithmic voltage, which can be applied, for example, as an input to a VMM memory array (e.g., a WL line or a CG line). The current-logarithmic voltage converter 4000 includes a memory cell 4001, a switch 4002 (selectively connecting the word line terminal of the memory cell 4001 to the node generating Vout), and a current source 4003 representing the neuron current IiN here. During operation, switch 4002 is closed and the output Vout increases in amplitude in proportion to the magnitude of the neuron current Ii N.

[0205] FIG. 41 shows a current-logarithmic voltage converter 4100 that can be optionally used to convert a neuron output current into a logarithmic voltage, which can be applied, for example, as an input to a VMM memory array (e.g., a WL line or a CG line). The current-logarithmic voltage converter 4100 includes a memory cell 4101, a switch 4102 (selectively connecting the control gate terminal of the memory cell 4101 to the node generating Vout), and a current source 4103 that represents the neuron current IiN here. During operation, the switch 4102 is closed, and the output Vout increases in amplitude in proportion to the magnitude of the neuron current IiN.

[0206] FIG. 42 shows a digital data-voltage converter 4200 that can be optionally used to convert digital data (i.e., data of 0 and 1) into a voltage, which can be applied, for example, as an input to a VMM memory array (e.g., a WL line or a CG line). The digital data-voltage converter 4200 includes a capacitor 4201, an adjustable current source 4202 (which is the current from a reference array of memory cells here), and a switch 4203. The digital data controls the switch 4203. For example, the switch 4203 can be closed when the digital data is "1" and opened when the digital data is "0". The voltage accumulated in the capacitor 4201 is the output OUT, corresponding to the value of the digital data. Optionally, the capacitor can be a word line or a control gate capacitance.

[0207] FIG. 43 shows a digital data-voltage converter 4300 that can optionally be used to convert digital data (i.e., 0s and 1s) into a voltage that can be applied, for example, as an input to a VMM memory array (e.g., a WL line or a CG line). The digital data-voltage converter 4300 includes a variable resistor 4301, an adjustable current source 4302 (here, the current from a reference array of memory cells), and a switch 4303. The digital data controls the switch 4303. For example, the switch 4303 can be closed when the digital data is "1" and opened when the digital data is "0". The output voltage corresponds to the value of the digital data.

[0208] FIG. 44 shows a reference array 4400 that can be used to provide the reference currents for the adjustable current sources 4202 and 4302 in FIGS. 42 and 43.

[0209] FIGS. 45-47 show components for verifying that after a programming operation, the flash memory cells in the VMM contain the appropriate charge corresponding to the W value intended to be stored in those flash memory cells.

[0210] FIG. 45 shows a digital comparator 4500 that receives a reference set of W values as a digital input and receives the W digital values sensed from a number of programmed flash memory cells. The digital comparator 4500 generates a flag if there is an inconsistency indicating that one or more of the flash memory cells are not programmed with the correct value.

[0211] FIG. 46 shows the digital comparator 4500 of FIG. 45 in cooperation with a converter 4600. The sensed W values are provided by a plurality of instantiations of the converter 4600. The converter 4600 receives the cell current ICELL from the flash memory cells and converts that cell current into digital data that can be provided to the digital comparator 4500 using one or more of the aforementioned converters such as ADC3400 or 3500.

[0212] Figure 47 shows an analog comparator 4700 that receives a reference set of W values as analog inputs and receives W analog values sensed from several programmed flash memory cells. The analog comparator 4700 generates a flag if there is an inconsistency indicating that one or more flash memory cells are not programmed with the correct value.

[0213] Figure 48 shows the analog comparator 4700 of FIG. 47 cooperating with a converter 4800. The sensed W value is provided by the converter 4800. The converter 4800 receives the digital value of the sensed W value and converts it to an analog signal that can be provided to the analog comparator 4700 using one or more of the converters described above (such as the pulse - voltage converter 3700, the digital data - voltage converter 4200, or the digital data - voltage converter 4300).

[0214] Figure 49 shows an output circuit 4900. It should be understood that there may still be a need to perform an activation function on the neuron output when the neuron output is digitized (such as by using the aforementioned integrating dual - slope ADCs 3400 or 3500). FIG. 49 shows an embodiment in which activation occurs before the neuron output is converted to a variable - width pulse or series of pulses. The output circuit 4900 includes an activation circuit 4901 and a current - pulse converter 4902. The activation circuit receives Ineuron values from various flash memory cells and generates Ineuron_act, which is the sum of the received Ineuron values. Next, the current - pulse converter 4902 converts Ineuron_act to a series of digital pulses and / or digital data representing the count of a series of digital pulses. Other converters described above (such as the integrating dual - slope ADCs 3400 or 3500) can be used instead of the converter 4902.

[0215] In another embodiment, activation may occur after the digital pulse is generated. In this embodiment, the digital output bits are mapped to a new set of digital bits using the functionality implemented by the activation mapping table or activation mapping unit 5010. Examples of such mapping are shown graphically in FIGS. 50 and 51. The activation digital mapping can simulate sigmoid, tanh, ReLu, or any activation function. Further, the activation digital mapping can quantize the output neurons.

[0216] FIG. 52 shows an example of a charge summing device 5200 that can be used to sum the outputs of the VMM during the verification operation after the programming operation to obtain a single analog value representing the output, and this single analog value can optionally be converted to a digital bit value. The charge summing device 5200 includes a current source 5201 and a sample-and-hold circuit including a switch 5202 and a sample-and-hold (S / H) capacitor 5203. As shown by way of example of a 4-bit digital value, there are 4 S / H circuits for holding the values from 4 evaluation pulses, and these values are summed at the end of the process. The S / H capacitor 5203 is selected at a ratio associated with the DINn bit position, for example, C_DIN3 = x8 Cu, C_DIN2 = x4 Cu, C_DIN1 = x2 Cu, DIN0 = x1 Cu. The current source 5201 is also multiplied by the corresponding ratio accordingly. * DINn bit position is selected at a ratio associated with the DINn bit position, for example, C_DIN3 = x8 Cu, C_DIN2 = x4 Cu, C_DIN1 = x2 Cu, DIN0 = x1 Cu. The current source 5201 is also multiplied by the corresponding ratio accordingly.

[0217] FIG. 53 shows a current summing device 5300 that can be used to sum the outputs of the VMM during the verification operation after the programming operation. The current summing device 5300 includes a current source 5301, switches 5302, 5303 and 5304, and a switch 5305. As shown by way of example of a 4-bit digital value, there are current source circuits for holding the values from 4 evaluation pulses, and these values are summed at the end of the process. The current source is 2^n *The ratio is multiplied based on the DINn bit position. For example, I_DIN3 = x8 Icell units, I_DIN2 = x4 Icell units, I_DIN1 = x2 Icell units, and I_DIN0 = x1 Icell unit.

[0218] Figure 54 shows a digital adder 5400 that receives a plurality of digital values, sums them together, and generates an output DOUT representing the sum of the inputs. The digital adder 5400 can be used during the verification operation after the programming operation. As shown in the example of 4-bit digital values, there are digital output bits for holding the values from four evaluation pulses, and these values are summed at the end of the process. The digital output is 2^n * It is digitally scaled based on the DINn bit position. For example, DOUT3 = x8 DOUT0, _DOUT2 = x4 DOUT1, I_DOUT1 = x2 DOUT0, and I_DOUT0 = DOUT0.

[0219] Figures 55A and 55B show a digital bit-pulse width converter 5500 used within an input block, row decoder, or output block. The pulse width output from the digital bit-pulse width converter 5500 is proportional to its value, as described above in relation to Figure 36B. The digital bit-pulse width converter includes a binary counter 5501. The state Q[N:0] of the binary counter 5501 can be loaded by serial data or parallel data within the load sequence. The row control logic 5510 outputs a voltage pulse having a pulse width proportional to the value of the digital data input provided from blocks such as the integrating ADCs of Figures 34 and 35.

[0220] Figure 55B shows the waveform of the output pulse width having a width proportional to the digital bit value. First, the data within the received digital bit is inverted, and the inverted digit bits are loaded into the counter 5501 either serially or in parallel. Then, the row pulse width is generated by the row control logic 5510 as shown in waveform 5520 by counting in binary until the maximum counter value is reached.

[0221] Optionally, a pulse train-pulse converter can be used to convert an output including a pulse train (such as signal 3411 or 3413 in FIG. 34B and signal 3513 or 3516 in FIG. 35B) into a single pulse (such as signals WL0, WL1, and WL in FIG. 36B) whose width varies in proportion to the number of pulses in the pulse train and is used as an input to the VMM array applied to the word line or control gate in the VMM array. An example of a pulse train-pulse converter is a binary counter having control logic. 2 An example is as shown in Table 9 for a 4-bit digital input.

[0222] An example is as shown in Table 9 for a 4-bit digital input. Table 9: Digital Input Bits to Output Pulse Width [Table 9]

[0223] Another embodiment uses an up binary counter and digital comparison logic. That is, the output pulse width is generated by counting up the up binary counter until the digital output of the binary counter becomes the same as the digital input bits.

[0224] Another embodiment uses a down binary counter. First, the down binary counter is loaded serially or in parallel with a digital data input pattern. Then, the output pulse width is generated by counting down the down binary counter until the digital output of the binary counter reaches the minimum value, i.e., the "0" logic state.

[0225] In another embodiment, the resolution of the analog-to-digital converter can be configured by a control signal. FIG. 60 shows a programmable ADC6000. The programmable ADC6000 receives an analog signal such as the output neuron current Ineu and converts it to an output 6002 with a set of digital bits. The programmable ADC600 receives a configuration 6001 that can be a set of analog control signals or digital control bits. In one example, the resolution of the output 6002 is determined by the configuration 6001. For example, if the configuration 6001 has a first value, the output can be a set of 4 bits, but if the configuration 6001 has a second value, the output can be a set of 8 bits.

[0226] A coarse level sensing circuit (not shown) can be used to sample the plurality of array output currents, and based on the value of this current, a gain (scaling factor) can be configured.

[0227] The gain can be configured for each specific neural network, and the gain can be set during neural network training for optimal performance.

[0228] In another embodiment, the ADC can be a hybrid of the architectures described above. For example, the first ADC can be a hybrid of a SAR ADC and a slope ADC. The second ADC may be a hybrid of a SAR ADC and a ramp ADC, and the third ADC may be a hybrid of an algorithmic ADC and a ramp ADC, without limitation.

[0229] FIG. 61A shows a hybrid output conversion block 6100. The output block 6100 receives differential signals IW+ and IW-. The successive approximation register ADC 6101 receives the differential signals IW+ and IW- and determines the upper digital bits (e.g., the most significant bits B7 to B4 in an 8-bit digital representation) that best correspond to the analog values represented by IW+ and IW-. When the SAR ADC 6101 determines those upper bits, an analog signal representing the result of subtracting the value represented by the upper bits from the signals IW+ and IW- is provided to a serial ADC block 6102 (such as a slope ADC or a ramp ADC), and then the serial ADC block 6102 determines the lower bits (e.g., the least significant bits B3 to B0 in an 8-bit digital representation) corresponding to that difference. Next, the upper bits and the lower bits are concatenated in series to generate a digital output representing the input signals IW+ and IW-.

[0230] FIG. 61B shows an output block 6110. The output block 6110 receives differential signals IW+ and IW-. The algorithmic ADC 6103 determines the upper bits (e.g., bits B7 to B4 in an 8-bit digital representation) corresponding to IW+ and IW-, and then the serial ADC block 6104 determines the lower bits (e.g., bits B3 to B0 in the case of an 8-bit digital representation).

[0231] FIG. 61C shows an output block 6120. The output block 6120 receives differential signals IW+ and IW-. The output block 6120 includes a hybrid ADC that converts the differential signals IW+ and IW- into digital bits by combining different conversion methods (such as those shown in FIGS. 61A and 61B) in one circuit.

[0232] FIG. 62 shows a configurable serial ADC 6200. The serial ADC 6200 includes an integrator 6270 that integrates the output neuron current Ineu onto an integration capacitor 6202 (Cint).

[0233] In one embodiment, VRAMP6250 is provided to the inverting input of comparator 6204. In this case, IREF6251 is off. The digital output (count value) 6221 is generated by ramping VRAMP6250 until comparator 6204 switches polarity, and counter 6220 counts clock pulses from the start of the ramp until comparator 6204 switches polarity, at which point counter 6220 provides the digital output (count value) 6221.

[0234] In another embodiment, VREF6255 is provided to the inverting input of comparator 6204. VOUT6203 is ramped down by ramp current 6251 (IREF) until VOUT6203 reaches VREF6255, at which point the EC6205 signal disables the count of counter 6220, at which point counter 6220 provides the digital output (count value) 6221. The (n-bit) ADC6200 can be configured to have lower accuracy (less than n bits) or higher accuracy (more than n bits) depending on the target application. The configurability of the accuracy is done by configuring, without limitation, the capacitance of capacitor 6202, current 6251 (IREF), the ramping speed of VRAMP6250, or the clock frequency of clock 6241.

[0235] In another embodiment, the ADC circuit of one VMM array is configured to have lower accuracy than n bits, and the ADC circuit of another VMM array is configured to have higher accuracy than n bits.

[0236] In another embodiment, one instance of the serial ADC circuit 6200 of one neuron circuit is combined with another instance of the serial ADC circuit 6200 of the next neuron circuit, such as by combining the integrating capacitors 6202 of the two instances of the serial ADC circuit 6200, to generate an ADC circuit having higher accuracy than n bits.

[0237] Figure 63 shows a configurable neuron SAR (successive approximation register) ADC6300. This circuit is a successive approximation converter based on charge redistribution using binary capacitors that converts the voltage input Vin to a digital output 6306 based on a reference voltage VREF. The ADC6300 includes a binary capacitor DAC (capacitor DAC, CDAC) 6301, an op-amp / comparator 6302, and SAR logic and registers 6303. As shown, GndV6304 is a low voltage reference level, for example, ground level. The SAR logic and registers 6303 provide the digital output 6306. Other non-binary capacitor structures can be implemented using weighted reference voltages or corrections by the output.

[0238] Figure 64 shows a pipelined SAR ADC circuit 6400 that can be used to increase the number of bits in a pipelined fashion in combination with the following SAR ADC. The SAR ADC circuit 6400 includes a binary CDAC6401, an op-amp / comparator 6402, an op-amp / comparator 6403, and SAR logic and registers 6404. As shown, GndV is a low voltage reference level, for example, ground level. The SAR logic and registers 6404 provide the digital output 6406. Vin is the input voltage and VREF is the reference voltage. V residue is generated by capacitor 6405 and provided as an input to the next stage of the SAR ADC conversion sequence.

[0239] Figure 65 shows a hybrid SAR+serial ADC circuit 6500 that can be used to increase the number of bits in a hybrid fashion. The SAR ADC circuit 6500 includes a binary CDAC6501, an op-amp / comparator 6502, and SAR logic and registers 6503. As shown, GndV is a low voltage reference level, for example, ground level during SAR ADC operation. The SAR logic and registers 6503 provide a digital output. Vin is the input voltage and VREF is the reference voltage. VREFRAMP is used as a reference ramp voltage during serial ADC operation instead of the GndV input to the op-amp / comparator 6502.

[0240] Other examples of hybrid ADC architectures include SAR ADC + sigma-delta ADC, flash ADC + serial ADC, pipelined ADC + serial ADC, serial ADC + SAR ADC, and other architectures.

[0241] FIG. 66 shows an algorithmic ADC output block 6600. The output block 6600 includes a sample-and-hold circuit 6601 configured as shown and 、a an analog-to-digital converter 6602, a digital-to-analog converter 6603, an integrator 6604, an operational amplifier 6605, and control switches 6606 and 6607.

[0242] In another embodiment, the sample-and-hold circuit is used for input to each row in the VMM array. For example, if the input includes a DAC, the DAC can include a sample-and-hold circuit.

[0243] As used herein, it should be noted that both the terms "over" and "on" include both "directly on" (with no intervening material, element, or gap therebetween) and "indirectly on" (with an intervening material, element, or gap therebetween). Similarly, the term "adjacent" includes "directly adjacent" (with no intervening material, element, or gap therebetween) and "indirectly adjacent" (with an intervening material, element, or gap therebetween), "attached to" includes "directly attached to" (with no intervening material, element, or gap therebetween) and "indirectly attached to" (with an intervening material, element, or gap therebetween), and "electrically coupled" includes "directly electrically coupled" (with no intervening material or element electrically connecting the elements together therebetween) and "indirectly electrically coupled" (with an intervening material or element electrically connecting the elements together therebetween). For example, forming an element "over a substrate" can include forming the element directly on the substrate without an intervening material / element therebetween, and forming the element indirectly on the substrate with one or more intervening material / elements therebetween.

Claims

1. A programmable neuron output block for generating an output of a neural network memory array, comprising: one or more input nodes for receiving current from the neural network memory array; a gain configuration circuit for receiving a gain configuration signal and applying a gain factor to the received current in response to the gain configuration signal to generate an output; wherein the gain configuration signal depends on (i) the number of rows enabled in the neural network memory array where the current is received, or (ii) all values input to the rows enabled in the neural network memory array that generated the current, the programmable neuron output block.

2. The programmable neuron output block according to claim 1, wherein the gain configuration signal includes an analog signal.

3. The programmable neuron output block according to claim 1, wherein the gain configuration signal includes digital bits.

4. The programmable neuron output block according to claim 1, wherein the gain configuration circuit includes a variable resistor controlled by the gain configuration signal.

5. The programmable neuron output block according to claim 1, wherein the gain configuration circuit includes a variable capacitor controlled by the gain configuration signal.

6. A programmable neuron output block for generating an output of a neural network memory array, comprising: an analog-to-digital converter for receiving current and a control signal from the neural network memory array and generating a digital output, wherein the resolution of the digital output is determined by the control signal, the analog-to-digital converter, the programmable neuron output block.

7. The programmable neuron output block according to claim 6, wherein the control signal includes an analog signal.

8. The programmable neuron output block according to claim 6, wherein the control signal includes digital bits.

9. A programmable neuron output block for generating an output of a neural network memory array, comprising: a hybrid analog-to-digital converter for converting the output of the neural network memory array into a digital output. A first analog-to-digital converter for generating a first portion of the digital output, wherein the first analog-to-digital converter includes an algorithmic analog-to-digital converter, the first analog-to-digital converter; A second analog-to-digital converter for generating a second portion of the digital output, a hybrid analog-to-digital converter comprising: A programmable neuron output block comprising: **Claim 10** The programmable neuron output block according to claim 9, wherein the second analog-to-digital converter includes a serial analog-to-digital converter.

Citation Information

Patent Citations

  • Object control system using neural circuit network arithmetic unit and evaluating method using neural circuit network arithmetic unit for controlled system

    JP1996194678A

  • Programmable neurons for analog non-volatile memory in deep learning artificial neural networks

    JP2021509514A

  • Methods of operating memory

    US20160358661A1

  • Artificial neural networks

    US20200327402A1

  • Configurable input blocks and output blocks and physical layout for analog neural memory in deep learning artificial neural network

    US20200349421A1