Sigma-delta analog-to-digital converter for generating digital output from a vector-matrix multiplication array

CN122804235APending Publication Date: 2026-09-22SILICON STORAGE TECHNOLOGY INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202480085167.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-01-29
Filing Date
2024-02-21
Publication Date
2026-09-22

Smart Images

  • Figure CN122804235A_ABST
    Figure CN122804235A_ABST
Patent Text Reader

Abstract

In one example, a system includes: a vector-matrix multiplication array comprising an array of non-volatile memory cells arranged in rows and columns; and a Σ-Δ analog-to-digital converter for receiving current from columns of the vector-matrix multiplication array and generating digital outputs in response to the current.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Priority Statement

[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 601,049, filed November 20, 2023, entitled “Output Circuit for a Vector-by-Matrix Multiplication Array,” and U.S. Patent Application No. 18 / 426,071, filed January 29, 2024, entitled “Sigma-Delta Analog-to-Digital Converter to Generate Digital Output from Vector-by-Matrix Multiplication Array.” Technical Field

[0003] Numerous examples of Σ-Δ analog-to-digital converters and associated methods for generating digital outputs from vector-matrix multiplication arrays are disclosed. Background Technology

[0004] Artificial neural networks mimic biological neural networks (the central nervous system of animals, especially the brain) and are used to estimate or approximate functions that can depend on a large number of inputs and are often unknown. Artificial neural networks typically consist of interconnected layers of "neurons" that exchange messages with each other.

[0005] Figure 1 An artificial neural network is illustrated, where circles represent the inputs or layers of neurons. Connections (called synapses) are indicated by arrows and have numerical weights that can be tuned empirically. This allows the neural network to adapt to its inputs and learn. Typically, a neural network consists of layers with multiple inputs. There are usually one or more intermediate layers of neurons, and an output layer of neurons that provide the output of the neural network. Neurons at each level make decisions individually or collectively based on data received from the synapses.

[0006] One of the major challenges in developing artificial neural networks for high-performance information processing is the lack of sufficient hardware technology. Real-world neural networks rely on a large number of synapses to achieve high connectivity between neurons, i.e., very high computational parallelism. In principle, this complexity can be achieved using digital supercomputers or dedicated clusters of graphics processing units. However, compared to biological networks, these methods are generally energy inefficient, in addition to being costly, as biological networks consume less energy primarily due to their ability to perform low-precision analog calculations. CMOS analog circuits have been used in artificial neural networks, but given the large number of neurons and synapses, most CMOS-implemented synapses are excessively large.

[0007] The applicant previously disclosed an artificial (simulated) neural network utilizing one or more non-volatile memory arrays as synapses in U.S. Patent Application Publication 2017 / 0337466A1, which is incorporated herein by reference. The non-volatile memory array operates as a simulated neural memory and includes non-volatile memory cells arranged in rows and columns. The neural network includes a plurality of first synapses configured to receive a first plurality of inputs and generate a first plurality of outputs therefrom, and a plurality of first neurons configured to receive the first plurality of outputs. The first plurality of synapses includes a plurality of memory cells, wherein each memory cell includes: spaced-apart source and drain regions formed in a semiconductor substrate, wherein a channel region extends between the source and drain regions; a floating gate disposed over and insulated from a first portion of the channel region; and a non-floating gate disposed over and insulated from a second portion of the channel region. Each memory cell stores weight values ​​corresponding to a plurality of electrons on the floating gate. The plurality of memory cells multiply the first plurality of inputs by the stored weight values ​​to generate the first plurality of outputs.

[0008] Non-volatile memory cells

[0009] Non-volatile memory is well known. For example, U.S. Patent 5,029,130 ​​(“the '130 Patent”), which is incorporated herein by reference, discloses an array of split-gate non-volatile memory cells, which is a type of flash memory cell. Figure 2The image shows such a memory cell 210. Each memory cell 210 includes a source region 14 and a drain region 16 formed in a semiconductor substrate 12, with a channel region 18 between the source and drain regions. A floating gate 20 is formed over and insulated from (and controls the conductivity of) a first portion of the channel region 18, and is formed over a portion of the source region 14. A word line terminal 22 (which is typically coupled to a word line) has a first portion disposed over and insulated from (and controlling the conductivity of) a second portion of the channel region 18, and a second portion extending upward and located over the floating gate 20. The floating gate 20 and the word line terminal 22 are insulated from the substrate 12 by a gate oxide. A bit line 24 is coupled to the drain region 16.

[0010] The memory cell 210 is erased by applying a high positive voltage to the word line terminal 22 (where electrons are removed from the floating gate), which causes electrons on the floating gate 20 to tunnel from the floating gate 20 to the word line terminal 22 through the intermediate insulator via the Fowler-Nordheim (FN).

[0011] Memory cell 210 is programmed via source-side injection (SSI) with hot electrons (where electrons are placed on the floating gate) by applying a positive voltage to word line terminal 22 and a positive voltage to source region 14. Electron flow occurs from drain region 16 to source region 14. As electrons reach the gap between word line terminal 22 and floating gate 20, they accelerate and become hot. Due to electrostatic attraction from floating gate 20, some of the heated electrons are injected onto floating gate 20 through the gate oxide.

[0012] Memory cell 210 is read by applying a positive read voltage to the drain region 16 and word line terminal 22 (which conducts the portion of channel region 18 below the word line terminal). If the floating gate 20 is positively charged (i.e., electrons are erased), the portion of channel region 18 below the floating gate 20 is also conducted, and current flows through channel region 18, which is sensed as an erased state or a "1" state. If the floating gate 20 is negatively charged (i.e., programmed electronically), the portion of channel region below the floating gate 20 is mostly or completely turned off, and current does not flow (or very little current) through channel region 18, which is sensed as a programmed state or a "0" state.

[0013] Table 1 depicts the typical voltage and current ranges that can be applied to the terminals of memory cell 210 to perform read, erase, and program operations:

[0014] Table 1: Figure 2 Operation of flash memory cell 210

[0015]

[0016] Other split-gate memory cell configurations as other types of flash memory cells are known. For example, Figure 3 A four-gate memory cell 310 is depicted, comprising a source region 14, a drain region 16, a floating gate 20 over a first portion of a channel region 18, a select gate 22 (typically coupled to a word line WL) over a second portion of the channel region 18, a control gate 28 over the floating gate 20, and an erase gate 30 over the source region 14. This configuration is described in U.S. Patent 6,747,310, which is incorporated herein by reference for all purposes. Here, all gates except the floating gate 20 are non-floating gates, meaning they are electrically connected to or can be electrically connected to a voltage source. Programming is performed by heated electrons from the channel region 18 that inject themselves into the floating gate 20. Erasing is performed by electrons tunneling from the floating gate 20 to the erase gate 30.

[0017] Table 2 depicts the typical voltage and current ranges that can be applied to the terminals of memory cell 310 to perform read, erase, and program operations:

[0018] Table 2: Figure 3 Operation of flash memory cell 310

[0019]

[0020] Figure 4 A tri-gate memory cell 410 is depicted, which is another type of flash memory cell. Memory cell 410 and... Figure 3 The memory cell 310 is the same as the memory cell 410, except that the memory cell 410 does not have a separate control gate. Except that no control gate bias is applied, the erase operation (thus erasing is performed using an erase gate) and read operation are the same as... Figure 3 The operation is similar. Programming is also performed without a control gate bias, and therefore, a higher voltage is applied to the source line during programming to compensate for the lack of a control gate bias.

[0021] Table 3 depicts the typical voltage and current ranges that can be applied to the terminals of memory cell 410 to perform read, erase, and program operations:

[0022] Table 3: Figure 4 Operation of flash memory cell 410

[0023]

[0024] Figure 5A stacked-gate memory cell 510 is depicted, which is another type of flash memory cell. Except that the floating gate 20 extends over the entire channel region 18 and the control gate 22 (which here will be coupled to the word line) extends over the floating gate 20 and is separated by an insulating layer (not shown), the memory cell 510... Figure 2 The memory cell 210 is similar. Erasure is performed by electron tunneling from FG to FN in the substrate, and programming is performed by channel hot electron (CHE) injection in the region between the channel 18 and the drain region 16, by electron flow from the source region 14 to the drain region 16, and by a read operation similar to that used for a read operation for a memory cell 210 with a higher control gate voltage.

[0025] Table 4 depicts the typical voltage range that can be applied to the terminals of memory cell 510 and substrate 12 to perform read, erase, and program operations:

[0026] Table 4: Figure 5 Operation of flash memory cell 510

[0027]

[0028] The methods and components described herein can be applied to other non-volatile memory technologies, such as, but not limited to, FINFET split-gate flash or stacked-gate flash memory, NAND flash memory, SONOS (silicon-oxide-nitride-oxide-silicon with charge trapped in nitride), MONOS (metal-oxide-nitride-oxide-silicon with metal charge trapped in nitride), ReRAM (resistive RAM), PCM (phase-change memory), MRAM (magnetic RAM), FeRAM (ferroelectric RAM), CT (charge-trapping) memory, CN (carbon nanotube) memory, OTP (two-level or multi-level one-time programmable) and CeRAM (associated electron RAM), etc.

[0029] To utilize memory arrays comprising one of the aforementioned types of non-volatile memory cells in artificial neural networks, two modifications were made. First, the circuitry was configured such that each memory cell could be individually programmed, erased, and read without adversely affecting the memory state of other memory cells in the array, as explained further below. Second, continuous (simulated) programming of the memory cells was provided.

[0030] Specifically, the memory state (i.e., the charge on the floating gate) of each memory cell in the array can be continuously changed from a fully erased state to a fully programmed state and vice versa, independently and with minimal interference to other memory cells. This means that the cell storage device is essentially analog, or at least can store one discrete value from many discrete values ​​(such as 16 or 64 different values). This allows for very precise and individual tuning of all memory cells in the memory array, making the memory array ideal for both storage and fine-tuning of synaptic weights in neural networks.

[0031] Neural networks using non-volatile memory cell arrays

[0032] Figure 6 This conceptual example illustrates a non-limiting example of a neural network utilizing a non-volatile memory array. This example uses a non-volatile memory array neural network for a facial recognition application, but any other suitable application can also be implemented using a neural network based on a non-volatile memory array.

[0033] In this example, S0 is the input layer, which is a 32x32 pixel RGB image with 5-bit precision (i.e., three 32x32 pixel arrays, one for each color R, G, and B, with 5-bit precision per pixel). The synapse CB1 from the input layer S0 to layer C1 applies different sets of weights in some cases and shared weights in others, and scans the input image with a 3x3 pixel overlapping filter (kernel), shifting the filter by one pixel (or more than one pixel as indicated by the model). Specifically, the values ​​of nine pixels in a 3x3 portion of the image (i.e., called the filter or kernel) are provided to synapse CB1, where these nine input values ​​are multiplied by appropriate weights, and after summing the output of this multiplication, a single output value is determined by the first synapse of CB1 and provided for generating one of the pixels in the feature map of layer C1. The 3x3 filter is then shifted one pixel to the right within the input layer S0 (i.e., adding a column of three pixels to the right and releasing a column of three pixels to the left), thereby providing the nine pixel values ​​from this newly positioned filter to synapse CB1, where they are multiplied by the same weights and a second single output value is determined by the associated synapse. This process continues until the 3x3 filter scans all three colors and all bits (precision values) across the entire 32x32 pixel image of the input layer S0. This process is then repeated using different sets of weights to generate different feature maps for layer C1 until all feature maps for layer C1 are computed.

[0034] At layer C1, in this example, there are 16 feature maps, each with 30x30 pixels. Each pixel is a new feature pixel extracted from the product of the input and the kernel, and therefore each feature map is a two-dimensional array, and thus in this example, layer C1 consists of a 16-layer two-dimensional array (remember that the layers and arrays referred to in this article are logical relationships and may not correspond to physical relationships—that is, the array may not be oriented in a physical two-dimensional array). Each of the 16 feature maps in layer C1 is generated by a set of sixteen different groups of synaptic weights applied to the filter scan. The C1 feature maps may all relate to different aspects of the same image feature, such as boundary identification. For example, the first map (generated using a first weight recombination, shared for all scans used to generate the first map) may identify circular edges, the second map (generated using a second weight recombination different from the first weight recombination) may identify rectangular edges, or the aspect ratio of some feature, and so on.

[0035] Before transitioning from layer C1 to layer S1, activation function P1 (pooling) is applied, which pools the values ​​from consecutive non-overlapping 2x2 regions in each feature map. The purpose of pooling function P1 is to average the neighboring locations (or, alternatively, use a max function) to reduce, for example, the dependence on edge locations and reduce the data size before moving to the next stage. At layer S1, there are 16 15x15 feature maps (i.e., sixteen distinct arrays, each with 15x15 pixels). The synapse CB2 from layer S1 to layer C2 scans the map in layer S1 using a 4x4 filter, where the filter is shifted by 1 pixel. At layer C2, there are 22 12x12 feature maps. Before transitioning from layer C2 to layer S2, activation function P2 (pooling) is applied, which pools the values ​​from consecutive non-overlapping 2x2 regions in each feature map. At layer S2, there are 22 6x6 feature maps. An activation function (pooling) is applied to the synapse CB3 from layer S2 to layer C3, where each neuron in layer C3 is connected to each mapping in layer S2 via a corresponding synapse in CB3. There are 64 neurons in layer C3. The synapse CB4 from layer C3 to the output layer S3 completely connects C3 to S3, meaning each neuron in layer C3 is connected to every neuron in layer S3. The output at S3 comprises 10 neurons, with the highest-output neuron determining the class. For example, this output could indicate an identification or classification of the content of the original image.

[0036] Synapses for each layer are implemented using an array or a portion of an array of non-volatile memory cells.

[0037] Figure 7 This is a block diagram of an array that could be used for this purpose. The vector-matrix multiplication (VMM) array 32 includes non-volatile memory cells and serves as synapses between layers (such as...) Figure 6 (CB1, CB2, CB3, and CB4 in the original text). Specifically, the VMM array 32 includes an array 33 of non-volatile memory cells, an erase gate and word line gate decoder 34, a control gate decoder 35, a bit line decoder 36, and a source line decoder 37, which decode the corresponding inputs of the non-volatile memory cell array 33. The inputs to the VMM array 32 may come from the erase gate and word line gate decoder 34 or from the control gate decoder 35. In this example, the source line decoder 37 also decodes the outputs of the non-volatile memory cell array 33. Alternatively, the bit line decoder 36 may decode the outputs of the non-volatile memory cell array 33.

[0038] The non-volatile memory cell array 33 serves two purposes. First, it stores weights that will be used by the VMM array 32. Second, the non-volatile memory cell array 33 efficiently multiplies the inputs with the weights stored in the non-volatile memory cell array 33 and adds them together on each output line (source line or bit line) to produce an output that will be used as the input to the next layer or the final layer. By performing multiplication and addition functions, the non-volatile memory cell array 33 eliminates the need for separate multiplication and addition logic circuits and is also highly efficient due to its in-situ memory computation.

[0039] The output of the non-volatile memory cell array 33 is provided to a differential summer (such as a summing operational amplifier or a summing current mirror) 38, which sums the output of the non-volatile memory cell array 33 to create a single value for the convolution. The differential summer 38 is arranged to perform the summation of positive and negative weights.

[0040] The summed output of the difference summer 38 is then provided to the activation function block 39, which modifies the output. Activation function block 39 can provide a sigmoid, tanh, or ReLU function. The modified output of activation function block 39 becomes the next layer's output (e.g., ...). Figure 6 The elements of the feature map of layer C1 are then applied to the next synapse to produce the next feature map layer or the final layer. Thus, in this example, the non-volatile memory cell array 33 constitutes multiple synapses (which receive their input from existing neuron layers or from input layers such as an image database), and the summing operational amplifier 38 and the activation function block 39 constitute multiple neurons.

[0041] Figure 7The inputs to the VMM array 32 (WLx, EGx, CGx and optional BLx and SLx) can be analog, binary or digital (in which case a DAC is provided to convert the digital bits to the appropriate input analog level), and the outputs can be analog, binary or digital (in which case an output ADC is provided to convert the output analog level to digital bits).

[0042] Figure 8 A block diagram illustrating the use of the multilayer VMM array 32, labeled VMM arrays 32a, 32b, 32c, 32d, and 32e. (See diagram for reference.) Figure 8 As shown, the input (denoted as Inputx) is converted from digital to analog by a digital-to-analog converter 31 and provided to the input VMM array 32a. The converted analog input can be voltage or current. The first-level input D / A conversion can be accomplished by using a function or LUT (lookup table) of appropriate analog level to map Inputx to the input VMM array 32a matrix multiplier. Input conversion can also be accomplished by an analog-to-analog (A / A) converter to convert the external analog input into a mapped analog input to the input VMM array 32a.

[0043] The output generated by the input VMM array 32a is provided as input to the next VMM array (hidden level 1) 32b, which in turn generates the output provided as input to the next VMM array (hidden level 2) 32c, and so on. The various layers of the VMM array 32 serve as different layers of synapses and neurons in a convolutional neural network (CNN). Each VMM array 32a, 32b, 32c, 32d, and 32e can be an independent physical non-volatile memory array, or multiple VMM arrays can utilize different portions of the same physical non-volatile memory array, or multiple VMM arrays can utilize overlapping portions of the same physical non-volatile memory array. Figure 8 The example shown contains five layers (32a, 32b, 32c, 32d, 32e): one input layer (32a), two hidden layers (32b, 32c), and two fully connected layers (32d, 32e). Those skilled in the art will understand that this is merely an example, and conversely, a system may include more than two hidden layers and more than two fully connected layers.

[0044] Vector-matrix multiplication (VMM) array

[0045] Figure 9 A neuronal VMM array 900 is depicted, which is particularly suitable for... Figure 3The memory cell 310 shown serves as a synapse and partial neuron between the input layer and the next layer. The VMM array 900 includes a memory array 901 of non-volatile memory cells and a reference array 902 of non-volatile reference memory cells (at the top of the array). Alternatively, another reference array may be placed at the bottom.

[0046] In the VMM array 900, control gate lines (such as control gate line 903) extend vertically (therefore, reference array 902 is orthogonal to control gate line 903 in the row direction), and erase gate lines (such as erase gate line 904) extend horizontally. Here, the inputs to the VMM array 900 are set on control gate lines (CG0, CG1, CG2, CG3), and the outputs of the VMM array 900 appear on source lines (SL0, SL1). In one example, even rows are used, and in another example, odd rows are used. The currents placed on each source line (SL0, SL1, respectively) perform a summation function of all currents from the memory cells connected to that particular source line.

[0047] As described herein with respect to neural networks, the non-volatile memory cells of the VMM array 900 (i.e., memory cells 310 of the VMM array 900) can be configured to operate in a region below a threshold.

[0048] Bias the non-volatile reference memory cell and non-volatile memory cell described in this paper in the weak inversion (below the threshold region):

[0049] Ids = Io e (Vg-Vth) / nVt = w Io e (Vg) / nVt ,

[0050] Where w = e (- Vth) / nVt

[0051] Where Ids is the drain-to-source current; Vg is the gate voltage on the memory cell; Vth is the threshold voltage of the memory cell; and Vt is the thermal voltage = k T / q, where k is Boltzmann's constant, T is the temperature in Kelvin, and q is the electron charge; n is the slope factor = 1 + (Cdep / Cox), where Cdep = the capacitance of the depletion layer, and Cox is the capacitance of the gate oxide layer; Io is the memory cell current at the gate voltage equal to the threshold voltage, and Io is related to (Wt / L). u Cox (n-1) Vt 2Proportional, where u is the carrier mobility, and Wt and L are the width and length of the memory cell, respectively.

[0052] For I-to-V logarithmic converters that use memory cells (such as reference memory cells or peripheral memory cells) or transistors to convert input current into input voltage:

[0053] Vg= n Vt log [Ids / wp Io]

[0054] Where wp represents the w in the reference memory cell or the peripheral memory cell.

[0055] For a memory array used as a vector matrix multiplier (VMM) array with current input, the output current is:

[0056] Iout = wa Io e (Vg)nVt ,Right now

[0057] Iout = (wa / wp) Iin = W Iin

[0058] W = e (Vthp - Vtha) / nVt

[0059] Here, wa = w for each memory cell in the memory array.

[0060] Vthp is the effective threshold voltage of the peripheral memory cell, and Vtha is the effective threshold voltage of the primary (data) memory cell. Note that the threshold voltage of the transistor is a function of the substrate bulk bias voltage, and the substrate bulk bias voltage, denoted as Vsb, can be modulated to compensate for various conditions at this temperature. The threshold voltage Vth can be expressed as:

[0061] Vth = Vth0 + γ(SQRT |Vsb – 2 F) - SQRT |2 F |)

[0062] Where Vth0 is the threshold voltage with zero substrate bias. F is the surface potential, and γ is the host effect parameter.

[0063] Word lines or control gates can be used as inputs to memory cells that accept input voltages.

[0064] Alternatively, the flash memory cells of the VMM array described herein can be configured to operate in a linear region:

[0065] Ids = β (Vgs-Vth) Vds;β = u Cox Wt / L

[0066] W = α (Vgs-Vth)

[0067] This means that the weight W in the linear region is proportional to (Vgs-Vth).

[0068] Word lines, control gates, bit lines, or source lines can be used as inputs to memory cells operating in a linear region. Bit lines or source lines can be used as outputs to memory cells.

[0069] For an IV linear converter, memory cells (such as reference memory cells or peripheral memory cells) or transistors operating in the linear region can be used to linearly convert input / output current into input / output voltage.

[0070] Alternatively, the memory cells of the VMM array described herein can be configured to operate in a saturation region:

[0071] Ids = ½ β (Vgs-Vth) 2 ;β=u Cox Wt / L

[0072] Wα (Vgs-Vth) 2 This means that the weight W is related to (Vgs-Vth). 2 proportional

[0073] Word lines, control gates, or erase gates can be used as inputs to memory cells operating in saturation regions. Bit lines or source lines can be used as outputs of output neurons.

[0074] Alternatively, the memory cells of the VMM array described herein can be used in all regions or combinations thereof (below threshold, linear or saturated regions) of each or more layers of a neural network.

[0075] It is described in U.S. Patent No. 10,748,630 Figure 7 Other examples of VMM array 32 are described herein by reference. As described in this application, source lines or bit lines can be used as neuron outputs (current summation outputs).

[0076] Figure 10A neuronal VMM array 1000 is depicted, which is particularly suitable for... Figure 2 The memory cell 210 shown serves as a synapse between the input layer and the next layer. The VMM array 1000 includes a memory array 1003 of non-volatile memory cells, a reference array 1001 of first non-volatile reference memory cells, and a reference array 1002 of second non-volatile reference memory cells. The reference arrays 1001 and 1002, arranged in the column direction of the array, are used to convert current inputs flowing into terminals BLR0, BLR1, BLR2, and BLR3 into voltage inputs WL0, WL1, WL2, and WL3. In practice, the first and second non-volatile reference memory cells are diode-connected via a multiplexer 1014 (partially depicted), into which current inputs flow. The reference cells are tuned (e.g., programmed) to a target reference level. The target reference level is provided by a reference microarray matrix (not shown).

[0077] Memory array 1003 serves two purposes. First, it stores the weights used by VMM array 1000 on their respective memory cells. Second, memory array 1003 efficiently multiplies the inputs (i.e., the current inputs provided in terminals BLR0, BLR1, BLR2, and BLR3, which reference arrays 1001 and 1002 convert into input voltages to supply word lines WL0, WL1, WL2, and WL3) by the weights stored in memory array 1003, and then adds all the results (memory cell currents) to produce an output on the corresponding bit lines (BL0 to BLN), which will be the input to the next layer or to the final layer. By performing multiplication and addition functions, memory array 1003 eliminates the need for separate multiplication and addition logic circuits and is also highly efficient. Here, voltage inputs are provided on word lines WL0, WL1, WL2, and WL3, and the outputs appear on the corresponding bit lines BL0 to BLN during read (inference) operations. The current placed on each of the bit lines BL0 to BLN performs a summation function of the currents from all non-volatile memory cells connected to that particular bit line.

[0078] Table 5 depicts the operating voltages and currents used for the VMM array 1000. The columns in the table indicate the voltages applied to the word lines for selected cells, word lines for unselected cells, bit lines for selected cells, bit lines for unselected cells, source lines for selected cells, and source lines for unselected cells. The rows indicate read, erase, and program operations.

[0079] Table 5: Figure 10 Operation of VMM array 1000 :

[0080]

[0081] Figure 11 A neuronal VMM array 1100 is depicted, which is particularly suitable for... Figure 2 The memory cell 210 shown serves as a synapse and partial neuron between the input layer and the next layer. The VMM array 1100 includes a memory array 1103 of non-volatile memory cells, a reference array 1101 of first non-volatile reference memory cells, and a reference array 1102 of second non-volatile reference memory cells. Reference arrays 1101 and 1102 extend in the row direction of the VMM array 1100. The VMM array is similar to VMM 1000, except that in VMM array 1100, word lines extend in the vertical direction. Here, inputs are set on word lines (WLA0, WLB0, WLA1, WLB2, WLA2, WLB2, WLA3, WLB3), and outputs appear on source lines (SL0, SL1) during read operations. The currents placed on each source line perform a summation function of all currents from the memory cells connected to that particular source line.

[0082] Table 6 depicts the operating voltages and currents used for the VMM array 1100. The columns in the table indicate the voltages applied to the word lines for selected cells, word lines for unselected cells, bit lines for selected cells, bit lines for unselected cells, source lines for selected cells, and source lines for unselected cells. The rows indicate read, erase, and program operations.

[0083] Table 6: Figure 11 Operation of VMM array 1100

[0084]

[0085] Figure 12 A neuronal VMM array 1200 was depicted, which is particularly suitable for... Figure 3The memory cell 310 shown serves as a synapse and partial neuron between the input layer and the next layer. The VMM array 1200 includes a memory array 1203 of non-volatile memory cells, a reference array 1201 of first non-volatile reference memory cells, and a reference array 1202 of second non-volatile reference memory cells. Reference arrays 1201 and 1202 are used to convert current inputs flowing into terminals BLR0, BLR1, BLR2, and BLR3 into voltage inputs CG0, CG1, CG2, and CG3. In practice, the first and second non-volatile reference memory cells are diode-connected via a multiplexer 1212 (partially shown), through which current inputs flow via BLR0, BLR1, BLR2, and BLR3. Each multiplexer 1212 includes a corresponding multiplexer 1205 and a common-source cascode transistor 1204 to ensure that the voltage on the bit lines (such as BLR0) of each of the first and second non-volatile reference memory cells remains constant during read operations. The reference cells are tuned to a target reference level.

[0086] Memory array 1203 serves two purposes. First, it stores weights that will be used by VMM array 1200. Second, memory array 1203 efficiently multiplies the inputs (current inputs provided to terminals BLR0, BLR1, BLR2, and BLR3, which reference arrays 1201 and 1202 convert into input voltages to be provided to the control gates (CG0, CG1, CG2, and CG3)) by the weights stored in the memory array, and then sums all the results (cell currents) to produce an output that appears on BL0-BLN and will be the input to the next layer or the final layer. By performing multiplication and addition functions, the memory array eliminates the need for separate multiplication and addition logic circuits and is also highly efficient. Here, the inputs are provided on the control gate lines (CG0, CG1, CG2, and CG3), and the outputs appear on the bit lines (BL0-BLN) during read operations. The currents placed on each bit line perform a summation function of all the currents from the memory cells connected to that particular bit line.

[0087] The VMM array 1200 performs unidirectional tuning of the non-volatile memory cells in the memory array 1203. That is, each non-volatile memory cell is erased and then partially programmed until the desired charge is reached on the floating gate. If too much charge is placed on the floating gate (causing an incorrect value to be stored in the cell), the cell is erased, and the sequence of partial programming operations restarts. As shown, two rows sharing the same erase gate (such as EG0 or EG1) are erased together (this may be referred to as page erasure), and thereafter, each cell is partially programmed until the desired charge is reached on the floating gate.

[0088] Table 7 depicts the operating voltages and currents used for the VMM array 1200. The columns in the table indicate the voltages applied to the word lines for selected cells, word lines for unselected cells, bit lines for selected cells, bit lines for unselected cells, control gates for selected cells, control gates for unselected cells in the same sector as the selected cell, control gates for unselected cells in different sectors from the selected cell, erase gates for selected cells, erase gates for unselected cells, source lines for selected cells, and source lines for unselected cells. Rows indicate read, erase, and program operations.

[0089] Table 7: Figure 12 Operation of VMM array 1200

[0090]

[0091] Figure 13 A neuronal VMM array 1300 was described, which is particularly suitable for... Figure 3 The memory cell 310 shown serves as a synapse and partial neuron between the input layer and the next layer. The VMM array 1300 includes a memory array 1303 of non-volatile memory cells, a reference array 1301 of first non-volatile reference memory cells, and a reference array 1302 of second non-volatile reference memory cells. EG lines EGR0, EG0, EG1, and EGR1 extend vertically, while CG lines CG0, CG1, CG2, and CG3 and SL lines WL0, WL1, WL2, and WL3 extend horizontally. The VMM array 1300 is similar to the VMM array 1400, except that the VMM array 1300 implements bidirectional tuning, where each individual cell can be completely erased, partially programmed, and partially erased as needed due to the use of separate EG lines to achieve a desired amount of charge on the floating gate. As shown in the figure, reference arrays 1301 and 1302 convert the input currents in terminals BLR0, BLR1, BLR2, and BLR3 into control gate voltages CG0, CG1, CG2, and CG3 to be applied to memory cells in the row direction (through the operation of reference cells connected via diodes of multiplexer 1314). The current outputs (neurons) are in bit lines BL0-BLN, where each bit line sums all currents from non-volatile memory cells connected to that particular bit line.

[0092] Table 8 depicts the operating voltages and currents used for the VMM array 1300. The columns in the table indicate the voltages applied to the word lines for selected cells, word lines for unselected cells, bit lines for selected cells, bit lines for unselected cells, control gates for selected cells, control gates for unselected cells in the same sector as the selected cell, control gates for unselected cells in different sectors from the selected cell, erase gates for selected cells, erase gates for unselected cells, source lines for selected cells, and source lines for unselected cells. Rows indicate read, erase, and program operations.

[0093] Table 8: Figure 13 Operation of VMM array 1300

[0094]

[0095] Figure 22 A neuronal VMM array 2200 was depicted, which is particularly suitable for... Figure 2 The memory unit 210 shown serves as a synapse and partial neuron between the input layer and the next layer. In the VMM array 2200, inputs INPUT0, ..., INPUT... N On bit lines BL0, ... BL respectively N The receiver receives data and outputs OUTPUT1, OUTPUT2, OUTPUT3 and OUTPUT4, which are generated on source lines SL0, SL1, SL2 and SL3, respectively.

[0096] Figure 23 A neuronal VMM array 2300 was depicted, which is particularly suitable for... Figure 2 The memory unit 210 shown serves as a synapse and partial neuron between the input layer and the next layer. In this example, the inputs are INPUT0 and INPUT... 1、 INPUT2 and INPUT3 receive signals on source lines SL0, SL1, SL2, and SL3 respectively, and output OUTPUT0, ..., OUTPUT. N In position lines BL0, ..., BL N Generate above.

[0097] Figure 24 The neuronal VMM array 2400 is depicted, which is particularly suitable for... Figure 2 The memory unit 210 shown serves as a synapse and partial neuron between the input layer and the next layer. In this example, the inputs are INPUT0..., INPUT... M On the word lines WL0... and WL respectively M The data is received and outputs OUTPUT0, ..., OUTPUT. N In position lines BL0, ..., BLN Generate above.

[0098] Figure 25 A neuronal VMM array 2500 was depicted, which is particularly suitable for... Figure 3 The memory unit 310 shown serves as a synapse and partial neuron between the input layer and the next layer. In this example, the inputs are INPUT0, ..., INPUT0. M On the word lines WL0, ..., WL respectively M The data is received and outputs OUTPUT0, ..., OUTPUT. N In position lines BL0, ..., BL N Generate above.

[0099] Figure 26 The neuronal VMM array 2600 is depicted, which is particularly suitable for... Figure 4 The memory unit 410 shown serves as a synapse and partial neuron between the input layer and the next layer. In this example, the inputs are INPUT0, ..., INPUT0. n On the vertical control grid lines CG0, ..., CG respectively N The data is received, and outputs OUTPUT1 and OUTPUT2 are generated on source lines SL0 and SL1.

[0100] Figure 27 The neuronal VMM array 2700 is depicted, which is particularly suitable for... Figure 4 The memory unit 410 shown serves as a synapse and partial neuron between the input layer and the next layer. In this example, the inputs are INPUT0, ..., INPUT0. N They are received on the gates of bit line control gates 2701-1, 2701-2, ..., 2701-(N-1) and 2701-N, respectively, and these gates are coupled to bit lines BL0, ..., BL0, respectively. N Example outputs OUTPUT1 and OUTPUT2 are generated on source lines SL0 and SL1.

[0101] Figure 28 A neuronal VMM array 2800 was depicted, which is particularly suitable for... Figure 3 The memory unit 310 shown Figure 5 The memory cell 510 shown and Figure 7 The memory unit 710 shown serves as a synapse and partial neuron between the input layer and the next layer. In this example, the inputs are INPUT0, ..., INPUT0. M In the word lines WL0, ..., WL MThe data is received and outputs OUTPUT0, ..., OUTPUT. N On bit lines BL0, ..., BL respectively N The above is generated.

[0102] Figure 29 The neuronal VMM array 2900 is depicted, which is particularly suitable for... Figure 3 The memory unit 310 shown Figure 5 The memory cell 510 shown and Figure 7 The memory unit 710 shown serves as a synapse and partial neuron between the input layer and the next layer. In this example, the inputs are INPUT0, ..., INPUT0. M In the control grid lines CG0, ..., CG M The data is received. Outputs are OUTPUT0, ..., OUTPUT. N On the vertical source lines SL0, ..., SL respectively N The above is generated, where each source line SL i The source line coupled to all memory cells in column i.

[0103] Figure 30 A neuronal VMM array 3000 is depicted, which is particularly suitable for... Figure 3 The memory unit 310 shown Figure 5 The memory cell 510 shown and Figure 7 The memory unit 710 shown serves as a synapse and partial neuron between the input layer and the next layer. In this example, the inputs are INPUT0, ..., INPUT0. M In the control grid lines CG0, ..., CG M The data is received. Outputs are OUTPUT0, ..., OUTPUT. N On the vertical lines BL0, ..., BL respectively N The above is generated, where each bit line BL i Bit lines coupled to all memory cells in column i.

[0104] Long Short-Term Memory

[0105] Existing technologies include a concept known as Long Short-Term Memory (LSTM). LSTM cells are commonly used in neural networks. LSTM allows neural networks to remember information for predetermined arbitrary time intervals and use that information in subsequent operations. A typical LSTM cell includes a cell, an input gate, an output gate, and a forget gate. These three gates regulate the flow of information into and out of the cell and the time interval at which information is remembered in the LSTM. Virtual Memory Models (VMMs) are particularly useful in LSTM cells.

[0106] Figure 14 An example LSTM 1400 is depicted. This example LSTM 1400 includes cells 1401, 1402, 1403, and 1404. Cell 1401 receives the input vector x0 and generates the output vector h0 and the cell state vector c0. Cell 1402 receives the input vector x1, the output vector (hidden state) h0 from cell 1401, and the cell state c0 from cell 1401, and generates the output vector h1 and the cell state vector c1. Cell 1403 receives the input vector x2, the output vector (hidden state) h1 from cell 1402, and the cell state c1 from cell 1402, and generates the output vector h2 and the cell state vector c2. Cell 1404 receives the input vector x3, the output vector (hidden state) h2 from cell 1403, and the cell state c2 from cell 1403, and generates the output vector h3. Additional cells can be used, and this four-cell LSTM is merely an example.

[0107] Figure 15 Depicting what can be used Figure 14 The following is a specific example implementation of LSTM cell 1500, comprising cells 1401, 1402, 1403, and 1404. LSTM cell 1500 receives an input vector x(t), a cell state vector c(t-1) from the previous cell, and an output vector h(t-1) from the previous cell, and generates the cell state vector c(t) and the output vector h(t).

[0108] LSTM unit 1500 includes sigmoid function devices 1501, 1502, and 1503, each applying a number between 0 and 1 to control how much of each component in the input vector is allowed to pass through to the output vector. LSTM unit 1500 also includes tanh devices 1504 and 1505 for applying a hyperbolic tangent function to the input vector, multiplier devices 1506, 1507, and 1508 for multiplying two vectors together, and adder device 1509 for adding two vectors together. The output vector h(t) can be provided to the next LSTM unit in the system, or it can be accessed for other purposes.

[0109] Figure 16LSTM unit 1600 is depicted, which is an example of a specific implementation of LSTM unit 1500. For the reader's convenience, LSTM unit 1600 uses the same numbering as LSTM unit 1500. Sigmoid function devices 1501, 1502, and 1503, and tanh device 1504 each include multiple VMM arrays 1601 and activation function blocks 1602. Thus, it can be seen that VMM arrays are particularly useful in LSTM units used in certain neural network systems. Multiplier devices 1506, 1507, and 1508, and adder device 1509 are implemented digitally or analogically. Activation function block 1602 can be implemented digitally or analogically.

[0110] exist Figure 17 The diagram shows an alternative form of the LSTM unit 1600 (and another example of a specific implementation of the LSTM unit 1500). Figure 17 In this LSTM unit, sigmoid function devices 1501, 1502, and 1503, and tanh device 1504 share the same physical hardware (VMM array 1701 and activation function block 1702) in a time-division multiplexing manner. The LSTM unit 1700 also includes a multiplier device 1703 for multiplying two vectors together, an adder device 1708 for adding two vectors together, a tanh device 1505 (which includes activation function block 1702), a register 1707 for storing the value i(t) when it is output from sigmoid function block 1702, and a register 1707 for storing the value f(t). When c(t-1) is output from multiplier device 1703 via multiplexer 1710, register 1704 stores the value of i(t). When u(t) is output from multiplier device 1703 via multiplexer 1710, register 1705 stores the value of u(t), and register 1705 is used to store the value of o(t). When c~(t) is output from multiplier device 1703 via multiplexer 1710, register 1706 and multiplexer 1709 store the value.

[0111] LSTM unit 1600 contains multiple VMM arrays 1601 and corresponding activation function blocks 1602, while LSTM unit 1700 contains a set of VMM arrays 1701 and activation function blocks 1702, which are used to represent multiple layers in the example of LSTM unit 1700. LSTM unit 1700 will require less space than LSTM 1600 because LSTM unit 1700 only needs 1 / 4 of its space for VMMs and activation function blocks compared to LSTM unit 1600.

[0112] It is also understood that an LSTM cell will typically comprise multiple VMM arrays, each using functionality provided by certain circuit blocks outside the VMM array itself, such as summer and activation function blocks, and high-voltage generation blocks. Providing a separate circuit block for each VMM array would require a significant amount of space within the semiconductor device and would be inefficient to some extent. Therefore, the examples described below reduce the circuitry used outside the VMM arrays themselves.

[0113] Gate control recursive unit

[0114] A simulated VMM implementation can be used in gated recurrent unit (GRU) systems. A GRU is a gated mechanism in a recurrent neural network. A GRU is similar to an LSTM, but a GRU unit generally contains fewer components than an LSTM unit.

[0115] Figure 18 An example GRU 1800 is depicted. This example GRU 1800 includes units 1801, 1802, 1803, and 1804. Unit 1801 receives an input vector x0 and generates an output vector h0. Unit 1802 receives an input vector x1, an output vector h0 from unit 1801, and generates an output vector h1. Unit 1803 receives an input vector x2 and an output vector (in a hidden state) h1 from unit 1802 and generates an output vector h2. Unit 1804 receives an input vector x3 and an output vector (in a hidden state) h2 from unit 1803 and generates an output vector h3. Additional units can be used, and this four-unit GRU is merely an example.

[0116] Figure 19 Depicting what can be used Figure 18 Examples of specific implementations of GRU unit 1900 are provided for units 1801, 1802, 1803, and 1804. GRU unit 1900 receives an input vector x(t) and an output vector h(t-1) from a previous GRU unit, and generates an output vector h(t). GRU unit 1900 includes sigmoid function devices 1901 and 1902, each applying a number between 0 and 1 to the components from the output vector h(t-1) and the input vector x(t). GRU unit 1900 also includes a tanh device 1903 for applying a hyperbolic tangent function to the input vector, multiple multiplier devices 1904, 1905, and 1906 for multiplying two vectors together, an adder device 1907 for adding two vectors together, and a complementary device 1908 for subtracting the input from 1 to generate the output.

[0117] Figure 20GRU unit 2000 is depicted, which is an example of a specific implementation of GRU unit 1900. For the reader's convenience, GRU unit 2000 uses the same numbering as GRU unit 1900. Figure 20 As can be seen, sigmoid function devices 1901 and 1902, and tanh device 1903 each include multiple VMM arrays 2001 and activation function blocks 2002. Therefore, it can be seen that VMM arrays are particularly useful in GRU units used in some neural network systems. Multiplier devices 1904, 1905, and 1906, adder device 1907, and complement device 1908 are implemented digitally or analogically. Activation function blocks 2002 can be implemented digitally or analogously.

[0118] Alternative forms of the GRU unit 2000 (and another example of a specific implementation of the GRU unit 1900) in Figure 21 As shown in [the image]. Figure 21 In this configuration, the GRU unit 2100 utilizes a VMM array 2101 and an activation function block 2102, which, when configured as a sigmoid function, applies numbers between 0 and 1 to control how much of each component in the input vector is allowed to pass through to the output vector. Figure 21 In this configuration, sigmoid function devices 1901 and 1902 and tanh device 1903 share the same physical hardware (VMM array 2101 and activation function block 2102) in a time-division multiplexing manner. The GRU unit 2100 also includes a multiplier device 2103 for multiplying two vectors together, an adder device 2105 for adding two vectors together, a complement device 2109 for subtracting the input from 1 to generate the output, a multiplexer 2104, and a multiplexer for... (the sentence is incomplete and requires further context to translate accurately). The register 2106, which holds the value of r(t) when it is output from the multiplier device 2103 via the multiplexer 2104, is used when the value h(t-1) is... The register 2107 holds the value of z(t) when it is output from the multiplier device 2103 via the multiplexer 2104, and the register 2107 is used to hold the value when the value h is... ^ (t) (1-z(t)) is output from the multiplier device 2103 via the multiplexer 2104 and is held in register 2108.

[0119] GRU unit 2000 contains multiple sets of VMM arrays 2001 and activation function blocks 2002, while GRU unit 2100 contains a set of VMM arrays 2101 and activation function blocks 2102, which are used to represent multiple layers in the example of GRU unit 2100. GRU unit 2100 will require less space than GRU unit 2000 because GRU unit 2100 only needs 1 / 3 of its space for VMMs and activation function blocks compared to GRU unit 2000.

[0120] It is also understandable that a GRU system would typically include multiple VMM arrays, each using functionality provided by certain circuit blocks outside the VMM array itself, such as summer and activation function blocks, and high-voltage generation blocks. Providing a separate circuit block for each VMM array would require a significant amount of space within the semiconductor device and would be inefficient to some extent. Therefore, the examples described below reduce the circuitry used outside the VMM arrays themselves.

[0121] The input to a VMM array can be analog level, binary level, pulse, time-modulated pulse, or digital bit (in which case a DAC is used to convert the digital bit to the appropriate input analog level), and the output can be analog level, binary level, timed pulse, pulse, or digital bit (in which case an output ADC is used to convert the output analog level to digital bit).

[0122] Generally, for each memory cell in a VMM array, each weight W can be implemented by a single memory cell, a differential cell, or a hybrid memory cell (the average of two cells). In the case of differential cells, two memory cells are used to implement the weight W as a differential weight (W = W+ - W-). In the case of two hybrid memory cells, two memory cells are used to implement the weight W as the average of the two cells.

[0123] Figure 31A VMM system 3100 is depicted. In some examples, the weights W stored in the VMM array are stored as differential pairs W+ (positive weights) and W- (negative weights), where W = (W+) - (W-). In VMM system 3100, half of the bit lines are designated as W+ lines, i.e., bit lines connected to memory cells that will store positive weights W+, and the other half of the bit lines are designated as W- lines, i.e., bit lines connected to memory cells that implement negative weights W-. The W- lines are distributed alternately among the W+ lines. The subtraction operation is performed by summing circuits (such as summing circuits 3101 and 3102), which receive current from the W+ and W- lines. The outputs of the W+ and W- lines are combined to effectively give W = W+ - W- for each (W+, W-) cell pair of all (W+, W-) line pairs. Although the W- lines, which are alternately distributed between the W+ lines, have been described above, in other examples, the W+ and W- lines can be located anywhere in the array.

[0124] Figure 32 Another example is depicted. In the VMM system 3210, positive weights W+ are implemented in a first array 3211 and negative weights W- are implemented in a second array 3212, which is separate from the first array, and the resulting weights are appropriately combined by a summing circuit 3213.

[0125] Figure 33 A VMM system 3300 is depicted, in which weights W stored in the VMM array are stored as differential pairs W+ (positive weights) and W- (negative weights), where W = (W+) - (W-). The VMM system 3300 includes arrays 3301 and 3302. Half of the bit lines in each of arrays 3301 and 3302 are designated as W+ lines, i.e., bit lines connected to memory cells that will store positive weights W+, and the other half of the bit lines in each of arrays 3301 and 3302 are designated as W- lines, i.e., bit lines connected to memory cells that implement negative weights W-. The W- lines are distributed alternately among the W+ lines. Subtraction operations are performed by summing circuits (such as summing circuits 3303, 3304, 3305, and 3306) that receive current from the W+ and W- lines. The outputs of the W+ line and the W- line from each array 3301, 3302 are combined to effectively give W = W+ - W- for each (W+, W-) cell pair of all (W+, W-) line pairs. Furthermore, the W values ​​from each array 3301 and array 3302 can be further combined by summing circuits 3307 and 3308 such that each W value is the result of subtracting the W value from array 3302 from the W value from array 3301. This means that the final result from summing circuits 3307 and 3308 is one of two differences.

[0126] Each non-volatile memory cell used in an analog neural memory system needs to be erased and programmed to maintain a very specific and precise amount of charge (i.e., the number of electrons) in a floating gate. For example, each floating gate can hold one of N different values, where N is the number of different weights that can be indicated by each cell. Examples of N include 16, 32, 64, 128, and 256.

[0127] Existing systems require significant area and involve substantial delays at the output stage. For example, it takes multiple clock cycles to convert analog current received from a VMM array into digital output data.

[0128] The goal is to reduce latency at the output to improve the overall operating speed of the system, which represents some or all of the artificial neural networks. Summary of the Invention

[0129] Numerous examples of output circuits and associated methods for neural network arrays are disclosed. Attached Figure Description

[0130] Figure 1 This is a diagram illustrating an artificial neural network.

[0131] Figure 2 The prior art split-gate flash memory cell is described.

[0132] Figure 3 Another prior art split-gate flash memory cell is described.

[0133] Figure 4 Another prior art split-gate flash memory cell is described.

[0134] Figure 5 Another prior art split-gate flash memory cell is described.

[0135] Figure 6 The following diagram illustrates different levels of an artificial neural network utilizing one or more non-volatile memory arrays.

[0136] Figure 7 Here is a block diagram of an example VMM system.

[0137] Figure 8 Here is a block diagram illustrating an example artificial neural network utilizing one or more VMM systems.

[0138] Figure 9 Another example of a VMM system is described.

[0139] Figure 10 Another example of a VMM system is described.

[0140] Figure 11 Another example of a VMM system is described.

[0141] Figure 12 Another example of a VMM system is described.

[0142] Figure 13 Another example of a VMM system is described.

[0143] Figure 14 The existing long short-term memory system is described.

[0144] Figure 15 An example cell used in a long short-term memory system is depicted.

[0145] Figure 16 Depicting Figure 15 The specific implementation of the unit is an example.

[0146] Figure 17 Depicting Figure 15 Another example of a specific implementation of the unit.

[0147] Figure 18 A prior art gate-controlled recursive cell system is described.

[0148] Figure 19 An example cell used in a gate-controlled recursive cell system is depicted.

[0149] Figure 20 Depicting Figure 19 The specific implementation of the unit is an example.

[0150] Figure 21 Depicting Figure 19 Another example of a specific implementation of the unit.

[0151] Figure 22 Another example of a VMM system is described.

[0152] Figure 23 Another example of a VMM system is described.

[0153] Figure 24 Another example of a VMM system is described.

[0154] Figure 25 Another example of a VMM system is described.

[0155] Figure 26 Another example of a VMM system is described.

[0156] Figure 27 Another example of a VMM system is described.

[0157] Figure 28Another example of a VMM system is described.

[0158] Figure 29 Another example of a VMM system is described.

[0159] Figure 30 Another example of a VMM system is described.

[0160] Figure 31 Another example of a VMM system is described.

[0161] Figure 32 Another example of a VMM system is described.

[0162] Figure 33 Another example of a VMM system is described.

[0163] Figure 34 Another example of a VMM system is described.

[0164] Figure 35 The output circuit is described.

[0165] Figure 36 A Σ-Δ analog-to-digital converter is described.

[0166] Figure 37 Another Σ-Δ analog-to-digital converter is described.

[0167] Figure 38 Another Σ-Δ analog-to-digital converter is described.

[0168] Figure 39 Another Σ-Δ analog-to-digital converter is described.

[0169] Figure 40 Another Σ-Δ analog-to-digital converter is described.

[0170] Figure 41 Another Σ-Δ analog-to-digital converter is described.

[0171] Figure 42 Another Σ-Δ analog-to-digital converter is described.

[0172] Figure 43 Another Σ-Δ analog-to-digital converter is described.

[0173] Figure 44A and Figure 44B The waveforms for the Σ-Δ analog-to-digital converter are depicted.

[0174] Figure 45 A graph depicting the relationship between bit line voltage and bit line current.

[0175] Figure 46A graph depicting the relationship between bit line voltage and bit line current for various values ​​of Qref is presented.

[0176] Figure 47 A graph depicting the relationship between bit line voltage and bit line current.

[0177] Figure 48 The output circuit is described.

[0178] Figure 49 The output circuit is shown.

[0179] Figure 50 The output circuit is described. Detailed Implementation

[0180] VMM system architecture

[0181] Figure 34 A block diagram of a VMM system 3400 is depicted. The VMM system 3400 includes a VMM array 3401, a row decoder 3402, a high-voltage decoder 3403, a column decoder 3404, bit line drivers 3405 (such as bit line control circuitry for programming), input circuitry 3406, output circuitry 3407, control logic components 3408, and a bias generator 3409. The VMM system 3400 also includes a high-voltage generation block 3410, which includes a charge pump 3411, a charge pump regulator 3412, and a high-voltage level generator 3413. The VMM system 3400 also includes an algorithm controller 3414, analog circuitry 3415, a control engine 3416 (which may include, but is not limited to, functions such as arithmetic functions, activation functions, embedded microcontroller logic), test control logic unit 3417, and static random access memory (SRAM) blocks 3418 (which are programmed / erased or weighted) to store intermediate data such as activation data for input circuits or intermediate data for output circuits (neuron output data, partial and output neuron data), or data inputs for programming (such as data inputs for a whole line or for multiple lines).

[0182] Input circuitry 3406 may include circuitry such as a DAC (digital-to-analog converter), a DPC (digital-to-pulse converter, digital-to-time modulated pulse converter), an AAC (analog-to-analog converter, such as a current-to-voltage converter, logarithmic converter), a PAC (pulse-to-analog level converter), or any other type of converter. Input circuitry 3406 may implement one or more of a normalization, linear, or nonlinear up / down scaling function or arithmetic function. Input circuitry 3406 may implement a temperature compensation function for the input level. Input circuitry 3406 may implement activation functions such as ReLU or sigmoid. Input circuitry 3406 may store digital activation data that will be applied as an input signal or combined with an input signal during programming or read operations. The digital activation data may be stored in a register. Input circuitry 3406 may include circuitry for driving array terminals, such as CG, WL, EG, and SL lines, which may include sample-and-hold circuitry and buffers. The DAC may be used to convert the digital activation data into an analog input voltage to be applied to the array.

[0183] Output circuitry 3407 may include circuitry such as ITV (current-to-voltage circuitry), ADC (analog-to-digital converter for converting analog neuron outputs into digital bits), AAC (analog-to-analog converter, such as current-to-voltage converter, logarithmic converter), APC (analog-to-pulse converter, analog-to-time-modulated pulse converter), or any other type of converter. Output circuitry 3407 may convert array outputs into activation data. Output circuitry 3407 may implement activation functions such as a modified linear activation function (ReLU) or sigmoid. Output circuitry 3407 may implement one or more of the following for neuron outputs: statistical normalization, regularization, up / down scaling / gain functions, statistical rounding, or arithmetic functions (e.g., addition, subtraction, division, multiplication, shifting, logarithmic). Output circuitry 3407 may implement temperature compensation functions for neuron outputs or array outputs (such as bitline outputs) to keep the array's power consumption approximately constant, or to improve the accuracy of the array (neuron) outputs, such as by keeping the IV slope approximately the same with temperature. The output circuit 3407 may include a register for storing output data.

[0184] Figure 35 The output circuit 3500 is described. The output circuit 3500 is... Figure 34An example implementation of the output circuit 3407 is shown below. The output circuit 3500 includes column multiplexers 3501-1, 3501-2…3501-(j-1), 3501-j and analog-to-digital converters (ADCs) 3502-1, 3502-2…3502-(j-1), 3502-j. Each column multiplexer 3501 receives current, such as bit line current, from one or more columns in the VMM array 3401. If more than one column is connected to a corresponding column multiplexer 3501, the column multiplexer 3501 selects the column and provides the current from that column to the corresponding ADC 3502, which converts the received current into a digital output DOUTx[n: 1], where x is the column number and DOUT comprises n bits. If j equals the number of columns in the VMM array 3401, which means that each column has its own ADC 3502, then the column multiplexer 3501 is optional, and each column in the VMM array 3401 can be directly connected to its associated ADC 3502.

[0185] Figure 36 The Σ-ΔADC 3600 is described, which can be used for... Figure 35 The ADC3502 is located in the output circuit 3500. The Σ-ΔADC 3600 receives the current IBL, enable signal EN, and clock signal CLK from a column of the VMM array 3401 (illustrated as a single memory cell for simplicity). CLK will typically have a higher frequency than the clock frequency used by other circuits in the VMM system because the Σ-ΔADC 3600 uses CLK when sampling the current received from the VMM array 3401. The Σ-ΔADC 3600 output includes an n-bit digital output DOUT[n:1]. The Σ-ΔADC 3600 directly converts the array current IBL into digital output bits without using a current-to-voltage converter (ITV).

[0186] Figure 37 The Σ-ΔADC 3700 is described, which is... Figure 36an example of the Σ-Δ ADC 3600 in . The Σ-Δ ADC 3700 includes a comparator 3701, a state machine (SM) 3702, a current source 3703, and a switch 3704. The Σ-Δ ADC 3700 is coupled to a column in the VMM array 3401 (illustrated as a single memory cell 3711 for simplicity), and the column draws current from the Σ-Δ ADC 3700. The SM 3702 may be implemented using discrete logic, a programmable device, a processor, a controller, or other mechanisms. The Σ-Δ ADC 3700 is coupled to a column in the VMM array 3401 (illustrated as a single memory cell for simplicity), and the column draws current from the Σ-Δ ADC 3700. The Σ-Δ ADC 3700 further receives: an enable signal EN; a clock signal CLK; and a reference voltage VREF. The current source 3703 provides a reference current IREF and is an example of an injection circuit. The switch 3704 is controlled by the SM 3702. The comparator 3701 includes a first terminal (here, an inverting terminal) coupled to the column (illustrated as the memory cell 3711) and the current source 3703 via the switch 3704, and a second terminal (here, a non-inverting terminal) coupled to the reference voltage VREF. The SM 3702 receives the output of the comparator 3701. When the voltage on the bit line (BL) node of the memory cell 3711 is < VREF, the clock signal CLK is transmitted by the SM 3702, or converted by the SM 3702 into a converted clock signal, which is used as the signal CLK_REF 3710. The signal CLK_REF 3710 thus has clock pulses of the same or opposite polarity to the clock signal CLK. In a case where the clock signal CLK has been converted by the SM 3702, the clock pulses of the signal CLK_REF 3710 may have a different duty cycle or polarity from the clock pulses of the clock signal CLK, and the signal CLK_REF 3710 controls the switch 3704. When the signal CLK_REF is asserted during each clock cycle, the switch 3704 is closed, and Iref from the current source 3703 is enabled to flow into the BL node of the memory cell 3711, which increases the voltage on the BL node or the memory cell 3711. The signal CLK_REF continues timing in response to the clock signal CLK, and at each assertion of the signal CLK_REF, the current IREF is injected into the BL node of the memory cell 3711, that is, the current is related to IREF A charge Qref proportional to T is injected into the BL node or the memory cell 3711, where T is the assertion time of the signal CLK_REF, until the voltage on the BL node of the memory cell 3711 is > VREF, then the SM 3702 deasserts the hold signal CLK_REF, which means that the CLK pulse is not transmitted to SM 3702 or converted by the SM for transmission to the signal CLK_REF. In response to the start of a read operation, the VMM array 3401 draws current to lower the BL node of the memory cell 3711 to a low level. When there is no additional charge, since the signal CLK_REF is deasserted, the VMM array 3401 draws current to lower the BL node of the memory cell 3711 to a low level until it is < VREF, and then the process repeats. SM 3702 tracks the timing on the signal CLK_REF and generates the digital output bit DOUT[n:] in response. In this example, DOUT[n:1] from SM 3702 is a digital representation of the count value of the number of times the switch 3704 is closed during a certain number of cycles of the clock signal CLK, and reflects the amount of current drawn by the VMM array 3401. For example, an 8-bit digital output DOUT[7:0] can be output for 512 CLK cycles. A high value indicates that the switch 3704 is opened and closed a relatively large number of times, which means the current drawn by the VMM array 3401 is relatively large, because the amount of charge provided by the current source 3703 to the BL node of the memory cell 3711 to match the current drawn by the VMM array 3401 is accordingly relatively large, while a low value indicates that the switch 3704 is opened and closed a relatively small number of times, which means the current drawn by the VMM array 3401 is relatively small. The digital output value reflects the total charge (i.e., the amount of Qref) injected into BL to balance the array current (to keep BL constant), where a smaller current produces a smaller digital output value, while a larger current produces a larger digital output value. DOUT[n:1] indicates the value of the current drawn by a column (exemplified as memory cell 3711).

[0187] Figure 38 depicts a Σ-ΔADC 3800, which is Figure 36An example of a Σ-Δ ADC 3600 is shown below. The Σ-Δ ADC 3800 includes a comparator 3801, an SM 3802, a capacitor CREF 3803, a switch 3804, and a switch 3805. The Σ-Δ ADC 3800 is coupled to a column in the VMM array 3401 (illustrated as a single memory cell 3811 for simplicity), which draws current from the Σ-Δ ADC 3700. The SM 3802 provides control signals CLK_REFB and CLK_REF to switches 3804 and 3805, respectively. CLK_REFB is the logical inverse of CLK_REF. Switch 3804 is used to charge capacitor 3803 to a specific voltage (VDD in this case) when it is closed (when CLK_REFB is asserted), and switch 3805 discharges the charge from capacitor 3803 into the bit line (BL) node of the single memory cell 3811 when it is closed (when CLK_REF is asserted), where the reference charge Qref on capacitor 3803 will be equal to (VDD - voltage of the BL node of cell 3811). The capacitance of capacitor 3803 is reduced proportionally. That is, for each clock pulse in signals CLK_REF and CLK_REFB, a reference charge Qref is injected into the BL node of cell 3811, where the reference charge Qref flows as a current through memory cell 3811 to ground. Capacitor 3803 is an example of the injection circuit. The Σ-ΔADC 3800 is coupled to a column in the VMM array 3401, which draws current from the Σ-ΔADC 3800. Comparator 3801 includes a first terminal (here, the inverting terminal) coupled to the column (exemplified as memory cell 3811) and capacitor 3803 via switch 3805, and a second terminal (here, the non-inverting terminal) coupled to the reference voltage VREF. SM 3802 receives the output of comparator 3801. The Σ-ΔADC 3800 also receives: an enable signal EN; a clock signal CLK; and the reference voltage VREF. By injecting a reference charge instead of a reference current into the BL node or memory cell 3811, capacitor 3803 will... Figure 37 The current source 3703 in the Σ-ΔADC 3800 is equivalent to this. Figure 37 The Σ-Δ ADC3700 operates in a similar manner, the main difference being that it uses a reference charge instead of a reference current. The output bit DOUT[n:1] from the SM 3802 is a digital count of the number of times switch 3805 closes during a certain number of cycles within the clock signal CLK. Higher values ​​indicate a relatively large number of times switch 3805 is opened and closed, representing a larger current drawn by the VMM array 3401, while lower values ​​indicate a relatively small number of times switch 3805 is opened and closed, representing a smaller current drawn by the VMM array 3401.

[0188] Figure 39 The Σ-ΔADC 3900 is described, which is... Figure 36 The example of the Σ-ΔADC 3600 is shown. The Σ-ΔADC 3900 is similar. Figure 37 The Σ-Δ ADC 3700 differs in that it uses an adjustable current source 3903 (which is an example of an injection circuit) to generate a variable reference current IREFV instead of a fixed reference current. Furthermore, the variable reference current source 3903 is controlled by SM 3902 or by a global control signal (not shown). The variable reference current source 3903 can be adjusted to provide different amounts of reference current to compensate for different array current ranges, different bit line capacitances, or any variations with PVT (process, temperature, and voltage). Comparator 3901 includes a first terminal (here, the inverting terminal) coupled to the column (exemplified as memory cell 3911) and the adjustable current source 3903 via switch 3904, and a second terminal (here, the non-inverting terminal) coupled to the reference voltage VREF. SM 3902 receives the output of comparator 3901.

[0189] Figure 40 The Σ-ΔADC 4000 is described, which is... Figure 36 The example is the Σ-ΔADC 3600. The Σ-ΔADC 4000 is similar. Figure 38 The Σ-Δ ADC 3800 differs from the one described above in that the voltage VREFSUP supplied to the reference capacitor 4003 is provided by a variable reference power supply 4006. Capacitor 4003 is an example of an injection circuit. Furthermore, the variable reference power supply 4006 generates a variable voltage in response to SM 4002 or via a global control signal (not shown); that is, the amount of voltage provided by the variable reference power supply 4006 can be adjusted in response to the signal CFG_VREFSUP, which can be provided by SM 4002. The voltage provided by the variable reference power supply 4006 can be adjusted to compensate for different array current ranges, different bit line capacitances, or any variations with PVT (process, temperature, and voltage). Comparator 4001 includes a first terminal (here, the inverting terminal) coupled to a column (e.g., memory cell 4011) and capacitor 4003 via switch 4005, and a second terminal (here, the non-inverting terminal) coupled to a reference voltage VREF. SM 4002 receives the output of comparator 4001.

[0190] Figure 41 The Σ-ΔADC 4100 is described, which is... Figure 36The example of the Σ-ΔADC 3600 is shown below. The Σ-ΔADC 4100 includes a comparator 4101, an SM 4102, a capacitor CREF 4103, switches 4104 and 4105, a transistor 4106, and a gate driver 4107 that drives the gate of transistor 4106. The Σ-ΔADC 4100 is similar to... Figure 38 The Σ-Δ ADC 3800 differs in that it uses transistor 4106, whose gate is controlled by a reference voltage provided by gate driver 4107. Comparator 4101 includes a first terminal (here, the inverting terminal) coupled to a column (exemplified as memory cell 4111) and capacitor 4103 (which is an example of an injection circuit) via switch 4105, and a second terminal (here, the non-inverting terminal) coupled to a reference voltage VREF. SM 4102 receives the output of comparator 4101. Transistor 4106 is used to control how much charge is transferred (injected) into the BL node of memory cell 4111. The injected charge is related to (Vdd - (VREFX + Vt_4106)). The capacitance of capacitor 4103 is proportional, where Vt_4106 is the threshold voltage of transistor 4106. VREFx can be controlled by SM 4102 or by a global control signal (not shown). It can be used to compensate for different array current ranges, different bit line capacitances, or any variations that occur with PVT (process, temperature, and voltage).

[0191] Figure 42 The Σ-ΔADC 4200 is described, which is... Figure 36 The example shown is the Σ-Δ ADC 3600. The Σ-Δ ADC 4200 includes a comparator 4201, an SM 4202, an adjustable capacitor 4203, a switch 4204, and a switch 4205. The Σ-Δ ADC 4200 is similar to... Figure 40 The Σ-ΔADC 4000 differs in that it uses an adjustable capacitor 4203 instead of a variable reference power supply. The adjustable capacitor 4203 generates a variable charge in response to SM 3902 or via a global control signal (not shown); that is, the capacitance provided by the adjustable capacitor 4203 is in response to the signal CFG_CREF, which can be provided by SM 4202. The signal CFG_CREF and the adjustable capacitor 4203 can be used to compensate for different array current ranges, different bit line capacitances, or any variations with PVT (process, temperature, and voltage). Comparator 4201 includes a first terminal (here, the inverting terminal) coupled to a column (exemplified as memory cell 4211) and the adjustable capacitor 4203 (which is an example of an injection circuit) via switch 4205, and a second terminal (here, the non-inverting terminal) coupled to a reference voltage VREF. SM 4202 receives the output of comparator 4201.

[0192] Figure 43 The Σ-ΔADC 4300 is described, which is... Figure 36 An example of the Σ-ΔADC 3600 is shown. The Σ-ΔADC 4300 includes a comparator 4301, an SM 4302, a variable capacitor 4303, a variable dummy BL capacitor 4306, a switch 4304, and a switch 4305. The Σ-ΔADC 4300 is similar to the Σ-ΔADC 4200, except that a variable dummy BL capacitor 4306 is added in parallel with the VMM array 3401. The variable dummy BL capacitor 4306 generates a variable charge in response to the SM 4302 or via a global control signal (not shown). That is, the capacitance provided by the variable dummy BL capacitor 4306 can be changed in response to the SM 4302 or via a global control signal (not shown). This can be used to compensate for different array current ranges, different bit line capacitances, or any variations that occur with PVT (process, temperature, and voltage). Comparator 4301 includes a first terminal (here, the inverting terminal) coupled via switch 4305 to a column (exemplified as memory cell 4311) and adjustable capacitors 4306 and 4303 (which together form an example of an injection circuit), and a second terminal (here, the non-inverting terminal) coupled to a reference voltage VREF. SM 4302 receives the output of comparator 4301.

[0193] Figure 44A and Figure 44BExample waveforms 4400 and 4410 are depicted, illustrating the operation of Σ-Δ ADCs 3600, 3700, 3800, 3900, 4000, 4100, 4200, and 4300. In waveform 4400, VMM array 3401 draws current I1, while in waveform 4410, VMM array 3401 draws half that amount, I1 / 2. VBL is the voltage on the bit line of VMM array 3401 coupled to the inverting inputs of the corresponding comparators 3701, 3801, 3901, 4001, 4101, 4201, and 4301, and VREF is the voltage supplied to the non-inverting inputs of the corresponding comparators 3701, 3801, 3901, 4001, 4101, 4201, and 4301. Clock signals CLK and CLK_REF are shown, with the outputs of the corresponding comparators 3701, 3801, 3901, 4001, 4101, 4201, and 4301 illustrated as the signal COMP_OUT. VBL rises when the corresponding switches 3704, 3805, 3904, 4005, 4105, 4205, and 4305 are closed, and subsequently falls when those switches are open. The number of times the corresponding switches 3704, 3805, 3904, 4005, 4105, 4205, and 4305 are closed (or opened) is counted, and the count value within a specific time period is output as the digital output DOUT[n:1]. In waveform 4400, the count is 4 during the 16 cycles of the clock signal CLK shown. In waveform 4410, the count is 2 during the 16 cycles of the clock signal CLK shown, because the array current is... Figure 44A The waveform 4400 represents half of the array current. The correlation period for the measurement count can be greater than or less than 16 cycles of the clock signal CLK.

[0194] Figure 45 Graph 4500 illustrates the operation of Σ-Δ ADCs 3600, 3700, 3800, 3900, 4000, 4100, 4200, and 4300, respectively. Graph 4500 shows the relationship between VBL and IBL, where VBL is the voltage at the BL node of memory cells 3711, 3811, 3911, 4011, and 4111, and IBL is the current drawn from the BL node by the VMM array 3401, which is coupled to the inputs of comparators 3701, 3801, 3901, 4001, 4101, 4201, and 4301, respectively. As shown, a large VBL variation occurs across the array current range, therefore improvements are desired to reduce the voltage variation across the shown current range.

[0195] Figure 46Graph 4600 illustrates the operation of Σ-Δ ADCs 3600, 3700, 3800, 3900, 4000, 4100, 4200, and 4300. Graph 4600 shows the relationship between VBL and IBL, where VBL is the voltage at the BL node of memory cells 3711, 3811, 3911, 4011, and 4111, and IBL is the current drawn from the BL node by the VMM array 3401, which is coupled to the inputs of comparators 3701, 3801, 3901, 4001, 4101, 4201, and 4301.

[0196] Figure 4600 depicts the traces of VBL and IBL for different Qref values, where Qref is the reference charge injected into the bitline. Here, 21 example traces of Qref are shown: QREF1, QREF2, ..., QREF20, QREF21. These are examples, and any number of different Qref values ​​can be used. SMs 3702, 3802, 3902, 4002, 4102, 4202, and 4302 can be programmed to modify Qref based on the range of the input current IBL being received or expected. The following algorithm can be used to select a Qref value for each current range.

[0197] First, the possible voltage range of VBL is divided into several ranges. Here, three example voltage ranges are shown: voltage ranges 4601, 4602, and 4603. Within each voltage range, the amount of acceptable error is determined. For example, for voltage range 4602 (VBL), the voltage changes from approximately 599mV to 601.5mV, and the acceptable error is 2.5mV.

[0198] Secondly, considering this acceptable error, the possible current range of the IBL is divided into several ranges. Here, six example current ranges are shown: current ranges 4604, 4605, 4606, 4607, 4608, and 4609. Within each current range, one or more Qref values ​​are identified that would cause the voltage trace to fall within the acceptable error amount for the voltage range; or alternatively, unacceptable Qref values ​​are identified.

[0199] For example, QREF1 through QREF4 may be acceptable in the current range 4604 because their traces are fairly linear, but some other QREFs may be unacceptable because they exhibit nonlinearity near the beginning or end of the range or outside the acceptable range 4602.

[0200] In the current range of 4605, QREF1, QREF2, and QREF3 may be considered unacceptable because their traces exhibit non-linearity (e.g., large drops) within this current range. QREF4 through QREF9 are acceptable, while all other QREFs may be considered unacceptable because their traces exhibit non-linearity or exceed the voltage range of 4602.

[0201] In the current range of 4606, QREF1 through QREF8 may be considered unacceptable because they do not even operate in the voltage range of 4602, while QREF10 through QREF13 are acceptable in the 4602 range and have fairly linear traces.

[0202] Similarly, within the current range 4607, QREF1 to QREF12 may be considered unacceptable, while QREF13 to QREF16 are acceptable.

[0203] Similarly, within the current range of 4608, QREF1 through QREF16 may be considered unacceptable, while QREF17 through QREF20 are acceptable.

[0204] Similarly, within the current range 4609, QREF1 to QREF19 may be considered unacceptable, while QREF20 to QREF21 are acceptable.

[0205] Because smaller QREFs are easier and faster to generate (e.g., due to capacitive charging time), SM3702, 3802, 3902, 4002, 4102, 4202, and 4302 can select the smallest acceptable QREF for each range. For example, SM3702, 3802, 3902, 4002, 4102, 4202, and 4302 can select QREF3 for current range 4604, QREF8 for current range 4605, QREF12 for current range 4606, QREF16 for current range 4607, QREF20 for current range 4608, and QREF21 for current range 4609.

[0206] Figure 47Graph 4700 illustrates the operation of Σ-Δ ADCs 3600, 3700, 3800, 3900, 4000, 4100, 4200, and 4300. Graph 4700 shows the relationship between VBL and IBL, where VBL is the voltage at the BL node of memory cells 3711, 3811, 3911, 4011, and 4111, and IBL is the current drawn from the BL node by the VMM array 3401, which is coupled to the inputs of comparators 3701, 3801, 3901, 4001, 4101, 4201, and 4301, respectively. As shown, the current range is divided into six regions, and different QREFs are used in each region by adjusting the values ​​of CREF, IREF, and VSUPREFs based on the Σ-Δ ADC design. SMs 3702, 3802, 3902, 4002, 4102, 4202, and 4302 can select specific CREF / IREF / VSUPREF values ​​based on the measured current range of the VMM array, for example, by monitoring the digital output bits to determine the current range based on the bit values, in order to achieve a more linear range among the various options illustrated in the traces of graphs 4500 and 4600, or options not shown. Optionally, these values ​​can be changed by the SM during the operation of the ADC.

[0207] Figure 48 Output circuit 4800 is depicted, which includes a current-to-voltage converter 4801 and a Σ-Δ ADC 4802. The current-to-voltage converter 4801 receives bit line current from VMM array 3401 (not shown) and converts the current into a voltage ITV_Ox, which is provided to the SD ADC 4802, which converts the voltage into a digital output [n:1].

[0208] Figure 49 Output circuitry 4900 is depicted, comprising a current-to-voltage converter 4901 and a Σ-Δ ADC 4902. The current-to-voltage converter 4901 receives different bit-line currents IBL+ and IBL- from a VMM array 3401 (not shown) and converts the differential currents into differential voltages ITV_O+ and ITV_O-, which are provided to the SD ADC 4902, which converts the voltages into digital outputs [n:1]. In this case, the Σ-Δ ADC (not shown) operates on voltages rather than currents.

[0209] Figure 50A differential output circuit 5000 is depicted, where the digital output bits represent the differential weights, such as W=(W+)-(W-) or IBL=(IBL+)-(IBL-). Σ-ΔADCs 5001 and 5002 can be any of the Σ-ΔADCs 3600, 3700, 3800, 3900, 4000, 4100, 4200, and 4300. During operation, IBL+ is converted to digital bits by Σ-ΔADC 5001, and the result is stored in an up / down counter 5003 by counting up from the intermediate value. Then, IBL- is converted to digital bits by Σ-ΔADC 5002, and the result is used to count down the value in counter 5003. Therefore, the final value DOUT[n:1] output by the up / down counter 5003 is the differential weight represented by (IBL+)-(IBL-). For example, if the up / down counter 5003 is an 8-bit counter, the intermediate value can be 127 (01000000). Counting up on the output of the Σ-ΔADC 5001 will produce a value between 127 and 255, and counting down will produce a final value between 0 and 255, where 128 to 255 represent positive weights and 0 to 127 represent negative weights.

[0210] As used herein, the terms “above” and “on” both encompass “directly on” (without intermediate material, elements, or spaces between) and “indirectly on” (with intermediate material, elements, or spaces between). Similarly, the term “adjacent” includes “directly adjacent” (without intermediate material, elements, or spaces between) and “indirectly adjacent” (with intermediate material, elements, or spaces between), “mounted to” includes “directly mounted to” (without intermediate material, elements, or spaces between) and “indirectly mounted to” (with intermediate material, elements, or spaces between), and “electrically coupled to” includes “directly electrically coupled to” (without intermediate material or elements electrically connecting the elements together) and “indirectly electrically coupled to” (with intermediate material or elements electrically connecting the elements together). For example, forming an element “above a substrate” can include forming an element directly on the substrate without intermediate material / elements between them, and forming an element indirectly on the substrate with one or more intermediate materials / elements between them.

Claims

1. A system comprising: Vector-matrix multiplication array, the vector-matrix multiplication array comprising an array of non-volatile memory cells arranged in rows and columns; and A Σ-Δ analog-to-digital converter, wherein the Σ-Δ analog-to-digital converter is used to receive current from columns of the vector-matrix multiplication array and generate a digital output in response to the current.

2. The system according to claim 1, wherein the Σ-Δ analog-to-digital converter comprises: Current source; switch; The comparator includes a first input terminal coupled to the column and the current source via the switch, a second input terminal coupled to a reference voltage, and an output terminal for providing an output; and A state machine is provided for receiving the output from the comparator, providing a control signal to the switch, and generating a digital output in response to the current, the digital output indicating the value of the current drawn by the column.

3. The system according to claim 1, wherein the Σ-Δ analog-to-digital converter comprises: Capacitor; switch; The comparator includes a first input terminal coupled to the column and the capacitor via the switch, a second input terminal coupled to a reference voltage, and an output terminal for providing an output; and A state machine is provided for receiving the output from the comparator, providing a control signal to the switch, and generating a digital output in response to the current, the digital output indicating the value of the current drawn by the column.

4. The system according to claim 1, wherein the Σ-Δ analog-to-digital converter comprises: Adjustable current source; switch; The comparator includes a first input terminal coupled to the column and the adjustable current source via the switch, a second input terminal coupled to a reference voltage, and an output terminal for providing an output; and A state machine is provided for receiving the output from the comparator, providing a control signal to the switch, and generating a digital output in response to the current, the digital output indicating the value of the current drawn by the column.

5. The system according to claim 1, wherein the Σ-Δ analog-to-digital converter comprises: Variable voltage source; Capacitor; switch; The comparator includes a first input terminal coupled to the column and the capacitor via the switch, a second input terminal coupled to a reference voltage, and an output terminal for providing an output; and A state machine is configured to receive the output from the comparator, provide a first control signal to the switch, provide a second control signal to the variable voltage source, and generate the digital output in response to the current, the digital output indicating the value of the current drawn by the column.

6. The system of claim 1, wherein the Σ-Δ analog-to-digital converter comprises: Capacitor; switch; transistor; Control circuit; The comparator includes a first input terminal coupled to the column and the capacitor via the switch and the transistor, a second input terminal coupled to a reference voltage, and an output terminal for providing an output; and A state machine is configured to receive the output from the comparator, provide a first control signal to the switch, provide a second control signal to the control circuit, and generate the digital output in response to the current, the digital output indicating the value of the current drawn by the column.

7. The system of claim 1, wherein the Σ-Δ analog-to-digital converter comprises: Adjustable capacitor; switch; The comparator includes a first input terminal coupled to the column and the adjustable capacitor via the switch, a second input terminal coupled to a reference voltage, and an output terminal for providing an output; and A state machine is configured to receive the output from the comparator, provide a first control signal to the switch, provide a second control signal to the adjustable capacitor, and generate the digital output in response to the current, the digital output indicating the value of the current drawn by the column.

8. The system of claim 7, wherein the Σ-Δ analog-to-digital converter comprises: A second adjustable capacitor is coupled to the first input terminal of the comparator.

9. A system comprising: Vector-matrix multiplication array, the vector-matrix multiplication array comprising an array of non-volatile memory cells arranged in rows and columns; and An analog-to-digital converter (ADC) is used to receive current from columns of the vector-matrix multiplication array and generate a digital output in response to the current, wherein variable charge or variable current is injected by an injection circuit to maintain the voltage of the columns approximately constant during operation.

10. The system of claim 9, wherein the variable charge or the variable current is selected based on a measured current range of the array.

11. The system of claim 10, wherein the variable charge or the variable current is changed during operation of the analog-to-digital converter.

12. The system of claim 9, wherein the variable charge is provided by a variable voltage source coupled to the capacitor.

13. The system of claim 9, wherein the variable charge is varied by changing the capacitance value of the variable capacitor.

14. The system of claim 9, wherein the system includes an up / down counter for determining differential weights provided by two columns in the vector-matrix multiplication array.

15. The system of claim 14, wherein the system includes a second analog-to-digital converter.

16. The system of claim 15, wherein the analog-to-digital converter generates a first value for the up / down counter to count upwards, and the second analog-to-digital converter generates a second value for the up / down counter to count downwards.

17. The system of claim 9, wherein the analog-to-digital converter is a Σ-Δ analog-to-digital converter.

18. A method, the method comprising: The current from an array of non-volatile memory cells arranged in rows and columns is converted into digital output bits, the conversion including injecting variable charge or variable current until the voltage of the column is equal to a reference voltage within a target voltage range.

19. The method of claim 18, wherein the array is used for vector-matrix multiplication.

20. The method of claim 18, wherein the variable charge is provided by an adjustable capacitor.

21. The method of claim 18, wherein the variable charge is provided by a capacitor, wherein a variable supply reference voltage is supplied at the terminals of the capacitor.

22. The method of claim 18, wherein the variable charge is provided by a plurality of capacitors.

23. The method of claim 18, wherein the conversion is performed by a Σ-Δ analog-to-digital converter.

24. The method of claim 18, wherein the method comprises: The magnitude of the variable charge or the magnitude of the variable current is selected based on the measurement of the current or the expected current range.

25. The method of claim 18, wherein the method comprises: The amount of the variable charge is selected based on the expected voltage range of the memory cell.

26. The method of claim 18, wherein the method comprises: The amount of the variable charge is selected based on the expected current range of the memory cell.

Citation Information

Patent Citations

  • High precision and highly efficient tuning mechanisms and algorithms for analog neuromorphic memory in artificial neural networks

    US10748630B2

  • Deep Learning Neural Network Classifier Using Non-volatile Memory Array

    US20170337466A1

  • Single transistor non-valatile electrically alterable semiconductor memory device

    US5029130A

  • Flash memory cells with separated self-aligned select and erase gates, and process of fabrication

    US6747310B2