Analog Neural Memory Array in an Artificial Neural Network with Adaptive Weight Mapping and Substantially Constant Array Source Impedance with Distributed Power

The analog neural memory array addresses the issue of varying source impedance and power consumption by employing a design with constant source impedance and power consumption, utilizing non-volatile memory cells and adaptive weight mapping for improved accuracy and noise resistance.

JP7690623B2Active Publication Date: 2025-06-10SILICON STORAGE TECHNOLOGY INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024009484
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-08-06
Filing Date
2024-01-25
Publication Date
2025-06-10
Estimated Expiration
2040-09-03

AI Technical Summary

Technical Problem

Existing analog neural memory arrays face challenges with varying source impedance and power consumption across the array, leading to inaccuracies and noise susceptibility during read, program, and erase operations.

Method used

The proposed analog neural memory array design features a substantially constant source impedance and power consumption across the array, achieved through the use of non-volatile memory cells arranged in rows and columns with dummy bit lines and bit line transistors, as well as adaptive weight mapping for optimal performance.

Benefits of technology

This design ensures consistent accuracy and reduced noise susceptibility across the array, maintaining constant power consumption and source impedance regardless of the selected cells during operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007690623000019
    Figure 0007690623000019
  • Figure 0007690623000020
    Figure 0007690623000020
  • Figure 0007690623000021
    Figure 0007690623000021
Patent Text Reader

Abstract

To constitute a bit line so that each memory cell can be individually programmed, erased and read out without giving an adverse effect to a memory state of other memory cells in the array.SOLUTION: In a VMM system 1900, each memory cell 1901 in a memory array has an approximately constant source impedance when the cell 1901 is operated. In a specific embodiment, power consumption is substantially constant from a bit line to a bit line in the array when the cell 1901 is read. In a specific embodiment, weight mapping is adaptively executed for optimum performance in power and noise.SELECTED DRAWING: Figure 19A
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] (Claim of Priority) This application claims priority to U.S. Provisional Patent Application No. 62 / 985,826, filed on March 5, 2020, entitled "Analog Neural Memory Array in Artificial Neural Network With Accurate Array Source Impedance With Adaptive Weight Mapping and Distributed Power", and U.S. Patent Application No. 16 / 986,812, filed on August 6, 2020, entitled "Analog Neural Memory Array in Artificial Neural Network with Substantially Constant Array Source Impedance with Adaptive Weight Mapping and Distributed Power".

[0002] (Field of the Invention) Numerous embodiments of analog neural memory arrays are disclosed. In certain embodiments, each memory cell within the array has a substantially constant source impedance when the cell is being operated. In certain embodiments, power consumption is substantially constant from bit line to bit line within the array when the cell is being read. In certain embodiments, weight mapping is performed adaptively for optimal performance in power and noise.

Background Art

[0003] An artificial neural network mimics a biological neural network (the central nervous system of an animal, particularly the brain), may rely on a number of inputs, and is used to estimate or approximate a generally unknown function. An artificial neural network generally includes layers of interconnected "neurons" that exchange messages.

[0004] Figure 1 shows an artificial neural network, in which the circles represent input or neuron layers. The connections (referred to as synapses) are represented by arrows and have numerical weights that can be tuned based on experience. This enables the artificial neural network to adapt to the input and become learnable. Typically, an artificial neural network includes multiple input layers. Typically, there is one or more intermediate layers of neurons and an output layer of neurons that provides the output of the neural network. At each level, the neurons make decisions individually or collectively based on the data received from the synapses.

[0005] One of the main challenges in the development of artificial neural networks for high-performance information processing is the lack of appropriate hardware technologies. In practice, practical artificial neural networks rely on a very large number of synapses, which enables a high connectivity between neurons, that is, a very high degree of parallelization of computational processing. In principle, such complexity can be realized by digital supercomputers or dedicated graphics processing unit clusters. However, in addition to high costs, these approaches also suffer from poor energy efficiency compared to biological networks, which mainly perform low-precision analog calculations and consume far less energy. CMOS analog circuits have been used in artificial neural networks, but most CMOS-implemented synapses have been too bulky assuming a large number of neurons and synapses.

[0006] The applicant has previously disclosed in U.S. Patent Application No. 15 / 594,439, published as U.S. Patent Publication No. 2017 / 0337466, which is incorporated by reference, an artificial (analog) neural network that utilizes one or more non-volatile memory arrays as synapses. The non-volatile memory arrays operate as analog neuromorphic memories. As used herein, the term neuromorphic means a circuit that implements a model of the nervous system. An analog neuromorphic memory includes a first plurality of synapses configured to receive a first plurality of inputs and then generate a first plurality of outputs, and a first plurality of neurons configured to receive the first plurality of outputs. The first plurality of synapses includes a plurality of memory cells, and each memory cell includes a spaced-apart source region and drain region formed in a semiconductor substrate with a channel region extending therebetween, a floating gate disposed above a first portion of the channel region and insulated from the first portion of the channel region, and a non-floating gate disposed above a second portion of the channel region and insulated from the second portion of the channel region. Each of the plurality of memory cells is configured to store a weight value corresponding to a plurality of electrons on the floating gate. The plurality of memory cells are configured to multiply the stored weight values by the first plurality of inputs to generate the first plurality of outputs. An array of memory cells arranged in this manner can be referred to as a vector by matrix multiplication (VMM) array.

[0007] Here, examples of different non-volatile memory cells that can be used in VMM are considered. <<Non-volatile memory cell>>

[0008] Various types of known non-volatile memory cells can be used in the VMM array. For example, U.S. Patent No. 5,029,130 (the “’130 patent”), which is incorporated herein by reference, discloses an array of split-gate non-volatile memory cells, which are a type of flash memory cell. Such a memory cell 210 is shown in FIG. 2. Each memory cell 210 includes a source region 14 and a drain region 16 formed in a semiconductor substrate 12, and there is a channel region 18 between the source region 14 and the drain region 16. The floating gate 20 is formed insulated above a first portion of the channel region 18 (and controls the conductivity of the first portion of the channel region 18), and is formed over a portion of the source region 14. The word line terminal 22 (typically coupled to a word line) is disposed above a second portion of the channel region 18 and is insulated from the second portion of the channel region 18, and has a first portion (which controls the conductivity of the second portion of the channel region 18) and a second portion that extends upward above the floating gate 20. The floating gate 20 and the word line terminal 22 are insulated from the substrate 12 by a gate oxide. The bit line terminal 24 is coupled to the drain region 16.

[0009] By applying a high positive voltage to the word line terminal 22, the memory cell 210 is erased (electrons are removed from the floating gate), whereby the electrons in the floating gate 20 pass through the insulator therebetween from the floating gate 20 to the word line terminal 22 via Fowler-Nordheim tunneling.

[0010] The memory cell 210 is programmed (electrons are applied to the floating gate) by applying a positive voltage to the word line terminal 22 and a positive voltage to the source region 14. The electron current flows from the drain region 16 toward the source region 14 (source line terminal). The electrons are accelerated and, when they reach the gap between the word line terminal 22 and the floating gate 20, they are energized (heated). A portion of the heated electrons is injected into the floating gate 20 through the gate oxide due to the electrostatic attraction from the floating gate 20.

[0011] Memory cell 210 is read by applying a positive read voltage to drain region 16 and word line terminal 22 (turning on the portion of channel region 18 below the word line terminal). When floating gate 20 becomes positively charged (i.e., electrons are erased), the portion of channel region 18 below floating gate 20 also turns on similarly, and current flows through channel region 18, which is sensed as the erased state, i.e., the "1" state. When floating gate 20 becomes negatively charged (i.e., programmed with electrons), the portion of the channel region below floating gate 20 becomes almost or completely off, and current does not (or hardly) flow through channel region 18, which is sensed as the programmed state, i.e., the "0" state.

[0012] Table 1 shows typical voltage ranges that can be applied to the terminals of memory cell 110 to perform read, erase, and program operations. Table 1: Operation of Flash Memory Cell 210 in FIG. 2 [Table 1] "Read 1" is a read mode in which the cell current is output to the bit line. "Read 2" is a read mode in which the cell current is output to the source line terminal.

[0013] FIG. 3 shows a memory cell 310 similar to memory cell 210 of FIG. 2 with an additional control gate (CG) terminal 28. Control gate terminal 28 is biased at a high voltage (e.g., 10V) during programming, a low or negative voltage (e.g., 0V / -8V) during erase, and a low or medium voltage (e.g., 0V / 2.5V) during read. The other terminals are biased in the same manner as the terminals of FIG. 2.

[0014] FIG. 4 shows a four-gate memory cell 410 including a source region 14, a drain region 16, a floating gate 20 above a first portion of a channel region 18, a select gate 22 (typically coupled to a word line, WL) above a second portion of the channel region 18, a control gate 28 above the floating gate 20, and an erase gate 30 above the source region 14. This configuration is described in U.S. Patent No. 6,747,310, which is incorporated herein by reference for all purposes. Here, all gates are non-floating gates except for the floating gate 20, that is, they are electrically connected or connectable to a voltage source. Programming is performed by injecting hot electrons themselves from the channel region 18 into the floating gate 20. Erasure is performed by electrons tunneling from the floating gate 20 to the erase gate 30.

[0015] Table 2 shows typical voltage ranges that can be applied to the terminals of the memory cell 410 to perform read, erase, and program operations. Table 2: Operation of the Flash Memory Cell 410 of FIG. 4

Table 2

[0016] FIG. 5 shows a memory cell 510 similar to the memory cell 410 of FIG. 4 except that the memory cell 510 does not include an erase gate (EG) terminal. Erasure is performed by biasing the substrate 18 to a high voltage and biasing the control gate CG terminal 28 to a low voltage or a negative voltage. Alternatively, erasure is performed by biasing the word line terminal 22 to a positive voltage and biasing the control gate terminal 28 to a negative voltage. Programming and reading are the same as those in FIG. 4.

[0017] Figure 6 shows a three-gate memory cell 610, which is another type of flash memory cell. The memory cell 610 is identical to the memory cell 410 of FIG. 4, except that the memory cell 610 does not have a separate control gate terminal. (Erasure occurs through the use of an erase gate terminal) The erase operation and the read operation are the same as those of FIG. 4, except that no control gate bias is applied. The programming operation is also performed without a control gate bias. As a result, during the programming operation, a higher voltage must be applied to the source line terminal to compensate for the lack of control gate bias.

[0018] Table 3 shows the typical voltage ranges that can be applied to the terminals of the memory cell 610 to perform read, erase, and program operations. Table 3: Operation of the flash memory cell 610 of FIG. 6

Table 3

[0019] Figure 7 shows a stacked gate memory cell 710, which is another type of flash memory cell. The memory cell 710 is the same as the memory cell 210 of FIG. 2, except that the floating gate 20 extends over the entire channel region 18 and the control gate terminal 22 (coupled to the word line) is separated by an insulating layer (not shown) and extends above the floating gate 20. Programming is performed using hot electron injection from the channel 18 to the floating gate 20 in the channel region adjacent to the drain region 16, and erasure is performed by Fowler-Nordheim electron tunneling from the floating gate 20 to the substrate 12. The read operation operates in a manner similar to that described above for the memory cell 210.

[0020] Table 4 shows the typical voltage ranges that can be applied to the terminals of the memory cell 710 and the substrate 12 to perform read, erase, and program operations. Table 4: Operations of Flash Memory Cell 710 in FIG. 7 [Table 4]

[0021] "Read 1" is a read mode in which the cell current is output to the bit line. "Read 2" is a read mode in which the cell current is output to the source line terminal. Optionally, in an array including rows and columns of memory cells 210, 310, 410, 510, 610, or 710, the source line can be coupled to one row of memory cells or two adjacent rows of memory cells. That is, the source line terminal can be shared by adjacent rows of memory cells.

[0022] FIG. 8 shows a twin split gate memory cell 810. The twin split gate memory cell 810 includes a pair of memory cells (A on the left and B on the right), and each of the memory cells includes a floating gate (FGA, FGB) 20 disposed above and insulated from the substrate 12, a control gate 28 (CGA, CGB) disposed above and insulated from the floating gate 20, and an erase gate 30 (EG) disposed adjacent to and insulated from the floating gate and the control gate 20 / 28 and also disposed above and insulated from the substrate 12. The erase gate is formed in a T shape, and as a result, each Floating gate's F GA, F GB's upper corner faces the inner corner of each of the T-shaped erase gates to improve the erase efficiency, and a drain region 16 (DRA, DRB) in the substrate adjacent to the floating gate 20 (each drain regionIt includes bit line contacts 24 (BLA, BLB) connected to domains 16 (DRA, DRB). The memory cells are formed as pairs of memory cells that share a common erase gate 30. This cell design is different from the memory cells discussed above with reference to FIGS. 2 to 7 in that it lacks at least the source region under the erase gate EG, lacks the select gate (also referred to as the word line), and lacks the channel region of each memory cell. Instead, a single continuous channel region 18 extends under both memory cells (i.e., from the drain region 16 of one memory cell to the drain region 16 of the other memory cell). To read or program one memory cell, the control gate 28 of the other memory cell is raised to a sufficient voltage, and the voltage coupling to the floating gate 20 therebetween activates the underlying channel region portion (e.g., to read or program cell A, the voltage on FGB is raised by voltage coupling from CGB to activate the channel region under FGB). Erasure is performed using Fowler-Nordheim electron tunneling from the floating gate 20A and / or the floating gate 20B to the erase gate 30. Programming is performed using hot electron injection from the channel 18 to the floating gate 20.

[0023] Table 5 shows the typical voltage ranges that can be applied to the terminals of the memory cell 810 to perform read, erase, and program operations. Cell A (FG, CGA, BLA) is selected for read, program, and erase operations. Table 5: Operations of the flash memory cell 810 in FIG. 8

Table 5

[0024] To utilize a memory array comprising one of the above-described types of non-volatile memory cells in an artificial neural network, in one embodiment, two modifications are made. First, as further described below, the lines are configured so that each memory cell can be individually programmed, erased, and read without adversely affecting the memory states of other memory cells in the array. Second, continuous (analog) programming of the memory cells is provided.

[0025] Specifically, the memory state of each memory cell in the array (i.e., the charge on the floating gate) can be changed continuously, independently, and with minimal disturbance to other memory cells, from a completely erased state to a completely programmed state. In another embodiment, the memory state of each memory cell in the array (i.e., the charge on the floating gate) can be changed continuously, independently, and with minimal disturbance to other memory cells, from a completely programmed state to a completely erased state, or from a completely erased state to a completely programmed state. This means that the cell memory is analog or can store at least one of a number of discontinuous values (such as 16 or 256 different values), which allows for very precise and individual tuning of all cells in the memory array, and which makes the memory array ideal for storing and creating finely tuned synaptic weights of a neural network.

[0026] The methods and means described herein can be applied, without limitation, to other non-volatile memory technologies such as FINFET split-gate flash or stacked-gate flash, SONOS (silicon-oxide-nitride-oxide-silicon, charge traps in nitride), MONOS (metal-oxide-nitride-oxide-silicon, metal charge traps in nitride), ReRAM (resistance change memory), PCM (phase change memory), MRAM (magnetoresistive memory), FeRAM (ferroelectric memory), OTP (one-time programmable, binary or multi-level), and CeRAM (strongly correlated electron memory). The methods and means described herein can be applied, without limitation, to volatile memory technologies used in neural networks such as SRAM, DRAM, and other volatile synaptic cells. <<Neural Network Using a Non-Volatile Memory Cell Array>>

[0027] FIG. 9 conceptually shows a non-limiting example of a neural network utilizing the non-volatile memory array of this embodiment. This example uses a non-volatile memory array neural network for a face recognition application, but it is also possible to implement other suitable applications using a non-volatile memory array-based neural network.

[0028] S0 is the input layer, which in this example is a 32×32 pixel RGB image with 5-bit precision (i.e., three 32×32 pixel arrays, one for each of the colors R, G, and B, and each pixel has 5-bit precision). The synapses CB1 going from the input layer S0 to layer C1 apply different sets of weights to some instances and shared weights to other instances, scan the input image with overlapping 3×3 pixel filters (kernels), and shift the filter one pixel (or more than two pixels depending on the model) at a time. Specifically, the 9 pixel values in the 3×3 portion of the image (i.e., what is referred to as the filter or kernel) are provided to the synapses CB1, where these 9 input values are multiplied by appropriate weights, and after summing the outputs of that multiplication, a single output value is determined and provided by the first synapse of CB1 to generate one pixel of the feature map of layer C1. The 3×3 filter is then shifted one pixel to the right within the input layer S0 (i.e., a 3-pixel column is added on the right and a 3-pixel column is dropped on the left), and the 9 pixel values of this newly positioned filter are provided to the synapses CB1, where they are multiplied by the same weights as above and a second single output value is determined by the relevant synapse. This process is continued until the 3×3 filter has scanned over the entire 32×32 pixel image of the input layer S0 for all three colors and all bits (precision values). The process is then repeated using different sets of weights until all of the feature maps of layer C1 are calculated, generating different feature maps of C1.

[0029] In this example, in layer C1, there are 16 feature maps each having 30×30 pixels. Each pixel is a new feature pixel extracted from the multiplication of the input and the kernel. Thus, each feature map is a two-dimensional array. Therefore, in this example, layer C1 consists of 16 layers of two-dimensional arrays (note that the layers and arrays referred to in this specification are logical relationships rather than necessarily physical relationships, that is, the arrays are not necessarily oriented in a physical two-dimensional array). Each of the 16 feature maps in layer C1 is generated by one of 16 different sets of synaptic weights applied to the filter scan. All of the C1 feature maps can target different aspects of the same image feature, such as edge discrimination. For example, the first map (generated using a first set of weights shared by all scans used to generate this first map) can identify circular edges, and the second map (generated using a second set of weights different from the first set of weights) can identify rectangular edges or the aspect ratio of a specific feature, etc.

[0030] Before going from layer C1 to layer S1, an activation function P1 (pooling) is applied that pools values from non-overlapping, consecutive 2×2 regions within each feature map. The purpose of the pooling function P1 is to average neighboring positions (or it is also possible to use the max function), for example, to reduce the dependence on edge positions, and to reduce the data size before going to the next stage. In layer S1, there are 16 15×15 feature maps (i.e., 16 different arrays of 15×15 pixels each). The synapses CB2 going from layer S1 to layer C2 scan the maps within layer S1 with a 4×4 filter with a 1-pixel filter shift. In layer C2, there are 22 12×12 feature maps. Before going from layer C2 to layer S2, an activation function P2 (pooling) is applied that pools values from non-overlapping, consecutive 2×2 regions within each feature map. In layer S2, there are 22 6×6 feature maps. In the synapses CB3 going from layer S2 to layer C3, an activation function (pooling) is applied, where all neurons within layer C3 are connected to all maps within layer S2 via their respective synapses of CB3. In layer C3, there are 64 neurons. The synapses CB4 going from layer C3 to the output layer S3 fully connect C3 to S3, i.e., all neurons within layer C3 are connected to all neurons within layer S3. The output at S3 contains 10 neurons, where the neuron with the highest output determines the class (classification). This output can, for example, indicate the identification or classification of the content of the original image.

[0031] Each layer of synapses is implemented using an array or a part of an array of non-volatile memory cells.

[0032] Figure 10 is a block diagram of a system that can be used for that purpose. The VMM system 32 includes non-volatile memory cells and is utilized as synapses (such as CB1, CB2, CB3, and CB4 in FIG. 6) between one layer and the next. Specifically, the VMM system 32 includes a VMM array 33 comprising non-volatile memory cells arranged in rows and columns, an erase gate and word line gate decoder 34, a control gate decoder 35, a bit line decoder 36, and a source line decoder 37, and these decoders decode respective inputs to the non-volatile memory cell array 33. Inputs to the VMM array 33 can be made from the erase gate and word line gate decoder 34 or from the control gate decoder 35. The source line decoder 37 in this example also decodes the output of the VMM array 33. Alternatively, the bit line decoder 36 can decode the output of the VMM array 33.

[0033] The VMM array 33 serves two purposes. First, it stores the weights used by the VMM system 32. Second, the VMM array 33 effectively multiplies the inputs by the weights stored in the VMM array 33 and sums them for each output line (source line or bit line) to generate an output, which becomes the input to the next layer or the input to the last layer. By performing the multiplication and addition functions, the VMM array 33 eliminates the need for separate multiplication and addition logic circuits and is also power-efficient due to in-memory computing.

[0034] The output of the VMM array 33 is supplied to a differential summer (such as a summing operational amplifier or a summing current mirror) 38 that sums the outputs of the non-volatile memory cell array 33 to create a single value for convolution. The differential summer 38 is arranged to perform the summation of both the positive and negative weight inputs and output a single value.

[0035] The combined output value of the differential combiner 38 is then supplied to an activation function circuit 39 that rectifies the output. The activation function circuit 39 can provide a sigmoid function, a tanh function, a ReLU function, or any other non-linear function. The rectified output value of the activation function circuit 39 becomes an element of the feature map of the next layer (e.g., C1 in FIG. 8) and is then applied to the next synapse to generate the next feature map layer or the last layer. Thus, in this example, the VMM array 33 constitutes a plurality of synapses (which receive inputs from the previous layer of neurons or from an input layer such as an image database), and the combiner 38 and the activation function circuit 39 constitute a plurality of neurons.

[0036] The inputs (WLx, EGx, CGx, and optionally BLx and SLx) to the VMM system 32 of FIG. 10 can be at an analog level, a binary level, a digital pulse (in which case a pulse - analog converter PAC may be required to convert the pulse to an appropriate input analog level), or a digital bit (in which case a DAC is provided to convert the digital bit to an appropriate input analog level), and the output can be at an analog level (e.g., current, voltage, or charge), a binary level, a digital pulse, or a digital bit (in which case an output ADC is provided to convert the output analog level to a digital bit).

[0037] FIG. 11 is a block diagram showing the use of multiple layers of the VMM system 32, labeled as VMM systems 32a, 32b, 32c, 32d, and 32e in the figure. As shown in FIG. 11, an input (denoted as Inputx) is converted from digital to analog by a digital-to-analog converter 31 and provided to the input VMM system 32a. The converted analog input can be a voltage or a current. The input D / A conversion of the first layer can be performed by using a function or a LUT (look-up table) that maps the input Inputx to an appropriate analog level of the matrix multiplier of the input VMM system 32a. The input conversion can also be performed by an analog-to-analog (A / A) converter to convert an external analog input to the mapped analog input to the input VMM system 32a. The input conversion can also be performed by a digital-to-digital pulse (D / P) converter to convert an external digital input to the mapped digital pulse(s) to the input VMM system 32a.

[0038] The output generated by the input VMM system 32a is then provided as input to the next VMM system (hidden level 1) 32b, which then generates an output that is in turn provided as input to the next input VMM system (hidden level 2) 32c, and so on. The various layers of the VMM system 32 function as the respective layers of synapses and neurons of a convolutional neural network (CNN). Each of the VMM systems 32a, 32b, 32c, 32d, and 32e can be a stand-alone physical system with its corresponding non-volatile memory array, or multiple VMM systems can utilize different portions of the same physical non-volatile memory array, or multiple VMM systems can utilize overlapping portions of the same physical non-volatile memory array. Each VMM system 32a, 32b, 32c, 32d, and 32e can also be time-multiplexed with respect to the various portions of its array or neurons. The example shown in FIG. 11 includes five layers (32a, 32b, 32c, 32d, 32e), namely, one input layer (32a), two hidden layers (32b, 32c), and two fully connected layers (32d, 32e). One skilled in the art will understand that this is merely exemplary and that the system can alternatively include more than two hidden layers and more than two fully connected layers. <<VMM Array>>

[0039] FIG. 12 shows a neuron VMM array 1200 that is particularly suitable for the memory cell 310 shown in FIG. 3 and is utilized as a portion of synapses and neurons between the input layer and the next layer. The VMM array 1200 includes a memory array 1201 of non-volatile memory cells and a reference array 1202 of non-volatile reference memory cells (located at the top of the array). Alternatively, another reference array can be located at the bottom.

[0040] In the VMM array 1200, control gate lines such as control gate line 1203 extend vertically (thus, the reference array 1202 in the row direction is orthogonal to the control gate line 1203), and erase gate lines such as erase gate line 1204 extend horizontally. Here, the input to the VMM array 1200 is provided to the control gate lines (CG0, CG1, CG2, CG3), and the output of the VMM array 1200 appears on the source lines (SL0, SL1). In one embodiment, only even rows are used, and in another embodiment, only odd rows are used. The current applied to each source line (SL0 and SL1 respectively) performs a summation function of all the currents from the memory cells connected to that particular source line.

[0041] As described herein for the neural network, the non-volatile memory cells of the VMM array 1200 are preferably configured to operate in the subthreshold region.

[0042] The non-volatile reference memory cells and non-volatile memory cells described herein are biased in the subthreshold region: Ids = Io * e (Vg-Vth) / nVt = w * Io * e (Vg) / nVt where w = e (-Vth) / nVt and

[0043] where Ids is the drain-source current, Vg is the gate voltage of the memory cell, Vth is the threshold voltage of the memory cell, Vt is the thermal voltage = k * T / q, where k is the Boltzmann constant, T is the Kelvin temperature, q is the electron charge, n is the slope factor = 1+(Cdep / Cox), Cdep = the capacitance of the depletion layer, and Cox is the capacitance of the gate oxide layer, Io is the memory cell current at a gate voltage equal to the threshold voltage, and Io is (Wt / L) * u * Cox * (n - 1) * Vt 2proportional to, where u is the carrier mobility, and Wt and L are the width and length of the memory cell, respectively.

[0044] When using an I-V log converter that converts the input current Ids to an input voltage Vg using a memory cell (such as a reference memory cell or a peripheral memory cell) or a transistor: Vg = n * Vt * log[Ids / wp * Io]

[0045] where wp is the w of the reference or peripheral memory cell.

[0046] When using an I-V log converter that converts the input current Ids to an input voltage Vg using a memory cell (such as a reference memory cell or a peripheral memory cell) or a transistor: Vg = n * Vt * log[Ids / wp * Io]

[0047] where wp is the w of the reference or peripheral memory cell.

[0048] For the memory array used as a vector matrix multiplier VMM array, the output current is as follows: Iout = wa * Io * e (Vg) / nVt i.e., Iout = (wa / wp) * Iin = W * Iin W = e (Vthp-Vtha) / nVt Iin = wp * Io * e (Vg) / nVt where for each memory cell in the memory array, wa = w, and wp is the w of the reference or peripheral memory cell.

[0049] The word line or control gate can be used as the input of the memory cell for the input voltage.

[0050] Alternatively, the non-volatile memory cells of the VMM array described herein can be configured to operate in the linear region. Ids = beta * (Vgs - Vth) * Vds; beta = u * Cox * Wt / L W α (Vgs - Vth) That is, the weight W in the linear region is proportional to (Vgs - Vth).

[0051] The word line or control gate or bit line or source line can be used as an input to the memory cell operating in the linear region. The bit line or source line can be used as an output of the memory cell.

[0052] For an I-V linear converter, a memory cell (such as a reference memory cell or a peripheral memory cell) or a transistor or a resistor operating in the linear region can be used to linearly convert the input / output current into an input / output voltage.

[0053] Alternatively, the memory cells of the VMM array described herein can be configured to operate in the saturation region. Ids = 1 / 2 * beta * (Vgs - Vth) 2 ; beta = u * Cox * Wt / L W α (Vgs - Vth) 2 , that is, the weight W is (Vgs - Vth) 2 proportional to

[0054] The word line, control gate, or erase gate can be used as an input to the memory cell operating in the saturation region. The bit line or source line can be used as an output of the output neuron.

[0055] Alternatively, the memory cells of the VMM arrays described herein can be used in all regions or combinations thereof (subthreshold, linear, or saturation) for each layer or layers of a neural network.

[0056] FIG. 13 shows a neuron VMM array 1300, particularly suitable for the memory cells 210 shown in FIG. 2, used as synapses between an input layer and the next layer. The VMM array 1300 comprises a memory array 1303 of non-volatile memory cells, a reference array 1301 of a first non-volatile reference memory cell, and a reference array 1302 of a second non-volatile reference memory cell. The reference arrays 1301 and 1302 arranged in the columns of the array serve to convert current inputs flowing into terminals BLR0, BLR1, BLR2, and BLR3 into voltage inputs WL0, WL1, WL2, and WL3. In practice, the first and second non-volatile reference memory cells are diode-connected through a multiplexer 1314 (only some of which are shown) with current inputs flowing into them. The reference cells are tuned (e.g., programmed) to a target reference level. The target reference level is provided by a reference mini-array matrix (not shown).

[0057] The memory array 1303 serves two purposes. First, it stores the weights used by the VMM array 1300 in each memory cell. Second, it effectively multiplies the inputs (i.e., the current inputs provided to terminals BLR0, BLR1, BLR2, and BLR3, which the reference arrays 1301 and 1302 convert to input voltages and provide to word lines WL0, WL1, WL2, and WL3) by the weights stored in the memory cell array 1303, and then adds all the results (memory cell currents) to generate outputs for each bit line (BL0-BLN), which are the inputs to the next layer or the last layer. With the memory array 1303 performing the multiplication and addition functions, the need for separate multiplication and addition logic is eliminated and is also power efficient. Here, voltage inputs are provided to word lines WL0, WL1, WL2, and WL3, and outputs appear on the respective bit lines BL0-BLN during read (inference) operations. The current applied to each bit line BL0-BLN performs a summation function of the currents from all the non-volatile memory cells connected to that particular bit line.

[0058] Table 6 shows the operating voltages for VMM array 1300. The columns in the table show the voltages applied to the word line of the selected cell, the word lines of the unselected cells, the bit lines of the selected cells, the bit lines of the unselected cells, the source lines of the selected cells, and the source lines of the unselected cells, and FLT indicates floating, i.e., no voltage is applied. The rows show the operations of read, erase, and program. Table 6: Operation of VMM Array 1300 in FIG. 13 [Table 6]

[0059] FIG. 14 shows a neuron VMM array 1400 that is particularly suited for the memory cells 210 shown in FIG. 2 and is utilized as part of the synapses and neurons between the input layer and the next layer. The VMM array 1400 comprises a memory array 1403 of non-volatile memory cells, a reference array 1401 of first non-volatile reference memory cells, and a reference array 1402 of second non-volatile reference memory cells. The reference arrays 1401 and 1402 run in the row direction of the VMM array 1400. The VMM array is similar to the VMM 1300, except that the word lines run vertically in the VMM array 1400. Here, inputs are provided to the word lines (WLA0, WLB0, WLA1, WLB2, WLA2, WLB2, WLA3, WLB3) and outputs appear on the source lines (SL0, SL1) during a read operation. The current applied to each source line performs a summation function of all the currents from the memory cells connected to that particular source line.

[0060] Table 7 shows the operating voltages for the VMM array 1400. The columns in the table show the voltages applied to the word line of the selected cell, the word lines of the unselected cells, the bit lines of the selected cells, the bit lines of the unselected cells, the source lines of the selected cells, and the source lines of the unselected cells. The rows show the operations of read, erase, and program. Table 7: Operation of VMM Array 1400 in FIG. 14 [Table 7]

[0061] 15 shows a neuron VMM array 1500, which is particularly suited for the memory cells 310 shown in FIG. 3 and is used as part of synapses and neurons between the input layer and the next layer. The VMM array 1500 comprises a memory array 1503 of non-volatile memory cells, a reference array 1501 of a first non-volatile reference memory cell, and a reference array 1502 of a second non-volatile reference memory cell. The reference arrays 1501 and 1502 function to convert current inputs flowing into terminals BLR0, BLR1, BLR2, and BLR3 into voltage inputs CG0, CG1, CG2, and CG3. In practice, the first and second non-volatile reference memory cells are diode-connected through a multiplexer 1512 (only a portion of which is shown), with the current inputs flowing through BLR0, BLR1, BLR2, and BLR3. Multiplexer 1512 each includes a respective multiplexer 1505 and cascoding transistor 1504 to ensure a constant voltage on the bit line (e.g., BLR0) of each of the first and second non-volatile reference memory cells during a read operation. The reference cells are tuned to a target reference level.

[0062] The memory array 1503 serves two purposes. First, it stores the weights used by the VMM array 1500. Second, the memory array 1503 multiplies the inputs (current inputs provided to terminals BLR0, BLR1, BLR2, and BLR3, which the reference arrays 1501 and 1502 convert to input voltages and provide to the control gates (CG0, CG1, CG2, and CG3)) by the weights stored in the memory cell array, and then adds all the results (cell currents) to generate an output that appears on BL0-BLN and is the input to the next layer or the input to the last layer. Having the memory array perform the multiplication and addition functions eliminates the need for separate multiplication and addition logic circuits and is also power efficient, where the inputs are provided to the control gate lines (CG0, CG1, CG2, and CG3) and the outputs appear on the bit lines (BL0-BLN) during read operations. The current applied to each bit line performs a summation function of all the currents from the memory cells connected to that particular bit line.

[0063] The VMM array 1500 implements one-way tuning of the non-volatile memory cells in the memory array 1503. That is, each non-volatile memory cell is erased and then partially programmed until the desired charge on the floating gate is reached. If too much charge is added to the floating gate (in which case an incorrect value will be stored in the cell), the cell is erased and the series of partial programming operations is restarted. As shown, two rows that share the same erase gate (such as EG0 or EG1) must be erased together (known as a page erase), and then each cell is partially programmed until the desired charge on the floating gate is reached.

[0064] Table 8 shows the operating voltages for the VMM array 1500. The columns in the table show the voltages applied to the word line of the selected cell, the word lines of the unselected cells, the bit line of the selected cell, the bit lines of the unselected cells, the control gate of the selected cell, the control gates of the unselected cells in the same sector as the selected cell, the control gates of the unselected cells in a different sector than the selected cell, the erase gate of the selected cell, the erase gates of the unselected cells, the source line of the selected cell, and the source lines of the unselected cells. The rows show the read, erase, and program operations. Table 8: Operation of VMM Array 1500 in FIG. 15 [Table 8]

[0065] 16 shows a neuron VMM array 1600 that is particularly suited for memory cells 310 shown in FIG. 3 and is used as part of synapses and neurons between an input layer and the next layer. The VMM array 1600 includes a memory array 1603 of non-volatile memory cells and 、the 1 non-volatile reference memory cell reference array 1601and a second reference array 1602 of non-volatile reference memory cells. EG lines EGR0, EG0, EG1, and EGR1 run vertically, and CG lines CG0, CG1, CG2, and CG3 and SL lines WL0, WL1, WL2, and WL3 run horizontally. VMM array 1600 is similar to VMM array 1600 except that VMM array 1600 implements bidirectional tuning, where each individual cell can be fully erased, partially programmed, and partially erased as needed to reach a desired amount of charge on the floating gate through the use of individual EG lines. As shown, reference arrays 1601 and 1602 convert input currents in terminals BLR0, BLR1, BLR2, and BLR3 into control gate voltages CG0, CG1, CG2, and CG3 (through the action of diode-connected reference cells via multiplexer 1614) that are applied to the memory cells in the row direction. The current outputs (neurons) are in bit lines BL0-BLN, each of which sums all the currents from the non-volatile memory cells connected to that particular bit line.

[0066] Table 9 shows the operating voltages for the VMM array 1600. The columns in the table show the voltages applied to the word line of the selected cell, the word lines of the unselected cells, the bit line of the selected cell, the bit lines of the unselected cells, the control gate of the selected cell, the control gates of the unselected cells in the same sector as the selected cell, the control gates of the unselected cells in a different sector than the selected cell, the erase gate of the selected cell, the erase gates of the unselected cells, the source line of the selected cell, and the source lines of the unselected cells. The rows show the read, erase, and program operations. Table 9: Operation of VMM Array 1600 in FIG. 16 [Table 9]

[0067] The input to the VMM array can be analog levels, binary levels, timing pulses, or digital bits, and the output can be analog levels, binary levels, timing pulses, or digital bits (in this case an output ADC is required to convert the output analog level current or voltage to digital bits).

[0068] For each memory cell in the VMM array, each weight W can be implemented by a single memory cell, or by a differential cell, or by two blended memory cells (an average of two or more cells). In the case of differential cells, two memory cells are required to implement weight W as a differential weight (W=W+-W-). In the case of two blended memory cells, two memory cells are required to implement weight W as an average of two cells.

[0069] One drawback of prior art arrays of non-volatile memory cells is that there is a large variation in the source impedance of the array and along the output lines (such as bit lines) of the array, with a resulting variation in accuracy and power consumption depending on which cell and its state is selected for a read, program, or erase operation. Another drawback is that they can be susceptible to noise.

[0070] What is needed is an improved VMM system that has low susceptibility to noise.

[0071] What is further needed is an improved VMM system that has a substantially constant source impedance of the array during operation (read, program, or erase) regardless of which cell or cells are selected.

[0072] What is further needed is an improved VMM system that has substantially constant power consumption during an operation (read, program, or erase) regardless of which cell or cells are selected. Summary of the Invention

[0073] Numerous embodiments of an analog neural memory array are disclosed. In certain embodiments, each memory cell within the array has a substantially constant source impedance while the cell is operating. In certain embodiments, power consumption is substantially constant from bit line to bit line within the array when the cell is read. In certain embodiments, weight mapping is performed adaptively for optimal performance in power and noise.

[0074] In one embodiment, an analog neural memory system is an array of non-volatile memory cells, where the cells are arranged in rows and columns, a first plurality of columns of cells are connected to bit lines, and a second plurality of columns of cells are connected to dummy bit lines, an array of non-volatile memory cells, and a first set of bit line transistors disposed at a first end of the array, each of the first set of bit line transistors being coupled to one of the dummy bit lines, a first set of bit line transistors, and a second set of bit line transistors disposed at a second end of the array opposite the first end of the array, each of the second set of bit line transistors being coupled to one of the bit lines, a second set of bit line transistors.

[0075] In another embodiment, an analog neural memory system comprises an array of non-volatile memory cells, where the cells are arranged in rows and columns, the columns are arranged in physically adjacent pairs, one column of the pair stores a W+ value, and one column of the pair stores a W- value.

[0076] In another embodiment, an analog neural memory system is a first array of non-volatile memory cells, where the cells are arranged in rows and columns, and the non-volatile memory cells within one or more of the columns store W+ values, a first array of non-volatile memory cells, and a second array of non-volatile memory cells, where the cells are arranged in rows and columns, and the non-volatile memory cells within one or more of the columns store W- values, a second array of non-volatile memory cells.

[0077] In another embodiment, an analog neural memory system includes a first array of non-volatile memory cells, the cells in the first array arranged in rows and columns, a first plurality of columns of cells in the first array connected to bit lines and a second plurality of columns of cells in the first array connected to dummy bit lines; and a second array of non-volatile memory cells, the cells in the second array arranged in rows and columns, a third plurality of columns of cells in the second array connected to bit lines and a fourth plurality of columns of cells in the second array connected to dummy bit lines. a set of bit line transistors disposed at a first end of the first array, each of the first set of bit line transistors being coupled to one of the bit lines; and a set of multiplexers disposed at a second end of the first array and coupled to ground, each of the first set of dummy bit lines being coupled to one of the multiplexers in the set of multiplexers, wherein impedances of the first array and the second array remain substantially constant when cells attached to the bit lines are selected for operation.

[0078] In another embodiment, an analog neural memory system comprises an array of non-volatile memory cells, the cells arranged in rows and columns, a first plurality of columns of cells connected to bit lines and a second plurality of columns of cells connected to dummy bit lines; a first set of bit line transistors disposed at a first end of the array, each of the first set of bit line transistors coupled to one of the dummy bit lines and configured to pull the coupled dummy bit line to ground; and a second set of bit line transistors disposed at a second end of the array opposite the first end of the array, each of the second set of bit line transistors coupled to one of the bit lines.

[0079] In another embodiment, an analog neural memory system includes an array of non-volatile memory cells, the cells arranged in rows and columns, a first plurality of columns of cells connected to bit lines and a second plurality of columns of cells connected to dummy bit lines; a first set of bit line transistors arranged at a first end of the array, each of the first set of bit line transistors coupled to one of the dummy bit lines; a second set of bit line transistors arranged at a second end of the array opposite the first end of the array, each of the second set of bit line transistors coupled to one of the bit lines; a source line coupled to a source line terminal of the row of non-volatile memory cells; and a source line transistor configured to pull the source line to ground.

[0080] In another embodiment, an analog neural memory system comprises an array of non-volatile memory cells, the cells arranged in rows and columns, the columns comprising a plurality of subsets of columns, each subset comprising columns storing W+ values, columns storing W− values, and redundant columns, and values ​​stored in the W+ or W− columns are remapped to the redundant columns.

[0081] In another embodiment, an analog neural memory system comprises an array of non-volatile memory cells, the cells arranged in rows and columns, the columns comprising a plurality of subsets of columns, each subset comprising a column storing a weight W value and a redundant column, and values ​​stored in the W column are remapped to the redundant column.

[0082] In another embodiment, a method of operating an analog neural memory system comprising a first array of non-volatile memory cells arranged in rows and columns, a second array of non-volatile memory cells arranged in rows and columns, a column multiplexer, local bit lines, and global bit lines, the method comprising: selecting, by the column multiplexer, a local bit line of the first array or a local bit line of the second array and coupling the selected bit line to the global bit line; and multiplexing adjacent global bit lines onto the current global bit line to effectively increase the width of the global bit line.

[0083] In another embodiment, a method of operating an analog neural memory system comprising an array of non-volatile memory cells arranged in rows and columns, bit lines, and an array output circuit, the method comprising combining a first analog-to-digital converter of a first array output with a second analog-to-digital converter of a second array output to form a third analog-to-digital converter.

[0084] In another embodiment, a method of operating an analog neural memory system comprising an array of non-volatile memory cells arranged in rows and columns, bit lines, an array output circuit, and a configurable n-bit analog-to-digital converter is disclosed, the method comprising configuring the n-bit analog-to-digital converter to have a lower accuracy than n bits or a higher accuracy than n bits.

[0085]

[0086]

[0087]

[0088]

[0089]

[0090]

[0091]

[0092]

[0093]

[0094]

[0095]

[0096]

[0097]

[0098]

[0099]

[0100]

[0101]

[0102]

[0103]

[0104]

[0105]

[0106]

[0107]

[0108]

[0109]

[0110]

[0111]

[0112]

[0113]

[0114]

[0115]

[0116]

[0117]

[0118]

[0119] [Brief description of the drawings]

[0120]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18A

Figure 18B

Figure 18C

Figure 19A

Figure 19B

Figure 19C

Figure 20

Figure 21A

Figure 21B

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27A

Figure 27B

Figure 28A

Figure 28B

Figure 28C

Figure 29

Figure 30

Figure 31

Figure 32

Figure 33

Embodiments for Carrying Out the Invention

[0121] The artificial neural network of the present invention utilizes a combination of CMOS technology and a non-volatile memory array. <<Embodiment of an Improved VMM System>>

[0122] FIG. 17 shows a block diagram of a VMM system 1700. The VMM system 1700 includes a VMM array 1701, a row decoder 1702, a high-voltage decoder 1703, a column decoder 1704, a bit-line driver 1705, an input circuit 1706, an output circuit 1707, control logic 1708, and a bias generator 1709. The VMM system 1700 further includes a high-voltage generation block 1710 including a charge pump 1711, a charge pump regulator 1712, and a high-voltage level generator 1713. The VMM system 1700 further includes an algorithm controller 1714, an analog circuit 1715, control logic 1716, and test control logic 1717. The systems and methods described below can be implemented in the VMM system 1700.

[0123] The input circuit 1706 may include circuits such as a DAC (digital-to-analog converter), a DPC (digital-to-pulse converter), a DTC (digital-to-time converter), an AAC (analog-to-analog converter such as a current-to-voltage converter), a PAC (pulse-to-analog level converter), or any other type of converter. The input circuit 1706 may implement a normalization function, a scaling function, or an arithmetic function. The input circuit 1706 may implement a temperature compensation function for the input, such as modulating the output voltage / current / time / pulse as a function of temperature. The input circuit 1706 may implement an activation function such as ReLU or sigmoid.

[0124] The output circuit 1707 may include circuits such as an ADC (analog-to-digital converter for converting neuron analog outputs to digital bits), an AAC (analog-to-analog converter such as a current-to-voltage converter), an ATC (analog-to-time converter), an APC (analog-to-pulse converter), or any other type of converter. The output circuit 1707 may implement an activation function such as ReLU or sigmoid. The output circuit 1707 may implement a statistical normalization, regularization, up / down scaling function, statistical rounding, or arithmetic function (e.g., addition, subtraction, division, multiplication, shift, logarithm) on the neuron output that is the output of the VMM array 1701. The output circuit 1707 may implement a temperature compensation function on the neuron output (voltage / current / time / pulse, etc.) or the array output (bit line output, etc.) in order to improve the accuracy of the VMM array 1701 (neuron) output, such as by keeping the power consumption of the VMM array 1701 substantially constant or by keeping the I-V gradient substantially the same.

[0125] FIG. 18A shows a prior art VMM system 1800. The VMM system 1800 includes exemplary cells 1801 and 1802, exemplary bit line switches 1803a, 1803b, 1803c, 1803d (connecting the bit lines to a sense circuit), exemplary dummy bit line switches 1804a, 1804b, 1804c, 1004d (coupled to a low bias level such as a ground (or near ground) level during read), and exemplary dummy cells 1805 and 1806 (source line pull-down cells). Bit line switch 1803a is coupled to a column of cells including cells 1801 and 1802 that are used to store data in the VMM system 1800. Dummy bit line switches 1804a, 1804b, 1804c, 1804 are each coupled to a column of cells (bit lines) that are dummy cells and are not used to store data in the VMM system 1800. This dummy bit line, which may also be referred to as a source line pull-down bit line, is used as a source line pull-down during a read operation, which means it is used to pull the source line to a low bias level such as ground (or near ground) through the dummy cells of the dummy bit line. Note that both the dummy bit line switches 1804a, 1804b, 1804c, 1804, and the bit line switches 1803a, 1803b, 1803c, 1803d appear at the same end of the array, i.e., they all appear at a common end of the column of cells to which they are coupled and are thus arranged in a single row.

[0126] One drawback of the VMM system 1800 is that the input impedance of each cell varies greatly due to the length of the electrical path through the associated bit line switch, the cell itself, and the associated dummy bit line switch. For example, FIG. 18B shows the electrical path through bit line switch 1803, cell 1801, dummy cell 1805, and dummy bit line switch 1804. Similarly, FIG. 18C shows the electrical path through bit line switch 1803, vertical metal bit line 1807, cell 1802, dummy cell 1806, vertical metal bit line 1808, and dummy bit line switch 1804. As can be seen, the path through cell 1802 traverses significantly longer bit lines and dummy bit lines, which is associated with higher capacitance and higher resistance. This results in cell 1802 having a greater parasitic impedance in the bit line or source line than cell 1801 in FIG. 18B. This variability is a drawback because it causes variations in the accuracy of the cell output applied to cell readout or verification (for program / erase tuning cycles), for example, depending on the position of the cell within the array.

[0127] FIG. 19A shows a VMM system 1900 that improves upon the prior art VMM system 1800. The VMM system 1900 includes exemplary cells 1901 and 1902, exemplary bit line switches 1903a, 1903b, 1903c, and 1903d that connect the bit lines to a sense circuit, exemplary dummy cells 1905 and 1906 that can function as source line pull-down cells, and exemplary dummy bit line switches 1904a, 1904b, 1904c, and 1904d. As an example, during a read operation, one end of dummy bit line switch 1904a is connected to a low voltage level such as ground, and the other end is connected to dummy cells 1905 and 1906 that are used as source line pull-downs. As can be seen, exemplary dummy bit line switch 1904a and the other dummy bit line switches are located at the opposite end of the array from bit line switches 1903a and the other bit line switches.

[0128] The advantages of this design can be seen in FIGS. 19B and 19C. In FIG. 19B, cell 1901 is selected for reading, and in FIG. 19C, cell 1902 is selected for reading.

[0129] FIG. 19B shows the electrical path through bit line switch 1903, cell 1901, dummy cell 1905 (source line pull-down cell), vertical metal bit line 1908, and dummy bit line switch 1904 (coupled to a low level such as ground during the read operation). FIG. 19C shows the electrical path through bit line switch 1903, vertical metal line 1907, cell 1902, dummy cell 1906 (source line pull-down cell), and dummy bit line switch 1904. The paths are substantially the same with respect to the interconnect length, which holds for all cells of the VMM system 1900. As a result, the impedance of the bit line impedance + source line impedance of each cell is substantially the same, which means that the variation in the amount of parasitic voltage drop drawn during the read or verification of the operation of each cell in the array is substantially the same.

[0130] FIG. 20 shows a VMM system 2000 with global source line pull-down bit lines. The VMM system 2000 is the same as the VMM system 1900 except that dummy bit lines 2005a - 2005n or 2007a - 2007n are connected together (to act as global source line pull-down lines for pulling the memory cell source lines to the ground level during read or verification), dummy bit line switches such as dummy bit line switches 2001 and 2002 are connected or coupled to a common ground labeled ARYGND (array ground), and the source lines are coupled together to a source line switch 2004 that selectively pulls the source lines to ground. These changes further reduce the variation in the parasitic impedance of each cell in the array during the read or verification operation.

[0131] In an alternative embodiment, one or more dummy bit lines and one or more dummy bit line switches can be used in place of the source line switch 2004 to pull the source line to ground.

[0132] In another embodiment, dummy rows can be utilized between rows as a physical barrier to avoid FG-FG coupling (between two adjacent cells).

[0133] FIG. 21A shows a VMM system 2100. In some embodiments, the weights W stored in the VMM are stored as a differential pair, W+ (positive weight) and W− (negative weight), where W = (W+)−(W−). In the VMM system 2100, half of the bit lines are designated as W+ lines, i.e., the bit lines connected to the memory cells storing the positive weight W+, and the other half of the bit lines are designated as W− lines, i.e., the bit lines connected to the memory cells implementing the negative weight W−. The W− lines are interspersed alternately between the W+ lines. The subtraction operation is performed by summing circuits that receive current from the W+ and W− lines, such as summing circuits 2101 and 2102. The outputs of the W+ lines and the outputs of the W− lines are combined together to effectively give W = W+−W− for each pair of (W+, W−) cells for all pairs of (W+, W−) lines. Optionally, dummy bit lines and source line pull-down bit lines, such as those shown in FIGS. 19 and 20, can be used in the VMM system 2100 to avoid FG-FG coupling (between two adjacent cells) and / or reduce the IR voltage drop of the source line during read or verify operations.

[0134] FIG. 21B shows another embodiment. In the VMM system 2110, the positive weight W+ is implemented in the first array 2111, the negative weight W- is implemented in a second array 2112 separate from the first array, and the resulting weights are appropriately combined together by the summing circuit 2113. Optionally, dummy bit lines and source line pull-down bit lines such as those shown in FIGS. 19 and 20 can be used in the VMM system 2110 to avoid FG-FG coupling and / or reduce the IR voltage drop of the source line during a read or verify operation.

[0135] The VMM system can be designed such that pairs of W+ and W- are arranged within the array in a manner that reduces FG-to-FG coupling or distributes power consumption in a more uniform manner across the array and output circuitry. This is described below with reference to Tables 10 and 11. Additional details regarding the FG-to-FG coupling phenomenon can be found in U.S. Patent Provisional Application No. 62 / 981,757, filed on February 26, 2020, by the same assignee and entitled "Ultra-Precise Tuning of Analog Neural Memory Cells in a Deep Learning Artificial Neural Network", which is incorporated herein by reference.

[0136] Table 10A shows an exemplary physical layout of the placement of two pairs of (W+, W-) bit lines. One pair is BL0 and BL1, and the second pair is BL2 and BL3. In this example, four rows are coupled to the source line pull-down bit line BLPWDN. BLPWDN is placed between each pair of (W+, W-) bit lines to prevent coupling (e.g., FG-to-FG coupling) between one pair of (W+, W-) bit lines and another pair of (W+, W-) bit lines. Thus, BLPWDN functions as a physical barrier between pairs of (W+, W-) bit lines. Table 10A: Exemplary Layout of W, W- Pairs

Table 10A

[0137] Table 10B shows different exemplary combinations of weights. A "1" means that the cell is used and has an actual output value, and a "0" means that the cell is not used and has no value or no significant output value. Table 10B: Exemplary Weight Combinations for W+, W- Pairs [Table 10B]

[0138] Table 11A shows another array embodiment of the physical arrangement of the (w+, w-) pair lines BL0 / 1 and BL2 / 3. The array includes redundant lines BL01 and BL23 and a source line pull-down bit line BLPWDN. The redundant bit line BL01 is used to remap a value from the pair BL0 / 1, and the redundant bit line BL23 is used to remap a value from the pair BL2 / 3, which is shown in a later table. Table 11A: Exemplary Layout for W+, W- Pairs [Table 11A]

[0139] Table 11B shows an example where the distributed weight values do not require remapping, basically there are no adjacent "1"s between adjacent bit lines. Table 11B: Exemplary Weight Combinations for W+, W- Pairs [Table 11B]

[0140] Table 11C shows an example where the distributed weights need to be remapped. Here, there are "1"s adjacent to BL1 and BL3, which causes adjacent bit line couplings. Therefore, this value is remapped as shown in Table 11D, resulting in no adjacent "1" values between adjacent bit lines. Additionally, the remapping reduces the total current along the bit line, leading to a more accurate value in that bit line, which also leads to a more distributed power consumption along the bit line. Optionally, additional bit lines (BL01, BL23) can be used, optionally acting as redundant columns. Table 11C: Exemplary weight combinations of w+, w-

Table 11C

Table 11D

[0141] Tables 11E and 11F show another embodiment of remapping noisy cells (or defective cells) to redundant columns such as BL01, BL23 in Table 11E or BL0B and BL1B in Table 11F. Table 11E: Remapped weight combinations of w+, w-

Table 11E

Table 11F

[0142] Table 11G shows an embodiment of the physical layout of an array suitable for FIG. 21B. Since each bit line has either a positive or negative weight, a dummy bit line acting as a source line pull-down, or an actual dummy bit line (not used, e.g., deeply or partially programmed, or partially erased), and a physical barrier to avoid FG-FG coupling are required for each bit line. Table 11G: Exemplary layout of w+, w- pairs

Table 11G

[0143] In another embodiment, a tuning bit line coupled to a column of cells is adjacent to a target bit line coupled to the column of cells, and the tuning bit line cells are used to tune the target bit line cells to a desired target value during a programming operation using FG-FG coupling between adjacent cells. Optionally, the source line pull-down bit line can be used on the side of the target bit line opposite the side adjacent to the tuning bit line.

[0144] Alternative embodiments for mapping noisy or defective cells can be implemented when such cells are designated as unused cells, meaning that they are (deeply) programmed so that they do not contribute any value to the neuron output.

[0145] Alternative embodiments for identifying fast cells (cells that can be programmed to reach a particular value faster than typical cells) can be implemented, where the fast cells are identified and receive a more accurate tuning algorithm so as not to overshoot the target during the programming operation.

[0146] FIG. 22 shows a VMM system 2200. The VMM system 2200 includes a redundant array 2201 that can be included in any of the VMM arrays previously considered. The redundant array 2201 can be used as redundancy for replacing a defective column if any of the columns attached to the bit line switches are considered defective. The redundant array can have its own redundant array (neuron) output (e.g., bit line), and / or redundant write and verification circuitry, and / or ADC circuitry for redundancy purposes. For example, if redundancy is required, the output of the redundant ADC replaces the output of the ADC for the bad bit line. The redundant array 2201 can also be used for weight mapping as described in Tables 10A and 10B to achieve a relatively uniform power distribution among the bit lines.

[0147] FIG. 23 shows a VMM system 2300 including an array 2301, an array 2302, a column multiplexer 2303, local bit lines LBL2305a-2305d, global bit lines GBL2308 and 2309, and dummy bit line switches 2305a-2305d. The column multiplexer 2303 is used to select each top local bit line 2305 of the array 2301 or the bottom local bit line 2305 of the array 2302 to the global bit line 2308. In one embodiment, the (metal) global bit line 2308 has the same number of lines as the number of local bit lines, for example, 8 or 16 lines. In another embodiment, the global bit line 2308 has only one (metal) line per N local bit lines, such as one global bit line per 8 or 16 local bit lines. The column multiplexer 2303 can multiplex adjacent global bit lines (such as GBL2309) to the target global bit line (such as GBL2308) to effectively increase the width of the current global bit line. This reduces the voltage drop across the target global bit line (GBL2308).

[0148] Here, various output circuits that can be used in conjunction with any of the VMM systems described herein are described.

[0149] FIG. 24 shows a VMM system 2400. The VMM system 2400 includes an array 2410, a shift register (SR) 2401, a digital-to-analog converter (DAC) 2402 that receives an input from the SR 2401 and outputs an equivalent (analog or pseudo-analog) level or information (e.g., voltage / timing), an adder circuit 2403, an analog-to-digital converter (ADC) 2404, and a bit line switch (not shown). There are dummy bit lines and dummy bit line switches, but they are not shown. As shown, the ADC circuits can be combined together to create a single ADC with higher accuracy (i.e., a larger number of bits).

[0150] The adder circuit 2403 can include the circuits shown in FIGS. 25 to 27. It can include, without limitation, circuits for normalization, scaling, arithmetic operations (e.g., addition, subtraction), activation, or statistical rounding.

[0151] FIG. 25 shows a variable resistor-adjustable current-voltage adder circuit 2500 including current sources 2501-1, ..., 2501-n that respectively draw currents Ineu(1), ..., Ineu(n) (which are currents received from the bit lines of the VMM array), an operational amplifier 2502, a variable holding capacitor 2504, and a variable resistor 2503. The operational amplifier 2502 outputs a voltage Vneuout = R2503 * (Ineu(1)+...+Ineu(n)), which is proportional to the sum of the currents Ineu(1), ..., Ineu(n). The holding capacitor 2504 is used to hold the output voltage when the switch 2506 is open. This held output voltage is used, for example, to be converted to digital bits by an ADC circuit.

[0152] FIG. 26 shows a variable capacitor (basically an integrator)-adjustable current-voltage adder circuit 2600 including current sources 2601-1, ..., 2601-n that respectively draw currents Ineu(1), ..., Ineu(n) (which are currents received from the bit lines of the VMM array), an operational amplifier 2602, a variable capacitor 2603, and a switch 2604. The operational amplifier 2602 outputs a voltage Vneuout2605 = (Ineu(1)+,...,+Ineu(n)) * Integral time / C2603, which is proportional to the sum of the currents Ineu(1), ..., Ineu(n).

[0153] Figure 27A shows a voltage adder 2700 that can be adjusted by variable capacitors (i.e., switch-cap SC circuits), which includes switches 2701 and 2702, variable capacitors 2703 and 2704, an operational amplifier 2705, a variable capacitor 2706, and a switch S1 2707. When switch 2701 is closed, input Vin0 is provided to the operational amplifier 2705. When switch 2702 is closed, input Vin1 is provided to the operational amplifier 2705. Optionally, switches 2701 and 2702 are not closed simultaneously. The operational amplifier 2705 generates an output Vout that is an amplified version of the input (either Vin0 and / or Vin1 depending on which of switches 2701 and 2702 is closed). That is, Vout = Cin / Cout * (Vin), where Cin is C2703 or C2704, and Cout is C2706. For example, Vout = Cin / Cout * Σ(Vinx), Cin = C2703 = C2704, where Vinx can be either Vin0 or Vin1. In one embodiment, Vin0 is the W+ voltage and Vin1 is the W- voltage, and the voltage adder 2700 adds them together (by enabling the appropriate polarity of the switches) to generate the output voltage Vout.

[0154] Figure 27B shows a voltage adder 2750 that includes switches 2751 (S1), 2752 (S3), 2753 (S2), and 2754 (S4), a variable input capacitor 2758, an operational amplifier 2755, a variable feedback capacitor 2756, and a switch 2757 (S5). In one embodiment, Vin0 is the W+ voltage and Vin1 is the W- voltage, and the voltage adder 2750 adds them together to generate the output voltage Vout (W+ - W-, by enabling the appropriate polarity of the switches).

[0155] When the input is Vin0: When switches 2754 and 2751 are closed and switches 2753, 2752, and 2757 are open, the input Vin0 is provided to the top terminal of capacitor 2758, and its bottom terminal is connected to VREF. Then, switch 2751 is opened and switch 2753 is closed to transfer charge from capacitor 2758 to feedback capacitor 2756. Basically, thereafter, the output VOUT = (C2758 / C2756) * Vin0 (for example, when VREF = 0).

[0156] When the input is Vin1: When switches 2753, 2754, and 2757 are closed and switches 2751, 2752, and 2757 are open, both terminals of capacitor 2758 are discharged to VREF. Then, switch 2754 is opened, switch 2752 is closed, the bottom terminal of capacitor 2758 is charged to Vin1, and then the feedback capacitor 2756 is charged to VOUT = -(C2758 / C2756) * Vin1 (when VREF = 0).

[0157] Therefore, when the sequence described above for Vin0 is implemented and then the sequence described above for the Vin1 input is implemented, for example, when VREF = 0, VOUT = (C2758 / C2756) * (Vin0 - Vin1). This is used, for example, to implement W = W+ - W-.

[0158] Each ADC shown in FIG. 27 can be configured to be combined with the next ADC for higher bit implementation using an appropriate design of the ADC.

[0159] Referring again to FIG. 17, the inputs to and outputs from the VMM array 1701 can be in digital or analog format. For example: · Sequential input IN[0:q] to the DAC: · In one embodiment, the input circuit 1706 receives digital inputs in a sequence starting from IN0, then IN1, ..., and then INq. All input bits have the same VCGin. The input bits are provided to a DAC and then an analog signal is applied as an input to the VMM array 1701. All bit line (neuron) outputs are summed by a scaled binary index multiplier either before or after the ADC. · In another embodiment, a scaled neuron (bit line) binary index multiplier method is used. As shown in FIG. 20, an exemplary summer has two bit lines BL0 and Bln. The weights are distributed across a plurality of bit lines BL0 to BLn. For example, there are four bit lines BL0, BL1, BL2, BL3. The output from bit line BL0 is multiplied by 2^0 = 1. The output from bit line Bln representing the nth binary bit position is multiplied by 2^n, e.g., for n = 3, 2^3 = 8. Then, the outputs from all bit lines after being appropriately multiplied by the binary bit position 2^n are summed together. This is then digitized by an ADC. This method means that all cells have only a binary range and the multi-level range (n bits) is achieved by peripheral circuitry (meaning "by the summer circuit"). Thus, the voltage drop across all bit lines is approximately the same for the highest bias level of the memory cell. · In another embodiment, digital inputs IN0, IN1, ..., and then INq are applied in a sequential manner. Each input bit has a corresponding analog value VCGin. All neuron outputs are summed for all input bit evaluations either before or after the ADC. · Parallel input to the DAC: · In another embodiment, inputs IN0, ... INq are provided to the DAC in a parallel manner. Each input IN[0:q] has a corresponding analog value VCGin. All neuron outputs are summed by a scaled binary index multiplier method either before or after the ADC.

[0160] In an embodiment involving sequential operation of the array, power is more evenly distributed.

[0161] In an embodiment using the neuron (bit line) binary index method, since each cell coupled to the bit line contains only binary levels, power consumption is reduced within the array, and the 2^n levels are achieved by an adder circuit.

[0162]

[0163] Figures 28A, 28B, and 28C show output circuits that can be used for the adder circuit 2403 and the analog-to-digital converter 2404 of FIG. 24.

[0164] FIG. 28A shows an output circuit 2800 including an analog-to-digital converter 2802 that receives a neuron output 2801 and outputs an output digital bit 2803.

[0165] FIG. 28B shows an output circuit 2810 including a neuron output circuit 2811 and an analog-to-digital converter 2812, which together receive a neuron output 2801 and generate an output 2813.

[0166] FIG. 28C shows an output circuit 2820 including a neuron output circuit 2821 and a converter 2822, which together receive a neuron output 2801 and generate an output 2823.

[0167] The neuron output circuit 2811 or 2821 can perform, without limitation, for example, addition, scaling, normalization, or arithmetic operations. The converter 2822 can perform, without limitation, for example, ADC, PDC, AAC, or APC operations.

[0168] FIG. 29 includes an adjustable (scaling) current source 2901 and an adjustable (scaling) current source 2902, which together are the output i that is the neuron output OUTshows the neuron output circuit 2900 that generates. This circuit can perform the summation of the positive weight W+ and the negative weight W-, i.e., W = W+ - W-, and can simultaneously perform up or down scaling of the output neuron current (through the adjustment of adjustable current sources 2901 and 2902). That is, I W+ is the scaled version of W+, and I W- is the scaled version of W-.

[0169] Figure 30 shows the configurable serial analog-to-digital converter 3000. It includes an integrator 3070 that integrates the neuron output current into the integration capacitor 3002 (Cint).

[0170] In one embodiment, VRAMP3050 is provided to the inverting input of the comparator 3004. The digital output (count value) 3021 is generated by ramping VRAMP3050 until the comparator 3004 switches polarity, and the counter 3020 counts the clock pulses from the start of the ramp.

[0171] In another embodiment, VREF3055 is provided to the inverting input of the comparator 3004. VC3010 is ramped down by the ramp current 3051 (IREF) until VOUT3003 reaches VREF3055, at which point the EC3005 signal disables the count of the counter 3020. The (n-bit) ADC3000 can be configured to have lower accuracy (less than n bits) or higher accuracy (more than n bits) depending on the target application. The configurability of the accuracy is achieved by configuring, without limitation, the capacitance of the capacitor 3002, the current 3051 (IREF), the ramping speed of VRAMP3050, or the clock frequency of the clock 3041.

[0172] In another embodiment, the ADC circuit of the VMM array is configured to have a lower accuracy than n bits, and the ADC circuit of another VMM array is configured to have a higher accuracy than n bits.

[0173] In another embodiment, one instance of the serial ADC circuit 3000 of one neuron circuit is combined with another instance of the serial ADC circuit 3000 of the next neuron circuit to generate an ADC circuit having a higher accuracy than n bits, such as by combining the integrating capacitors 3002 of the two instances of the serial ADC circuit 3000.

[0174] FIG. 31 shows a configurable neuron SAR (successive approximation register) analog-to-digital converter 3100. This circuit is a successive approximation converter based on charge redistribution using binary capacitors. It includes a binary CDAC (capacitor-based DAC) 3101, an operational amplifier / comparator 3102, and SAR logic and register 3103. As shown, GndV3104 is a low voltage reference level, for example, ground level. The SAR logic and register 3103 provide a digital output 3106.

[0175] Figure 32 shows a configurable neuron-combo SAR analog-to-digital converter circuit 3200. This circuit combines two n-bit ADCs from two neuron circuits into one to achieve higher accuracy than n bits. For example, in the case of a 4-bit ADC of one neuron circuit, this circuit can achieve an accuracy exceeding 4 bits, such as 8-bit ADC accuracy, by combining two 4-bit ADCs. The combo circuit topology is equivalent to a split-cap (bridge capacitor (cap) or attenuation cap) SAR ADC circuit. For example, an 8-bit 4C-4C SAR ADC is brought about by combining two adjacent 4-bit 4C SAR ADC circuits. To achieve this, a bridge circuit 3204 (C split) is used, and the capacitance of this circuit = (total number of CDAC cap units / total number of CDAC cap units - 1).

[0176] Figure 33 shows a pipelined SAR ADC circuit 3300 that can be used to increase the number of bits in a pipelined manner in combination with the following SAR ADC. The SAR ADC circuit 3300 includes a binary CDAC (capacitor-based DAC) 3301, an operational amplifier / comparator 3302, an operational amplifier / comparator 3303, SAR logic, and a register 3304. As shown, GndV3104 is a low voltage reference level, for example, the ground level. The SAR logic and register 3103 provide a digital output 3106. Vin is at the input voltage, VREF is the reference voltage, and GndV is the ground voltage. V residue is generated by capacitor 3305 and provided as the input to the next stage of the SAR ADC.

[0177] Additional implementation details regarding circuits of configurable output neurons (such as configurable neuron ADCs) can be found in U.S. Patent Application No. 16 / 449,201, filed on June 21, 2019, by the same assignee and entitled "Configurable Input Blocks and Output Blocks and Physical Layout for Analog Neural Memory in a Deep Learning Artificial Neural Network", which is incorporated herein by reference.

[0178] As used herein, it should be noted that both the terms "over" and "on" comprehensively include "directly over" (with no intervening material, element, or gap therebetween) and "indirectly over" (with an intervening material, element, or gap therebetween). Similarly, the term "adjacent" includes "directly adjacent" (with no intervening material, element, or gap therebetween) and "indirectly adjacent" (with an intervening material, element, or gap therebetween), "attached to" includes "directly attached to" (with no intervening material, element, or gap therebetween) and "indirectly attached to" (with an intervening material, element, or gap therebetween), and "electrically coupled" includes "directly electrically coupled" (with no intervening material or element electrically connecting the elements together therebetween) and "indirectly electrically coupled" (with an intervening material or element electrically connecting the elements together therebetween). For example, forming an element "over a substrate" can include forming the element directly on the substrate without an intervening material / element therebetween and forming the element indirectly over the substrate with one or more intervening materials / elements therebetween.

Claims

1. 1. A method of operating an analog neural memory system comprising a first array of non-volatile memory cells arranged in rows and columns, a second array of non-volatile memory cells arranged in rows and columns, a column multiplexer, local bit lines, and global bit lines, the method comprising: selecting, by the column multiplexer, a local bit line of the first array or a local bit line of the second array and coupling the selected bit line to a global bit line; multiplexing the global bit line with a second global bit line adjacent to the global bit line, thereby forming a current global bit line having an effective width formed by a width of the global bit line and a width of the second global bit line.

2. 2. The method of claim 1, further comprising summing outputs from one or more of the global bit lines.

Citation Information

Patent Citations

  • Nonvolatile storage device, semiconductor device, and electronic apparatus

    JP2018195358A

  • In-memory multiply and accumulate with global charge-sharing

    US20190043560A1