Method and apparatus for programming analog neural memories in deep learning artificial neural networks

The integration of CMOS technology and non-volatile memory arrays in neural networks addresses the challenge of precise programming in VMM arrays, improving computational efficiency and power efficiency by enabling precise tuning of memory states.

JP7730398B2Active Publication Date: 2025-08-27SILICON STORAGE TECHNOLOGY INC
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2024081446
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-05-25
Filing Date
2024-05-20
Publication Date
2025-08-27
Estimated Expiration
2039-01-18

AI Technical Summary

Technical Problem

Existing artificial neural networks face challenges in achieving high-performance information processing due to the lack of suitable hardware technology, particularly in programming non-volatile memory arrays with the required precision and granularity for analog neuromorphic memory systems.

Method used

A programming system and method for vector matrix multiplication (VMM) arrays in artificial neural networks that utilize CMOS technology and non-volatile memory arrays, enabling precise programming of memory cells to store one of N different values, allowing for continuous and independent tuning of memory states with minimal disturbance to other cells.

Benefits of technology

The system achieves precise and efficient programming of memory cells, reducing energy consumption and eliminating the need for separate multiplication logic, thereby enhancing the computational efficiency and power efficiency of neural networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007730398000005
    Figure 0007730398000005
  • Figure 0007730398000006
    Figure 0007730398000006
  • Figure 0007730398000007
    Figure 0007730398000007
Patent Text Reader

Abstract

To provide programming systems and methods for use with a vector-by-matrix multiplication (VMM) array in an artificial neural network.SOLUTION: In a method for carrying out a given appropriate application by use of a non-volatile memory array-based neural network, a synapse CB1 which goes from an input S0 to a feature map C1 scans an input with a filter to shift the filter, multiplies them with the same weight to determine a second single output value by a related neuron, uses a set of different weights to repeatedly generate a different feature map of C1, applies an activation function P1, pools a value from a continuous but non overlapping 2×2 region in each feature map, applies an activation function P2 (pooling) before going from C2 to S2, pools the value from the continuous but non-overlapping 2×2 region in each feature map, and completely connects together S3 and C3 with a synapse CB4 which goes from C3 to output S3.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] (Priority Claim) This application claims priority to U.S. Provisional Patent Application No. 62 / 642,878, filed March 14, 2018, entitled "Method and Apparatus for Programming Analog Neuromorphic Memory in an Artificial Neural Network," and U.S. Patent Application No. 15 / 990,395, filed May 25, 2018, entitled "Method And Apparatus For Programming Analog Neural Memory In A Deep Learning Artificial Neural Network."

[0002] FIELD OF THE INVENTION Numerous embodiments of a programming apparatus and method for use with vector-by-matrix multiplication (VMM) arrays in artificial neural networks are disclosed. [Background technology]

[0003] Artificial neural networks mimic biological neural networks (the central nervous systems of animals, particularly the brain), which can rely on a large number of inputs and are used to estimate or approximate functions that are largely unknown. Artificial neural networks generally contain layers of interconnected "neurons" that exchange messages.

[0004] FIG. 1 illustrates an artificial neural network, where circles represent layers of inputs or neurons. Connections (called synapses) are represented by arrows and have numerical weights that can be adjusted based on experience. This allows the neural network to adapt to the inputs and learn. Typically, a neural network contains multiple layers of inputs. There are typically one or more hidden layers of neurons, and an output layer of neurons that provide the neural network's output. At each level, neurons make decisions, individually or collectively, based on the data received from the synapses.

[0005] One of the major challenges in developing artificial neural networks for high-performance information processing is the lack of suitable hardware technology. In practice, practical neural networks rely on a very large number of synapses, enabling high connectivity between neurons and therefore a very high degree of computational parallelism. In principle, such complexity could be achieved by digital supercomputers or specialized graphics processing unit clusters. However, in addition to high costs, these approaches also suffer from poor energy efficiency compared to biological networks, which primarily perform low-precision analog computations and therefore consume much less energy. CMOS analog circuits have been used in artificial neural networks, but most CMOS-implemented synapses are too bulky for the large number of neurons and synapses required.

[0006] The applicant previously disclosed an artificial (analog) neural network utilizing one or more non-volatile memory arrays as synapses in U.S. Patent Application No. 15 / 594,439, which is incorporated by reference. The non-volatile memory array operates as an analog neuromorphic memory. The neural network device includes a first plurality of synapses configured to receive a first plurality of inputs and generate a first plurality of outputs therefrom, and a first plurality of neurons configured to receive the first plurality of outputs. The first plurality of synapses includes a plurality of memory cells, each including spaced apart source and drain regions formed in a semiconductor substrate with a channel region extending therebetween, a floating gate disposed above and insulated from a first portion of the channel region, and a non-floating gate disposed above and insulated from a second portion of the channel region. Each of the plurality of memory cells is configured to store a weight value corresponding to a number of electrons in the floating gate. The plurality of memory cells are configured to multiply the first plurality of inputs by the stored weight values ​​to generate the first plurality of outputs.

[0007] Each non-volatile memory cell used in an analog neuromorphic memory system must be erased and programmed to hold a very specific and precise amount of charge in its floating gate. For example, each floating gate must hold one of N different values, where N is the number of different weights that can be exhibited by each cell. Examples of N include 16, 32, and 64.

[0008] One challenge in VMM systems is the ability to program selected cells with the precision and granularity required for different values ​​of N. For example, if a selected cell can contain one of 64 different values, then a very high degree of precision is required in the program operation.

[0009] What is needed is an improved programming system and method suitable for use with a VMM in an analog neuromorphic memory system. Summary of the Invention

[0010] Numerous embodiments of programming systems and methods are disclosed for use with vector matrix multiplication (VMM) arrays in artificial neural networks, whereby selected cells can be programmed with great precision to hold one of N different values.

[0011]

[0012]

[0013]

[0014]

[0015]

[0016]

[0017]

[0018]

[0019]

[0020]

[0021]

[0022]

[0023]

[0024]

[0025]

[0026]

[0027]

[0028]

[0029]

[0030]

[0031]

[0032]

[0033]

[0034]

[0035]

[0036]

[0037]

[0038]

[0039]

[0040]

[0041]

[0042]

[0043]

[0044]

[0045] [Brief explanation of the drawings]

[0046] [Figure 1] FIG. 1 illustrates an artificial neural network. [Figure 2] 1 is a cross-sectional view of a conventional two-gate nonvolatile memory cell. [Figure 3] 1 is a cross-sectional view of a conventional four-gate nonvolatile memory cell. [Figure 4] 1 is a cross-sectional view of a conventional three-gate nonvolatile memory cell. [Figure 5] FIG. 2 is a cross-sectional view of another conventional two-gate nonvolatile memory cell. [Figure 6]FIG. 1 illustrates different levels of an exemplary artificial neural network utilizing a non-volatile memory array. [Figure 7] FIG. 2 is a block diagram illustrating a vector multiplier matrix. [Figure 8] FIG. 2 is a block diagram illustrating various levels of a vector multiplier matrix. [Figure 9] 10 illustrates another embodiment of a vector multiplier matrix. [Figure 10] 10 illustrates another embodiment of a vector multiplier matrix. [Figure 11] 11 shows the operating voltages for performing the operations on the vector multiplier matrix of FIG. [Figure 12] 10 illustrates another embodiment of a vector multiplier matrix. [Figure 13] 13 shows the operating voltages for performing the operations on the vector multiplier matrix of FIG. 12. [Figure 14] 10 illustrates another embodiment of a vector multiplier matrix. [Figure 15] 15 shows the operating voltages for performing the operations on the vector multiplier matrix of FIG. 14. [Figure 16] 10 illustrates another embodiment of a vector multiplier matrix. [Figure 17] The operating voltages for performing the operations in the vector multiplier matrix of FIG. 216 are shown. [Figure 18A] 1 shows how to program the vector multiplier matrix. [Figure 18B] 1 shows how to program the vector multiplier matrix. [Figure 19] 18A and 18B show waveforms for programming. [Figure 20] 18A and 18B show waveforms for programming. [Figure 21] 18A and 18B show waveforms for programming. [Figure 22] 1 shows a vector multiplier matrix system. [Figure 23] Column drivers are shown. [Figure 24] 1 shows multiple reference matrices. [Figure 25] A single reference matrix is ​​shown. [Figure 26] The reference matrix is ​​shown. [Figure 27] 1 shows another reference matrix. [Figure 28] A comparison circuit is shown. [Figure 29] 2 shows another comparison circuit. [Figure 30] 2 shows another comparison circuit. [Figure 31] 1 shows a current-to-digital bit circuit. [Figure 32] The waveforms for the circuit of FIG. 31 are shown. [Figure 33] A current-slope circuit is shown. [Figure 34] The waveforms for the circuit of FIG. 33 are shown. [Figure 35] A current-slope circuit is shown. [Figure 36] The waveforms for the circuit of FIG. 35 are shown. DETAILED DESCRIPTION OF THE INVENTION

[0047] The artificial neural network of the present invention utilizes a combination of CMOS technology and non-volatile memory arrays. Nonvolatile Memory Cell

[0048] Digital nonvolatile memories are well known. For example, U.S. Pat. No. 5,029,130 ​​(the "'130 patent") discloses an array of split-gate nonvolatile memory cells and is incorporated herein by reference for all purposes. Such a memory cell is shown in FIG. 2. Each memory cell 210 is formed in a semiconductor substrate 12 and includes a source region 14 and a drain region 16 having a channel region 18 therebetween. A floating gate 20 is formed above and insulated from (and controls the conductivity of) a first portion of the channel region 18, and is also formed above a portion of the source region 16. A word line terminal 22 (typically coupled to a word line) has a first portion disposed above and insulated from (and controls the conductivity of) a second portion of the channel region 18, and a second portion extending above the floating gate 20. A floating gate 20 and a wordline terminal 22 are insulated from the substrate 12 by a gate oxide. A bitline 24 is coupled to the drain region 16.

[0049] The memory cell 210 is erased (where electrons are removed from the floating gate) by applying a high positive voltage to the word line terminal 22, thereby causing electrons in the floating gate 20 to tunnel through the intermediate insulator from the floating gate 20 to the word line terminal 22 by Fowler-Nordheim tunneling.

[0050] The memory cell 210 is programmed by applying a positive voltage to the word line terminal 22 and a positive voltage to the source 16 (where electrons are applied to the floating gate). An electron current will flow from the source 16 toward the drain 14. When the electrons reach the gap between the word line terminal 22 and the floating gate 20, they accelerate and heat up. Some of the heated electrons are injected into the floating gate 20 through the gate oxide 26 due to electrostatic attraction from the floating gate 20.

[0051] The memory cell 210 is read by applying a positive read voltage to the drain 14 and word line terminal 22 (turning on the channel region under the word line terminal). If the floating gate 20 is positively charged (i.e., erased of electrons and positively coupled to the drain 16), then the portion of the channel region under the floating gate 20 will also be turned on and current will flow through the channel region 18, which is sensed as an erased or "1" state. If the floating gate 20 is negatively charged (i.e., programmed with electrons), then the portion of the channel region under the floating gate 20 will be mostly or completely off and no (or only a small) current will flow through the channel region 18, which is sensed as a programmed or "0" state.

[0052] Table 1 shows typical voltage ranges that may be applied to the terminals of memory cell 210 to perform read, erase, and program operations. Table 1: Operation of flash memory cell 210 of FIG. 2 [Table 1]

[0053] Other split-gate memory cell configurations are known. For example, FIG. 3 shows a four-gate memory cell 310 including a source region 14, a drain region 16, a floating gate 20 above a first portion of the channel region 18, a select gate 28 (typically coupled to a word line) above a second portion of the channel region 18, a control gate 22 above the floating gate 20, and an erase gate 30 above the source region 14. This configuration is described in U.S. Pat. No. 6,747,310, which is incorporated herein by reference for all purposes. Here, all gates, except for the floating gate 20, are non-floating gates, meaning they are electrically connected or connectable to a voltage source. Programming is indicated by heated electrons from the channel region 18 injecting themselves into the floating gate 20. Erasing is indicated by electrons tunneling from the floating gate 20 to the erase gate 30.

[0054] Table 2 shows typical voltage ranges that may be applied to the terminals of memory cell 310 to perform read, erase, and program operations. Table 2: Operation of the flash memory cell 310 of FIG. 3 [Table 2]

[0055] Figure 4 shows a split-gate, triple-gate memory cell 410. Memory cell 410 is identical to memory cell 310 of Figure 3, except that memory cell 410 does not have a separate control gate. Erase operations (erase through the erase gate) and read operations are similar to those of Figure 3, except that there is no control gate bias. Programming operations are also performed without a control gate bias, so the program voltage on the source line is higher to compensate for the lack of control gate bias.

[0056] Table 3 shows typical voltage ranges that may be applied to the terminals of memory cell 410 to perform read, erase, and program operations. Table 3: Operation of flash memory cell 410 of FIG. 4 [Table 3]

[0057] Figure 5 shows a stacked gate memory cell 510. Memory cell 510 is similar to memory cell 210 of Figure 2, except that the floating gate 20 extends over the entire channel region 18, and the control gate 22 extends over the floating gate 20 separated by an insulating layer. Erase, programming, and read operations operate in a similar manner as described above for memory cell 210.

[0058] Table 4 shows typical voltage ranges that may be applied to the terminals of memory cell 510 to perform read, erase, and program operations. Table 4: Operation of flash memory cell 510 of FIG. 5 [Table 4]

[0059] To utilize a memory array containing one of the types of non-volatile memory cells in an artificial neural network, two modifications are made: First, as explained further below, the lines are configured so that each memory cell can be individually programmed, erased, and read without adversely affecting the memory state of other memory cells in the array; Second, continuous (analog) programming of the memory cells is provided.

[0060] Specifically, the memory state (i.e., the charge on the floating gate) of each memory cell in the array can be changed from a fully erased state to a fully programmed state independently and continuously with minimal disturbance to other memory cells. In another embodiment, the memory state (i.e., the charge on the floating gate) of each memory cell in the array can be changed from a fully programmed state to a fully erased state, and vice versa, independently and continuously with minimal disturbance to other memory cells. This means that the cell storage is analog, or at a minimum, can store one of many distinct values ​​(such as 16 or 64 different values), which allows for very precise and individual tuning of every cell in the memory array and makes the memory array ideal for storing and fine-tuning the synaptic weights of a neural network. Neural networks using nonvolatile memory cell arrays

[0061] 6 conceptually illustrates a non-limiting example of a neural network utilizing a non-volatile memory array. This example uses a non-volatile memory array neural net for a face recognition application, although any other suitable application can be implemented using a non-volatile memory array-based neural network.

[0062] S0 is the input, which in this example is a 32x32 pixel RGB image with 5-bit precision (i.e., three 32x32 pixel arrays, one for each color R, G, and B, with each pixel having 5-bit precision). Synapse CB1 going from S0 to C1 has both a set of distinct and shared weights, scans the input image with overlapping 3x3 pixel filters (kernels), and shifts the filters by one pixel (or two or more pixels, as determined by the model). Specifically, the values ​​of nine pixels in the 3x3 portion of the image (i.e., called the filter or kernel) are provided to synapse CB1, which multiplies these nine input values ​​by the appropriate weights, and after summing the outputs of that multiplication, a single output value is determined and provided by the first neuron of CB1 to generate one pixel of the layer of feature map C1. The 3x3 filter is then shifted one pixel to the right (i.e., adding a column of 3 pixels to the right and dropping a column of 3 pixels on the left), so that the 9 pixel values ​​of this newly positioned filter are provided to synapse CB1, which multiplies them by the same weight to determine a second single output value by the associated neuron. This process continues until the 3x3 filter has scanned the entire 32x32 pixel image for all three colors and all bits (the precision value). The process is then repeated using different sets of weights to generate different feature maps for C1 until all of layer C1's feature maps have been calculated.

[0063] In C1, in this example, there are 16 feature maps, each with 30x30 pixels. Each pixel is a new feature pixel extracted from the multiplication of the input and the kernel, and therefore each feature map is a two-dimensional array. Thus, in this example, synapses CB1 constitute 16 layers of two-dimensional arrays. (Note that the neuron layers and arrays referred to herein are logical, not necessarily physical, relationships; i.e., the arrays are not necessarily oriented in a physical two-dimensional array.) Each of the 16 feature maps is generated by one of 16 different sets of synaptic weights applied to the filter scans. The C1 feature maps can all target different aspects of the same image feature, such as boundary identification. For example, a first map (generated using a first set of weights shared by all scans used to generate this first map) can identify circular edges, while a second map (generated using a second set of weights different from the first set of weights) can identify rectangular edges or the aspect ratio of a particular feature, etc.

[0064] Activation function P1 (pooling) is applied before going from C1 to S1 to pool values ​​from consecutive, non-overlapping 2x2 regions in each feature map. The purpose of the pooling step is to average to nearby locations (or a max function can also be used), e.g., to reduce dependency on edge locations, and to reduce data size before going to the next stage. In S1, there are 16 15x15 feature maps (i.e., 16 different arrays of 15x15 pixels each). The synapses and associated neurons in CB2 going from S1 to C2 scan the maps in S1 with a 4x4 filter using a 1-pixel filter shift. In C2, there are 22 12x12 feature maps. Activation function P2 (pooling) is applied before going from C2 to S2 to pool values ​​from consecutive, non-overlapping 2x2 regions in each feature map. In S2, there are 22 6x6 feature maps. The activation function is applied at synapse CB3 going from S2 to C3, where all neurons in C3 connect to all maps in S2. There are 64 neurons in C3. Synapse CB4 going from C3 to output S3 fully connects S3 to C3. The output at S3 contains 10 neurons, where the highest output neuron determines the class. This output can indicate, for example, the identity or classification of the content of the original image.

[0065] Each level of synapses is implemented using an array or portion of an array of non-volatile memory cells. FIG. 7 is a block diagram of a vector matrix multiplication (VMM) array containing non-volatile memory cells and utilized as a synapse between an input layer and the next layer. Specifically, the VMM 32 includes an array 33 of non-volatile memory cells, an erase gate and word line gate decoder 34, a control gate decoder 35, a bit line decoder 36, and a source line decoder 37, which decode the inputs to the memory array 33. The source line decoder 37 in this example also decodes the output of the memory cell array. Alternatively, the bit line decoder 36 can decode the output of the memory array. The memory array serves two purposes. First, it stores the weights used by the VMM. Second, the memory array effectively multiplies the inputs by the weights stored in the memory array and sums them for each output line (source line or bit line) to generate an output, which becomes the input to the next layer or the input to the last layer. By performing multiplication and addition functions, the memory array eliminates the need for separate multiplication and addition logic and is also more power efficient due to in-place memory computation.

[0066] The output of the memory array is fed to a differential adder (e.g., summing op-amp) 38, which sums the outputs of the memory cell array to generate a single value for the convolution. The differential adder is such that a positive input implements the sum of positive and negative weights. The summed output value is then fed to an activation function circuit 39, which rectifies the output. Activation functions may include sigmoid, tanh, or ReLU functions. The rectified output value becomes an element of a feature map for the next layer (e.g., C1 in the above description) and is then applied to the next synapse to generate the next feature map layer or the final layer. Thus, in this example, the memory array constitutes multiple synapses (receiving input from a previous layer of neurons or from an input layer such as an image database), and the summing op-amp 38 and activation function circuit 39 constitute multiple neurons.

[0067] FIG. 8 is a block diagram of the various levels of VMMs. As shown in FIG. 8, input is converted from digital to analog by a digital-to-analog converter 31 and provided to an input VMM 32a. The output generated by the input VMM 32a is provided as input to the next VMM (hidden level 1) 32b, which in turn generates an output that is provided as input to the next VMM (hidden level 2) 32b, and so on. The 32 various layers of VMMs function as different layers of synapses and neurons of a convolutional neural network (CNN). Each VMM can be a standalone non-volatile memory array, or multiple VMMs can utilize different portions of the same non-volatile memory array, or multiple VMMs can utilize overlapping portions of the same non-volatile memory array. 8 includes five layers (32a, 32b, 32c, 32d, 32e): one input layer (32a), two hidden layers (32b, 32c), and two fully connected layers (32d, 32e). Those skilled in the art will appreciate that this is merely an example and that the system may alternatively include more than two hidden layers and more than two fully connected layers. Vector Matrix Multiplication (VMM) Array

[0068] FIG. 9 shows a neuron VMM 900 that is particularly suited for memory cells of the type shown in FIG. 3 , utilized as part of synapses and neurons between an input layer and the next layer. VMM 900 comprises a memory array 901 of non-volatile memory cells and a reference array 902 (at the top of the array). Alternatively, a separate reference array can be located at the bottom. In VMM 900, control gate lines, such as control gate line 903, run vertically (thus the row-oriented reference array 902 is orthogonal to the input control gate lines), and erase gate lines, such as erase gate line 904, run horizontally. Here, inputs are provided to the control gate lines and outputs appear on source lines. In one embodiment, only even rows are used, and in another embodiment, only odd rows are used. The current in the source line is a function of the sum of all currents from the memory cells connected to the source line.

[0069] As described herein for neural networks, the flash cells are preferably configured to operate in the sub-threshold region.

[0070] The memory cells described herein are biased in weak inversion. Ids=Io * e (Vg-Vth) / kVt =w * Io * e (Vg) / kVt w=e (-Vth) / kVt

[0071] Regarding the IV-log converter, which uses a memory cell to convert an input current to an input voltage: Vg=k * Vt * log[Ids / wp * Io]

[0072] For a memory array used as a vector matrix multiplier VMM, the output current is: Iout=wa * Io * e (Vg) / kVt , i.e. Iout=(wa / wp) *Iin=W * Iin W=e (Vthp-Vtha) / kVt

[0073] The word line or control gate can be used as the input of the memory cell for the input voltage.

[0074] Alternatively, the flash memory cells can be configured to operate in the linear region. Ids=β * (Vgs-Vth) * Vds;β=u * Cox * W / L W α(Vgs-Vth)

[0075] In an IV linear converter, memory cells operating in the linear region can be used to linearly convert input / output currents to input / output voltages.

[0076] Other embodiments of the ESF vector matrix multiplier are described in U.S. Patent Application No. 15 / 826,345, which is incorporated herein by reference. The source lines or bit lines can be used as neuron outputs (current sum outputs).

[0077] FIG. 10 shows a neuron VMM1000 that is particularly suited to memory cells of the type shown in FIG. 2 and used as synapses between the input layer and the next layer. VMM1000 comprises a memory array 1003 of non-volatile memory cells, a reference array 1001, and a reference array 1002. Reference arrays 1001 and 1002 serve to convert current inputs flowing into terminals BLR0-3 into voltage inputs WL0-3 in the column direction of the array. In practice, the reference memory cells are diode-connected to the current inputs flowing into them through multiplexers. The reference cells are adjusted (e.g., programmed) to a target reference level, which is provided by a reference mini-array matrix. Memory array 1003 serves two purposes. First, it stores the weights used by VMM1000. Second, memory array 1003 effectively multiplies its inputs (current inputs provided to terminals BLR0-3, which reference arrays 1001 and 1002 convert to input voltages and provide to word lines WL0-3) with weights stored in the memory array to generate outputs that are inputs to the next layer or to the final layer. By performing a multiplication function, the memory array eliminates the need for separate multiplication logic and is also power efficient. Here, voltage inputs are provided to word lines and outputs appear on bit lines during read (inference) operations. The current on the bit lines performs a function of the sum of all the currents from the memory cells connected to the bit lines.

[0078] 11 shows the operating voltages of the VMM 1000. The columns in the table show the voltages applied to the word line of the selected cell, the word lines of the unselected cells, the bit lines of the selected cell, the bit lines of the unselected cells, the source lines of the selected cell, and the source lines of the unselected cells. The rows show the read, erase, and program operations.

[0079] FIG. 12 shows a neuron VMM 1200 that is particularly suited for memory cells of the type shown in FIG. 2, used as part of synapses and neurons between an input layer and the next layer. VMM 1200 comprises a memory array 1203 of non-volatile memory cells, a reference array 1201, and a reference array 1202. Reference arrays 1201 and 1202, which run along the rows of array VMM 1200, are similar to VMM 1000, except that word lines run vertically in VMM 1200. Here, inputs are provided on the word lines, and outputs appear on source lines during read operations. The current in the source line performs a function of the sum of all currents from memory cells connected to the source line.

[0080] Figure 13 shows the operating voltages of VMM1200. The columns in the table show the voltages applied to the word line of the selected cell, the word lines of the unselected cells, the bit lines of the selected cell, the bit lines of the unselected cells, the source line of the selected cell, and the source lines of the unselected cells. The rows show the read, erase, and program operations.

[0081] FIG. 14 shows a neuron VMM 1400 that is particularly suited for memory cells of the type shown in FIG. 3, utilized as part of synapses and neurons between the input layer and the next layer. VMM 1400 comprises a memory array 1403 of non-volatile memory cells, a reference array 1401, and a reference array 1402. Reference arrays 1401 and 1402 serve to convert current inputs flowing into terminals BLR0-3 into voltage inputs CG0-3. In practice, the reference memory cells are diode-connected with their current inputs flowing into them via cascoding mux 1414. Mux 1414 includes mux 1405 and cascoding transistor 1404 to ensure a constant voltage on the bit line of the reference cell during read. The reference cell is adjusted to a target reference level. Memory array 1403 serves two purposes. First, it stores the weights used by VMM 1400. Second, memory array 1403 effectively multiplies its inputs (current inputs provided to terminals BLR0-3, which reference arrays 1401 and 1402 convert to input voltages provided to control gates CG0-3) by weights stored in the memory array to produce an output that serves as the input to the next layer or to the final layer. By performing a multiplication function, the memory array eliminates the need for separate multiplication logic and is also power efficient. Here, the inputs are provided to word lines and the outputs appear on bit lines during read operations. The currents on the bit lines perform a function of the sum of all the currents from the memory cells connected to the bit lines.

[0082] The VMM 1400 performs one-way conditioning of the memory cells in the memory array 1403. That is, each cell is erased and then partially programmed until the desired charge on the floating gate is reached. If too much charge is present on the floating gate (such as an incorrect value being stored in the cell), the cell must be erased and a series of partial programming operations must be redone. As shown, two rows that share the same erase gate must be erased together (known as a page erase), and then each cell is partially programmed until the desired charge on the floating gate is reached.

[0083] 15 shows the operating voltages of the VMM 1400. The columns in the table indicate the voltages applied to the word line of the selected cell, the word lines of the unselected cells, the bit line of the selected cell, the bit lines of the unselected cells, the control gate of the selected cell, the control gates of the unselected cells in the same sector as the selected cell, the control gates of the unselected cells in a different sector from the selected cell, the erase gate of the selected cell, the erase gates of the unselected cells, the source line of the selected cell, and the source lines of the unselected cells. The rows indicate the read, erase, and program operations.

[0084] FIG. 16 shows a neuron VMM 1600 that is particularly suited for memory cells of the type shown in FIG. 3, utilized as part of synapses and neurons between the input layer and the next layer. VMM 1600 comprises a memory array 1603 of nonvolatile memory cells, a reference array 1601, and a reference array 1602. The EG lines run vertically, while the CG and SL lines run horizontally. VMM 1600 is similar to VMM 1400, except that VMM 1600 implements bidirectional tuning: each individual cell can be fully erased, partially programmed, and optionally partially erased to achieve a desired amount of charge on its floating gate. As shown, reference arrays 1601 and 1602 convert input currents in terminals BLR 0-3 into control gate voltages CG 0-3 (through the action of diode-connected reference cells via multiplexers) that are applied to the memory cells in the row direction. The current output (neuron) is in a bit line that sums all the currents from the memory cells connected to the bit line.

[0085] 17 shows the operating voltages of the VMM 1600. The columns in the table indicate the voltages applied to the word line of the selected cell, the word lines of the unselected cells, the bit line of the selected cell, the bit lines of the unselected cells, the control gate of the selected cell, the control gates of the unselected cells in the same sector as the selected cell, the control gates of the unselected cells in a different sector from the selected cell, the erase gate of the selected cell, the erase gates of the unselected cells, the source line of the selected cell, and the source lines of the unselected cells. The rows indicate the read, erase, and program operations.

[0086] 18A and 18B illustrate a programming method 1800. Initially, the method typically begins in response to a received program command (step 1801). Next, a mass program operation programs all cells to a "0" state (step 1802). A soft erase then erases all cells to an intermediate, weakly erased level of approximately 3-5 μA (step 1803). This contrasts with a deep erase, where all cells are fully erased in digital applications, e.g., a cell current of approximately 20-30 μA. Next, a hard program is performed on all unselected cells to remove charge from the cells to a very deep program state for unused cells to ensure that the cells are truly off, i.e., these memory cells provide no significant current. Next, a soft program is performed on selected cells to remove some charge from the cells to an intermediate, weakly programmed level of approximately 0.7-1.5 μA using a coarse-grained algorithm (steps 1805, 1806, 1807). A coarse step program cycle followed by a verify operation is performed, in which the charge of the selected cell is compared to various thresholds in a coarse, iterative manner (steps 1806 and 1807). The coarse step program cycle includes coarse voltage increments, and / or coarse program times, and / or coarse program currents, resulting in a coarse cell current change from one program step to the next. Next, fine programming is performed (steps 1808-1813), in which all cells are programmed to a target level within a range of 100 pA to 20 nA, depending on the desired level, using a fine step program algorithm. A fine step program cycle followed by a verify operation is performed (steps 1809 and 1810). The fine step program cycle may include a combination of coarse and fine resolution voltage increments, and / or program times, and / or program currents, resulting in a fine cell current change from one program step to the next. Once the selected cell reaches the desired target, the programming operation is complete (step 1811). If not, the fine programming operation is repeated until the desired level is reached.However, if the number of attempts exceeds a threshold number (step 1812), the programming operation stops and the selected cell is deemed a bad cell (step 1813).

[0087] FIG. 19 shows an example waveform 1900 for performing a programming operation using pulse modulation. Signal 1901 is a program cycle enable signal consisting of multiple program and verify cycles. Signal 1902 is an individual pulse program cycle enable signal (signal 1902 = logic "1" enables programming). A verify cycle follows each pulse program cycle (signal 1902 = logic "0" enables verification). Signal 1903 is an individual pulse program cycle enable signal for a specific bit line. Signal 1904 is an individual pulse program cycle enable signal for another specific bit line. As shown, the width of signal 1903 is narrower than the width of signal 1904. Thus, signal 1903 allows a smaller charge to be changed on the floating gate of a memory cell, resulting in smaller current accuracy in programming. Within program cycle 1901, different program pulse widths 1903 can be used to achieve the desired programming accuracy for a particular memory cell.

[0088] 20 shows an example waveform 2000 for performing a programming operation using high voltage level modulation. Signal 2003 is an individual pulse program cycle enable signal for a specific bit line. Signal 2004 is an individual pulse program cycle enable signal for another specific bit line. The program pulse widths of signals 2003 and 2004 are the same in this waveform. Signal 2005 is a high voltage increment for a source line or control gate, etc. for programming, incrementing from one program pulse to the next.

[0089] FIG. 21 shows an example waveform 2100 for performing a programming operation using high voltage level modulation. Signal 2103 is an individual pulse program cycle enable signal for a specific bit line. Signal 2104 is an individual pulse program cycle enable signal for another specific bit line. Signal 2105 is a program high voltage increment for a source line or control gate, etc., for programming. It can be the same from one program pulse to the next, or it increments. The program pulse width of signal 2103 is the same across multiple pulses with incremental high voltages. The program pulse width of signal 2104 varies across multiple pulses of the same high voltage increment, e.g., narrower for the first pulse.

[0090] FIG. 22 shows a VMM system 2200 including a VMM matrix 2201 , a column decoder 2202 , and a column driver 2203 .

[0091] FIG. 23 shows an exemplary column driver 2300 that can be used as column driver 2203 of FIG. 22. Column driver 2300 includes latch 2301, inverter 2302, NOR gate 2303, PMOS transistor 2304, and NMOS transistors 2305 and 2306, configured as shown. Latch 2301 receives a data signal DIN and an enable signal EN. A node BLIO between PMOS transistor 2304 and NMOS transistor 2305 includes a bit line input or output signal, which is selectively connected to a bit line via a column decoder, such as column decoder 2202 of FIG. 22. A sense amplifier SA2310 is coupled to the bit line via node BLIO to read the cell current of a selected memory cell. Sense amplifier SA2310 is used to verify a desired current level in a selected memory cell, such as after an erase or program operation. PMOS transistor 2304 serves to inhibit the bit line from programming in response to inhibit control circuits 2301 / 2302 / 2303. NMOS transistor 2306 provides a bias program current to the bit line. NMOS transistor 2305 enables the bias program current to the bit line, thus enabling programming of the selected memory cell. Thus, a program control signal (such as signals 1902 / 1903 in FIG. 19, signals 2003 / 2004 in FIG. 20, or signals 2103 / 2104 in FIG. 21) to the gate of NMOS transistor 2305 enables programming of the selected cell.

[0092] 24 shows an exemplary VMM system 2400, including reference array matrices 2401a, 2401b, 2401c, and 2401d and VMM matrices 2402a, 2402b, 2402c, and 2402d. Each VMM matrix has its own reference array matrix.

[0093] 25 shows an exemplary VMM system 2500 including a single reference array matrix 2501 and VMM matrices 2502a, 2502b, 2502c, and 2502d. The single reference array matrix 2501 is shared across multiple VMM matrices.

[0094] FIG. 26 illustrates an exemplary reference matrix 2600 that can be used as reference matrices 2401a-d of FIG. 24, matrix 2501 of FIG. 25, or reference matrices 2801 or 2901 of FIGS. 28 and 29. Reference matrix 2600 includes reference memory cells 2602a, 2602b, 2602c, and 2602x coupled to a common control gate signal 2603 and a source line reference signal 2604, and a bit line reference decoder 2601 that provides multiple bit line reference signals for use in read or verify operations. For example, reference memory cells 2602a-x can provide incremental current levels of 100 pA / 200 pA / 300 pA / ... / 30 nA. Alternatively, reference memory cells 2602a-x can provide a constant current level of 100 pA per reference memory cell. In this case, combinations of reference memory cells 2602a-x are used to generate different reference current levels, such as generating 100 pA / 200 pA / 300 pA / etc. in a thermometer coded manner. Other combinations of constant and / or incremental reference current levels can be used to generate desired reference current levels. Furthermore, the difference current between two reference cells can be used to generate reference memory cell currents such as 100 pA reference current = 250 pA reference current - 150 pA reference current. This can be used, for example, to generate a reference current that is compensated for temperature.

[0095] Figure 27 shows an exemplary reference matrix 2700 that can be used as reference matrices 2401a-d of Figure 24, matrix 2501 of Figure 25, or reference matrices 2801 or 2901 of Figures 28 and 29. Reference matrix 2700 includes reference memory cells 2702a, 2702b, 2702c, and 2702x coupled to a common control gate reference signal 2703, erase gate reference signal 2704, and source line reference signal 2705, and a bit line reference decoder 2701 that provides multiple bit line reference signals for use in read or verify operations. A combination of constant and / or incremental reference current levels can be used to generate the desired reference current levels, as in Figure 26.

[0096] FIG. 28 shows an Icell PMOS compare circuit 2800, which includes PMOS transistors 2801 and 2804, NMOS cascoding transistors 2802 and 2805, a selected memory cell 2803 from a VMM memory array 2820, and a reference matrix 2806 (such as reference matrix 2500 or 2600), arranged as shown. NMOS transistors 2802 and 2805 are used to bias the reference bit lines to the desired voltage levels. An output current Iout is a current value indicative of the value stored in the selected memory cell 2803. The voltage level at output node 2810 indicates the result of a comparison of the current in the selected memory cell 2803 with the reference current from reference matrix 2806. The voltage at node 2801 rises to Vdd (or falls to ground) if the current in the selected memory cell 2803 is greater (or less) than the reference current from reference matrix 2806.

[0097] FIG. 29 shows an Icell PMOS compare circuit 2900, which includes, arranged as shown, PMOS transistor 2901, switches 2902, 2903, and 2904, NMOS cascoding transistors 2905 and 2907, a selected memory cell 2908 from a VMM memory array 2920, a reference matrix 2906 (such as reference matrix 2600 or 2700), and a comparator 2909. The output COMP_OUT is a current value indicative of the value stored in the selected memory cell 2908 compared to a reference current. The Icell compare circuit functions by using it as a time-multiplexed single PMOS current mirror to eliminate mismatch between the two mirror PMOS transistors. During a first period, S0 and S1 are closed and S2 is open. The current from the reference memory matrix 2906 is stored (held) in PMOS transistor 2901. During a second period, S0 and S1 are open and S2 is closed. The stored reference current is compared with the current from memory cell 2908 and the result of the comparison is indicated at output node 2910. Optionally, the comparator may compare the voltage at node 2901 with a reference voltage VREF to indicate the result of the comparison, where the reference current is sampled and held, and the current from the selected memory cell 2902 is sampled and held in PMOS transistor 2901 and compared to the reference current.

[0098] In another embodiment, an array leakage compensation circuit 3051 as shown in FIG. 30 can be used with a single PMOS mirror circuit to sample the leakage of the VMM array 3020 (S3 closed) (all word lines are off and the bit line leakage current is sampled into the retention PMOS) and hold the leakage in the retention transistor (S3 open). This leakage is then subtracted from the current in the selected memory cell during the comparison period to obtain the actual memory cell current from the VMM array 3020 for comparison. This can be used for all the comparison circuits described herein. This can be used for reference array leakage compensation.

[0099] 31 shows Icell-to-digital data circuit 3100, which includes current source 3101, switch 3102, capacitor 3103, comparator 3104, and counter 3105. At the start of a comparison period, signal 3110 is pulled down to ground. Signal 3110 then begins to ramp up according to cell current 3101 (derived from a VMM memory array with array leakage compensation as described above). The ramp rate is proportional to cell current 3101 and capacitor 3103. The output 3112 of comparator 3104 then enables counter 3105 to begin counting digitally. When the voltage at node 3110 reaches voltage level VREF3111, comparator 3104 switches polarity and stops counter 3105. Digital output Q <n:0>The value of 3113 indicates the value of the cell current 3101 .

[0100] Figure 32 shows waveforms 3200 for the operation of Icell to digital data circuit 3100. Signal 3201 is the ramp voltage (corresponding to signal 3110 in Figure 31). The ramp rate of signal 3201 is shown for different cell current levels. Signals 3205 and 3207 are the comparator outputs for two different cell currents (corresponding to signal 3121 in Figure 31). Signals 3206 and 3208 are the digital outputs Q for the two different cell currents. <n:0>is.

[0101] FIG. 33 shows an Icell-slope circuit 3300, including a memory cell current source 3301, a switch 3302, a capacitor 3303, and a comparator 3304. The memory cell current is extracted from the VMM memory array with array leakage compensation as described above. At the start of a comparison period, signal 3310 is pulled down to ground. Signal 3310 then begins to ramp up in response to cell current 3301 (extracted from the VMM memory array). The ramp rate is proportional to cell current 3301 and capacitor 3303. After a fixed comparison period, the voltage at node 3310 is compared by comparator 3304 with a reference voltage VREFx 3311. VREFx 3311 is, for example, 0.1V, 0.2V, 0.3V, . . . , 1.5V, 1.6V for 16 reference levels. Thus, each level corresponds to a current level for 16 different current levels. The output of comparator 3304 indicates the value of cell current 3301. To compare the voltage at node 3310 (which may be held on capacitor 3303 by shutting off S1 after a fixed comparison period) with 16 reference levels, either 16 comparators with 16 reference levels are used, or one comparator with reference levels multiplexed by 16 for the 16 reference levels.

[0102] Figure 34 shows waveforms 3400 for the operation of the Icell-slope circuit 3300. Signal 3401 shows different ramp rates with different voltage levels (Vcellx) at the rising edge of enable signal 3402. Voltage Vcellx is compared to a reference voltage to indicate the value of the cell current (reference voltage VREFx 3311 in Figure 33).

[0103] Figure 35 shows an Icell-slope conversion circuit 3500, including a memory cell current source 3504, a switch 3502, a capacitor 3501, an NMOS cascoding transistor 3503, and a comparator 3505. The memory cell current 3501 is extracted from a VMM memory array with array leakage compensation as described above. NMOS 3503 is used to bias the voltage on the bit line of a selected memory cell (shown as Icell 3504). Operation is similar to that of Figure 31, except that the ramp direction is ramp down instead of ramp up.

[0104] FIG. 36 shows waveforms 3600 for the operation of the Icell-slope circuit 3500.

[0105] It should be noted that, as used herein, both the terms "over" and "on" are inclusive of "directly" (without any intermediate material, element, or space disposed between them) and "indirectly on" (with an intermediate material, element, or space disposed between them). Similarly, the term "adjacent" includes "directly adjacent" (without any intermediate material, element, or space disposed between them) and "indirectly adjacent" (with an intermediate material, element, or space disposed between them); "attached" includes "directly attached" (without any intermediate material, element, or space disposed between them) and "indirectly attached to" (with an intermediate material, element, or space disposed between them); and "electrically coupled" includes "directly electrically coupled" (without any intermediate material or element electrically coupling the elements together) and "indirectly electrically coupled to" (with an intermediate material or element electrically coupling the elements together). For example, forming an element "over a substrate" can include forming the element directly on the substrate without any intermediate materials / elements therebetween, as well as forming the element indirectly on the substrate with one or more intermediate materials / elements therebetween.

Claims

1. 1. A circuit for converting memory cell current signals for a vector matrix multiplier into a set of digital bits, the circuit comprising: a memory cell that produces a read current during a read operation, the read current being based on one of N possible values ​​stored in the memory cell, where N is greater than two; a capacitor having a first terminal coupled to the memory cell and a second terminal coupled to ground; a switch coupled in parallel with the capacitor between the memory cell and ground, the switch being closed before the read operation to pull down the first terminal of the capacitor to ground and being open during the read operation; a comparator including a first input coupled to the memory cell, the switch, and the first terminal of the capacitor, a second input coupled to a voltage reference, and an output; a counter coupled to receive the output from the comparator as an enable signal, to digitally count when the enable signal is asserted, to stop counting when the enable signal is deasserted, and to output a count value indicative of a value of the read current.

2. The circuit of claim 1 , wherein the read current is maintained in a transistor.

3. The circuit of claim 1 further comprising an array leakage compensation circuit.

4. 2. The circuit of claim 1, wherein the memory cells are split-gate memory cells.

5. The circuit of claim 1 , wherein the memory cells are stacked gate memory cells.

6. 1. A circuit for converting memory cell current signals to slopes for a vector matrix multiplier, the circuit comprising: a memory cell from which a read current is drawn; a transistor including a first terminal and a second terminal coupled to the memory cell; a capacitor having a first terminal coupled to a power supply and a second terminal coupled to the first terminal of the transistor; a switch coupled in parallel with the capacitor between the power supply and the first terminal of the transistor, the switch being closed before a read operation to charge the second terminal of the capacitor to the voltage of the power supply, and being opened during a read operation; a comparator including a first input coupled to the first terminal of the transistor, the switch, and the second terminal of the capacitor, a second input coupled to a voltage reference, and an output indicative of a value stored in the memory cell.

7. The circuit of claim 6 further comprising an array leakage compensation circuit.

8. 7. The circuit of claim 6, wherein the memory cells are split-gate memory cells.

9. The circuit of claim 6 wherein the memory cells are stacked gate memory cells.

Citation Information

Patent Citations

  • Multi-valued sense amplifier circuit

    JP1997069293A

  • Current integrated sense amplifier for memory modules in rfid

    JP2006504211A

  • Product arithmetic unit including resistance change type variable resistor element, product sum arithmetic unit, neural network providing each neuron element with these arithmetic units, and product arithmetic method

    JP2009282782A

  • Non-volatile storage device

    WO2014119329A1

  • Low power operation for flash memory system

    WO2016195845A1