Neural Network Classifier Using Tri-Gate Non-Volatile Memory Cell Array

By combining CMOS technology with nonvolatile memory arrays, synaptic memory units that are independently programmed and read are realized, solving the problem of low energy efficiency of artificial neural network hardware and improving computing parallelism and energy efficiency.

CN113330461BActive Publication Date: 2025-08-08SILICON STORAGE TECHNOLOGY INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN201980089312.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-04-11
Filing Date
2019-08-29
Publication Date
2025-08-08
Estimated Expiration
2039-08-29

AI Technical Summary

Technical Problem

In the prior art, the hardware implementation of artificial neural networks has problems such as low energy efficiency and excessive synaptic implementation, making it difficult to achieve a balance between high computing parallelism and high-efficiency energy consumption.

Method used

Combined with nonvolatile memory arrays, the memory cells are configured to store weight values, and efficient independent programming and reading of synapses are realized, and the memory cells are connected in lines arranged in rows and columns, and inputs are received and outputs are generated.

Benefits of technology

It realizes efficient synaptic weight regulation and independent control, reduces energy consumption, improves computing parallelism, and is suitable for high-performance information processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113330461B_ABST
    Figure CN113330461B_ABST
Patent Text Reader

Abstract

The present invention relates to a neural network device having a synapse, the synapse having memory cells, each memory cell having a floating gate disposed above a first portion of a channel region between a source region and a drain region, a first gate disposed above a second portion of the channel region, and a second gate disposed above either the floating gate or the source region. A first line is electrically connected to the first gates in one of the rows of memory cells, a second line is electrically connected to the second gates in one of the rows of memory cells, a third line is electrically connected to the source regions in one of the columns of memory cells, and a fourth line is electrically connected to the drain regions in one of the columns of memory cells. The synapse receives a first plurality of inputs as voltages on the first line or the second line, and provides a first plurality of outputs as currents on the third line or the fourth line.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related patent applications

[0002] This application claims the benefit of U.S. Application No. 16 / 382,045, filed April 11, 2019, which claims priority to Provisional Application No. 62 / 794,492, filed January 18, 2019, and Provisional Application No. 62 / 798,417, filed January 29, 2019. Technical Field

[0003] The present invention relates to neural networks. Background Art

[0004] Artificial neural networks simulate biological neural networks (the central nervous system of animals, especially the brain) and are used to estimate or approximate functions that can depend on a large number of inputs and are generally known. Artificial neural networks typically consist of layers of interconnected "neurons" that exchange messages with each other. Figure 1 An artificial neural network is shown, where circles represent layers of inputs or neurons. Connections (called synapses) are represented by arrows and have numerical weights that can be adjusted based on experience. This allows the neural network to adapt to the input and learn. Typically, a neural network includes multiple layers of inputs. There are typically one or more intermediate layers of neurons, and an output layer of neurons that provide the output of the neural network. Neurons at each level make decisions based on the data received from the synapses, either individually or collectively.

[0005] One of the main challenges in developing artificial neural networks for high-performance information processing is the lack of adequate hardware technology. In fact, practical neural networks rely on a large number of synapses to achieve high connectivity between neurons, that is, very high computational parallelism. In principle, such complexity can be achieved using digital supercomputers or clusters of dedicated graphics processing units. However, in addition to being high-cost, these approaches are also mediocre in energy efficiency compared to biological networks, which consume less energy mainly due to the low-precision analog calculations they perform. CMOS analog circuits have been used in artificial neural networks, but given the large number of neurons and synapses, the synapses of most CMOS implementations are too large. Summary of the Invention

[0006] The aforementioned problems and needs are addressed by a neural network device comprising a first plurality of synapses configured to receive a first plurality of inputs and thereby generate a first plurality of outputs. The first plurality of synapses comprises a plurality of memory cells, wherein each of the memory cells comprises: a source region and a drain region formed in a semiconductor substrate, wherein a channel region extends between the source region and the drain region; a floating gate disposed over and insulated from a first portion of the channel region; a first gate disposed over and insulated from a second portion of the channel region; and a second gate disposed over and insulated from the floating gate or over and insulated from the source region. Each of the plurality of memory cells is configured to store a weight value corresponding to a plurality of electrons on the floating gate. The plurality of memory cells is configured to generate the first plurality of outputs based on the first plurality of inputs and the stored weight values. The memory cells of the first plurality of synapses are arranged in rows and columns. The first plurality of synapses includes: a plurality of first lines, each of which electrically connects first gates in one of the rows of memory cells; a plurality of second lines, each of which electrically connects second gates in one of the rows of memory cells; a plurality of third lines, each of which electrically connects source regions in one of the columns of memory cells; and a plurality of fourth lines, each of which electrically connects drain regions in one of the columns of memory cells. The first plurality of synapses is configured to receive the first plurality of inputs as voltages on the first plurality of lines or the second plurality of lines, and to provide the first plurality of outputs as currents on the third plurality of lines or the fourth plurality of lines.

[0007] A neural network device may include a first plurality of synapses configured to receive a first plurality of inputs and thereby generate a first plurality of outputs. The first plurality of synapses includes a plurality of memory cells, wherein each of the memory cells includes: a source region and a drain region formed in a semiconductor substrate, wherein a channel region extends between the source region and the drain region; a floating gate disposed over and insulated from a first portion of the channel region; a first gate disposed over and insulated from a second portion of the channel region; and a second gate disposed over and insulated from the floating gate or over and insulated from the source region. Each of the plurality of memory cells is configured to store a weight value corresponding to a plurality of electrons on the floating gate. The plurality of memory cells is configured to generate the first plurality of outputs based on the first plurality of inputs and the stored weight values. The memory cells of the first plurality of synapses are arranged in rows and columns. The first plurality of synapses includes: a plurality of first lines, each of which electrically connects the first gates of one of the rows of memory cells together; a plurality of second lines, each of which electrically connects the second gates of one of the rows of memory cells together; a plurality of third lines, each of which electrically connects the source regions of one of the rows of memory cells together; and a plurality of fourth lines, each of which electrically connects the drain regions of one of the columns of memory cells together. The first plurality of synapses is configured to receive the first plurality of inputs as voltages on the plurality of first lines, the plurality of second lines, or the plurality of third lines, and to provide the first plurality of outputs as currents on the plurality of fourth lines.

[0008] A neural network device may include a first plurality of synapses configured to receive a first plurality of inputs and thereby generate a first plurality of outputs. The first plurality of synapses includes a plurality of memory cells, wherein each of the memory cells includes: a source region and a drain region formed in a semiconductor substrate, wherein a channel region extends between the source region and the drain region; a floating gate disposed over and insulated from a first portion of the channel region; a first gate disposed over and insulated from a second portion of the channel region; and a second gate disposed over and insulated from the floating gate or over and insulated from the source region. Each of the plurality of memory cells is configured to store a weight value corresponding to a plurality of electrons on the floating gate. The plurality of memory cells is configured to generate the first plurality of outputs based on the first plurality of inputs and the stored weight values. The memory cells of the first plurality of synapses are arranged in rows and columns. The first plurality of synapses includes: a plurality of first lines, each of which electrically connects the first gates of one of the rows of memory cells together; a plurality of second lines, each of which electrically connects the second gates of one of the rows of memory cells together; a plurality of third lines, each of which electrically connects the source regions of one of the rows of memory cells together; and a plurality of fourth lines, each of which electrically connects the drain regions of one of the columns of memory cells together. The first plurality of synapses is configured to receive the first plurality of inputs as voltages on the fourth plurality of lines and to provide the first plurality of outputs as currents on the third plurality of lines.

[0009] A neural network device may include a first plurality of synapses configured to receive a first plurality of inputs and thereby generate a first plurality of outputs. The first plurality of synapses includes a plurality of memory cells, wherein each of the memory cells includes: a source region and a drain region spaced apart formed in a semiconductor substrate, wherein a channel region extends between the source region and the drain region; a floating gate disposed over and insulated from a first portion of the channel region; a first gate disposed over and insulated from a second portion of the channel region; and a second gate disposed over and insulated from the floating gate or over and insulated from the source region. Each of the plurality of memory cells is configured to store a weight value corresponding to a plurality of electrons on the floating gate. The plurality of memory cells is configured to generate the first plurality of outputs based on the first plurality of inputs and the stored weight values. The memory cells of the first plurality of synapses are arranged in rows and columns, and the first plurality of synapses include: a plurality of first lines, each of which electrically connects first gates of the memory cells in one of the rows; a plurality of second lines, each of which electrically connects second gates of the memory cells in one of the rows; a plurality of third lines, each of which electrically connects source regions of the memory cells in one of the rows; a plurality of fourth lines, each of which electrically connects drain regions of the memory cells in one of the columns; and a plurality of transistors, each of which is electrically connected in series with one of the fourth lines. The first plurality of synapses are configured to receive the first plurality of inputs as voltages on gates of the plurality of transistors and provide the first plurality of outputs as currents on the plurality of third lines.

[0010] A neural network device may include a first plurality of synapses configured to receive a first plurality of inputs and thereby generate a first plurality of outputs. The first plurality of synapses includes a plurality of memory cells, wherein each of the memory cells includes: a source region and a drain region spaced apart formed in a semiconductor substrate, wherein a channel region extends between the source region and the drain region; a floating gate disposed over and insulated from a first portion of the channel region; a first gate disposed over and insulated from a second portion of the channel region; and a second gate disposed over and insulated from the floating gate or over and insulated from the source region. Each of the plurality of memory cells is configured to store a weight value corresponding to a plurality of electrons on the floating gate. The plurality of memory cells is configured to generate the first plurality of outputs based on the first plurality of inputs and the stored weight values. The memory cells of the first plurality of synapses are arranged in rows and columns, and the first plurality of synapses include: a plurality of first lines, each of which electrically connects first gates in one of the rows of memory cells; a plurality of second lines, each of which electrically connects second gates in one of the columns of memory cells; a plurality of third lines, each of which electrically connects source regions in one of the rows of memory cells; and a plurality of fourth lines, each of which electrically connects drain regions in one of the columns of memory cells. The first plurality of synapses are configured to receive the first plurality of inputs as voltages on the second plurality of lines or the fourth plurality of lines, and to provide the first plurality of outputs as currents on the third plurality of lines.

[0011] A neural network device may include a first plurality of synapses configured to receive a first plurality of inputs and thereby generate a first plurality of outputs. The first plurality of synapses includes a plurality of memory cells, wherein each of the memory cells includes: a source region and a drain region formed in a semiconductor substrate, wherein a channel region extends between the source region and the drain region; a floating gate disposed over and insulated from a first portion of the channel region; a first gate disposed over and insulated from a second portion of the channel region; and a second gate disposed over and insulated from the floating gate or over and insulated from the source region. Each of the plurality of memory cells is configured to store a weight value corresponding to a plurality of electrons on the floating gate. The plurality of memory cells is configured to generate the first plurality of outputs based on the first plurality of inputs and the stored weight values. The memory cells of the first plurality of synapses are arranged in rows and columns. The first plurality of synapses includes: a plurality of first lines, each of which electrically connects the first gates of one of the columns of memory cells together; a plurality of second lines, each of which electrically connects the second gates of one of the rows of memory cells together; a plurality of third lines, each of which electrically connects the source regions of one of the rows of memory cells together; and a plurality of fourth lines, each of which electrically connects the drain regions of one of the columns of memory cells together. The first plurality of synapses is configured to receive the first plurality of inputs as voltages on the first plurality of lines or the fourth plurality of lines, and to provide the first plurality of outputs as currents on the third plurality of lines.

[0012] Other objects and features of the present invention will become apparent from a review of the specification, claims and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 A schematic diagram showing an artificial neural network.

[0014] Figure 2 FIG. 1 is a side cross-sectional view of a conventional 2-gate non-volatile memory cell.

[0015] Figure 3 To show Figure 2 Schematic diagram of a conventional array architecture of memory cells.

[0016] Figure 4 FIG. 1 is a side cross-sectional view of a conventional 2-gate non-volatile memory cell.

[0017] Figure 5 To show Figure 4 Schematic diagram of a conventional array architecture of memory cells.

[0018] Figure 6FIG. 1 is a side cross-sectional view of a conventional 4-gate non-volatile memory cell.

[0019] Figure 7 To show Figure 6 Schematic diagram of a conventional array architecture of memory cells.

[0020] Figure 8A A diagram showing the distribution of evenly spaced neural network weight levels.

[0021] Figure 8B A diagram showing the non-uniformly spaced distribution of neural network weight levels.

[0022] Figure 9 FIG. 1 is a flow chart showing a bidirectional tuning algorithm.

[0023] Figure 10 is a block diagram illustrating weight mapping using current comparison.

[0024] Figure 11 is a block diagram illustrating weight mapping using voltage comparison.

[0025] Figure 12 A schematic diagram showing different levels of an exemplary neural network utilizing a non-volatile memory array.

[0026] Figure 13 is a block diagram showing a vector multiplier matrix.

[0027] Figure 14 is a block diagram showing the various levels of the vector multiplier matrix.

[0028] Figure 15 A side cross-sectional view of a 3-gate nonvolatile memory cell.

[0029] Figure 16 To illustrate the arrangement of a source sum matrix multiplier Figure 15 Schematic diagram of the array architecture of the tri-gate memory cell.

[0030] Figure 17 FIG. 1 is a schematic diagram illustrating a current-to-voltage converter using a tri-gate memory cell.

[0031] Figures 18 to 29 To illustrate the arrangement of a source or drain sum matrix multiplier Figure 15 Schematic diagram of the array architecture of the tri-gate memory cell.

[0032] Figure 30 is a side cross-sectional view of a 3-gate nonvolatile memory cell.

[0033] Figures 31 to 39 To illustrate the arrangement of a drain or source sum matrix multiplier Figure 30Schematic diagram of the array architecture of the tri-gate memory cell.

[0034] Figure 40 The diagram is a schematic diagram showing a controller on the same chip as the memory array for implementing operations of the memory array. DETAILED DESCRIPTION

[0035] The artificial neural network of the present invention utilizes a combination of CMOS technology and a non-volatile memory array. Digital non-volatile memories are well known. For example, U.S. Patent 5,029,130 (the '130 patent) discloses an array of split-gate non-volatile memory cells. The memory cells disclosed in the '130 patent are Figure 2 1 and 2 are shown as memory cells 10. Each memory cell 10 includes a source region 14 and a drain region 16 formed in a semiconductor substrate 12, with a channel region 18 therebetween. A floating gate 20 is formed over and insulated from a first portion of the channel region 18 (and controls its electrical conductivity), and is formed over a portion of the drain region 16. A control gate 22 (i.e., a second channel control gate) has a first portion 22b disposed over and insulated from (and controls the electrical conductivity of) the second portion of the channel region 18, and a second portion 22c extending upwardly along and over the floating gate 20. The floating gate 20 and the control gate 22 are insulated from the substrate 12 by a gate oxide 26.

[0036] Memory cell 10 is erased (wherein electrons are removed from floating gate 20) by placing a high positive voltage on control gate 22, causing the electrons on floating gate 20 to tunnel from floating gate 20 through intervening insulator 24 to control gate 22 via Fowler-Nordheim tunneling.

[0037] The memory cell 10 is programmed by placing a positive voltage on the control gate 22 and a positive voltage on the drain 16 (which places electrons on the floating gate 20). Electron current will flow from the source 14 to the drain 16. When the electrons reach the gap between the control gate 22 and the floating gate 20, they will accelerate and become heated. Due to the electrostatic attraction from the floating gate 20, some of the heated electrons will be injected through the gate oxide 26 onto the floating gate 20.

[0038] Memory cell 10 is read by placing a positive read voltage on drain 16 and control gate 22 (which turns on the portion of the channel region below the control gate). If floating gate 20 is positively charged (i.e., electrons are erased and capacitively coupled to the positive voltage on drain 16), then the portion of channel region 18 below floating gate 20 is also turned on, and current will flow through channel region 18, which is sensed as an erased state or a "1" state. If floating gate 20 is negatively charged (i.e., programmed by electrons), then the portion of channel region 18 below floating gate 20 is mostly or completely turned off, and no current (or very little current) will flow through channel region 18, which is sensed as a programmed state or a "0" state.

[0039] Figure 3 The architecture of a conventional array architecture of memory cells 10 is shown. The memory cells 10 are arranged in rows and columns. In each column, the memory cells are arranged end-to-end in a mirrored manner so that they form pairs of memory cells, each memory cell pair shares a common source region 14 (S), and each adjacent group of memory cell pairs shares a common drain region 16 (D). All source regions 14 of any given row of memory cells are electrically connected together by a source line 14a. All drain regions 16 of any given column of memory cells are electrically connected together by a bit line 16a. All control gates 22 of any given row of memory cells are electrically connected together by a control gate line 22a. Therefore, although memory cells can be programmed and read individually, memory cell erasure is performed row by row (each row of memory cells is erased together by applying a high voltage on the control gate line 22a). If a specific memory cell is to be erased, all memory cells in the same row are also erased.

[0040] Those skilled in the art will appreciate that the source and drain may be interchangeable, wherein the floating gate 20 may partially extend over the source 14 instead of the drain 16, as shown in FIG. Figure 4 shown. Figure 5The corresponding memory cell architecture is best illustrated, including memory cells 10, source lines 14a, bit lines 16a, and control gate lines 22a. As is apparent from the figure, memory cells 10 in the same row share the same source line 14a and the same control gate line 22a, while the drain regions of all cells in the same column are electrically connected to the same bit line 16a. The array design is optimized for digital applications and allows selected cells to be individually programmed, for example, by applying 1.6V and 7.6V to the selected control gate line 22a and source line 14a, respectively, and grounding the selected bit line 16a. By applying a voltage greater than 2 volts to the unselected bit line 16a and grounding the remaining lines, disturbance of unselected memory cells in the same pair is avoided. Memory cells 10 cannot be erased individually because the process responsible for erasure (Fowler-Nordheim tunneling of electrons from floating gate 20 to control gate 22) is only weakly affected by the drain voltage (i.e., the only voltage that may be different for two adjacent cells in the row direction sharing the same source line 14a). Non-limiting examples of operating voltages may include:

[0041] Table 1

[0042] CG 22a BL 16a SL 14a Read 1 0.5-3V 0.1-2V 0V Read 2 0.5-3V 0-2V 2-0.1V Erase About 11-13V 0V 0V programming 1-2V 1-3uA 9-10V

[0043] Read 1 is a read mode in which the cell current is output on the bit line. Read 2 is a read mode in which the cell current is output on the source line.

[0044] Split-gate memory cells having more than two gates are also known. For example, a memory cell having a source region 14, a drain region 16, a floating gate 20 located above a first portion of a channel region 18, a select gate 28 (i.e., a second channel control gate) located above a second portion of the channel region 18, a control gate 22 located above the floating gate 20, and an erase gate 30 located above the source region 14 is known. Figure 6 6,747,310). Here, except for floating gate 20, all gates are non-floating, meaning they are electrically connected or capable of being connected to a voltage source or current source. Programming is illustrated by heated electrons from channel region 18 injecting themselves onto floating gate 20. Erasing is illustrated by electrons tunneling from floating gate 20 to erase gate 30.

[0045] The architecture of the quad-gate memory cell array can be as follows Figure 71 and 2. The memory cells are configured as shown in FIG. 1 . In this embodiment, each horizontal select gate line 28a electrically connects all select gates 28 of the row memory cells together. Each horizontal control gate line 22a electrically connects all control gates 22 of the row memory cells together. Each horizontal source line 14a electrically connects all source regions 14 of two rows of memory cells sharing a source region 14 together. Each bit line 16a electrically connects all drain regions 16 of the column memory cells together. Each erase gate line 30a electrically connects all erase gates 30 of two rows of memory cells sharing an erase gate 30 together. As with the previous architecture, individual memory cells can be independently programmed and read. However, memory cells cannot be erased individually. Erasing is performed by placing a high positive voltage on the erase gate line 30a, which results in erasing two rows of memory cells sharing the same erase gate line 30a at the same time. Exemplary non-limiting operating voltages may include those in Table 2 below (in this embodiment, the select gate line 28a may be referred to as a word line WL):

[0046] Table 2

[0047] SG 28a BL 16a CG 22a EG 30a SL 14a Read 1 0.5-2V 0.1-2V 0-2.6V 0-2.6V 0V Read 2 0.5-2V 0-2V 0-2.6V 0-2.6V 2-0.1V Erase -0.5V / 0V 0V 0V / -8V 8-12V 0V programming 1V 1uA 8-11V 4.5-5V 4.5-5V

[0048] Read 1 is a read mode in which the cell current is output on the bit line. Read 2 is a read mode in which the cell current is output on the source line.

[0049] In order to utilize the above-described non-volatile memory array in a neural network, two modifications can be made. First, the circuitry can be reconfigured so that each memory cell can be individually programmed, erased, and read without adversely affecting the memory state of other memory cells in the array, as explained further below. Second, continuous (analog) programming of the memory cells can be provided. Specifically, the storage or programming state (i.e., the charge on the floating gate as reflected by the number of electrons on the floating gate) of each memory cell in the array can be continuously changed from a fully erased state to a fully programmed state and vice versa independently and with minimal interference with other memory cells. This means that the cell storage device is analog, or at least can store one discrete value among many, which allows very precise and individual tuning of all cells in the memory array, and makes the memory array ideal for storing and fine-tuning the synaptic weights of a neural network.

[0050] Memory cell programming and storage

[0051] The neural network weight levels stored in the memory cells may be evenly spaced (e.g. Figure 8A ), or unevenly spaced (as shown in Figure 8B ). You can use Figure 9The bidirectional tuning algorithm shown is used to program nonvolatile memory cells. Icell is the read current of the target cell being programmed, and Itarget is the desired read current when the cell is ideally programmed. The target cell read current Icell is read (step 1) and compared to the target read current Itarget (step 2). If the target cell read current Icell is greater than the target read current Itarget, a program tuning process is performed (step 3) to increase the number of electrons on floating gate 20 (where a lookup table or silicon-based approximation function can be used to determine the desired initial and incremental program voltages VCG on control gate 22) (steps 3a-3b), which can be repeated as needed (step 3c). If the target cell read current Icell is less than the target read current Itarget, an erase tuning process is performed (step 4) to decrease the number of electrons on floating gate 20 (where a lookup table or silicon-based approximation function can be used to determine the desired initial and incremental erase voltages VEG on erase gate 30) (steps 4a-4b), which can be repeated as needed (step 4c). If the program tuning process exceeds the target read current, the erase tuning process is performed (step 3d and starting from step 4a), and vice versa (step 4d and starting from step 3a), until the target read current is reached (within an acceptable delta value).

[0052] Instead, a unidirectional programming algorithm utilizing programming programming can be used to implement programming of non-volatile memory cells. Using this algorithm, the memory cell 10 is first completely erased and then a Figure 9 The programming tuning steps 3a-3c in the above method are repeated until the read current of the target memory cell 10 reaches the target threshold. Alternatively, a unidirectional tuning algorithm using erase tuning can be used to achieve the tuning of non-volatile memory cells. In this method, the memory cell is first fully programmed and then the Figure 9 The erase tuning steps 4a-4c in FIG. 5 are repeated until the read current of the target memory cell reaches the target threshold.

[0053] Figure 10Schematic diagram illustrating weight mapping using current comparison. The weight digital bits (e.g., a 5-bit weight for each synapse representing the target digital weight of the memory cell) are input to a digital-to-analog converter (DAC) 40, which converts the bits to a voltage Vout (e.g., 64 voltage levels - 5 bits). Vout is converted to a current Iout (e.g., 64 current levels - 5 bits) by a voltage-to-current converter V / I Conv 42. The current Iout is provided to a current comparator IComp 44. A program or erase algorithm enable is input to the memory cell 10 (e.g., erase: increase EG voltage; or program: increase CG voltage). The output memory cell current Icellout (i.e., from a read operation) is provided to the current comparator IComp 44. The current comparator IComp 44 compares the memory cell current Icellout with the current Iout derived from the weight digital bits to generate a signal indicative of the weight stored in the memory cell 10.

[0054] Figure 11 Schematic diagram illustrating weight mapping using voltage comparison. The weight digital bits (e.g., 5-bit weight for each synapse) are input to a digital-to-analog converter (DAC) 40, which converts the bits to a voltage Vout (e.g., 64 voltage levels - 5 bits). Vout is provided to a voltage comparator VComp 46. A program or erase algorithm enable is input to the memory cell 10 (e.g., erase: increase EG voltage; or program: increase CG voltage). The output memory cell current Icellout is provided to a current-to-voltage converter I / V Conv 48 to convert the current to a voltage V2out (e.g., 64 voltage levels - 5 bits). Voltage V2out is provided to a voltage comparator VComp 46. Voltage comparator VComp 46 compares voltages Vout and V2 to generate a signal indicative of the weight stored in the memory cell 10.

[0055] Another embodiment for weight map comparison uses variable pulse widths (i.e., pulse widths that are proportional or inversely proportional to the weight values) for the input weights and / or outputs of the memory cells. In yet another embodiment for weight map comparison, digital pulses (e.g., pulses generated by a clock, where the number of pulses is proportional or inversely proportional to the weight values) are used for the input weights and / or outputs of the memory cells.

[0056] Neural Networks Using Nonvolatile Memory Cell Arrays

[0057] Figure 12A non-limiting example of a neural network utilizing a non-volatile memory array is conceptually illustrated. This example utilizes a non-volatile memory array neural network for a facial recognition application, but any other suitable application can be implemented using a non-volatile memory array-based neural network. For this example, S0 is the input layer, which is a 32×32 pixel RGB image with 5-bit precision (i.e., three 32×32 pixel arrays, one for each color R, G, and B, with 5-bit precision per pixel). Synapse CB1 from S0 to C1 has both different sets of weights and shared weights, and scans the input image with a 3×3 pixel overlapping filter (kernel), shifting the filter by 1 pixel (or more than 1 pixel as dictated by the model). Specifically, the values of 9 pixels in a 3×3 portion of the image (i.e., referred to as the filter or kernel) are provided to synapse CB1, which multiplies these 9 input values by the appropriate weights. After summing the outputs of these multiplications, a single output value is determined and provided by the first synapse of CB1 for use in generating a pixel of one layer C1 of the feature map. The 3×3 filter is then shifted one pixel to the right (i.e., a column of three pixels on the right is added and a column of three pixels on the left is released), whereby the nine pixel values in this newly positioned filter are provided to synapse CB1, whereby they are multiplied by the same weights and a second single output value is determined by the associated synapse. This process continues until the 3×3 filter has scanned all three colors and all bits (precision values) across the entire 32×32 pixel image. This process is then repeated using different sets of weights to generate different feature maps for C1 until all feature maps for layer C1 are calculated.

[0058] At layer C1, in this example, there are 16 feature maps, each with 30×30 pixels. Each pixel is a new feature pixel extracted from the product of the input and the kernel, so each feature map is a two-dimensional array, so in this example, synapse CB1 is composed of a two-dimensional array of 16 layers (remember that the references to neuron layers and arrays in this article are logical, not necessarily physical, i.e., the arrays do not have to be oriented to physical two-dimensional arrays). Each of the 16 feature maps is generated by one of sixteen different sets of synaptic weights applied to the filter scans. The C1 feature maps can all relate to different aspects of the same image features, such as edge recognition. For example, a first map (generated using a first set of weights, shared for all scans used to generate the first map) can identify circular edges, a second map (generated using a second set of weights different from the first) can identify rectangular edges, or the aspect ratio of certain features, and so on.

[0059] Before passing from layer C1 to layer S1, an activation function P1 (pooling) is applied, which pools the values from consecutive non-overlapping 2×2 regions in each feature map. The purpose of the pooling stage is to average neighboring positions (or a max function can also be used) to, for example, reduce dependencies on edge positions and reduce the data size before entering the next stage. At layer S1, there are 16 15×15 feature maps (i.e., sixteen different arrays of 15×15 pixels per feature map). The synapses and associated neurons in CB2 from layer S1 to layer C2 scan the maps in S1 using a 4×4 filter, where the filter is shifted by 1 pixel. At layer C2, there are 22 12×12 feature maps. Before passing from layer C2 to layer S2, an activation function P2 (pooling) is applied, which pools the values from consecutive non-overlapping 2×2 regions in each feature map. At layer S2, there are 22 6×6 feature maps. An activation function is applied to synapse CB3 from layer S2 to layer C3, where every neuron in layer C3 is connected to every mapping in layer S2. At layer C3, there are 64 neurons. Synapse CB4 from layer C3 to output layer S3 fully connects S3 to C3. The output at layer S3 consists of 10 neurons, with the highest output neuron determining the class. For example, this output can indicate recognition or classification of the content of the original image.

[0060] Each level of synapses is implemented using an array or portion of an array of non-volatile memory cells. Figure 13 is a block diagram of a vector-matrix multiplication (VMM) array that includes nonvolatile memory cells and serves as a synapse between the input layer and the next layer. Specifically, VMM array 32 includes a nonvolatile memory cell array 33, an erase gate and word line gate decoder 34, a control gate decoder 35, a bit line decoder 36, and a source line decoder 37, which decode the inputs of memory cell array 33. In this example, source line decoder 37 also decodes the outputs of memory cell array 33. Alternatively, bit line decoder 36 can decode the outputs of nonvolatile memory cell array 33. The memory array serves two purposes. First, it stores weights to be used by VMM array 32. Second, the memory cell array efficiently multiplies the inputs with the weights stored in the memory cell array and adds the results along each output line to produce an output, which serves as the input to the next layer or the final layer. By performing multiplication and addition functions, the memory array eliminates the need for separate multiplication and addition logic circuits and is also highly power-efficient due to its in-situ memory calculations.

[0061] The output of the memory cell array is provided to a single or differential summation circuit 38, which sums the output of the memory cell array to create a single value for the convolution. The summed output value is then provided to an activation function circuit 39, which corrects the output. The activation function can be a sigmoid, tanh, or ReLu function. The corrected output value from circuit 39 becomes an element of the feature map of the next layer (e.g., C1 in the above description), and is then applied to the next synapse to produce the next feature map layer or the final layer. Thus, in this example, the memory cell array 33 constitutes a plurality of synapses (the plurality of synapses receive their inputs from an existing neuron layer or from an input layer such as an image database), and the summation circuit 38 and the activation function circuit 39 constitute a plurality of neurons.

[0062] Figure 14 FIG. 1 is a block diagram illustrating the use of multiple layers of VMM arrays 32 (labeled here as VMM arrays 32a, 32b, 32c, 32d, and 32e). Figure 14 As shown, the input (denoted as Inputx) is converted from digital to analog by a digital-to-analog converter 31 and provided to an input VMM array 32a. The output generated by the input VMM array 32a is provided as input to the next VMM array (hidden level 1) 32b, which in turn generates an output that is provided as input to the next VMM array (hidden level 2) 32c, and so on. The layers of the VMM array 32 serve as different layers of synapses and neurons of a convolutional neural network (CNN). Each VMM array 32a, 32b, 32c, 32d, and 32e can be an independent physical non-volatile memory array, or multiple VMM arrays can utilize different portions of the same non-volatile memory array, or multiple VMM arrays can utilize overlapping portions of the same physical non-volatile memory array. Figure 14 The example shown includes five layers (32a, 32b, 32c, 32d, and 32e): one input layer (32a), two hidden layers (32b, 32c), and two fully connected layers (32d, 32e). Those skilled in the art will appreciate that this is merely exemplary and that, on the contrary, the system may include more than two hidden layers and more than two fully connected layers.

[0063] Figures 15 and 16 The configuration of a tri-gate memory cell array arranged as a source-sum matrix multiplier is shown. Figure 15 shown in , and with Figure 6 The same as the tri-gate memory cell of FIG, except that there is no erase gate. In one embodiment, cell erase is performed by applying a positive voltage to the select gate 28, where electrons tunnel from the floating gate 20 to the select gate 28. Figure 15A table of exemplary non-limiting operating voltages for a tri-gate memory cell.

[0064] Table 3

[0065] SG 28a BL 16a CG 22a SL 14a Read 1 0.50-2V 0.1-2V 0-2.6V 0.1-2V Read 2 0.5-2V 0-2V 0-2.6V 2-0.1V Erase 3-12V 0V 0V to -13V 0V programming 1V 1uA 8-11V 4.5-8V

[0066] Read 1 is a read mode in which the cell current is output on the bit line. Read 2 is a read mode in which the cell current is output on the source line. Figure 15 An alternative erase operation for a tri-gate memory cell may place the p-type substrate 12 at a high voltage (eg, 10V to 20V) and the control gate 22 at a low or negative voltage (eg, -10V to 0V), whereby electrons will tunnel from the floating gate 20 to the substrate 12.

[0067] Used for Figure 15 The lines of the tri-gate memory cell array are Figure 16 shown in , and with Figure 7 The lines in the array are identical, except that there is no erase gate line 30a (because there is no erase gate) and the source line 14a extends vertically rather than horizontally, so that each memory cell can be independently programmed, erased, and read. Specifically, each column of memory cells includes a source line 14a that connects all the source regions 14 of the memory cells in the column together. After programming each memory cell with the appropriate weight value for that cell, the array acts as a source summation matrix multiplier. The matrix voltage inputs are Vin0-Vin3 and are arranged on the control gate line 22a. The matrix current outputs Iout0...Ioutn are generated on the source line 14a. For all cells in the column, each output Iout is the sum of the input current I multiplied by the weight W stored in the cell:

[0068] Iout=Σ(Ii*Wij)

[0069] Where "i" represents the row and "j" represents the column where the memory cell is located. Figure 16 In the case of Vin0-Vin3 shown in FIG, for all cells in the column, each output Iout is proportional to the sum of the input voltage multiplied by the weight W stored in the cell:

[0070] IoutαΣ(Vi*Wij)

[0071] Each memory cell acts as a single neuron with an additive weighted value represented as an output current, Iout, which is determined by the sum of the weight values stored in the memory cells in that column. The output of any given neuron is in the form of a current, which can then be used as the input current, Iin, for the next subsequent VMM array stage after being adjusted by the activation function circuit.

[0072] Given that the input is voltage and the output is current, Figure 16 In the embodiment of the present invention, each subsequent VMM stage after the first stage preferably includes a circuit for converting the input current from the previous VMM stage into a voltage to be used as the input voltage Vin. Figure 17 An example of such a current-to-voltage conversion circuit is shown, which converts the input current Iin0...Iin1 logarithmically into the input voltage Vin0..Vin1 for application to a modified memory cell row in a subsequent stage. The memory cells described herein are biased in weak inversion,

[0073] Ids=Io*e (Vg-Vth) / kVt =w*Io*e (Vg) / kVt

[0074] where w = e (-Vth) / kVt

[0075] For an I to V logarithmic converter that uses a memory cell to convert input current to input voltage:

[0076] Vg=k*Vt*log[Ids / wp*Io]

[0077] Here, wp is the w of the reference memory cell or the peripheral memory cell. For a memory array used as a vector matrix multiplier VMM, the output current is:

[0078] Iout=wa*Io*e (Vg) / kVt ,Right now

[0079] Iout=(wa / wp)*Iin=W*Iin

[0080] W=e (Vthp-Vtha) / kVt

[0081] Here, wa = w for each memory cell in the memory array. A select gate line 28a, which is connected to the bit line 16a through a switch BLR that is closed during current-to-voltage conversion, can be used as an input to the memory cell to obtain an input voltage.

[0082] Alternatively, the non-volatile memory cells of the VMM array described herein may be configured to operate in the linear region:

[0083] Ids=beta*(Vgs-Vth)*Vds; beta=u*Cox*Wt / L,

[0084] Where Wt and L are the width and length of the transistor respectively

[0085] Wα(Vgs-Vth), weight W is proportional to (Vgs-Vth)

[0086] The select gate line or control gate line or bit line or source line can be used as the input of the memory cell operating in the linear region.The bit line or source line can be used as the output of the output neuron.

[0087] For an IV linear converter, a memory cell (such as a reference memory cell or a peripheral memory cell) or a transistor or resistor operating in the linear region can be used to linearly convert the input / output current to the input / output voltage. Alternatively, the non-volatile memory cells of the VMM array described herein can be configured to operate in the saturation region:

[0088] Ids=α1 / 2*beta*(Vgs-Vth) 2 ; beta=u*Cox*Wt / L

[0089] Wα(Vgs-Vth) 2 , refers to the weight W and (Vgs-Vth) 2 Proportional

[0090] The select gate line or control gate can be used as the input of the memory cell operating in the saturation region. The bit line or source line can be used as the output of the output neuron. Alternatively, the non-volatile memory cells of the VMM array described herein can be used in all regions or combinations thereof (subthreshold, linear or saturation regions). Any of the above-mentioned current-to-voltage conversion circuits or techniques can be used with any of the embodiments herein so that the current output from any given neuron in the form of current can then be used as the input for the next subsequent VMM array stage after being adjusted by the activation function circuit.

[0091] Figure 18 Shown Figure 15 Another configuration of an array of tri-gate memory cells 10 arranged as a source-sum matrix multiplier is shown in FIG. Figure 18 The lines of the array are Figure 16. However, the matrix voltage inputs Vin0-Vin3 are arranged on select gate line 28a, and the matrix current outputs Iout0...IoutN are produced on source line 14a (i.e., for all cells in the column, each output Iout is the sum of the cell currents proportional to the weight W stored in the cell). As with the previous embodiment (and any other embodiment described herein), the output of any given neuron is in the form of a current, which can then be used as input for the next subsequent VMM array stage after being adjusted by the activation function circuit.

[0092] Figure 19 Shown Figure 15 Another configuration of an array of tri-gate memory cells 10 arranged as a drain-sum matrix multiplier is shown in FIG. Figure 19 The lines of the array are Figure 16 However, the matrix voltage inputs Vin0-Vin3 are arranged on control gate line 22a, and the matrix current outputs Iout0...IoutN are produced on bit lines 16a (i.e., for all cells in a column, each output Iout is the sum of the cell currents proportional to the weight W stored in the cell).

[0093] Figure 20 Shown Figure 15 Another configuration of an array of tri-gate memory cells 10 arranged as a drain-sum matrix multiplier is shown in FIG. Figure 20 The lines of the array are Figure 16 However, the matrix voltage inputs Vin0-Vin3 are arranged on select gate line 28a, and the matrix current outputs Iout0...IoutN are produced on bit line 16a (i.e., for all cells in a column, each output Iout is the sum of the cell currents proportional to the weight W stored in the cell).

[0094] Figure 21 Shown Figure 15 Another configuration of an array of tri-gate memory cells 10 arranged as a drain-sum matrix multiplier is shown in FIG. Figure 21 The lines of the array are Figure 16 1 . The same as the lines in the array of , except that the source lines 14a extend horizontally rather than vertically. Specifically, each row of memory cells includes a source line 14a that connects all the source regions 14 of the memory cells in that row together. The matrix voltage inputs Vin0-Vin3 are arranged on select gate line 28a, and the matrix current outputs Iout0...IoutN are generated on bit lines 16a (i.e., for all cells in a column, each output Iout is the sum of the cell currents proportional to the weight W stored in the cell).

[0095] Figure 22 Shown Figure 15 Another configuration of an array of tri-gate memory cells 10 arranged as a drain-sum matrix multiplier is shown in FIG. Figure 22 The lines of the array are Figure 21 Matrix voltage inputs Vin0-Vin3 are arranged on control gate line 22a, and matrix current outputs Iout0...IoutN are produced on bit lines 16a (i.e., for all cells in a column, each output Iout is the sum of the cell currents proportional to the weight W stored in the cell).

[0096] Figure 23 Shown Figure 15 Another configuration of an array of tri-gate memory cells 10 arranged as a drain-sum matrix multiplier is shown in FIG. Figure 23 The lines of the array are Figure 21 Matrix voltage inputs Vin0-Vin1 are arranged on source line 14a, and matrix current outputs Iout0...IoutN are produced on bit lines 16a (i.e., for all cells in a column, each output Iout is the sum of the cell currents proportional to the weight W stored in the cell).

[0097] Figure 24 Shown Figure 15 Another configuration of an array of tri-gate memory cells 10 arranged as a source-sum matrix multiplier is shown in FIG. Figure 24 The lines of the array are Figure 21 The matrix voltage inputs Vin0-Vinn are arranged on bit line 16a, and the matrix current outputs Iout0...Iout1 are produced on source line 14a (i.e., for all cells in a row, each output Iout is the sum of the cell currents proportional to the weight W stored in the cell).

[0098] Figure 25 Shown Figure 15 Another configuration of an array of tri-gate memory cells 10 arranged as a source-sum matrix multiplier is shown in FIG. Figure 25 The lines of the array are Figure 21The lines in the array of are identical except that each bit line includes a bit line buffer transistor 60 connected in series with the bit line (i.e., a transistor between whose source and drain any current on the bit line flows). Transistor 60 acts as a graduated switch that selectively and gradually turns the bit line on (i.e., the transistor couples the bit line to its current or voltage source) as the input voltage on the transistor gate terminal increases. Matrix voltage inputs Vin0...Vinn are provided to the gates of transistors 60, and matrix current outputs lout0...Iout1 are provided onto source line 14a. An advantage of this configuration is that the matrix inputs can be provided as voltages (to operate transistors 60) rather than providing inputs directly to the bit lines in the form of voltages. This enables the bit lines to be operated using a constant voltage source that is gradually coupled to the bit lines using transistors 60 in response to the input voltage Vin provided to the transistor gates, thereby eliminating the need to provide a voltage input to the memory array. For a memory array in which the bit lines receive inputs Figure 15 For any of the embodiments herein of memory cell 10, transistor 60 may be included on the bit line to receive input on its gate rather than applying the input to the bit line.

[0099] Figure 26 Shown Figure 15 Another configuration of an array of tri-gate memory cells 10 arranged as a source-sum matrix multiplier is shown in FIG. Figure 26 The lines of the array are Figure 21 The lines in the array are identical, except that the control gate line 22a extends vertically rather than horizontally. Specifically, each column of memory cells includes a control gate line 22a connecting all the control gates 22 in the column together. After programming each memory cell with the appropriate weight value for that cell, the array acts as a source-sum matrix multiplier. The matrix voltage inputs are Vin0-Vinn and are arranged on the control gate line 22a. The matrix current outputs Iout0...Iout1 are generated on the source line 14a. For all cells in the row, each output Iout is the sum of the cell currents proportional to the weight W stored in the cell.

[0100] Figure 27 Shown Figure 15 Another configuration of an array of tri-gate memory cells 10 arranged as a source-sum matrix multiplier is shown in FIG. Figure 27 The lines of the array are Figure 21The lines in the array are identical, except that the control gate line 22a extends vertically rather than horizontally. Specifically, each column of memory cells includes a control gate line 22a connecting all the control gates 22 in that column together. After programming each memory cell with the appropriate weight value for that cell, the array acts as a source-sum matrix multiplier. The matrix voltage inputs are Vin0-Vinn and are arranged on bit lines 16a. The matrix current outputs lout0...Iout1 are generated on source lines 14a. For all cells in the row, each output lout is the sum of the cell currents proportional to the weight W stored in the cell.

[0101] Figure 28 Shown Figure 15 Another configuration of an array of tri-gate memory cells 10 arranged as a source-sum matrix multiplier is shown in FIG. Figure 28 The lines of the array are Figure 21 The lines in the array are identical, except that the select gate line 28a extends vertically rather than horizontally. Specifically, each column of memory cells includes a select gate line 28a that connects all the select gates 28 in that column together. After programming each memory cell with the appropriate weight value for that cell, the array acts as a source-sum matrix multiplier. The matrix voltage inputs are Vin0-Vinn and are arranged on the select gate line 28a. The matrix current outputs Iout0...Iout1 are generated on the source line 14a. For all cells in the row, each output Iout is the sum of the cell currents proportional to the weight W stored in the cell.

[0102] Figure 29 Shown Figure 15 Another configuration of an array of tri-gate memory cells 10 arranged as a source-sum matrix multiplier is shown in FIG. Figure 29 The lines of the array are Figure 21 The lines in the array are identical, except that the select gate line 28a extends vertically rather than horizontally. Specifically, each column of memory cells includes a select gate line 28a connecting all the select gates 28 in that column together. After programming each memory cell with the appropriate weight value for that cell, the array acts as a source-sum matrix multiplier. The matrix voltage inputs are Vin0-Vinn and are arranged on bit lines 16a. The matrix current outputs lout0...Iout1 are generated on source lines 14a. For all cells in the row, each output lout is the sum of the cell currents proportional to the weight W stored in the cell.

[0103] Figures 30 to 31 Another configuration of a tri-gate memory cell array having a different configuration arranged as a drain-sum matrix multiplier is shown. Figure 30 shown in , and with Figure 6 The tri-gate memory cell is identical to that of FIG, except that there is no control gate 22. The select gate 28 as shown does not extend upwardly over any portion of the floating gate 20. However, the select gate 28 may include an additional portion that extends upwardly over a portion of the floating gate 20.

[0104] Reading, erasing, and programming are performed in a similar manner, but without any bias on the control gate (not included). Example non-limiting operating voltages may include those in Table 4 below:

[0105] Table 4

[0106] SG 28a BL 16a EG 30a SL 14a Read 1 0.5-2V 0.1-2V 0-2.6V 0V Read 2 0.5-2V 0-2V 0-2.6V 2-0.1V Erase -0.5V / 0V 0V 8-12V 0V programming 1V 1uA 4.5-7V 4.5-7V

[0107] Read 1 is a read mode in which the cell current is output on the bit line. Read 2 is a read mode in which the cell current is output on the source line.

[0108] Figure 31 Shows the use Figure 30 FIG2 is a memory cell array architecture of memory cells 10, wherein all lines except bit line 16a extend in the horizontal / row direction. After programming each memory cell with the appropriate weight value for that cell, the array functions as a bit line summing matrix multiplier. The matrix voltage inputs are Vin0-Vin3 and are arranged on select gate line 28a. The matrix outputs Iout0...Ioutn are generated on bit line 16a. For all cells in the column, each output Iout is the sum of the cell currents proportional to the weight W stored in the cell.

[0109] Figure 32 Shown Figure 30 Another configuration of an array of tri-gate memory cells 10 arranged as a drain-sum matrix multiplier is shown in FIG. Figure 32 The lines of the array are Figure 31 Matrix voltage inputs Vin0-Vin1 are arranged on erase gate line 30a, and matrix current outputs Iout0...IoutN are produced on bit lines 16a (i.e., for all cells in a column, each output Iout is the sum of the cell currents proportional to the weight W stored in the cell).

[0110] Figure 33 Shown Figure 30 Another configuration of an array of tri-gate memory cells 10 arranged as a drain-sum matrix multiplier is shown in FIG. Figure 33 The lines of the array are Figure 31Matrix voltage inputs Vin0-Vin1 are arranged on source line 14a, and matrix current outputs Iout0...IoutN are produced on bit lines 16a (i.e., for all cells in a column, each output Iout is the sum of the cell currents proportional to the weight W stored in the cell).

[0111] Figure 34 Shown Figure 30 Another configuration of an array of tri-gate memory cells 10 arranged as a source-sum matrix multiplier is shown in FIG. Figure 34 The lines of the array are Figure 31 The matrix voltage inputs Vin0-Vinn are arranged on bit line 16a, and the matrix current outputs Iout0...Iout1 are produced on source line 14a (i.e., for all cells in a row, each output Iout is the sum of the cell currents proportional to the weight W stored in the cell).

[0112] Figure 35 Shown Figure 30 Another configuration of an array of tri-gate memory cells 10 arranged as a source-sum matrix multiplier is shown in FIG. Figure 35 The lines of the array are Figure 31 The lines in the array of are identical except that each bit line includes a bit line buffer transistor 60 connected in series with the bit line (i.e., a transistor between whose source and drain any current on the bit line flows). Transistor 60 acts as a graduated switch that selectively and gradually turns the bit line on (i.e., the transistor couples the bit line to its current or voltage source) as the input voltage on the transistor gate terminal increases. Matrix voltage inputs Vin0...Vinn are provided to the gates of transistors 60, and matrix current outputs lout0...Iout1 are provided onto source line 14a. An advantage of this configuration is that the matrix inputs can be provided as voltages (to operate transistors 60) rather than providing inputs directly to the bit lines in the form of voltages. This enables the bit lines to be operated using a constant voltage source that is gradually coupled to the bit lines using transistors 60 in response to the input voltage Vin provided to the transistor gates, thereby eliminating the need to provide a voltage input to the memory array. For a memory array in which the bit lines receive inputs Figure 30 For any of the embodiments herein of memory cell 10, transistor 60 may be included on the bit line to receive input on its gate rather than applying the input to the bit line.

[0113] Figure 36 Shown Figure 30 Another configuration of an array of tri-gate memory cells 10 arranged as a source-sum matrix multiplier is shown in FIG. Figure 36 The lines of the array are Figure 31The lines in the array are identical, except that the select gate line 28a extends vertically rather than horizontally. Specifically, each column of memory cells includes a select gate line 28a that connects all the select gates 28 in that column together. After programming each memory cell with the appropriate weight value for that cell, the array acts as a source-sum matrix multiplier. The matrix voltage inputs are Vin0-Vinn and are arranged on the select gate line 28a. The matrix current outputs Iout0...Iout1 are generated on the source line 14a. For all cells in the row, each output Iout is the sum of the cell currents proportional to the weight W stored in the cell.

[0114] Figure 37 Shown Figure 30 Another configuration of an array of tri-gate memory cells 10 arranged as a source-sum matrix multiplier is shown in FIG. Figure 37 The lines of the array are Figure 31 The lines in the array are identical, except that the select gate line 28a extends vertically rather than horizontally. Specifically, each column of memory cells includes a select gate line 28a connecting all the select gates 28 in that column together. After programming each memory cell with the appropriate weight value for that cell, the array acts as a source-sum matrix multiplier. The matrix voltage inputs are Vin0-Vinn and are arranged on bit lines 16a. The matrix current outputs lout0...Iout1 are generated on source lines 14a. For all cells in the row, each output lout is the sum of the cell currents proportional to the weight W stored in the cell.

[0115] Figure 38 Shown Figure 30 Another configuration of an array of tri-gate memory cells 10 arranged as a source-sum matrix multiplier is shown in FIG. Figure 38 The lines of the array are Figure 31 The lines in the array are identical, except that the erase gate line 30a extends vertically rather than horizontally. Specifically, each column of memory cells includes an erase gate line 30a that connects all the erase gates 30 in the column together. After programming each memory cell with the appropriate weight value for that cell, the array acts as a source sum matrix multiplier. The matrix inputs are Vin0-Vinn and are arranged on the erase gate line 30a. The matrix current outputs Iout0...Iout1 are generated on the source line 14a. For all cells in the row, each output Iout is the sum of the cell currents proportional to the weight W stored in the cell.

[0116] Figure 39 Shown Figure 30 Another configuration of an array of tri-gate memory cells 10 arranged as a source-sum matrix multiplier is shown in FIG. Figure 39 The lines of the array are Figure 31 The lines in the array are identical, except that the erase gate line 30a extends vertically rather than horizontally. Specifically, each column of memory cells includes an erase gate line 30a that connects all the erase gates 30 in that column together. After programming each memory cell with the appropriate weight value for that cell, the array acts as a source-sum matrix multiplier. The matrix voltage inputs are Vin0-Vinn and are arranged on bit lines 16a. The matrix current outputs Iout0...Iout1 are generated on source lines 14a. For all cells in the row, each output Iout is the sum of the cell currents proportional to the weight W stored in the cell.

[0117] about Figures 15 to 20 implementation plan, namely Figure 15 An array of memory cells in which Figures 16 to 20 Each of the arrays in includes source lines 14a extending vertically (in the column direction) rather than horizontally (in the row direction), Figures 16 to 20 The array will also be applicable to an array having an erase gate 30 instead of a control gate 22. Figure 30 Specifically, Figures 16 to 20 The array will be Figure 30 Memory cell 10 instead of Figure 15 Memory cells 10 in the array, where in each array, the control gate line 22a will be replaced by an erase gate line 30a (ie, each of the erase gate lines 30a will connect all erase gates 30 of a row of memory cells 10 together). Figures 16 to 20 All other lines in will remain unchanged.

[0118] All of the above functions may be performed under the control of a controller 100 connected to the memory array of the above memory unit 10 for neural network functions. Figure 40 As shown, controller 100 is preferably on the same semiconductor chip or substrate 110 as memory array 120. However, controller 100 may also be located on a separate semiconductor chip or substrate, and may be a collection of multiple controllers disposed at different locations on or off semiconductor chip or substrate 110.

[0119] It should be understood that the present invention is not limited to the embodiments described above and shown herein, but encompasses any and all variations within the scope of any claims. For example, reference to the invention herein is not intended to limit the scope of any claim or claim term, but rather is merely a reference to one or more features that may be covered by one or more claims. The examples of materials, processes, and numerical values described above are merely exemplary and should not be construed as limiting the claims. A single layer of material may be formed as multiple layers of such material or similar materials, and vice versa. Although the outputs of each memory cell array are manipulated by filtering and condensing before being sent to the next neuron layer, they do not have to be so. Finally, for each of the above-described matrix multiplier array embodiments, for any lines that are not used for input voltage or output current, the nominal read voltages for that memory cell configuration disclosed in the table herein may (but need not) be applied to those lines during operation.

[0120] It should be noted that, as used herein, the terms "above" and "on" include inclusively "directly on" (without an intervening material, element, or space therebetween) and "indirectly on" (with an intervening material, element, or space therebetween). Similarly, the term "adjacent" includes "directly adjacent" (without an intervening material, element, or space therebetween) and "indirectly adjacent" (with an intervening material, element, or space therebetween), "mounted to" includes "directly mounted to" (without an intervening material, element, or space therebetween) and "indirectly mounted to" (with an intervening material, element, or space therebetween), and "electrically coupled to" includes "directly electrically coupled to" (without an intervening material or element electrically connecting the elements together) and "indirectly electrically coupled to" (with an intervening material or element electrically connecting the elements together). For example, forming an element "above a substrate" may include forming the element directly on the substrate without an intervening material / element therebetween, and forming the element indirectly on the substrate with one or more intervening materials / elements therebetween.

Claims

1. A neural network processing device, comprising: a first plurality of synapses configured to receive a first plurality of inputs and generate a first plurality of outputs therefrom, wherein the first plurality of synapses comprises: A plurality of memory cells, wherein each of the memory cells comprises: a spaced-apart source region and a drain region formed in a semiconductor substrate, wherein a channel region extends between the source region and the drain region; a floating gate disposed over and insulated from a first portion of the channel region; and a first gate disposed over and insulated from a second portion of the channel region. and a second gate disposed above the floating gate and insulated from the floating gate or disposed above the source region and insulated from the source region; Each of the plurality of memory cells is configured to store a weight value corresponding to a plurality of electrons on the floating gate; The plurality of memory cells are configured to generate the first plurality of outputs based on the first plurality of inputs and stored weight values; wherein the memory cells of the first plurality of synapses are arranged in rows and columns, and wherein the first plurality of synapses comprises: a plurality of first lines, each first line electrically connecting the first gates in one of the rows of the memory cells together; a plurality of second lines, each second line electrically connecting the second gates in one of the rows of the memory cells together; a plurality of third lines, each third line electrically connecting the source regions in one of the columns of the memory cells together; a plurality of fourth lines, each fourth line electrically connecting together the drain regions in one of the columns of the memory cells; wherein the first plurality of synapses are configured to receive the first plurality of inputs as voltages on the plurality of second lines and to provide the first plurality of outputs as currents on the plurality of third lines or the plurality of fourth lines.

2. The neural network processing device of claim 1 , wherein for each of the plurality of memory cells, the second gate is disposed above and insulated from the floating gate, and wherein the first plurality of synapses are configured to provide the first plurality of outputs as currents on the plurality of third lines.

3. The neural network processing device of claim 1 , wherein for each of the plurality of memory cells, the second gate is disposed above and insulated from the source region, and wherein the first plurality of synapses are configured to provide the first plurality of outputs as currents on the plurality of third lines.

4. The neural network processing device of claim 1 , wherein for each of the plurality of memory cells, the second gate is disposed above and insulated from the floating gate, and wherein the first plurality of synapses are configured to provide the first plurality of outputs as currents on the plurality of fourth lines.

5. The neural network processing device of claim 1 , wherein for each of the plurality of memory cells, the second gate is disposed above and insulated from the source region, and wherein the first plurality of synapses are configured to provide the first plurality of outputs as currents on the plurality of fourth lines.

6. The neural network processing device according to claim 1, further comprising: A first plurality of neurons is configured to receive the first plurality of outputs.

7. The neural network processing device according to claim 6, further comprising: a second plurality of synapses configured to receive a second plurality of inputs from the first plurality of neurons and generate a second plurality of outputs therefrom, wherein the second plurality of synapses comprises: a plurality of second memory cells, wherein each of the second memory cells comprises: a second source region and a second drain region spaced apart from each other formed in the semiconductor substrate, wherein a second channel region extends between the second source region and the second drain region; a second floating gate disposed over a first portion of the second channel region and insulated from the first portion; a third gate disposed over a second portion of the second channel region and insulated from the second portion; and a fourth gate disposed over the second floating gate and insulated from the second floating gate or disposed over the second source region and insulated from the second source region; Each of the plurality of second memory cells is configured to store a second weight value corresponding to a plurality of electrons on the second floating gate; the plurality of second memory cells being configured to generate the second plurality of outputs based on the second plurality of inputs and the stored second weight values; wherein the second memory cells of the second plurality of synapses are arranged in rows and columns, and wherein the second plurality of synapses comprises: a plurality of fifth lines, each fifth line electrically connecting the third gates in one of the rows of the second memory cells together; a plurality of sixth lines, each sixth line electrically connecting the fourth gates in one of the rows of the second memory cells together; a plurality of seventh lines, each seventh line electrically connecting the second source regions in one of the columns of the second memory cells together; a plurality of eighth lines, each eighth line electrically connecting the second drain regions in one of the columns of the second memory cells together; Wherein the second plurality of synapses are configured to receive the second plurality of inputs as voltages on the plurality of sixth lines and to provide the second plurality of outputs as currents on the plurality of seventh lines or the plurality of eighth lines.

8. The neural network processing device of claim 7 , wherein for each of the plurality of second memory cells, the fourth gate is disposed above and insulated from the second floating gate, and wherein the second plurality of synapses are configured to provide the second plurality of outputs as currents on the plurality of seventh lines.

9. The neural network processing device of claim 7 , wherein for each of the plurality of second memory cells, the fourth gate is disposed above and insulated from the second source region, and wherein the second plurality of synapses are configured to provide the second plurality of outputs as current on the plurality of seventh lines.

10. The neural network processing device of claim 7 , wherein for each of the plurality of second memory cells, the fourth gate is disposed above and insulated from the second floating gate, and wherein the second plurality of synapses are configured to provide the second plurality of outputs as currents on the plurality of eighth lines.

11. The neural network processing device according to claim 7 , wherein for each of the plurality of second memory cells, the fourth gate is set at the second The second source region is above and insulated from the second source region, and wherein the second plurality of synapses are configured to provide the second plurality of outputs as currents on the eighth plurality of lines.

12. A neural network processing device comprising: a first plurality of synapses configured to receive a first plurality of inputs and generate a first plurality of outputs therefrom, wherein the first plurality of synapses comprises: A plurality of memory cells, wherein each of the memory cells comprises: a spaced-apart source region and a drain region formed in a semiconductor substrate, wherein a channel region extends between the source region and the drain region; a floating gate disposed over and insulated from a first portion of the channel region; and a first gate disposed over and insulated from a second portion of the channel region. and a second gate disposed above the floating gate and insulated from the floating gate or disposed above the source region and insulated from the source region; Each of the plurality of memory cells is configured to store a weight value corresponding to a plurality of electrons on the floating gate; The plurality of memory cells are configured to generate the first plurality of outputs based on the first plurality of inputs and stored weight values; wherein the memory cells of the first plurality of synapses are arranged in rows and columns, and wherein the first plurality of synapses comprises: a plurality of first lines, each first line electrically connecting the first gates in one of the rows of the memory cells together; a plurality of second lines, each second line electrically connecting the second gates in one of the rows of the memory cells together; a plurality of third lines, each third line electrically connecting the source regions in one of the rows of the memory cells together; a plurality of fourth lines, each fourth line electrically connecting together the drain regions in one of the columns of the memory cells; wherein the first plurality of synapses are configured to receive the first plurality of inputs as voltages on the plurality of second lines or the plurality of third lines, and to provide the first plurality of outputs as currents on the plurality of fourth lines.

13. The neural network processing device of claim 12, wherein the first plurality of synapses are configured to receive the first plurality of inputs as voltages on the second plurality of lines.

14. The neural network processing device of claim 12, wherein the first plurality of synapses are configured to receive the first plurality of inputs as voltages on the third plurality of lines.

15. The neural network processing device according to claim 12, further comprising: A first plurality of neurons is configured to receive the first plurality of outputs.

16. The neural network processing device according to claim 15, further comprising: a second plurality of synapses configured to receive a second plurality of inputs from the first plurality of neurons and generate a second plurality of outputs therefrom, wherein the second plurality of synapses comprises: a plurality of second memory cells, wherein each of the second memory cells comprises: a second source region and a second drain region spaced apart from each other formed in the semiconductor substrate, wherein a second channel region extends between the second source region and the second drain region; a second floating gate disposed over a first portion of the second channel region and insulated from the first portion; a third gate disposed over a second portion of the second channel region and insulated from the second portion; and a fourth gate disposed over the second floating gate and insulated from the second floating gate or disposed over the second source region and insulated from the second source region; Each of the plurality of second memory cells is configured to store a second weight value corresponding to a plurality of electrons on the second floating gate; The plurality of second memory cells are configured to generate the second plurality of outputs based on the second plurality of inputs and the stored second weight values; wherein the second memory cells of the second plurality of synapses are arranged in rows and columns, and wherein the second plurality of synapses comprises: a plurality of fifth lines, each fifth line electrically connecting the third gates in one of the rows of the second memory cells together; a plurality of sixth lines, each sixth line electrically connecting the fourth gates in one of the rows of the second memory cells together; a plurality of seventh lines, each seventh line electrically connecting the second source regions in one of the rows of the second memory cells together; a plurality of eighth lines, each eighth line electrically connecting together the second drain regions in one of the columns of the second memory cells; wherein the second plurality of synapses are configured to receive the second plurality of inputs as voltages on the plurality of sixth lines or the plurality of seventh lines and to provide the second plurality of outputs as currents on the plurality of eighth lines. 17 . The neural network processing device of claim 16 , wherein the second plurality of synapses are configured to receive the second plurality of inputs as voltages on the plurality of sixth lines.

18. The neural network processing device of claim 16, wherein the second plurality of synapses are configured to receive the second plurality of inputs as voltages on the seventh plurality of lines.

19. A neural network processing device comprising: a first plurality of synapses configured to receive a first plurality of inputs and generate a first plurality of outputs therefrom, wherein the first plurality of synapses comprises: A plurality of memory cells, wherein each of the memory cells comprises: a spaced-apart source region and a drain region formed in a semiconductor substrate, wherein a channel region extends between the source region and the drain region; a floating gate disposed over and insulated from a first portion of the channel region; and a first gate disposed over and insulated from a second portion of the channel region. and a second gate disposed above the floating gate and insulated from the floating gate or disposed above the source region and insulated from the source region; Each of the plurality of memory cells is configured to store a weight value corresponding to a plurality of electrons on the floating gate; The plurality of memory cells are configured to generate the first plurality of outputs based on the first plurality of inputs and stored weight values; wherein the memory cells of the first plurality of synapses are arranged in rows and columns, and wherein the first plurality of synapses comprises: a plurality of first lines, each first line electrically connecting the first gates in one of the rows of the memory cells together; a plurality of second lines, each second line electrically connecting the second gates in one of the rows of the memory cells together; a plurality of third lines, each third line electrically connecting the source regions in one of the rows of the memory cells together; a plurality of fourth lines, each fourth line electrically connecting together the drain regions in one of the columns of the memory cells; wherein the first plurality of synapses are configured to receive the first plurality of inputs as voltages on the plurality of fourth lines and to provide the first plurality of outputs as currents on the plurality of third lines.

20. The neural network processing device according to claim 19, further comprising: A first plurality of neurons is configured to receive the first plurality of outputs.

21. The neural network processing device according to claim 20, further comprising: a second plurality of synapses configured to receive a second plurality of inputs from the first plurality of neurons and generate a second plurality of outputs therefrom, wherein the second plurality of synapses comprises: a plurality of second memory cells, wherein each of the second memory cells comprises: a second source region and a second drain region spaced apart from each other formed in the semiconductor substrate, wherein a second channel region extends between the second source region and the second drain region; a second floating gate disposed over a first portion of the second channel region and insulated from the first portion; a third gate disposed over a second portion of the second channel region and insulated from the second portion; and a fourth gate disposed over the second floating gate and insulated from the second floating gate or disposed over the second source region and insulated from the second source region; Each of the plurality of second memory cells is configured to store a second weight value corresponding to a plurality of electrons on the second floating gate; the plurality of second memory cells being configured to generate the second plurality of outputs based on the second plurality of inputs and the stored second weight values; wherein the second memory cells of the second plurality of synapses are arranged in rows and columns, and wherein the second plurality of synapses comprises: a plurality of fifth lines, each fifth line electrically connecting the third gates in one of the rows of the second memory cells together; a plurality of sixth lines, each sixth line electrically connecting the fourth gates in one of the rows of the second memory cells together; a plurality of seventh lines, each seventh line electrically connecting the second source regions in one of the rows of the second memory cells together; a plurality of eighth lines, each eighth line electrically connecting the second drain regions in one of the columns of the second memory cells together; Wherein the second plurality of synapses are configured to receive the second plurality of inputs as voltages on the plurality of eighth lines and to provide the second plurality of outputs as currents on the plurality of seventh lines.

22. A neural network processing device comprising: a first plurality of synapses configured to receive a first plurality of inputs and generate a first plurality of outputs therefrom, wherein the first plurality of synapses comprises: A plurality of memory cells, wherein each of the memory cells comprises: a spaced-apart source region and a drain region formed in a semiconductor substrate, wherein a channel region extends between the source region and the drain region; a floating gate disposed over and insulated from a first portion of the channel region; and a first gate disposed over and insulated from a second portion of the channel region. and a second gate disposed above the floating gate and insulated from the floating gate or disposed above the source region and insulated from the source region; Each of the plurality of memory cells is configured to store a weight value corresponding to a plurality of electrons on the floating gate; the plurality of memory cells being configured to generate the first plurality of outputs based on the first plurality of inputs and the stored weight values; wherein the memory cells of the first plurality of synapses are arranged in rows and columns, and wherein the first plurality of synapses comprises: a plurality of first lines, each first line electrically connecting the first gates in one of the rows of the memory cells together; a plurality of second lines, each second line electrically connecting the second gates in one of the rows of the memory cells together; a plurality of third lines, each third line electrically connecting the source regions in one of the rows of the memory cells together; a plurality of fourth lines, each fourth line electrically connecting the drain regions in one of the columns of the memory cells together; a plurality of transistors, each transistor being electrically connected in series with one of the fourth lines; The first plurality of synapses are configured to receive the first plurality of inputs as voltages on gates of the plurality of transistors and to provide the first plurality of outputs as currents on the plurality of third lines.

23. The neural network processing device according to claim 22, further comprising: A first plurality of neurons is configured to receive the first plurality of outputs.

24. The neural network processing device according to claim 23, further comprising: a second plurality of synapses configured to receive a second plurality of inputs from the first plurality of neurons and generate a second plurality of outputs therefrom, wherein the second plurality of synapses comprises: a plurality of second memory cells, wherein each of the second memory cells comprises: a second source region and a second drain region spaced apart from each other formed in the semiconductor substrate, wherein a second channel region extends between the second source region and the second drain region; a second floating gate disposed over a first portion of the second channel region and insulated from the first portion; a third gate disposed over a second portion of the second channel region and insulated from the second portion; and a fourth gate disposed over the second floating gate and insulated from the second floating gate or disposed over the second source region and insulated from the second source region; Each of the plurality of second memory cells is configured to store a second weight value corresponding to a plurality of electrons on the second floating gate; the plurality of second memory cells being configured to generate the second plurality of outputs based on the second plurality of inputs and the stored second weight values; wherein the second memory cells of the second plurality of synapses are arranged in rows and columns, and wherein the second plurality of synapses comprises: a plurality of fifth lines, each fifth line electrically connecting the third gates in one of the rows of the second memory cells together; a plurality of sixth lines, each sixth line electrically connecting the fourth gates in one of the rows of the second memory cells together; a plurality of seventh lines, each seventh line electrically connecting the second source regions in one of the rows of the second memory cells together; a plurality of eighth lines, each eighth line electrically connecting the second drain regions in one of the columns of the second memory cells together; a second plurality of transistors, each second transistor electrically connected in series with one of the eighth lines; Wherein the second plurality of synapses are configured to receive the second plurality of inputs as voltages on gates of the second plurality of transistors and to provide the second plurality of outputs as currents on the plurality of seventh lines.

25. A neural network processing device comprising: a first plurality of synapses configured to receive a first plurality of inputs and generate a first plurality of outputs therefrom, wherein the first plurality of synapses comprises: A plurality of memory cells, wherein each of the memory cells comprises: a spaced-apart source region and a drain region formed in a semiconductor substrate, wherein a channel region extends between the source region and the drain region; a floating gate disposed over and insulated from a first portion of the channel region; and a first gate disposed over and insulated from a second portion of the channel region. and a second gate disposed above the floating gate and insulated from the floating gate or disposed above the source region and insulated from the source region; Each of the plurality of memory cells is configured to store a weight value corresponding to a plurality of electrons on the floating gate; the plurality of memory cells being configured to generate the first plurality of outputs based on the first plurality of inputs and the stored weight values; wherein the memory cells of the first plurality of synapses are arranged in rows and columns, and wherein the first plurality of synapses comprises: a plurality of first lines, each first line electrically connecting the first gates in one of the rows of the memory cells together; a plurality of second lines, each second line electrically connecting the second gates in one of the columns of the memory cells together; a plurality of third lines, each third line electrically connecting the source regions in one of the rows of the memory cells together; a plurality of fourth lines, each fourth line electrically connecting the drain regions in one of the columns of the memory cells together; The first plurality of synapses are configured to receive the first plurality of inputs as voltages on the second plurality of lines or the fourth plurality of lines, and to provide the first plurality of outputs as currents on the third plurality of lines.

26. The neural network processing device of claim 25, wherein the first plurality of synapses are configured to receive the first plurality of inputs as voltages on the second plurality of lines.

27. The neural network processing device of claim 25, wherein the first plurality of synapses are configured to receive the first plurality of inputs as voltages on the fourth plurality of lines.

28. The neural network processing device according to claim 25, further comprising: A first plurality of neurons is configured to receive the first plurality of outputs.

29. The neural network processing device according to claim 28, further comprising: a second plurality of synapses configured to receive a second plurality of inputs from the first plurality of neurons and generate a second plurality of outputs therefrom, wherein the second plurality of synapses comprises: a plurality of second memory cells, wherein each of the second memory cells comprises: a second source region and a second drain region spaced apart from each other formed in the semiconductor substrate, wherein a second channel region extends between the second source region and the second drain region; a second floating gate disposed over a first portion of the second channel region and insulated from the first portion; a third gate disposed over a second portion of the second channel region and insulated from the second portion; and a fourth gate disposed over the second floating gate and insulated from the second floating gate or disposed over the second source region and insulated from the second source region; Each of the plurality of second memory cells is configured to store a second weight value corresponding to a plurality of electrons on the second floating gate; the plurality of second memory cells being configured to generate the second plurality of outputs based on the second plurality of inputs and the stored second weight values; wherein the second memory cells of the second plurality of synapses are arranged in rows and columns, and wherein the second plurality of synapses comprises: a plurality of fifth lines, each fifth line electrically connecting the third gates in one of the rows of the second memory cells together; a plurality of sixth lines, each sixth line electrically connecting the fourth gates in one of the columns of the second memory cells together; a plurality of seventh lines, each seventh line electrically connecting the second source regions in one of the rows of the second memory cells together; a plurality of eighth lines, each eighth line electrically connecting the second drain regions in one of the columns of the second memory cells together; Wherein the second plurality of synapses are configured to receive the second plurality of inputs as voltages on the plurality of sixth lines or the plurality of eighth lines and to provide the second plurality of outputs as currents on the plurality of seventh lines.

30. The neural network processing device of claim 29, wherein the second plurality of synapses are configured to receive the second plurality of inputs as voltages on the plurality of sixth lines.

31. The neural network processing device of claim 29, wherein the second plurality of synapses are configured to receive the second plurality of inputs as voltages on the plurality of eighth lines.

32. A neural network processing device comprising: a first plurality of synapses configured to receive a first plurality of inputs and generate a first plurality of outputs therefrom, wherein the first plurality of synapses comprises: A plurality of memory cells, wherein each of the memory cells comprises: a spaced-apart source region and a drain region formed in a semiconductor substrate, wherein a channel region extends between the source region and the drain region; a floating gate disposed over and insulated from a first portion of the channel region; and a first gate disposed over and insulated from a second portion of the channel region. and a second gate disposed above the floating gate and insulated from the floating gate or disposed above the source region and insulated from the source region; Each of the plurality of memory cells is configured to store a weight value corresponding to a plurality of electrons on the floating gate; the plurality of memory cells being configured to generate the first plurality of outputs based on the first plurality of inputs and the stored weight values; wherein the memory cells of the first plurality of synapses are arranged in rows and columns, and wherein the first plurality of synapses comprises: a plurality of first lines, each first line electrically connecting together the first gates in one of the columns of the memory cells; a plurality of second lines, each second line electrically connecting the second gates in one of the rows of the memory cells together; a plurality of third lines, each third line electrically connecting the source regions in one of the rows of the memory cells together; a plurality of fourth lines, each fourth line electrically connecting the drain regions in one of the columns of the memory cells together; Wherein the first plurality of synapses are configured to receive the first plurality of inputs as voltages on the fourth plurality of lines and to provide the first plurality of outputs as currents on the third plurality of lines.

33. The neural network processing device according to claim 32, further comprising: A first plurality of neurons is configured to receive the first plurality of outputs.

34. The neural network processing device according to claim 33, further comprising: a second plurality of synapses configured to receive a second plurality of inputs from the first plurality of neurons and generate a second plurality of outputs therefrom, wherein the second plurality of synapses comprises: a plurality of second memory cells, wherein each of the second memory cells comprises: a second source region and a second drain region spaced apart from each other formed in the semiconductor substrate, wherein a second channel region extends between the second source region and the second drain region; a second floating gate disposed over a first portion of the second channel region and insulated from the first portion; a third gate disposed over a second portion of the second channel region and insulated from the second portion; and a fourth gate disposed over the second floating gate and insulated from the second floating gate or disposed over the second source region and insulated from the second source region; Each of the plurality of second memory cells is configured to store a second weight value corresponding to a plurality of electrons on the second floating gate; the plurality of second memory cells being configured to generate the second plurality of outputs based on the second plurality of inputs and the stored second weight values; wherein the second memory cells of the second plurality of synapses are arranged in rows and columns, and wherein the second plurality of synapses comprises: a plurality of fifth lines, each fifth line electrically connecting the third gates in one of the columns of the second memory cells together; a plurality of sixth lines, each sixth line electrically connecting the fourth gates in one of the rows of the second memory cells together; a plurality of seventh lines, each seventh line electrically connecting the second source regions in one of the rows of the second memory cells together; a plurality of eighth lines, each eighth line electrically connecting the second drain regions in one of the columns of the second memory cells together; Wherein the second plurality of synapses are configured to receive the second plurality of inputs as voltages on the plurality of eighth lines and to provide the second plurality of outputs as currents on the plurality of seventh lines.

Citation Information

Patent Citations

  • Single transistor non-valatile electrically alterable semiconductor memory device

    US5029130A

  • Flash memory cells with separated self-aligned select and erase gates, and process of fabrication

    US6747310B2

  • Deep Learning Neural Network Classifier Using Non-volatile Memory Array

    US20170337466A1