Hyperdimensional stochastic compute-in-memory with charge-mode dram array
The hyperdimensional stochastic compute-in-memory system addresses memory bandwidth bottlenecks in AI applications by embedding logic within a charge-mode DRAM array, achieving efficient and energy-effective data processing.
Patent Information
- Application Number
- PCT/US2024/059041
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-08
- Filing Date
- 2024-12-06
- Publication Date
- 2025-06-12
AI Technical Summary
Current computing systems for large language models and AI applications face memory bandwidth bottlenecks due to separate compute logic and memory chips, leading to inefficient data processing.
The implementation of a hyperdimensional stochastic compute-in-memory system using a charge-mode DRAM array, which embeds logic functions within the memory, enabling efficient matrix-vector multiplication and reducing the need for separate compute and memory units.
This approach enhances energy efficiency and throughput by tightly coupling memory and computation, reducing refresh rates, and minimizing leakage currents, while supporting both deterministic and stochastic computing modes.
Smart Images

Figure US2024059041_12062025_PF_FP_ABST
Abstract
Description
HYPERDIMENSIONAL STOCHASTIC COMPUTE-IN-MEMORY WITH CHARGE-MODE DRAM ARRAYCROSS-REFERENCE TO RELATED APPLICATION
[0001] This application claims the benefit of priority under 35 U.S.C. § 119(e) to earlier-filed U.S. Provisional Patent Application No. 63 / 608,055, filed on December 8, 2023, the contents of which are herein incorporated by reference.STATEMENT OF GOVERNMENT SPONSORED SUPPORT
[0002] This invention was made with government support under N00014-20-1-2405 awarded by the Office of Naval Research. The government has certain rights in the invention.SUMMARY
[0002] In some example embodiments, there may be provided a method of performing a compute-in-memory operation on a eDRAM cell comprising a plurality of transistors, the plurality of transistors comprising at least: an input transistor, a charge pump transistor, a write select transistor, a share select transistor, and an output transistor, the plurality of transistors having a respective channel located underneath, the method comprising: performing a program operation on the plurality of transistors, the program operation comprising: coupling the input transistor and the output transistor to an analog supply, and sequentially pulsing the charge pump transistor, the write select transistor, and the share select transistor between ground and the analog supply while the channel underneath the output transistor is programmed between ground and a digital supply; and performing a compute operation on the plurality of transistors, the compute operation comprising: coupling the channel underneath the input transistor to the channel underneath the output transistor, and pulsing the input transistor from ground to the analog supply to transfer charge in the channel underneath the output transistor to the channel underneath the input transistor.
[0003] In some variations, the analog supply comprises a voltage of between about 1.5V and about 2. IV In further variations, the digital supply comprises a voltage between about 0.6V and about 0.8V. In certain variations, sequentially pulsing the charge pump transistor, the write selecttransistor, and the share select transistor comprises pulsing the charge pump transistor, the write select transistor, and the share select transistor over a plurality of cycles.
[0004] In some variations, during a first cycle of the plurality of cycles, the pulsing writes an equal amount of charge to the channel underneath the input transistor and to the channel underneath the output transistor. In further variations, a second cycle of the plurality of cycles comprises a first phase and a second phase, and wherein: during the first phase of the second cycle of the plurality of cycles, the charge pump transistor and the write select transistor are pulsed and, during the second phase of the cycle of the plurality of cycles, the share select transistor is pulsed.
[0005] In some variations, pulsing the input transistor comprises accomplishing a linear matrix-vector multiplication operation by inducing a transfer of charge onto the output transistor proportional to the charge in the channel underneath the input transistor. In certain variations, the method further comprises performing a non-destructive readout operation to determine a result of the linear matrix-vector multiplication. In further variations, the plurality of transistors are disposed within a memory cell, and a memory array comprises a plurality of memory cells arranged in a plurality of rows and a plurality of columns. In some variations, the memory array further comprises: a row driver to drive the plurality of rows of memory cells; and a column driver to drive the plurality of columns of memory cells.
[0006] In certain variations, the memory array is configured to perform fully row-parallel and column-parallel matrix-vector multiplication operations. In some variations, the row driver comprises a pseudo-random number generator. In further variations, the row driver comprises a binary -to-stochastic converter configured to convert a binary input into a stochastic representation.
[0007] In an embodiment, there is provided a system comprising a monolithic 3D silicon arrangement, the monolithic 3D silicon arrangement comprising: a plurality of eDRAM cells, each of the plurality of eDRAM cells comprising: a plurality of transistors, the plurality of transistors comprising at least: an input transistor, a write select transistor, and an output transistor, each of the plurality of transistors having a respective channel located underneath, wherein the plurality of eDRAM cells are disposed in a first direction, a plurality of input lines and a plurality of bitlines couple the plurality of eDRAM cells in a second direction, and a plurality of output lines and a plurality of write select lines couple the plurality of eDRAM cells in a third direction; a firstcontroller configured to drive the plurality of input lines and the plurality of bitlines of the monolithic 3D silicon arrangement; and a second controller configured to drive the plurality of output lines and the plurality of write select lines of the monolithic 3D silicon arrangement.
[0008] In some variations, the monolithic 3D silicon arrangement is configured to perform a program operation and a compute operation on the plurality of transistors. In certain variations, the program operation comprises: coupling the input transistor and the output transistor to an analog supply, and pulsing the write select transistor while the channel underneath the output transistor is programmed between a multi-level voltage between ground and a digital supply.
[0009] In further variations, the compute operation comprises: coupling the channel underneath the input transistor to the channel underneath the output transistor, and pulsing the input transistor from ground to the analog supply to transfer charge in the channel underneath the output transistor to the channel underneath the input transistor. In some variations, wherein the compute operation accomplishes a linear matrix-vector multiplication operation by inducing a transfer of charge onto the output transistor proportional to the charge in the channel underneath the input transistor. In certain variations, the first direction, second direction, and third direction are orthogonal. In some variations, the first controller comprises a binary-to-stochastic converter configured to convert a binary input into a stochastic representation.
[0010] The details of one or more variations of the subject matter described herein are set forth in the accompanying drawings and the description below. Other features and advantages of the subject matter described herein will be apparent from the description and drawings, and from the claims.BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The accompanying drawings, which are incorporated in and constitute a part of this specification, show certain aspects of the subject matter disclosed herein and, together with the description, help explain some of the principles associated with the disclosed implementations. In the drawings,
[0012] FIG. 1A shows an illustrative embodiment of a eDRAM cell structure in accordance with the systems and methods described herein;
[0013] FIG. IB shows an illustrative embodiment of a program operation of a eDRAM cell in accordance with the systems and methods described herein;
[0014] FIG. 1C shows an illustrative embodiment of a switched-capacitor structure equivalent to the programmed eDRAM cell in accordance with the systems and methods described herein;
[0015] FIG. ID shows the linear multiplication that is enabled by the eDRAM cell in accordance with the systems and methods described herein;
[0016] FIG. IE shows a graph of leakage of a eDRAM cell versus a negative reverse body bias voltage in accordance with the systems and methods described herein;
[0017] FIGs. 2A-2C shows a system architecture for performing the CIM operations in a eDRAM array accordance with the systems and methods described herein;
[0018] FIGs. 3A-3F illustrate embodiments of peripheral circuits in accordance with the systems and methods described herein;
[0019] FIG. 3G illustrates a timing diagram showing the pre-charge, COMPUTE, and readout operations performed on eDRAM cells in accordance with the systems and methods described herein;
[0020] FIG. 4 illustrates a table showing deterministic to stochastic conversion;
[0021] FIGs. 5A-5C illustrate relationships between linearity and input and weight vectors in deterministic and stochastic computing modes in accordance with the systems and methods described herein;
[0022] FIGs. 6A-6D illustrate relationships between eDRAM weight coefficient and programmed weight, measured at four instances in time;
[0023] FIG. 7A illustrates an energy recovery logic (ERL) circuit for CV2energy recovery savings achieved by the systems and methods in accordance with embodiments described herein;
[0024] FIG. 7B illustrates the energy recovery measured in accordance with embodiments described herein; and
[0025] FIGs. 8A and 8B illustrate monolithic 3D implementations of eDRAM arrays in accordance with embodiments described herein.DETAILED DESCRIPTION
[0026] In computing systems used to process data for large language model (LLM) and other artificial intelligence (Al) applications, significant compute power and memory capacity are required. However, compute logic and memory are traditionally realized with separate chips. The communication between such separate chips may be input / output- (I / O) limited. I / O-limited communications introduces memory bandwidth bottlenecks in such systems that maintain compute logic and memory on separate chips. Such systems include many current Al accelerators. Compute-in-memory (CIM) architectures address this problem by embedding logic functions into the memory.
[0027] A 768 by (x) 768 crossbar array of double-differential, 4 x 5T charge-mode DRAM compute-in-memory cells performs signed analog-digital matrix-vector multiplication with embedded DAC for 4b multibit programmable analog weights, and stochastic DAC binary encoding of lb-8b digital inputs and digital lb-8b encoding of stochastic ADC binary outputs for Al on the edge and cloud.
[0028] The charge-mode computational DRAM (eDRAM) CIM macro may present one or more of the following key advantages over other CIM crossbars for Al on the edge / Tiny ML: i) highly linear multiply-and-accumulate operations with signed weight and input encoding as well as doubled dynamic range; ii) non-destructive read operations with low leakage current leveraging FDSOI body isolation and thick-oxide transistors for extremely small gate and subthreshold leakage allowing lower refresh rates; iii) highly energy-efficient charge-mode sensing from the cell, quantized by a high input-impedance dynamic comparator, thus no static current and IR drop along the array; iv) low device-to-device variability in a standard 22nm FDSOI CMOS process scaling to more advanced technology nodes; v) high-density, high-resolution digitally encoded analog weights in the cell performing area-efficient multiplication; vi) supporting bothdeterministic and stochastic input encoding and output decoding; and vii) dynamically reconfiguring the array for reuse in model-level pipelining.
[0029] By tightly coupling memory and computation in massively parallel crossbar array architecture, compute-in-memory (CIM) promises energy -efficient and high-throughput alternatives to conventional von-Neumann topologies. Typical CIM processors perform multiply- accumulate (MAC) operations in the analog domain through accumulating output currents collected from bit cells subjected to digitally supplied input voltages and digitizing the outputs with a parallel bank of current-sensing or voltage-sensing ADCs. Fundamentally, charge-mode, rather than conductive CIM, elements offer potentially greater overall efficiency by conserving charge in storage and energy in computation. Charge-mode elements provide these benefits by allowing for non-destructive reversible charge transfer triggered by an applied input voltage and by registering a change in output voltage through sensing the charge capacitively coupled to the output line.
[0030] Other determining factors of scaling analog CIM to deployable systems include on- chip model capacity and area efficiency, i.e., throughput per unit area. CIM architectures involving SRAMs store multi-bit weights across separate bit cells that couple through capacitors or other gain summing elements to output lines for voltage / current readout, that are binary weighted and combined to obtain a composite MAC sum. SRAMs offer superior memory retention but are subject to static power due to subthreshold leakage, gate leakage, and junction leakage. Alternative embedded DRAM (1T1C) cell for CIM cores offers greater bit-cell density than SRAM and offers direct charge-mode readout, although readout operations are destructive requiring frequent refresh and incurring peripheral circuit design complexities even greater than SRAM. Multi-bit weights are preferred in fixed deployment scenarios for edge Al and tiny ML which has led designers in this application space to explore traditional non-volatile memories such as Flash and emerging nonvolatile memory cells such as RRAM. The relatively high write energy and low operational speed of Flash macros limit their practical use in demanding settings. Likewise, sneak paths, IR drop, and cycle-to-cycle and device-to-device variability of RRAM have impeded progress towards practical solutions. Production RRAMs have very low resistance and low RON / ROFF ratio which limits the operation of the array to activate only a small number of rows at a time, decreasing the degree of parallelism and the overall throughput of the system. RRAM drivers aretypically over-dimensioned to support large static currents, which increases the macro area and lowers the on-chip memory density. Both RRAMs and Flash have proven to be hard in scaling to the latest technology nodes because of the challenges in floating gate oxide thickness and memristive material integration.
[0031] The charge-mode computation DRAM (eDRAM) CIM macro for row-parallel and column-parallel signed analog-digital matrix-vector multiplication (MVM) presented here combines the non-destructive charge-conserving readout and adiabatic hot-clock energy recovery of a eDRAM CIM crossbar array, with cell-level embedded functionality for direct DAC multi -bit programming of lb-4b analog weights, and array-peripheral functionality for stochastic DAC binary encoding of lb-8b digital inputs and digital lb-8b encoding of stochastic ADC binary outputs, as a versatile macro for hyperdimensional probabilistic computing and Bayesian Al on the edge.
[0032] The eDRAM arrays and eDRAM systems presented herein can be dynamically reconfigured for reuse in in-memory / near-memory computations through fast reprogramming (similar to commercial DRAM). This ability of the systems described herein to be dynamically reconfigured contributes to device longevity.
[0033] FIG. 1A illustrates the principle of operation of the macro at the level of a single quadrant of the double-differential eDRAM cell 100, serving weight storage with in-built DAC and MAC functions with a plurality of nMOS transistors 102a, 102b, 102c, 102d, and 102e. While the example of FIG. 1A shows 5 nMOS transistors 102a-102e, any number of transistors could be used. In the example of FIG. 1A, the 5 nMOS transistors comprise a charge pump transistor 102a, a write select transistor 102b, an output transistor 102c, a share select transistor 102d, and an input transistor 102e. Each of the plurality of transistors 102a-102e has a respective channel located underneath. For example, the output transistor 102c has beneath it an output channel, and the input transistor 102 has beneath it an input channel. The channels located underneath each of the plurality transistors may function as a charge well as charge is transferred between or otherwise shared between different channels.
[0034] In some implementations, the output transistor 102c and the input transistor 102e may have respective lengths that are larger than lengths of the charge pump transistor 102a, the writeselect transistor 102b, and the share select transistor 102d. In certain implementations, the lengths of the output transistor 102c and the input transistor 102e are between about two and about six times larger than the lengths of the charge pump transistor 102a, the write select transistor 102b, and the share select transistor 102d. In some implementations, the lengths of the output transistor 102c and the input transistor 102e are between about three and about five times larger than the lengths of the charge pump transistor 102a, the write select transistor 102b, and the share select transistor 102d.
[0035] Row drivers supply the bit-line (BL) and input (VIN) horizontal lines 104, while column drivers provide charge-pump (CP), write (WR), and share (SH) vertical lines 106. The CP vertical line is used to add additional charge into the shared junction between the charge pump transistor 102a and the write select transistor 102b without fully overwriting the charge well underneath the output transistor 102c and the input transistor 102e. This additional charge can be used to slowly refresh the charge well underneath the output transistor 102c and the input transistor 102e using a subthreshold current mechanism. However, during PROGRAM and COMPUTE operations, the CP and WR vertical lines are driven to the same voltage.
[0036] The eDRAM cell is configured to be subject to a weight programming operation, also called a program operation or a programming operation. The program operation of the eDRAM cell 100 is illustrated in FIG. IB. As shown in FIG. IB, during weight programming, the input (VIN) and output (VOUT) vertical lines are tied to the analog supply voltage VoDa. As also shown in FIG. IB, the bitline voltage (BL) is varied between ground and the digital supply voltage (VDDD) modulated by the programming input. The analog supply voltage VoDa and the digital supply voltage VDDD may be configured to allow above subthreshold operations for linear programming of the eDRAM cell 100. For example, the analog supply voltage VoDa may be between about 1.5 V and about 2.1V In certain implementations, the analog supply voltage may be 1.8V. The digital supply voltage VDDD may be between about 0.6V and about 0.8V.
[0037] As also shown in FIG. IB, the CP, WR, and SH vertical lines are pulsed as part of the program operation. The CP, WR, and SH may be pulsed between ground (GND) and VoDa in synchrony with the sequence of bits Di presented from least significant bit (LSB) to most significant bit (MSB) on the bitline (BL) according to the waveforms shown. This pulsing of theCP, WR, and SH vertical lines configures the eDRAM cell as an algorithmic charge-division digital-to-analog converter (DAC) equivalent to the familiar switched-capacitor structure shown in FIG. 1C.
[0038] As shown in FIG. IB, the CP, WR, and SH vertical lines may be pulsed over a plurality of cycles 130a, 130b, .. . 130n. During a first cycle 130a of the plurality of cycles, the write select transistor 102b and share select transistor 102d are pulsed together. During the LSB Do, active high CP / WR and SH short the channel potentials underneath both VIN and VOUT to the BL potential, writing equal amounts of charge in both of the channel potentials that are either zero for Do = 0, or a nominal half-well charge for Do = 1. In some implementations, this writing of charge to both of the channel potentials underneath VI and VOUT occurs on a first pulsing cycle of the plurality of pulsing cycles. This shorting of the channel potentials underneath VIN and VOUT serves to reset the memory cell. After a predetermined amount of time, the write select transistor 102b and share select transistor 102d are driven to a lower voltage. The write select transistor 102b may be turned off prior to the share select transistor 102d being turned off to ensure a clean reset operation. Further, by turning off the write select transistor 102b prior to turning off the share select transistor 102d, the impact of charge injection on the charge stored in the channel underneath the input transistor 102e and the output transistor 102c caused by turning off the write select transistor 102b and the share select transistor 102d can be minimized.
[0039] In a second cycle 130b of the plurality of cycles, the charge pump transistor 102a and the write select transistor 102b are pulsed to a first voltage during a first phase of the second cycle 130b. The pulsing of the charge pump transistor 102a and the write select transistor 102b to the first voltage writes into the channel beneath the output transistor 102c the charge that was transferred to the channel underneath VOUT during the first cycle. After the charge pump transistor 102a and the write select transistor 102b are driven to the first voltage, they are driven to a second voltage, the second voltage being lower than the first voltage to disconnect the BL from charge well underneath VOUT. Subsequently, during a second phase of the second cycle, the share select transistor 102d is pulsed so that the charge stored in the channel underneath VIN from the first cycle and the charged stored in the channel underneath VOUT from the first phase of the second cycle are subject to a divide-by-two operation.
[0040] The pulsing patern of the second pulsing cycle of the plurality of cycles repeats as many times as needed for the input bits to be written to the eDRAM cell. Alternating activation of CP / WR and SH during subsequent DAC cycles with bit Di presented on BL as shown during cycles 130b and 130n of FIG. IB, writes the corresponding zero or nominal charge in the channel underneath VOUT followed by sharing the net charge equally across both channels resulting in consecutive addition and divide-by-two operations, establishing algorithmic n-bit DAC of a multibit input to multi-level charge residing in both channels after n cycles.
[0041] During a COMPUTE operation, CP / WR is deactivated low, and SH remains activated high to continuously couple the two channels underneath VIN and VOUT, sharing their charge in a common well as in a charge injection device. The distribution of the shared charge between the two channels underneath VIN and VOUT depends on the input voltage VIN. With VOUT left floating near mid-level potential, all the weight charge stored in the well resides underneath VOUT. If the input D is low during the computational cycle, VIN is maintained at GND potential, and no charge transfer takes place from the channel underneath VOUT to the channel underneath VIN. In contrast, if D is high, VIN is pulsed from GND to VoDa which causes the well charge to transfer from the channel underneath VOUT to the channel underneath VIN.
[0042] As illustrated in FIG. ID, this transfer induces an equal transfer of charge onto the VOUT line due to capacitive coupling across the gate oxide of the output transistor, which causes a rise in voltage VOUT proportional to the charge stored in the well, hence accomplishing linear matrixvector multiplication across the array of eDRAM cells. The eDRAM cell 100 of FIG. 1A allows for high-density, high-resolution digitally encoded analog weights in a single dynamic memory cell, allowing the eDRAM cell 100 to perform area-efficient matrix-vector multiplication.
[0043] FIG. ID illustrates the linear multiplication that is enabled by the eDRAM systems described herein. When the transistor (the input vertical line) VIN is subject to a voltage, an output signal on the output line VOUT is proportional to the charge stored in the eDRAM cell. This operation can be done adiabatically or non-adiabatically. FIG. ID illustrates adiabatic and non- adiabatic modes of operation of the eDRAM systems described herein. In the non-adiabatic case, the applied inputs are digital pulses (e.g., 0 or 1). In the adiabatic case, a sinusoidal input is applied such that, during the COMPUTE operation, CV2energy consumption is recovered.
[0044] Performing this COMPUTE operation provides linearity of the programmed charge and greater throughput and energy efficiency relative to other CIM technologies. In particular, refresh rates of the eDRAM systems described herein are reduced relative to conventional CIM systems because low body leakage currents are obtained via fully depleted silicon-on-insulator body isolation. Further, thick gate-oxide transistors provide extremely small gate and subthreshold leakage.
[0045] FIG. IE shows a graph illustrating the leakage of a eDRAM cell, such as the eDRAM cell 100 of FIG. 1A, versus a negative reverse body bias voltage (VRBB). The eDRAM cell of FIG. IE may be manufactured according to a fully depleted silicon-on-insulator process and configured to operate under zero reverse body bias voltage.
[0046] As shown by FIG. IE, there is only a contribution to leakage from gate-induced drain lowering (GIDL) mechanisms absent gate leakage and suppressed subthreshold leakage by carefully controlled transistor threshold voltages, write select transistor bias and bit line programming biases. The leakage due to the GIDL mechanisms is exponential in nature with respect to the difference in the programmed BL bias voltage and the BL bias at which the cell is maintained during the COMPUTE operation. The GIDL leakage can be further suppressed or linearized because fully depleted silicon-on-insulator processes have two gate controls. In addition to the gate of the transistor, the transistor body beneath the channel and the buried oxide layer of the transistor can act as an additional gate to control the GIDL leakage. The GIDL leakage can be further linearized by a factor of lOx by driving the body bias to -3V. The back bias can be selected between - IV to -5V to lower the GIDL leakage and provide superior retention.
[0047] FIG. 2A illustrates an embodiment of the architecture of a eDRAM system 200 with double-differential encoding in accordance with the systems and methods described herein. The eDRAM system 200 of FIG. 2A comprises a eDRAM crossbar array 202, a row controller 204, and a column controller 206. In some implementations, the row controller 204 and the column controller 206 may be the same controller.
[0048] The eDRAM crossbar array 202 may comprise a plurality of eDRAM cells 208 (i.e., a first eDRAM cell 208a, a second eDRAM cell 208b, a third eDRAM cell 208c, fourth eDRAM cell 208d, and so on). The eDRAM cells 208 may comprise a eDRAM cell having the structuredescribed in FIG. 1A. In some implementations, the eDRAM crossbar array 202 may be, for example, a 768 x 768 double-differential eDRAM crossbar array comprising an array with 768 rows of double-differential eDRAM cells and 768 columns of double-differential eDRAM cells. The eDRAM crossbar array 202 may comprise a symmetrical structure. The symmetrical structure of the eDRAM crossbar array 202 may facilitate linear MVM owing to better common-mode rejection and offset cancellation due to capacitive coupling.
[0049] The row controller 204 may comprise a plurality of modules configured to control various functions of the rows of the eDRAM crossbar array 202. For example, the row controller 204 may comprise a reconfigurable bit precision logic (RBPL) module 204a, a pseudo-random number generator (PRBS) module 204b, and a binary-to-stochastic converter (BTSC) module 204c. Similarly, the column controller 206 may comprise a reconfigurable bit precision logic (RBPL) module 206a and a binary-to-stochastic converter (BTSC) module 206b, both of which may control the various functions of the columns of the eDRAM crossbar array 202.
[0050] The eDRAM crossbar array 202 of the eDRAM system 200 may be configured to perform CIM by first being subject to a pre-charge operation to perform autozeroing. During such a pre-charge operation, the output transistor is driven to a pre-charge voltage, VPRE, while the input transistor is kept at ground.
[0051] After the pre-charge operation, the eDRAM crossbar array 202 is subject to a COMPUTE operation. Driving the output well to VPRE while maintaining the input well at Vss establishes a known output voltage for an input vector that is used as a reference during the COMPUTE operation. During the COMPUTE operation, the voltage applied to the input line is changed from Vss to VIN. In the case of a double-differential configuration, the pre-charging operation could be configured to double the output dynamic range during the COMPUTE operation. For example, if the output is pre-charged with an input of -1 (VIN+ is driven to Vss while VIN- is driven to VDDA), and the COMPUTE operation is done with an input of +1 (VIN+ is driven to VDDA while VIN- is driven to Vss), the sensed output is twice the output compared to a single bitcell operation.
[0052] The charge transfer that occurs during the COMPUTE operations to be carried out on the eDRAM system 200 described herein may be described mathematically as follows. It is firstassumed that a total charge QPROG is stored in the eDRAM well during the PROGRAM operation. Under this assumption, during the first step of a COMPUTE operation, the input transistor VIN may take a charge aQPR0 Gunder its well. After channel inversion, the channel voltage VCII,COMP is defined by
[0053] When the input transistor Vin has the charge aQPR 0Gunder its well, the charge under the output transistor well is given by (1 — a) QPR0 Gfollowing the principle of conservation of charge. It follows thatwhere equation (3) follows from plugging equation (1) into equation (2). Here, Cgsrepresents the gate-to-source capacitance of the transistor arrangement of the eDRAM cell. The capacitive load on the Vout line (COUT), shared by other cells in the array, induces a charge QPROG to compensate for the charge transfer:
[0054] VOu t=aQ pR 0G(4). VOUT is typically pre-charged to a known bias VPRE before the cOutCOMPUTE operation. Thus, after the COMPUTE operation, the absolute value of VOUT is given by
[0055] VPRE represents a known pre-charge bias that is selected for the eDRAM cell. In some implementations, VPRE equals VoDa / 2 to allow the voltage swing to go up to VoDa after the charge transfer takes place. In certain implementations, VPRE is between Vss and VDD3. Equations (l)-(5) can be manipulated to derive the equation for a, namely:This induces a multiplicative gain on the output line (proportional to VIN and QPROG):The parameter a depends on the applied input voltage VIN, and thus describes how much charge is transferred between the channel underneath the input transistor 102e and the channel underneath the output transistor 102c during the COMPUTE operation. In some implementations, VIN needs to be higher than a threshold voltage above the channel voltage VCH to allow an above-threshold linear charge transfer operation. In other words, in order for the above-threshold linear charge transfer operation to occur, VIN must satisfyT / IV =5 ^CH + ^THRESH -
[0056] FIG. 2C illustrates the double differential structure with differential inputs and differential outputs. As shown in FIG. 2C, double-differential encoding of the inputs, outputs, and weights across the eDRAM crossbar array 202 supports four-quadrant signed analog-digital MVM, for a four-fold increase in cell size occupying four identical complementary quadrants. In return, the symmetrical 4x5T structure offers greater common-mode noise rejection, as well as substantially improved linearity due to charge balancing with uniform capacitive loading along rows and columns across the eDRAM crossbar array 202. This improved linearity is also critically important in guaranteeing that resonance between the capacitive load of the eDRAM array and an external inductive tank for hot-clock adiabatic energy-recycling be maintained at constant hot- clock frequency, independent of the state of the inputs and outputs of the eDRAM crossbar array 202. This obviates elaborate stochastic modulation schemes to ensure balanced input statistics that have limited the efficiency gains achievable by adiabatic energy recovery in charge-domain CIM due to residual central-limit statistical variations in resonance frequency.
[0057] On the array level, the total MAC sum across each output column of the eDRAM crossbar array 202 is a normalized sum of the outputs of different multiply operations that are performed by every double-differential eDRAM cell of N rows in the eDRAM crossbar array 202 (weight stored in the double-differential cell represented by a differential conductance Wk = Gk+ — Gk-), which is given by,Vout+ - Vout-where the index k represents a given row of the eDRAM crossbar array 202 that contributes to the normalized sum of the output for a given column.
[0058] Via the STBC modules 204c and 206b, the eDRAM CIM system 200 provides for stochastic encoding of the inputs and outputs to offer greater versatility in input and output digital formats accommodating both probabilistic and deterministic activations of neural state variables.
[0059] FIG. 3A- illustrates peripheral circuit architecture in accordance with embodiments herein. Referring to FIG. 3A, the digital peripheral circuits 300 includes reconfigurable input and output bit width for computing at multiple precisions ranging up to 8-bits with support for multiplexing between both deterministic and stochastic encoding / decoding of inputs and outputs. As shown in FIG. 3A, the peripheral circuit 300 comprises a comparator 302, a multiplexer (MUXer) 304, an input chopper 306, energy recovery logic drivers 308 (discussed further below with respect to FIGs. 7A and 7B), a compute module 310, an output chopper 312, a double-tall latch-type dynamic comparator 314, a demultiplexer 316, and an output decoder 318.
[0060] The comparator 302 is a digital comparator that is configured to compare the multi-bit input and generated random number. The multiplexer 304 is configured to decide between the deterministic or stochastic datapath for the given inputs.
[0061] The input chopper 306 is configured to perform input bit flipping without changing the actual input register value, which is useful for double-dynamic range operation with complementary inputs during pre-charge and compute operations.
[0062] The energy recovery logic drivers 308 are configured to act as a simple level shifter (from digital supply voltage to analog supply voltage) in non-adiabatic mode of operation. However, in the case of adiabatic supply, the energy recovery logic drivers 308 are configured to act as a continuous analog driver to feed the sinusoidal supply voltage.
[0063] The compute module 310 is configured to execute the core charge-mode compute-inmemory multiply-and-accumulate operation described with respect to FIG. ID and 2A.
[0064] The output chopper 312, the demultiplexer 316, and the output decoder 318 are configured to act as a peripheral controller to decode results of MAC operations during stochastic and deterministic computing under different bit-precision compute constraints.
[0065] The dynamic comparator 314 is configured to serve as the building block for singleslope analog-to-digital conversion. As shown by the peripheral circuit 300 of FIG. 3A, computing is possible according to two modes of operation: the peripheral circuit 300 can support both deterministic and stochastic input encoding and output decoding for probabilistic CIM. In some implementations, it is possible to switch between the deterministic mode and the stochastic mode of operation.
[0066] The digital peripheral circuit 300, and integrated pseudo-random number generators (PRNGs) and digital comparators, one each provided per row, generate stochastically encoded binary input streams to the array from supplied digital inputs with unbiased mean and adjustable variance set by PRGN magnitude. Strictly independent and identically distributed (i.i.d) PRGN channels are generated employing no more than two-bit shift registers and one XOR gate per PRGN bit slice. The 1 -bit quantized column outputs resolved by the array of dynamic comparators are decimated accordingly to accumulate statistics producing sigmoidal activation functions of varying spread corresponding to the combined variance of the PRGN additive noise in the central limit. Fig. 3A additionally shows a timing diagram 305 that shows the pulsing that occurs during a COMPUTE operation.
[0067] Referring to FIGs. 3B-3F, the peripheral circuits 320, 330, 340, 350, and 360 can be used to drive the transistors of the eDRAM cells in accordance with embodiments herein. In particular, peripheral circuit 320 of FIG. 3B is configured to drive the bitlines of the eDRAM cells described herein. Peripheral circuit 320 is configured to drive BL+ and BL- to either VBLHI or VBLLO in a complementary fashion, using a user-defined digital-to-analog converter at the periphery. The choice of VBLHI and VBLLO is governed by compute-in-memory performance metrics such as the output dynamic range, average refresh time, and multiplication SNR in multilevel storage.
[0068] Peripheral circuit 330 of FIG. 3C is configured to drive the input transistors of the eDRAM cells described herein. Peripheral circuit 330 comprises level shifter 331, ERL drivers332, and control switches 333. The level shifter 331 and the ERL drivers 332 ensure the generation of analog bias voltage modulated by the digital input bits. Finally, the control switches 333 determine whether to bias the IN at a fixed voltage during the PROGRAM operation or to drive the IN lines dynamically based on the digital input bits.
[0069] Peripheral circuit 340 of FIG. 3D is configured to drive the write select transistors of the eDRAM cells described herein. Peripheral circuit 350 of FIG. 3E is configured to drive the share select transistors of the eDRAM cells described herein. During a PROGRAM operation, the write select transistors, such as write select transistor 102b of FIG. 1A and share select transistors, such as share select transistor 102d of FIG. 1A, are driven to Vnna / Hot Clock (HC) or ground (VSS) based on external control bits WR SEL and SH SEL, to control the switched-capacitor DAC operation of eDRAM. However, during a COMPUTE operation, the write select transistors, such as write select transistor 102b of FIG. 1A and share select transistors, such as share select transistor 102d of FIG. 1A, are driven to VWRLO and VSHHI respectively, set from another digital-to-analog converter at the periphery. The choice of VWRLO is governed by the desired leakage of eDRAM, while VSHHI is chosen to maximize the dynamic range of computation, as the share select transistor 102d governs the coupling of the channels beneath the input transistor and the output transistor. Peripheral circuit 360 of FIG. 3F is configured to drive the output transistors of the eDRAM cells described herein. During a pre-charge operation, they are driven to VPRE, set from another digital-to-analog converters at the periphery. Differential output lines maybe pre-charged to separate biases VpRE+_and VPRE- to set the threshold voltage for single-slope analog-to-digital conversion. However, during a COMPUTE operation, they are left floating to allow it to set the multiplication gain of the bitcell. The eDRAM cell having transistors driven by peripheral circuits 320, 330, 340, 350, and 360 may be, for example, eDRAM cell 100 of FIG. 1 A or any of the eDRAM cells 208a, 208b, 208c, and so on, of FIGs. 2A and 2B.
[0070] FIG. 3G illustrates a timing diagram for the pre-charge, COMPUTE, and readout operations performed using the eDRAM cells, such as eDRAM cells 208a, 208b, 208c, and 208d, of FIG. 2A. In a first segment 370a of the timing diagram of FIG. 3G, the pre-charge operation is performed on the eDRAM cells by pulsing the output transistor is driven to a pre-charge voltage, VPRE, while the gate of the input transistor is kept at ground to implement autozeroing. In a second segment 370b of the timing diagram of FIG. 3G, COMPUTE operations are performed on theeDRAM cell. As many compute operations in a bit-serial fashion may be performed as are required to perform a given logical operation in the eDRAM cell and the results are accumulated in a digital register. The eDRAM cells, such as eDRAM cells 208a, 208b, 208c, and 208d, of FIG. 2A, allow for non-destructive compute / readout operations across multiple cycles, without requiring refresh after every compute cycle. In a subsequent segment 370n of the timing diagram of FIG. 3G, readout of the result of the bit-serial COMPUTE operations is performed from a digital register.
[0071] FIG. 4 illustrates a table 400 showing deterministic to stochastic conversion of a bit in accordance with embodiments described herein. A deterministic binary or bipolar input is received by the system. The input has a resolution as described by column 402 of the table 400. The binary or bipolar input may have a representation according to column 404 of the table 400, which may be analogous to an internal representation of the input according to column 406. The input may be mapped to an 8-bit internal representation. The 8-bit representation of the input may be compared to an 8-bit random number generated by the PRBS module, such as PRBS module 204b of FIG. 2A. In general, the input may be mapped to a representation having as many bits as contained in the output of the PRBS module. The stochastic mapping of the binary or bipolar input is given in column 408 of the table 400. Column 410 of the table 400 displays the expected variance of the stochastic representation of the input. As shown in column 410 of the table 400, the variance of the stochastic representation increased with higher resolution inputs.
[0072] FIGs. 5A-5C illustrate linearity of multiplication achieved by the systems described herein relative to different input and weight vectors. The linearity of multiplication and accumulation (also referred to as multiply-accumulate or MAC) is a measure of the signal-to-noise ratio (SNR) of the eDRAM systems described herein. Noise may come from the eDRAM cell, the periphery of the eDRAM cell, or from device variations. In order to quantize the full-scale dynamic range of the MAC output of eDRAM array with a single dynamic comparator, a singleslope ramp analog-to-digital conversion (ADC) method is leveraged. In this method, the threshold voltage of the ADC is varied by pre-charging VOUT+ and VOUT- to separate biases, VPRE+ and VPRE, respectively. In some implementations, the threshold voltage is given by VThresh= VPRE+— VPRE_. The single-slope ramp ADC method allows for highly efficient charge-mode sensing from a eDRAM cell, quantized by a high input-impedance ADC. This eliminates any static current and current-resistance (IR) drops along the eDRAM array.
[0073] FIG. 5A demonstrates the MAC required to effect the flipping of a bit from 0 to 1 at different threshold voltages VThresh. In particular, FIG. 5 A shows the MAC error characterization at various random weights after offset cancellation using autozeroing, for a plurality of different threshold voltages.
[0074] FIG. 5B shows the ADC switching outputs (as describing by the cumulative spiking probability) versus the expected MAC transfer function of stochastic encoded inputs at different input precisions (the input precisions being shown by the legend within FIG. 5A). FIG. 5B illustrates that higher bit-precision has higher variance when 1-b inputs [-1, +1] are mapped into the center values in multi -bit representation. By mapping 1-b inputs in this manner, the desired noise-shaping on MAC transfer function can be achieved for different bit precision.
[0075] FIG. 5C shows the linearity of weight programming by plotting estimated analog weight value versus the programmed multi-bit digital weight value estimated weight coefficients versus the programmed weights at 4-bit precision.
[0076] FIGs. 5A-5C illustrate that when the entire array, such as eDRAM crossbar array 202 of FIG. 2A, is utilized, the double-differential sensing margin is measured to cover the +175 mV to -175 mV range, substantially larger than the dynamic comparator random threshold variations measured to be within 1 mV, thus allowing for 6b-8b output quantization. Furthermore, linearity and uniformity in the weights also shown in FIGs. 5A-5C are measured to be within 4-bit resolution across the array, minimizing impact of random variations on array-level function in general MVM-based Al tasks.
[0077] FIGs. 6A-6D show the estimated weight coefficients versus the programmed weights at 4-bit precision. FIGs. 6A-6D are measured at four sequential instances in time: FIG. 6A is measured at 0.90 ms, FIG. 6B is measured at 4.05 ms, FIG. 6C is measured at 7.20 ms, and FIG. 6D is measured at 10.12 ms.
[0078] FIG. 7A illustrates an energy recovery logic (ERL) circuit 700 in accordance with embodiments described herein. The ERL circuit 700 of FIG. 7A provides CV2energy recovery savings. In some implementations, the ERL circuit 700 may be configured to drive lines with a hot clock (HC) 702. In some implementations, the hot clock may be a sinusoidal hot clock. Bydriving lines with a sinusoidal HC, rather than a constant voltage supply, full CV2energy recovery savings can be realized across the eDRAM array, such as the eDRAM array 202 of FIGs. 2A-2C. FIG. 7B illustrates the measured energy recovery in graph 706, which plots both the instantaneous and average power measured versus time. Even accounting for constant losses in the periphery of the eDRAM array 202, including CMOS stages preceding the ERL drivers, the overall savings in energy per computational cycle are substantial, boosting the net energy efficiency of the eDRAM array tenfold relative to the non-adiabatic mode of operation achieved with traditional digital pulsing.
[0079] FIGs. 8A shows a first vertical implementation of a monolithic three-dimensional (3D) silicon arrangement 800. The arrangement 800 comprises a plurality of eDRAM cells 801a, 801b, 801c, ... The plurality of eDRAM cells 801a, 801b, 801c, ... may be oriented in a three- dimensional arrangement. The plurality of eDRAM cells 801a, 801b, 801c, ... may comprise a plurality of transistors, including an input transistor, an output transistor, and a write select transistor. In some implementations, the DAC function provided by a charge pump transistor and a share select transistor may be shared at the periphery of the 3D arrangement such that a charge pump transistor and a share select transistor need not be included in the plurality of eDRAM cells 801a, 801b, 801c. In certain implementations, the plurality of eDRAM cells 801a, 801b, 801c may comprise an input transistor, an output transistor, a write select transistor, a charge pump transistor, and a share select transistor, as is the case with the eDRAM cell 100 of FIG. 1A.
[0080] The input lines VIN1, VIN2, VIN3 of the arrangement are disposed a long a first direction, the output lines VOUT1, VOUT2, VOUT3 and write lines WR1, WR2, and WR3, are disposed along a second direction that is perpendicular to the first direction. The first direction, the second direction, and the third direction may be orthogonal. The plurality of eDRAM cells 801a, 801b, 801c are disposed in a third direction perpendicular to the first and second directions. In the example of FIG. 8 A, the first direction is the direction into and out of the page, the second direction runs horizontally along the page, and the third direction runs vertically along the page.
[0081] FIG. 8B shows an additional vertical implementation of a monolithic three-dimensional (3D) silicon arrangement 810 comprising a plurality of eDRAM cells 802a, 802b, 802c, .... The plurality of eDRAM cells 802a, 802b, 802c, ... may be oriented in a three-dimensionalarrangement. The plurality of eDRAM cells 802a, 802b, 802c, ... may comprise a plurality of transistors, including an input transistor, an output transistor, and a write select transistor. In some implementations, the DAC function provided by a charge pump transistor and a share select transistor may be shared at the periphery of the 3D arrangement such that a charge pump transistor and a share select transistor need not be included in the plurality of eDRAM cells 802a, 802b, 802c. In certain implementations, the plurality of eDRAM cells 802a, 802b, 802c may comprise an input transistor, an output transistor, a write select transistor, a charge pump transistor, and a share select transistor, as is the case with the eDRAM cell 100 of FIG. 1A.
[0082] The plurality of eDRAM cells 802a, 802b, 802c, ... may be oriented in a three- dimensional arrangement. The input lines VIN1, VIN2, VIN3 of the arrangement are disposed a long a first direction, the output lines VOUT1, VOUT2, VOUT3 and write lines WR1, WR2, and WR3, are disposed along a second direction that is perpendicular to the first direction. The plurality of eDRAM cells 802a, 802b, 802c are disposed in a third direction perpendicular to the first and second directions. In the example of FIG. 8B, the first direction is runs vertically along the page, the second direction runs into and out of the page, and the third direction runs horizontally along the page.
[0083] The monolithic 3D silicon arrangement 800 and 810 of FIGs. 8 A and 8B are manufactured using 3D vNAND-flash process, 3D vNAND-flash-like process, or 3D DRAM process technologies where the transistors in the eDRAM bit cell are implemented vertically or horizontally, respectively, using multi-tier silicon. The monolithic three-dimensional (3D) silicon arrangements 800 and 810 can be used to perform PROGRAM and COMPUTE operations with high throughput and energy efficiency. The monolithic 3D silicon arrangements 800 and 810 shown in FIGs. 8A and 8B have a manufacturability with monolithic 3D silicon processes that results in extremely high memory capacity, and compute density. Because the monolithic 3D silicon arrangements 800 and 810 of FIGs. 8 A and 8B can be manufactured with standard CMOS processes that are scalable, the CIM operations described herein can be realized with very low device-to-device variability in multi-level storage and compute.
[0084] One or more aspects or features of the subject matter described herein can be realized in digital electronic circuitry, integrated circuitry, specially designed ASICs, field programmablegate arrays (FPGAs) computer hardware, firmware, software, and / or combinations thereof. These various aspects or features can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device. The programmable system or computing system may include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.
[0085] These computer programs, which can also be referred to as programs, software, software applications, applications, components, or code, include machine instructions for a programmable processor, and can be implemented in a high-level procedural and / or object- oriented programming language, and / or in assembly / machine language. As used herein, the term “machine-readable medium” refers to any computer program product, apparatus and / or device, such as for example magnetic discs, optical disks, memory, and Programmable Logic Devices (PLDs), used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor. The machine-readable medium can store such machine instructions non-transitorily, such as for example as would a non-transient solid-state memory or a magnetic hard drive or any equivalent storage medium. The machine-readable medium can alternatively or additionally store such machine instructions in a transient manner, such as for example, as would a processor cache or other random-access memory associated with one or more physical processor cores.
[0086] To provide for interaction with a user, one or more aspects or features of the subject matter described herein can be implemented on a computer having a display device, such as for example a cathode ray tube (CRT) or a liquid crystal display (LCD) or a light emitting diode (LED) monitor for displaying information to the user and a keyboard and a pointing device, such as for example a mouse or a trackball, by which the user may provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well. For example, feedbackprovided to the user can be any form of sensory feedback, such as for example visual feedback, auditory feedback, or tactile feedback; and input from the user may be received in any form, including acoustic, speech, or tactile input. Other possible input devices include touch screens or other touch-sensitive devices such as single or multi-point resistive or capacitive track pads, voice recognition hardware and software, optical scanners, optical pointers, digital image capture devices and associated interpretation software, and the like.
[0087] In the descriptions above and in the claims, phrases such as “at least one of’ or “one or more of’ may occur followed by a conjunctive list of elements or features. The term “and / or” may also occur in a list of two or more elements or features. Unless otherwise implicitly or explicitly contradicted by the context in which it used, such a phrase is intended to mean any of the listed elements or features individually or any of the recited elements or features in combination with any of the other recited elements or features. For example, the phrases “at least one of A and B;” “one or more of A and B;” and “A and / or B” are each intended to mean “A alone, B alone, or A and B together.” A similar interpretation is also intended for lists including three or more items. For example, the phrases “at least one of A, B, and C;” “one or more of A, B, and C;” and “A, B, and / or C” are each intended to mean “A alone, B alone, C alone, A and B together, A and C together, B and C together, or A and B and C together.” Use of the term “based on,” above and in the claims is intended to mean, “based at least in part on,” such that an unrecited feature or element is also permissible.
[0088] The subject matter described herein can be embodied in systems, apparatus, methods, and / or articles depending on the desired configuration. The implementations set forth in the foregoing description do not represent all implementations consistent with the subject matter described herein. Instead, they are merely some examples consistent with aspects related to the described subject matter. Although a few variations have been described in detail above, other modifications or additions are possible. In particular, further features and / or variations can be provided in addition to those set forth herein. For example, the implementations described above can be directed to various combinations and subcombinations of the disclosed features and / or combinations and subcombinations of several further features disclosed above. In addition, the logic flows depicted in the accompanying figures and / or described herein do not necessarilyrequire the particular order shown, or sequential order, to achieve desirable results. Other implementations may be within the scope of the following claims.
Claims
CLAIMWhat is claimed:
1. A method of performing a compute-in-memory operation on a eDRAM cell comprising a plurality of transistors, the plurality of transistors comprising at least: an input transistor, a charge pump transistor, a write select transistor, a share select transistor, and an output transistor, the plurality of transistors having a respective channel located underneath, the method comprising: performing a program operation on the plurality of transistors, the program operation comprising: coupling the input transistor and the output transistor to an analog supply, and sequentially pulsing the charge pump transistor, the write select transistor, and the share select transistor between ground and the analog supply while the channel underneath the output transistor is programmed between ground and a digital supply; and performing a compute operation on the plurality of transistors, the compute operation comprising: coupling the channel underneath the input transistor to the channel underneath the output transistor, and pulsing the input transistor from ground to the analog supply to transfer charge in the channel underneath the output transistor to the channel underneath the input transistor.
2. The method of claim 1, wherein the analog supply comprises a voltage of between about 1.5V and about 2. IV.
3. The method of claim 1, wherein the digital supply comprises a voltage between about 0.6V and about 0.8V.
4. The method of claim 1, wherein: sequentially pulsing the charge pump transistor, the write select transistor, and the share select transistor comprises pulsing the charge pump transistor, the write select transistor, and the share select transistor over a plurality of cycles.
5. The method of claim 4, wherein: during a first cycle of the plurality of cycles, the pulsing writes an equal amount of charge to the channel underneath the input transistor and to the channel underneath the output transistor.
6. The method of claim 5, wherein: a second cycle of the plurality of cycles comprises a first phase and a second phase, and wherein: during the first phase of the second cycle of the plurality of cycles, the charge pump transistor and the write select transistor are pulsed and, during the second phase of the cycle of the plurality of cycles, the share select transistor is pulsed.
7. The method of claim 1, wherein pulsing the input transistor comprises accomplishing a linear matrix-vector multiplication operation by inducing a transfer of charge onto the output transistor proportional to the charge in the channel underneath the input transistor.
8. The method of claim 7, further comprising performing a non-destructive readout operation to determine a result of the linear matrix-vector multiplication.
9. The method of claim 1, wherein the plurality of transistors are disposed within a memory cell, and wherein a memory array comprises a plurality of memory cells arranged in a plurality of rows and a plurality of columns.
10. The method of claim 9, wherein the memory array further comprises: a row driver to drive the plurality of rows of memory cells; and a column driver to drive the plurality of columns of memory cells.
11. The method of claim 9, wherein the memory array is configured to perform fully row-parallel and column-parallel matrix-vector multiplication operations.
12. The method of claim 10, wherein the row driver comprises a pseudo-random number generator.
13. The method of claim 10, wherein the row driver comprises a binary -to-stochastic converter configured to convert a binary input into a stochastic representation.
14. A system comprising a monolithic 3D silicon arrangement, the monolithic 3D silicon arrangement comprising: a plurality of eDRAM cells, each of the plurality of eDRAM cells comprising: a plurality of transistors, the plurality of transistors comprising at least: an input transistor, a write select transistor, and an output transistor, each of the plurality of transistors having a respective channel located underneath, wherein the plurality of eDRAM cells are disposed in a first direction, a plurality of input lines and a plurality of bitlines couple the plurality of eDRAM cells in a second direction, and a plurality of output lines and a plurality of write select lines couple the plurality of eDRAM cells in a third direction; a first controller configured to drive the plurality of input lines and the plurality of bitlines of the monolithic 3D silicon arrangement; and a second controller configured to drive the plurality of output lines and the plurality of write select lines of the monolithic 3D silicon arrangement.
15. The system of claim 14, wherein the monolithic 3D silicon arrangement is configured to perform a program operation and a compute operation on the plurality of transistors.
16. The system of claim 15, wherein the program operation comprises: coupling the input transistor and the output transistor to an analog supply, and pulsing the write select transistor while the channel underneath the output transistor is programmed between a multi-level voltage between ground and a digital supply.
17. The system of claim 16, wherein the compute operation comprises:coupling the channel underneath the input transistor to the channel underneath the output transistor, and pulsing the input transistor from ground to the analog supply to transfer charge in the channel underneath the output transistor to the channel underneath the input transistor.
18. The system of claim 17, wherein the compute operation accomplishes a linear matrix-vector multiplication operation by inducing a transfer of charge onto the output transistor proportional to the charge in the channel underneath the input transistor.
19. The system of claim 14, wherein the first direction, second direction, and third direction are orthogonal.
20. The system of claim 14, wherein the first controller comprises a binary-to- stochastic converter configured to convert a binary input into a stochastic representation.
Citation Information
Patent Citations
Multi-layer vector-matrix multiplication apparatus for a deep neural network
US20190370639A1
Power-efficient compute-in-memory pooling
US20220012580A1
Ternary logic circuit device
US20220352893A1
Compute in memory three-dimensional non-volatile NAND memory for neural networks with weight and input level expansions
US20220398439A1
Memory device using semiconductor element
US20230012075A1