Charge domain in-DRAM computations for binary neural networks
By implementing analog calculations in DRAM and using multiplication and accumulation operations in the charge domain, the power consumption and throughput problems caused by frequent memory operations in deep neural networks are solved, and more efficient calculations are achieved.
Patent Information
- Application Number
- CN202411142780.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-07-01
- Filing Date
- 2024-08-20
- Publication Date
- 2025-06-27
AI Technical Summary
Deep neural networks reduce power consumption and throughput due to frequent memory reading and writing when performing tasks.
Analog calculation is implemented in dynamic random access memory (DRAM). By performing multiplication and accumulation operations in the charge domain, high parallelism calculation is performed using the capacitance characteristics of DRAM to reduce dependence on the processor.
Significantly reduces power consumption and improves throughput, avoiding irrelevant CPU reads and writes by performing calculations entirely in DRAM.
Smart Images

Figure CN120216448A_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims priority to U.S. Provisional Application No. 63 / 614,989, filed on December 27, 2023, which is incorporated herein by reference in its entirety. Technical Field
[0003] Embodiments of the present disclosure generally relate to memory devices. More specifically, embodiments of the present disclosure relate to in-memory computing within a dynamic random access memory in the charge domain. Background Art
[0004] Typically, a computing system may include electronic devices that communicate information via electrical signals during operation. For example, a computing system may include a processor communicatively coupled to a memory device such as a dynamic random access memory (DRAM) device. In this way, the processor may communicate with the memory device (e.g.) to retrieve executable instructions, retrieve data to be processed by the processor, and / or store data output from the processor.
[0005] Deep neural networks are becoming increasingly popular due to their excellent performance in performing machine learning tasks such as image classification, speech recognition, anomaly detection, and other tasks. The basic computational operations using deep neural networks are performed using multiply-accumulate operations that require frequent memory reads and writes. These frequent reads and writes from the DRAM by the processor greatly increase power consumption and greatly reduce the throughput of the task.
[0006] Embodiments of the present disclosure may address one or more of the problems set forth above. Summary of the Invention
[0007] In one aspect, the present application provides a method for computing in-memory computing of a dynamic random access memory (DRAM), comprising: loading input parameters into a first group of cells of the DRAM; loading inverted input parameters each complementary to a corresponding input parameter into a second group of cells of the DRAM; loading an indication of an offset voltage into an offset group of cells of the DRAM; performing an operation on weights with the corresponding stored input parameter or stored inverted input parameter; activating columns in the first group and the second group to perform an accumulation of the operations of the weights of the cells in the columns to store a sum; using the indication to generate an offset voltage in the columns; and generating an output based on the sum and the offset voltage and storing the output in an output group of cells of the DRAM.
[0008] In one aspect, the present application provides a system, comprising: a weight register configured to store weights; a plurality of word lines; a plurality of digit lines; a demultiplexer configured to route the weights through corresponding word lines among the plurality of word lines at least partially based on the value of the weights; a plurality of multi-bit cells, comprising: a first set of bits for storing input parameters; a second set of bits for storing inverted values of the input parameters, each inverted value being located in a corresponding bit corresponding to the corresponding bit of the input parameter, wherein the plurality of word lines are configured to combine the weights with the corresponding input parameters or inverted values to form combined values, and accumulate the combined values in the plurality of columns of the plurality of multi-bit cells on corresponding first digit lines among the plurality of digit lines; a third set of bits, wherein each column is configured to store an indication of an offset voltage to be generated on a corresponding second digit line among the plurality of digit lines; and a fourth set of bits configured to store an output of a comparison between the corresponding combined value on the first digit line and the corresponding offset voltage on the second digit line.
[0009] In yet another aspect, the present application provides a method for performing in-memory computing within a dynamic random access memory (DRAM), comprising: loading input parameters into a first set of bits in a first array bank of the DRAM; loading inverted input parameters complementary to the corresponding input parameters into a second set of bits in a second array bank of the DRAM; loading an indication of an offset voltage into an offset set of bits of the DRAM; multiplying the weights by the corresponding stored input parameters or inverted input parameters in a charge domain by selecting the first set when the weight has a first value and selecting the second set when the weight has a second value; activating columns comprising the first set and the second set to perform accumulation of operations of the weights of the cells in the columns in the charge domain to store a sum on a first digit line of the columns; generating an offset voltage in the columns by placing an amount of charge on a second digit line, wherein the amount of charge is at least partially based on the indication; performing a comparison in the charge domain between the sum on the first digit line and the offset voltage on the second digit line by activating a sense amplifier, generating an output based on the comparison; and storing the output in an output set of bits of the DRAM via the sense amplifier. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 A simplified block diagram illustrating certain features of a memory device having a sense amplifier, memory banks, and bank control in accordance with an embodiment of the present disclosure;
[0011] Figure 2 For a memory device in accordance with an embodiment of the present disclosure Figure 1 a block diagram of a portion of the memory device, showing portions of the sense amplifier, memory banks, and bank control for performing in-memory computing within the device in a charge domain;
[0012] Figure 3 For a memory device in accordance with an embodiment of the present disclosure Figure 1A block diagram of a portion of a portion of the above, showing a sense amplifier, a memory bank, and a portion of the bank control for determining a multiply-accumulate (MAC) sum in one or more columns of the memory bank;
[0013] Figure 4 is according to an embodiment of the present disclosure Figure 1 A block diagram of a portion of a portion of the above, showing a sense amplifier, a memory bank, and a portion of the bank control for generating an offset voltage and an output value;
[0014] Figure 5 is a timing diagram of offset voltage generation according to an embodiment of the present disclosure using nine columns and six-digit digital-to-analog conversion (DAC) codes Figure 4 for the offset voltage generation;
[0015] Figure 6 is a flowchart of a process for performing in-memory DRAM MAC calculations in the charge domain according to an embodiment of the present disclosure;
[0016] Figure 7 is according to an embodiment of the present disclosure Figure 1 A circuit diagram of a sense amplifier of the above, the sense amplifier configured to perform in-memory DRAM MAC calculations in the charge domain and including threshold voltage compensation (VTC) circuitry; and
[0017] Figure 8 is according to an embodiment of the present disclosure Figure 1 A circuit diagram of a sense amplifier of the above, the sense amplifier configured to perform in-memory DRAM MAC calculations in the charge domain without VTC circuitry. DETAILED DESCRIPTION
[0018] One or more specific embodiments will be described below. To provide a concise description of these embodiments, not all features of the actual implementation are described in the specification. It should be understood that in the development of any such actual implementation, as in any engineering or design project, many implementation-specific decisions must be made to achieve the developer's specific goals, such as compliance with system-related and business-related constraints, and the constraints of different implementations may be different. In addition, it should be understood that such development work may be complex and time-consuming, but for those of ordinary skill in the art who benefit from the present disclosure, these are routine tasks in design, construction, and manufacturing.
[0019] As previously discussed, deep neural networks (DNNs) can be used to perform machine learning tasks. However, using DNNs with traditional processor-based processing involves frequent reads and writes from memory, which has a negative impact on power consumption and throughput when executing tasks. In-memory computing can be applied, but in-memory computing is generally not applied to dynamic random access memory (DRAM) devices. The following presents an in-DRAM computing method for implementing analog computing in DRAM in the charge domain. In the computation of binary neural networks (BNNs), DRAM can implement bit-serial XNOR for multiply-accumulate operations.
[0020] BNNs simplify weights and input parameters to +1 or -1. BNNs can also use the sign (sgn) function to simplify the output of the result to +1 or -1 as well. For example, a BNN can use the following transformation:
[0021]
[0022] where w i,n is a weight parameter, x i,n is an input parameter, and a n is a bias parameter. The bias parameter can be a simplified parameter that combines one or more batch normalization parameters and filter biases. Using the transformation, if the multiply-accumulate (MAC) sum is greater than the bias, the output (z n ) will be +1. Otherwise, the output will be -1. As an example, the filter weights can include the following set: +1, +1, -1, +1, -1, +1, +1, -1, and +1, and the input parameters include the following set: +1, -1, -1, -1, -1, +1, -1, -1, and +1. When the BNN performs XNOR on these two values together, the MAC sum is +3, which provides an output when the sum is submitted to the sign function with a bias. If the bias is a value less than +3 (e.g., -2), the output will be +1. Otherwise, the output will be -1.
[0023] In memory, these +1 and -1 values can be implemented as a first value of +1 (e.g., 1) and a second value of -1 (e.g., 0). Additionally, word lines can be used to represent weights. Input parameters are stored in multi-bit cells. In-memory computing in the charge domain using such representations can be a generally high-parallelism computation. Such computations can also utilize high-density DRAM cell arrays to be compatible with various filter sizes and input feature sizes. Moreover, since the computation can be fully executed in DRAM to produce an output, irrelevant reads and writes of CPU-based computing are omitted, resulting in significantly better power efficiency and throughput when using in-DRAM computing.
[0024] Turning now to the figures, Figure 1Simplified block diagram for illustrating certain features of memory device 10. Specifically, Figure 1 The block diagram of is a functional block diagram for illustrating certain functionality of memory device 10. According to one embodiment, memory device 10 may be a double data rate type five synchronous dynamic random access memory (DDR5 SDRAM) device. The various features of DDR5 SDRAM allow for reduced power consumption, more bandwidth, and more storage capacity compared to previous generations of DDR SDRAM.
[0025] Memory device 10 may include several memory banks 12. For example, memory bank 12 may be a DDR5 SDRAM memory bank. Memory bank 12 may be disposed on one or more chips (e.g., SDRAM chips) arranged on a dual in-line memory module (DIMM). It should be understood that each DIMM may include several SDRAM memory chips (e.g., ×8 or ×16 memory chips). Each SDRAM memory chip may include one or more memory banks 12. Memory device 10 represents a portion of a single memory chip (e.g., SDRAM chip) having several memory banks 12. For DDR5, memory banks 12 may be further arranged to form bank groups. For example, for an 8 gigabyte (Gb) DDR5 SDRAM, the memory chip may include 16 memory banks 12 arranged in 8 bank groups, with each bank group including 2 memory banks. For example, for a 16 Gb DDR5 SDRAM, the memory chip may include 32 memory banks 12 arranged in 8 bank groups, with each bank group including 4 memory banks. Depending on the application and design of the overall system, various other configurations, organizations, and sizes of memory banks 12 on memory device 10 may be utilized.
[0026] Memory bank 12 and / or bank control block 22 includes sense amplifiers 13. As previously noted, sense amplifiers 13 are used by memory device 10 during a sense operation. Specifically, the sense circuit system of memory device 10 utilizes sense amplifiers 13 to receive low voltage (e.g., low differential) signals from the memory cells of memory bank 12 and amplify the small voltage difference so that memory device 10 can properly interpret the data.
[0027] Memory device 10 may include a command interface 14 and an input / output (I / O) interface 16. Command interface 14 is configured to provide several signals (e.g., signal 15) from an external (e.g., host) device (not shown), such as a processor or a controller. The processor or controller may provide various signals 15 to memory device 10 to facilitate the transmission and reception of data to be written to or read from memory device 10.
[0028] As should be understood, the command interface 14 may include several circuits, such as a clock input circuit 18 and a command address input circuit 20, for example, to ensure proper handling of the signal 15. The command interface 14 may receive one or more clock signals from an external device. Typically, a double data rate (DDR) memory utilizes a differential pair of system clock signals: a true clock signal Clk_t and an inverted / complementary clock signal Clk_c. The positive clock edge of DDR refers to the point where the rising true clock signal Clk_t crosses the falling complementary clock signal Clk_c, and the negative clock edge indicates the transition of the falling true clock signal Clk_t and the rising of the complementary clock signal Clk_c. Commands (e.g., read commands, write commands, activate commands, precharge commands, etc.) are typically input on the positive edge of the clock signal, and data is transmitted or received on both the positive and negative clock edges.
[0029] The clock input circuit 18 receives the true clock signal Clk_t and the complementary clock signal Clk_c, and generates an internal clock signal CLK. The internal clock signal CLK is supplied to an internal clock generator, such as a delay locked loop (DLL) circuit 30. The DLL circuit 30 generates a phase-controlled internal clock signal LCLK based on the received internal clock signal CLK. The phase-controlled internal clock signal LCLK is supplied to, for example, the I / O interface 16 and is used as a timing signal for determining the output timing of the read data. In some embodiments, the clock input circuit 18 may include circuitry that splits the clock signal into multiple (e.g., 4) phases. The clock input circuit 18 may also include phase detection circuitry that is used to detect which phase receives the first pulse when pulse clusters occur too frequently so that the clock input circuit 18 can reset between pulse clusters.
[0030] One or more internal clock signals / phases CLK may also be provided to various other components within the memory device 10 and may be used to generate various additional internal clock signals. For example, the internal clock signal CLK may be provided to the command decoder 32. The command decoder 32 may receive command signals from the command bus 34 and may decode the command signals to provide various internal commands. For example, the command decoder 32 may provide the command signals to the DLL circuit 30 via the bus 36 to coordinate the generation of the phase-controlled internal clock signal LCLK. The phase-controlled internal clock signal LCLK may be used, for example, to time data via the IO interface 16.
[0031] In addition, the command decoder 32 can decode commands such as read commands, write commands, mode register set commands, activate commands, precharge commands, etc., and provide access to a specific memory bank 12 corresponding to the command via the bus path 40. As should be understood, the memory device 10 can include various other decoders, such as row decoders and column decoders, to facilitate access to the memory bank 12. In one embodiment, each memory bank 12 includes a bank control block 22 that provides the necessary decoding (e.g., row decoder and column decoder) and other features, such as timing control and data control, to facilitate the execution of commands to and from the memory bank 12.
[0032] The memory device 10 performs operations such as read commands and write commands based on command / address signals received from an external device such as a processor. In one embodiment, the command / address bus can be a 14-bit bus (CA<13:0>) for accommodating command / address signals. The command / address signals are timed into the command interface 14 using clock signals (Clk_t and Clk_c). The command interface can include a command address input circuit 20 configured to receive and transmit commands through, for example, the command decoder 32 to provide access to the memory bank 12. In addition, the command interface 14 can receive a chip select signal (CS_n). The CS_n signal enables the memory device 10 to process commands on the incoming CA<13:0> bus. Access to a specific bank 12 within the memory device 10 is encoded by commands on the CA<13:0> bus.
[0033] In addition, the command interface 14 can be configured to receive several other command signals. For example, a command / address on-die termination (CA_ODT) signal can be provided to facilitate proper impedance matching within the memory device 10. For example, a reset command RESET_n can be used to reset the command interface 14, status register, state machine, etc. during power-up. The command interface 14 can also receive a command / address inversion (CAI) signal, and the command / address inversion signal can be provided to invert the state of the command / address signals CA<13:0> on the command / address bus, for example, depending on the command / address routing of the specific memory device 10. A mirror (MIR) signal can also be provided to facilitate the mirroring function. Based on the configuration of multiple memory devices in a specific application, the MIR signal can be used to multiplex signals so that they can be swapped for a certain routing of signals to the memory device 10. Various signals can also be provided to facilitate testing of the memory device 10, such as a test enable (TEN) signal. For example, the TEN signal can be used to place the memory device 10 into a test mode for connectivity testing.
[0034] The command interface 14 can also be used to provide a warning signal ALERT_n to the system processor or controller for certain detectable errors. For example, the warning signal ALERT_n can be emitted from the memory device 10 when a cyclic redundancy check (CRC) error is detected. Other warning signals can also be generated. Additionally, the bus and pins used to emit the warning signal ALERT_n from the memory device 10 can be used as input pins during some operations, such as the connectivity test mode performed using the TEN signal described above.
[0035] Using the commands and timing signals discussed above, data can be sent to and from the memory device 10 by transmitting and receiving data signals 44 via the I / O interface 16. More specifically, data can be sent to or retrieved from the memory bank 12 via a data path 46 that includes a plurality of bidirectional data buses. Data I / O signals, commonly referred to as DQ signals, are generally transmitted and received in one or more bidirectional data buses. For certain memory devices, such as DDR5 SDRAM memory devices, the IO signals can be divided into upper and lower bytes. For example, for a ×16 memory device, the IO signals can be divided into upper and lower IO signals (e.g., DQ<15:8> and DQ<7:0>) corresponding to the upper and lower bytes of the data signal.
[0036] To allow for higher data rates within the memory device 10, certain memory devices, such as DDR memory devices, can utilize data strobe signals, commonly referred to as DQS signals. The DQS signal is driven by an external processor or controller that transmits data (e.g., for a write command) or by the memory device 10 (e.g., for a read command). For a read command, the DQS signal effectively serves as an additional data output (DQ) signal with a predetermined pattern. For a write command, the DQS signal is used as a clock signal to capture the corresponding input data. Similar to the clock signals (Clk_t and Clk_c), the DQS signal can be provided as a differential pair (DQS_t and DQS_c) of data strobe signals to provide differential pair signaling during both reads and writes. For certain memory devices, such as DDR5 SDRAM memory devices, the differential pair of DQS signals can be divided into upper and lower data strobe signals (e.g., UDQS_t and UDQS_c; LDQS_t and LDQS_c), which correspond to the upper and lower bytes of the data transmitted to and from the memory device 10.
[0037] The impedance (ZQ) calibration signal can also be provided to the memory device 10 through the IO interface 16. The ZQ calibration signal can be provided to a reference pin and is used to tune the output driver and ODT values by adjusting the pull-up and pull-down resistors of the memory device 10 across process, voltage, and temperature (PVT) value variations. Since PVT characteristics can affect the ZQ resistor value, the ZQ calibration signal can be provided to the ZQ reference pin to adjust the resistance to calibrate the input impedance to a known value. As should be understood, a precision resistor is typically coupled between the ZQ pin on the memory device 10 and GND / VSS external to the memory device 10. This resistor serves as a reference for adjusting the drive strength of the internal ODT and IO pins.
[0038] Additionally, the loopback data signal LBDQ and the loopback strobe signal LBDQS can be provided to the memory device 10 through the IO interface 16. The loopback data signal and the loopback strobe signal can be used during the test or debug phase to set the memory device 10 into a mode where signals are looped back through the memory device 10 through the same pins. For example, the loopback signal can be used to set the memory device 10 to test the data output (DQ) of the memory device 10. The loopback can include both LBDQ and LBDQS or may only include the loopback data pins. This is typically intended for monitoring the data captured by the memory device 10 at the IO interface 16. LBDQ can indicate the data operation of the target memory device (e.g., the memory device 10), and thus can be analyzed to monitor the data operation of the target memory device (e.g., debug it and / or perform diagnostics on it). Additionally, LBDQS can indicate the strobe operation (e.g., the timing of the data operation) of the target memory device, such as the memory device 10, and thus can be analyzed to monitor the strobe operation of the target memory device (e.g., debug it and / or perform diagnostics on it).
[0039] As should be understood, various other components such as a power supply circuit (for receiving external VDD and VSS signals), a mode register (for defining various modes of programmable operations and configurations), a read / write amplifier (for amplifying signals during read / write operations), a temperature sensor (for sensing the temperature of the memory device 10), etc. can also be incorporated into the memory device 10. Therefore, it should be understood that only the Figure 1 block diagram is provided to highlight certain functional features of the memory device 10 to facilitate the subsequent detailed description. Additionally, although the memory device 10 has been described above as a DDR5 device, the memory device 10 can be any suitable device (e.g., a low-power double data rate (LPDDR) device, a double data rate type 4 DRAM (DDR4) device, another DRAM type, or a combination of different types of memory devices).
[0040] Figure 2 To include Figure 1Block diagram of a system 50 of portions of a memory bank 12, sense amplifiers 13, and bank control 22 of a memory device 10. As illustrated, bank control 22 includes a word line (WL) controller 52 that controls WL drivers 54, 56, 58, and 60. WL drivers 54 and 56 control access to the cells of memory bank 12 in a first array core 62, and WL drivers 58 and 60 control access to the cells of memory bank 12 in a second array core 64. Although WL drivers 54, 56, 58, and 60 are shown as monolithic drivers, each driver may include more than one WL driver. For example, WL drivers 54, 56, 58, and 60 may include WLs for each row or for groups of rows. WL drivers 54, 56, 58, and 60 can be used to read and write data from / to selected cells by asserting the WLs. As discussed below, WL drivers 54, 56, 58, and 60 can be used to load the weights of a BNN to perform the MACs previously discussed.
[0041] The first array core 62 may include a first group 66 of cells that can be used to load input parameters using a memory write. Similarly, the second array core 62 may include a second group 68 of cells that are used to load complementary input parameters. In some embodiments, by activating the WLs of the first group 66 using WL drivers 54 and / or 56, the input parameters stored in the first group 66 can be inverted and in the second group 68. The sense amplifier 13 is then activated by activating the WLs of the second group 68 to invert the data from the corresponding cells of the first group 66 by storing the inverted values from the corresponding digit lines (e.g., DLF) into the corresponding cells of the second group 68. In some embodiments, both the first group 66 and the second group 68 can be written using conventional memory writes.
[0042] The first array core 62 also includes an output group 72 of cells that are used to store the output of the MAC and compare it as the output of the BNN and / or for the next convolutional layer as part of machine learning. The output group 72 may include one or more rows of cells. The second array core 64 also includes an offset group 70 that can be used to store the offsets used when comparing the BNNs. As discussed more below, the offsets can be stored as digital codes that are used to control an analog offset charge to determine the result of the comparison (e.g., 0 or 1, also referred to as -1 and +1, respectively).
[0043] Figure 3Block diagram of a system 80 for performing a multiply-accumulate process, the system including a first set 66, a second set 68, and a weight register 74. The weight register 74 receives and stores weights 82. When a corresponding weight 82 (e.g., w2) is a first value (e.g., 1 indicating +1), the demultiplexer 84 activates a corresponding WL 86 (e.g., WL<2>) via WL drivers 88 (individually referred to as WL88A and 88B). Similarly, when a corresponding weight 82 (e.g., w3) is a second value (e.g., 0 indicating -1), the demultiplexer activates a corresponding WL 90 (e.g., WL<3>) via WL drivers 92 (individually referred to as WL 92A and 92B). In some embodiments, the WL drivers 88 and 92 can be triggered simultaneously. By leveraging the capacitive characteristics of DRAM, an XNOR operation can be performed, resulting in the following truth table:
[0044] <![CDATA[Weight w i > <![CDATA[Input parameter x i > Results 0 0 1 0 1 0 1 0 0 1 1 1
[0045] Table 1. Truth table with weight and input parameter values.
[0046] This principle holds because the weight of the first value adds the input parameter directly to column 94, while the weight of the second value adds the inverse of the input parameter directly to column 94. In other words, the cells of the column can be used to perform the XNOR operation of Table 1, and the column can be used to accumulate all the XNOR operations in column 94. For example, the accumulation can be performed on the digit lines (e.g., DLT and DLF) of column 94. The sense amplifier 13 can then store the MAC value of column 94 on a first digit line (e.g., DLT), while decoupling the other digit line and enabling the other digit line (e.g., DLF) to be precharged to another (e.g., lower) voltage to prepare to generate an offset voltage for evaluating the comparison of the MAC with the offset voltage, as previously discussed.
[0047] Figure 4 Block diagram of a system 100 for performing offset voltage generation and calculating a result based on the comparison of the offset voltage with the stored MAC value. As discussed with respect to Figure 3 The MAC sum is stored on the digit line (e.g., DLT) of column 94, while the other digit line (e.g., DLF) can be precharged to a voltage (e.g., VBLP). Then, the WL driver 102 uses a pulse 104 to open and close the cells in the offset group 70. In other words, the pulse 104 serially connects and disconnects each of the stored values in the DAC code 106 of the column. The DAC code 106 indicates how much charge is to be added to the precharged level of the other digit line to generate the offset voltage. For example, if the DAC code 106 is a six-bit number containing the value "011110", then for the four middle bits, the offset voltage increases as each corresponding charge is added to the other digit line (e.g., DLF).
[0048] In some embodiments, the DRAM may be able to synchronously calculate the magnitude of the length of the digital lines (e.g., 1,000 input parameters) and half of the WL length (e.g., 500 weights). This parallelism can be further increased by using multiple sense amplifiers 13 for calculation simultaneously. In addition, as the size of the input parameters increases, the width and number of bits in the offset group 70 and the output group 72 should also increase.
[0049] Figure 5 A timing diagram 120 showing the generation of offset voltages on nine columns of digital lines is presented. The timing diagram 120 includes a line 122 corresponding to the pulse 104, which pulses six different WLs in sequence, where each WL corresponds to the pulse and bit of the DAC code 106 for each of the nine columns. The line pairs 124, 126, 128, 130, 132, 134, 136, 138, and 140 each correspond to the voltages on a digital line pair: the true digital line (DLT) and the complementary digital line (DLF). The DLT is indicated by a solid line, and the DLF is indicated by a dashed line. The pulse activates the corresponding WLs at times 142, 144, 146, 148, 150, and 152. The line pair 124 corresponds to the value "011110", which causes the DLF of the line pair 124 to increase at times 144, 146, 148, and 150. The line pair 126 corresponds to the value "100011", which causes the DLF of the line pair 126 to increase at times 142, 150, and 152. The line pair 128 corresponds to the value "000001", which causes the DLF of the line pair 128 to increase only at time 152. The line pair 130 corresponds to the value "000000", which causes the DLF of the line pair 130 not to increase at all and only to decrease after each pulse. The line pair 132 corresponds to the value "011100", which causes the DLF of the line pair 132 to increase at times 144, 146, and 148. The line pair 134 corresponds to the value "010110", which causes the DLF of the line pair 134 to increase at times 144, 148, and 150. The line pair 136 corresponds to the value "000011", which causes the DLF of the line pair 136 to increase at times 150 and 152. The line pair 138 corresponds to the value "110110", which causes the DLF of the line pair 138 to increase at times 142, 144, 148, and 150. The line pair 140 corresponds to the value "111111", which causes the DLF of the line pair 140 to increase at times 142, 144, 146, 148, 150, and 152. In other words, when the bit of the DAC code 106 stores a logic high value, the activation of the corresponding WL causes the voltage of the DLF to increase. Similarly, when the bit of the DAC code 106 stores a logic low value, the activation of the corresponding WL causes the voltage of the DLF to decrease.
[0050] Return Figure 4Once the digital line has the stored MAC sum and another digital line has the stored offset voltage, activating the sense amplifier 13 results in a BNN calculation of the offset voltage minus the MAC sum. This value of 1 (+1) or 0 (-1) can then be stored in the output bank 108 by activating the corresponding WL (or WLs).
[0051] Figure 6 FIG. is a block diagram of a process 160 for performing in-memory computing in the charge domain. Input parameters are loaded into the DRAM memory bank 12 (block 162). In other words, the input parameters are written to the first bank 66 using a memory write. Similarly, an offset is loaded into the DRAM memory bank 12 as a DAC code 106 for each column to be computed (block 164). Similar to the input parameters, the offset can be loaded into the offset bank 72 instead of the first bank 66 using a memory write. The offset and the input parameters can be written simultaneously or at different times. Once the input parameters are loaded into the DRAM, the inverses of the input parameters can be stored in the second bank (block 166). As previously noted, these inverted values can be loaded by activating the WLs in the first bank 66, then enabling the sense amplifier 13, and then triggering the corresponding WLs in the second bank 68. Next, the weights 82 are used to activate the corresponding WL drivers according to the values of the weights 82 to cause an operation on the weights 82 and the input parameters in the first bank 66 or the inverted input parameters in the second bank 68 (block 168). For example, the demultiplexer 84 can be used to select the WL 86 of the first bank 66 for a first value (e.g., 1) and the WL 90 of the second bank for a second value (e.g., 0). Then the column 94 is activated to obtain the MAC result (block 170). This MAC result can be on the DLT and DLF of the column. Next, the system generates an offset voltage according to the offset (block 172). Generating the offset voltage can include precharging the DLF to a set voltage (e.g., VBLP), and then sequentially connecting the bits of the DAC code 106 to the column to generate the offset voltage. The column then generates the result of the BNN comparison in the charge domain and stores the result in the output bank 72 of the cell (block 174). For example, the sense amplifier 13 can be activated to compare the offset voltage with the MAC result. If the offset voltage is lower than the MAC result, the output is a first output value (e.g., 1, which corresponds to +1 in the BNN). Otherwise, the output is a second output value (e.g., 0, which corresponds to -1 in the BNN). This generation and storage of the output values can occur in each column and can be stored in one or more rows of the output bank 72. These output values can be output as the output of the BNN and / or can be used as the input to the next convolutional layer.
[0052] Figure 7 Can be implemented as Figure 1The circuit diagram of the sense amplifier 13 of the embodiment shown in [description], the sense amplifier can be used to perform voltage threshold compensation (VTC). Although only a single sense amplifier 13 is shown, multiple sense amplifiers 13 are included in the memory device 10, and the multiple sense amplifiers function similarly and can share at least some control signals and / or power supply voltages.
[0053] As illustrated, the sense amplifier 13 includes a PSA section 252, which includes PMOS transistors MPT 254 and MPB 256. The sense amplifier 13 also includes an NSA section 258, which includes NMOS transistors MNT 260 and MNB 262. MPT 254 and MPB 256 receive the ACT signal 264 at the terminals (e.g., source terminals) of MPT 254 and MPB 256 via a "top node". Although the illustrated embodiment shows both MPT 254 and MPB 256 coupled to the same ACT signal 264 and thus receiving the same voltage, some embodiments of the sense amplifier 13 may connect MPT 254 and MPB 256 to different ACT signals so that the source terminals of MPT 254 and MPB 256 can be driven at different voltage levels. The ACT signal 264 is generally used to control data movement and control of the sense amplifier 13. The ACT signal 264 can be driven using an array voltage (VARY) 266, which is selectively coupled and decoupled to MPT 254 and MPB 256 as the ACT signal 264 by a transistor 268 controlled by an SAP signal 270.
[0054] An isolation transistor (MN3b) 272 can be used to separate the transistor 268 (and VARY 266) from MPT 254 and MPB 256 based on the assertion of an isolation (ISOCS) signal 274, thereby substantially isolating the sense amplifier 13 from the SAP 270 signal. This avoids common source leakage during charge accumulation during the summation operation.
[0055] The terminal (e.g., drain) of MPT 254 is coupled to DLT 280 which is selectively coupled to gut node true (GUTT) 284, and the other terminal (e.g., drain) of MPB 256 is coupled to DLF 288 which is selectively coupled to gut node bar (GUTB) 283. The gate terminal of MPT 254 is also coupled to GUTB 283, and the gate terminal of MPB 256 is also coupled to GUTT 284. In other words, MPT 254 and MPB 256 are cross-coupled PMOS transistors coupled between the gut nodes and the ACT signal 264. As previously discussed, sense amplifier 13 receives signals from the memory cells and amplifies any difference. Sense amplifier 13 is selectively coupled to the memory cells via digital lines DLT 280 and DLF 288. DLT 280 carries the value (e.g., 1) indicating the value of the stored bit from the memory cell, while DLF 288 is complementary to the value (e.g., 0). Isolation ISOSA signal 278 can be used to selectively couple GUTT 284 to DLT 280 and decouple from DLT 280 via transistor MN1a 285, and selectively couple GUTB 283 to DLF 288 and decouple from DLF 288 via transistor MN1b 286. Transistors MN2a292 and MN2b 294 can be used to couple DLF 288 to GUTT 284 and DLT 280 to GUTB 283 respectively using BLCP signal 296.
[0056] MNT 260 has a terminal (e.g., source terminal) coupled to GUTT 284, and the gate terminal of MNT 260 is coupled to DLF 288. Similarly, MMB 262 has a terminal (e.g., source terminal) coupled to GUTB 283, and the gate terminal of MMB 262 is coupled to DLT 280. The other terminals of NMOS transistors MNT 260 and MNB 262 are coupled together to RNL signal 298. The RNL signal 298 (e.g., NMOS strobe signal) can be an optional voltage which can strobe MNT 260 and MMB 262 to a voltage level (e.g., ground / VSS 300) to complete the latching once the amplification in sense amplifier 13 has amplified a relatively low voltage from the memory cell. For example, when the SAN signal 304 is asserted, this RNL signal 298 can transition to VSS 300 to perform such latching via transistor 302. However, transistor MN3a 299 can isolate the sense amplifier 13 from the SAN signal 304 using the ISOCS signal 274 to avoid leakage during MAC accumulation operations.
[0057] In operation, Figure 7The sense amplifier 13 can operate using a VTC that compensates for the threshold voltage offsets of the combination of MPT 254, MPB 256, MNT 260, and MMB 262. Additionally, since both DLT 280 and DLF 288 will be used for accumulation in column 94, they are shorted together during accumulation by asserting the BLCP signal 296 and the ISOSA signal 278. They can be isolated after the MAC sum is generated so that DLF 288 can be precharged to another voltage.
[0058] The sense amplifier 13 also includes a precharge section 306 that can selectively enable DLT 280 to be charged to VBLP 310 using a transistor 312 with a precharge (BLPRU) signal 308. Similarly, the precharge section 306 can selectively enable DLF 288 to be precharged separately from DLT 280 to VBLP 310 using a transistor 316 with a precharge (BLPRD) signal 314. Thus, DLF 288 can be precharged as part of or in preparation for the offset voltage generation discussed previously.
[0059] Figure 8 Can be implemented as Figure 1 The circuit diagram of the sense amplifier 13 of the embodiment shown in, the sense amplifier may not include VTC circuitry. In other words, in addition to the absence of the ISOSA signal 278, transistor MN1a 285, or transistor MN1b 286, Figure 8 The sense amplifier 13 is similar to Figure 7Sense amplifier 13. Thus, to perform equalization, precharge section 306 includes transistor MN5 400 enabled by shorting DLT 280 and DLF 288 together during accumulation. Specifically, when the precharge signal (BLPR) 402 is not asserted, but the equalization (BLEQU) signal 404 and the equalization (BLEQD) signal 406 are asserted, DLT 280 and DLF 288 are shorted together. The BLEQU signal 404 can be used to enable DLT 280 to be shorted to DLF 288 and / or charged to VBLP 310. The BLEQD signal 406 can be used to enable DLF 288 to be shorted to DLT 280 and / or charged to VBLP 310. In other words, if the BLPR signal 402 is asserted, the assertion of the corresponding equalization signal (e.g., BLEQU signal 404 or BLEQD signal 406) can enable DLF 288 to be charged to VBLP 310. For example, after storing the MAC sum on DLT 280 and DLF 288, the BLEQU signal 404 can be deasserted to separate DLT 280 from DLF 288. Then, the BLPR signal 402 can be asserted to precharge DLF 288 to VBLP 310 to prepare for the offset voltage generation and / or a part thereof discussed previously.
[0060] While the present disclosure may admit of various modifications and alternative forms, specific embodiments have been shown by way of example in the drawings and have been described in detail herein. However, it should be understood that the present disclosure is not intended to be limited to the particular forms disclosed. Indeed, the present disclosure is intended to cover all modifications, equivalents, and alternatives falling within the spirit and scope of the present disclosure as defined by the appended claims.
[0061] The techniques presented and claimed herein are referenced and applied to substantial objects and specific instances with a practical nature, which substantially improve the technical field of the present invention in an arguable manner and are thus not abstract, intangible, or purely theoretical. Additionally, if any technical solution appended at the end of this specification contains one or more elements designated as "means for [performing] [function]..." or "steps for [performing] [function]...", such elements are intended to be construed in accordance with 35 U.S.C. 112(f). However, for any technical solution containing elements designated in any other way, such elements are not intended to be construed in accordance with 35 U.S.C. 112(f).
Claims
1. A method for calculating a calculation in a dynamic random access memory DRAM, comprising: loading input parameters to a first group of cells of the DRAM; loading a second group of cells of the DRAM with inverted input parameters each complementary to a corresponding input parameter; an instruction to load an offset voltage to an offset group of cells of the DRAM; performing an operation on the weights using corresponding stored input parameters or stored inverse input parameters; activating columns in the first group and the second group to perform accumulation on the operations of the weights of the cells in the columns to store a sum; generating an offset voltage in the column using the indication; and An output is generated based on the sum and the offset voltage and stored in an output group of cells of the DRAM. 2 . The method of claim 1 , wherein loading the first group comprises writing to a memory of the first group. The method of claim 1 , wherein loading the second group comprises writing to a memory of the second group.
4. The method of claim 1 , wherein loading into the second group comprises inverting the input parameters of the first group into the inverted input parameters of the second group by: activating the word lines of the first group, Activate the sense amplifier, and A word line of the second group is activated to store an inversion of a value to the second group.
5. The method of claim 1, wherein the indication comprises a digital-to-analog code indicating a number of pulses to add charge to generate the offset voltage.
6. The method according to claim 1, comprising: selecting a word line of the first group based on a first value of a first weight of the weights; and The word line of the second group is selected based on a second value of a second weight of the weights, wherein the first value and the second value are logical complements. The method of claim 1 , wherein the operation comprises an XNOR operation. The method of claim 1 , wherein storing the sum comprises storing the sum on two digit lines of the column.
9. The method of claim 8, wherein generating the offset voltage comprises: maintaining storage of the sum on a first digit line of the two digit lines; and The offset voltage is generated on a second digit line of the two digit lines.
10. The method of claim 9, comprising precharging the second digit line before generating the offset voltage on the second digit line.
11. The method of claim 1, comprising using the output in a binary neural network (BNN).
12. A system comprising: a weight register configured to store a weight; multiple word lines; Multiple digital lines; a demultiplexer configured to route the weights through corresponding word lines of the plurality of word lines based at least in part on values of the weights; A plurality of multi-bit cells comprising: A first group of bits, which are used to store input parameters; a second set of bits for storing inverted values of the input parameter, each inverted value being located in a respective bit corresponding to a respective bit of the input parameter, wherein the plurality of word lines are configured to combine the weights with the corresponding input parameter or inverted value to form a combined value, and to accumulate the combined values in the plurality of columns of the plurality of multi-bit cells on a respective first digit line of the plurality of digit lines; a third set of bits, wherein each column is configured to store an indication of an offset voltage to be generated on a corresponding second digit line of the plurality of digit lines; and A fourth set of bits is configured to store an output of a comparison between a corresponding combined value on the first digit line and a corresponding offset voltage on the second digit line.
13. The system of claim 12, comprising a convolutional neural network implemented in a memory device, the memory device comprising the weight register, the plurality of word lines, the demultiplexer, and the plurality of multi-bit cells.
14. The system of claim 12, wherein combining the weight with the corresponding input parameter or inversion value comprises an XNOR operation.
15. The system of claim 12, comprising a plurality of sense amplifiers configured to perform the comparison on corresponding columns of the first and second groups.
16. The system of claim 15 , wherein each of the plurality of sense amplifiers is configured to perform threshold voltage compensation, wherein accumulating the combined value of the corresponding column comprises shorting the corresponding first digit line and the second digit line together by asserting a pair of signals, wherein generating the offset voltage comprises ceasing to short the corresponding first digit line and the second digit line together after accumulation has been completed, and wherein generating the offset voltage further comprises precharging the corresponding second digit line and causing the charge on the corresponding second digit line to change by a number of pulses and a value indicated by a corresponding indication.
17. The system of claim 15, wherein as part of the offset voltage generation, each of the plurality of sense amplifiers is configured to: equalizing the respective first digit line and the respective second digit line by disabling precharging and enabling equalization signals for the respective first digit line and the respective second digit line; After the accumulation, disabling equalization of the corresponding second digital line; After disabling equalization, precharging the corresponding digit line; and The charge on the corresponding digit line is caused to change by a number of pulses and a value indicated by the corresponding indication.
18. A method for calculating a calculation in a dynamic random access memory DRAM, comprising: loading input parameters into a first set of bits in a first array core of the DRAM; loading a second group of bits in a second array core of the DRAM with an inverted input parameter complementary to the corresponding input parameter; an indication of loading an offset voltage to an offset group bit of the DRAM; multiplying the weights with corresponding stored input parameters or inverted input parameters in the charge domain by selecting the first group when the weights have a first value and selecting the second group when the weights have a second value; activating a column including the first group and the second group to perform an accumulation in the charge domain of the weights of the cells in the column to store a sum on a first digit line of the column; using the indication to generate an offset voltage in the column by placing an amount of charge on a second digit line, wherein the amount of charge is based at least in part on the indication; performing a comparison of the sum on the first digit line and the offset voltage on the second digit line in the charge domain by activating a sense amplifier, generating an output based on the comparison; and The output is stored in output group bits of the DRAM via the sense amplifier.
19. The method of claim 18, wherein the multiplication comprises an XNOR operation of the corresponding weight and the stored input parameter or the inverted input parameter.
20. The method of claim 18, wherein the offset group is in the second array core and the output group is in the first array core.