In-memory computation with ternary activations
By using switch pairs and capacitors in the memory calculation bit cell, combined with the configuration of the controller, efficient multiplication of signed input bits is achieved, which solves the problems of reduced calculation accuracy and high energy consumption in the prior art, and improves the performance of in-memory calculations.
Patent Information
- Application Number
- CN202280019411.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-03-17
- Filing Date
- 2022-03-10
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2042-03-10
AI Technical Summary
When processing signed input bits in existing memory, there are problems of reduced calculation accuracy and high energy consumption when processing signed input bits. Especially when accommodating symbols, binary 0 is mapped to -1, resulting in the input vectors that can only represent odd numbers, limiting the calculation accuracy.
A calculation bit unit in memory is provided, and multiplication operation of memory bits and signed input bits is realized through the combination of switch pairs and capacitors, combined with the configuration of the controller. In the first stage of the operation, the switch state is controlled according to the symbol value of the input bit; in the second stage, the switch state is reversed or maintained according to the amplitude value of the input bit to complete the multiplication operation.
Improves calculation accuracy, reduces energy consumption, can effectively process signed input bits, and enhances the performance of calculations in memory.
Smart Images

Figure CN116964675B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims the benefit of U.S. Application No. 17 / 204,649, which is hereby incorporated by reference in its entirety. Technical Field
[0003] The present application relates to in-memory computing, and more particularly, to in-memory computing with ternary activations. Background Art
[0004] Unlike conventional bit cells, compute-in-memory (CiM) bit cells not only store bits, but also include logic gates for multiplying the stored bits with input bits. CiM greatly speeds up computation time for applications such as artificial intelligence, because the resulting multiplication does not require fetching bits from memory to transfer to an arithmetic logic unit for subsequent multiplication, as would be performed in a classic von Neumann computer architecture. Instead, the multiplication occurs in the memory itself.
[0005] Although in-memory computational bit cells are advantageous over conventional bit cells for computationally intensive applications such as artificial intelligence, problems arise regarding accommodating the sign (positive or negative) of the input bit multiplied by the storage bit of the in-memory computational bit cell. In order to accommodate the sign, the binary 0 value of the input bit can be considered to represent -1. In this accommodation, a set of input bits forms an input vector. Since binary 0 is mapped to -1, each input vector represents an odd number. For example, -7 can be represented by an input vector [-1, -1, -1], while 7 can be represented by an input vector [1, 1, 1]. This restriction on the odd number of input vectors in a signed implementation reduces the computational accuracy. In addition, using this conventional signed implementation, the charging and discharging of the capacitors in the in-memory computational bit cells may consume a lot of energy. Summary of the invention
[0006] A memory is provided, comprising: a bit cell having a switch pair connected to an output node; a capacitor coupled to the output node; a first storage element and a plurality of additional storage elements; and a controller configured to: during a first phase of operation of the memory, select a first bit from the first storage element to control the switch pair in response to the first bit, and to: during a second phase of operation of the memory, select a second bit from the plurality of additional storage elements to control the switch pair in response to the second bit.
[0007] In addition, a method of controlling a bit cell to multiply a storage bit with a signed input bit is provided, the method comprising: during a first phase of operation and in response to the sign of the signed input bit having a first binary value, closing a first switch coupled between a node for storing the bit and an output node, and opening a second switch coupled between a node for storing the complement of the bit and the output node; during the first phase of operation and in response to the sign of the signed bit having a second binary value, opening the first switch and closing the second switch; during a second phase of operation and in response to the amplitude of the signed input bit having the first binary value, inverting the switch states of the first switch and the second switch established during the first phase of operation; and during the second phase of operation and in response to the amplitude of the signed input bit having the second binary value, maintaining the switch states of the first switch and the second switch established during the first phase of operation.
[0008] In addition, a memory is provided, the memory comprising: a bit cell configured to store a storage bit, the bit cell comprising a first switch coupled between a node for storing the bit and an output node and a second switch coupled between a node for storing the complement of the bit and the output node; a capacitor having a first plate connected to the output node; and a controller, the controller being configured to: in a first stage of operation, in response to a sign of an input word having a first binary value, open the second switch and close the first switch, and in response to a sign bit of the input word having a second binary value, open the first switch and close the second switch to control the switching states of the first switch and the second switch, wherein the second binary value is the complement of the first binary value.
[0009] Finally, a method of operation for in-memory computing is provided, the method comprising: during a first phase of operation, controlling a switch pair coupled between a bit cell and a plate of a capacitor in response to a sign bit; and during a second phase of operation, controlling the switch pair in response to a magnitude bit.
[0010] These and other advantageous features will be better understood from the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 is a circuit diagram of an in-memory computation bit cell including a switch pair according to one aspect of the present disclosure.
[0012] Figure 2 According to one aspect of the present disclosure, Figure 1 An implementation of an in-memory computational bit cell, wherein the switch pair comprises a transmission gate pair.
[0013] Figure 3 According to one aspect of the present disclosure, Figure 1An implementation of an in-memory computational bit cell, wherein the switch pair comprises a PMOS transistor pair.
[0014] Figure 4 Illustrated are aspects of a controller for selecting from an input buffer during ternary computation of a bit cell in memory according to one aspect of the present disclosure.
[0015] Figure 5 Some operational waveforms of an in-memory computation bit cell performing ternary computation according to one aspect of the present disclosure are illustrated.
[0016] Figure 6 Illustrated are columns of in-memory computation bit cells configured for ternary computation and organized to form a multiply and accumulate (MAC) circuit according to one aspect of the present disclosure.
[0017] Figure 7 A memory including an array having a plurality of columns, each column including multiply and accumulate circuits configured for ternary computations, according to one aspect of the disclosure is illustrated.
[0018] Figure 8 is a flow chart of an example ternary computation for computing a bit cell in memory according to one aspect of the present disclosure.
[0019] Fig. 9 Several example electronic systems are illustrated that each include an array of in-memory computation bit cells configured for ternary computation according to one aspect of the present disclosure.
[0020] Embodiments of the present disclosure and their advantages may be best understood by referring to the following detailed description.It should be appreciated that like reference numerals are used to identify like elements illustrated in one or more of the figures. DETAILED DESCRIPTION
[0021] In deep learning and other machine learning applications, convolutional layers are basic building blocks. A convolutional layer comprises a collection of nodes for multiplication of filter weight bits with an input vector from a previous layer (or from input data such as an image being analyzed). Nodes may also be designated as neurons. To increase processing speed, neurons or nodes are implemented using in-memory computational bit cells. To provide increased computational accuracy and reduce power consumption, a ternary computation technique is provided in which the input vector can have either odd or even sign values. The technique is denoted as a "ternary" computation technique because the resulting computation at the in-memory computational bit cell can cause the voltage of the read bit line to increase, remain unchanged, or decrease.
[0022] Ternary computing disclosed herein is also referred to as ternary activation. Ternary computing can be practiced using any suitable in-memory computing bit cell including a switch pair and a capacitor. Figure 1 An example computation in memory (CiM) bit cell 100 is shown in FIG. A pair of cross-coupled inverters 120 and 125 store bits on an output node wt. The stored bits are also referred to as filter weight bits. Therefore, the output node wt will also be represented as a filter weight bit node wt. The pair of cross-coupled inverters 120 and 125 also store a complement filter weight bit (the complement of the filter weight bit) on a complement filter weight bit node wtb. The filter weight bit node wt is the output node of inverter 120, and the complement filter weight bit node wtb is the output node of inverter 125. The logic gates in the computation in memory bit cell 100 are formed by a left (L) switch and a (R) switch. The L switch is coupled between the filter weight bit node wt and the output node 105. Similarly, the R switch is coupled between the complement filter weight bit node wtb and the output node 105. The capacitor C is coupled between the output node 105 and the read bit line (RBL). As used herein, the term "bit cell" for CiM applications will be understood to refer to inverters 120 and 125 and L and R switches, since these devices are formed from transistors implemented on a semiconductor die. In contrast, capacitor C is a passive device that can be shared by other CiM bit cells in alternative implementations.
[0023] The in-memory calculation bit unit 100 calculates the multiplication of the filter weight bit and the signed input bit. The signed input bit is a bit within a signed input vector. The signed input vector can also be represented as a signed input word. In order to better understand the advantageous ternary activation for the multiplication of the signed input bit and the filter weight bit disclosed herein, the signed multiplication with the input vector having only odd magnitudes will first be discussed. As previously mentioned, the signed implementation of the input vector usually limits the input bits of the signed input vector to be considered as representing -1 or 1. For example, the binary 0 of the input bit can be mapped to -1, and the binary 1 is mapped to 1. In this mapping, the signed input vector can then only represent odd numbers. For example, the input vector [-1, -1, -1] represents the signed value -7. Similarly, the input vector [1, 1, 1] represents the signed value 7. In this way, a 3-bit wide input vector in odd-only signed representation can represent the odd values -7, -5, -3, -1, 1, 3, 5, and 7, depending on the binary value of the individual input bits.
[0024] In one implementation, each of the L and R switches is implemented using a transmission gate. The transmission gate can pass strong 0 (passing ground through the transmission gate) and strong 1 (passing the supply voltage VDD through the transmission gate), but requires both p-type metal oxide semiconductor (PMOS) transistors and n-type metal oxide semiconductor (NMOS) transistors. A single transistor (such as a PMOS transistor) can also be used to implement each L and R switch, but the PMOS transistor cannot pass strong 0, but can only pass strong 1 and weak 0 (positive voltage instead of ground, due to transistor threshold voltage requirements).
[0025] The number of operation phases that the in-memory computation bit cell 100 uses to perform the multiplication of the input bit and the filter weight bit depends on whether both strong 0 and strong 1 can be passed by the switch implementation. In the multiplication, the final result will be that the output node 105 is grounded (passing a strong 0), or the output node 105 is charged to the supply voltage VDD (passing a strong 1). Since the PMOS implementation of the L switch and the R switch cannot pass a strong 0, the output node 105 is first grounded in the pre-charge phase. However, the transmission gate implementation of the L switch and the R switch can pass a strong 0, so the pre-charge phase is not required in the transmission gate implementation. The following discussion of odd-only signed multiplication will assume that the L switch and the R switch are transmission gates so that the multiplication is performed through the calculation phase of the operation and the accumulation phase of the operation. In contrast, if the L switch and the R switch are implemented using PMOS transistors, the pre-charge phase of the operation is necessary so that the output node 105 can be initially discharged (e.g., grounded) to a binary 0. The following computation phase may then discharge the output node 105 to represent a binary 0, or may charge the output node 105 to the supply voltage VDD to represent a binary 1.
[0026] Prior to the calculation phase, the read bit line is precharged by turning on the precharge PMOS transistor P1. The source of the precharge transistor P1 is connected to a node for a common mode voltage VCM. The common mode voltage VCM may be equal to the supply voltage VDD in some implementations, or may be a fraction of the supply voltage VDD in alternative implementations. Without loss of generality, the following discussion will assume that the common mode voltage VCM is equal to one-half of the supply voltage VDD. Regardless of whether there is a precharge phase, the precharge transistor P1 remains on during the calculation phase to keep the read bit line charged to the common mode voltage VCM.
[0027] In the calculation phase, if the input bit is a binary 0 (which maps to -1, as described above), the R switch is closed and the L switch is open. If the filter weight bit is a binary 1, the R switch then passes the 0 to the ground output node 105, so that the capacitor C is charged to the reference voltage VCM. Conversely, if the filter weight bit is a binary 0, the output node 105 is charged to the supply voltage VDD during the calculation phase, so that the capacitor C is discharged in the realization that the reference voltage VCM is equal to the supply voltage VDD.
[0028] If the input bit is a binary 1, the switches are controlled in a complementary manner. In this case, during the calculation phase, the L switch is closed and the R switch is open. If the filter weight bit is a binary 1, the output node 105 is charged to the supply voltage VDD, so that the capacitor C is charged to -VCM. If the filter weight bit is a 0, the output node 105 remains discharged, so that the capacitor C remains charged to VCM.
[0029] The input bit can also be represented as an activation bit as part of an activation vector. For an implementation where the input vector has only odd magnitudes, the resulting relationship between the activation (Act) bit, the filter weight bit (Wt), and the binary state of the output node (Out) in the calculation phase can be summarized in Table 1 below:
[0030] Table 1
[0031]
[0032] It can therefore be seen that the binary state of the output node voltage (0 for ground and 1 for supply voltage VDD) is the exclusive NOT OR (XNOR) of the activation bit and the filter weight bit, so that the resulting CiM calculation can be specified as an XNOR-based calculation.
[0033] The accumulation phase follows the calculation phase. In the accumulation phase, precharge transistor P1 is turned off to float the read bit line. Output node 105 is then grounded. If capacitor C is charged to -VCM in the calculation phase, the read bit line voltage is then pulled below the common mode voltage VCM. Conversely, if capacitor C remains charged in the calculation phase, the read bit line voltage is not affected in the accumulation phase. Note that the read bit line spans the columns of CiM bit cells ( Figure 1 The capacitors C for the column all have their bottom plates grounded by the grounding of the corresponding output node 105 during the accumulation phase. The capacitors C all have their top plates or terminals coupled to the read bit line. Thus, the read bit line accumulates the shared charge from all capacitors C during the accumulation phase.
[0034] Since the activation bit is either binary 0 or binary 1, the representation of the negative and positive signs of the corresponding only odd-magnitude activation vector (from which the activation bit is derived) may require mapping 0 to -1 (1 remains 1). Although the resulting XNOR-based computation is more efficient than using conventional von Neumann computer architectures, the odd-only restriction of the signed activation vector reduces computational accuracy. In addition, the charging and discharging of capacitor C when transitioning from the computation phase to the accumulation phase may consume significant energy.
[0035] In order to reduce power consumption and increase accuracy, a ternary computing scheme is provided for a capacitive CiM bit cell. The capacitive CiM bit cell can be arranged as previously discussed for the CiM bit cell 100 in an XNOR-based computing scheme. Therefore, there are L switches and R switches, which are controlled in the ternary computing. If these L switches and R switches are transmission gates, only a computing phase and an execution phase are required. If PMOS transistors are used to form the L switches and R switches instead, a pre-charge phase can be included, as will be discussed further below.
[0036] Figure 2 , a transmission gate implementation of an example CiM bit cell 200 is shown in FIG. Inverters 120 and 125, read bit line RBL, and precharge transistor P1 are arranged as discussed for bit cell 100. The filter weight bit node wt (the output of inverter 120) is coupled to output node 105 through transmission gate T1, which forms an L switch. Capacitor C is coupled between output node 105 and read bit line RBL, as discussed for bit cell 100. A PMOS transistor P2 arranged in parallel with NMOS transistor M1 forms transmission gate T1. An activation bit signal (Act) controls the gate of transistor M1, while the complement of the activation bit signal (ActB) controls the gate of transistor P2. Thus, when the activation bit signal is true (binary 1 in a high-active implementation), transmission gate T1 is closed, and when the activation bit signal is false (binary 0 in a high-active implementation), transmission gate T1 is open. It should be understood that a low-active activation bit signal may be used in an alternative implementation.
[0037] The transmission gate T2 that forms the R switch is similar in that it is also formed by a parallel combination of a PMOS transistor P3 and an NMOS transistor M2. The complement activation bit signal (ActB) controls the gate of transistor M2, while the activation bit signal (Act) controls the gate of transistor P3. Therefore, when the activation bit signal is false (binary 0 in a high-effective implementation), the transmission gate T2 is closed, and when the activation bit signal is true (binary 1 in a high-effective implementation), the transmission gate T2 is open. Since the transmission gates T1 and T2 can pass both strong 1s and strong 0s, when the read bit line is precharged by turning on transistor P1, there is no need for the precharge operation phase in which the bit cell 200 grounds the output node 105.
[0038] Reference now Figure 3 , an alternative in-memory computation bit cell 300 is shown in which the L switch and the R switch cannot pass a strong 0. Inverters 120 and 125, the read bit line RBL, and the precharge transistor P1 are arranged as discussed for the bit cell 100. PMOS transistor P2 forms an L switch, which is coupled between the filter weight bit node wt and the output node 105. PMOS transistor P3 forms an R switch, which is coupled between the complementary filter weight bit node wtb and the output node 105. Capacitor C is coupled between the output node 105 and the read bit line, as discussed for the bit cells 100 and 200. The complementary activation bit signal ActB controls the gate of transistor P2. Therefore, when the activation bit signal Act is true, the left switch is turned on. Similarly, the activation bit signal Act controls the gate of transistor P3 so that when the activation bit signal is false, transistor P3 is turned on.
[0039] Since neither transistor P2 nor P3 can pass a strong zero, output node 105 is initially discharged by NMOS reset transistor M3 during the precharge phase of operation. The source of transistor M3 is connected to ground, while its drain is connected to output node 105. A read word line (RWL) controls the gate of transistor M3. The read word line RWL is asserted during the initial precharge phase, during which precharge transistor P1 is also turned on. Therefore, transistor M3 is turned on during the precharge phase of operation so that capacitor C can be charged to the common mode voltage VCM. For a transmission gate implementation such as that discussed for bit cell 200, the precharge phase of operation is unnecessary.
[0040] Regardless of whether the precharge phase of the operation is used, the sign of the signed activation vector can be represented by a sign bit. The activation bit for the signed activation vector (which can also be represented as an amplitude bit) can be arranged from the least significant bit (LSB) to the most significant bit (MSB) as in a conventional binary word. For example, in a three-bit wide signed activation vector, the range of the activation bit can be from
[000] to
[111] . In this three-bit wide implementation, for example, the activation bit portion of the signed activation vector value 5 or -5 can therefore all be represented by
[011] . The sign bit is 0 or 1, to represent the negative or positive sign of the signed activation vector, respectively. If the sign bit is multiplied by the activation bit, the result can have one of the following four possible values: -1, -0, 0, and 1.
[0041] Given these four possible values of the sign bit multiplied by the activation bit, the calculation stage in the ternary calculation is completely different from the conventional XNOR-based calculation discussed previously. In the XNOR-based calculation, the control of the L switch and the R switch depends on the binary state of the activation bit. But in the ternary calculation stage, the control of the L switch and the R switch depends only on the sign bit, as shown in Table 2 below.
[0042] Table 2
[0043]
[0044] In the calculation phase, if the sign bit is positive (equal to 1 in a high-effective implementation), the L switch is closed and the R switch is open regardless of the value of the activation bit. The precharge transistor P1 remains on during the calculation phase. Conversely, if the sign bit is negative (equal to 0 in a high-effective implementation), the R switch is closed and the L switch is open during the calculation phase. Likewise, the control of the R and L switches by a negative sign bit is independent of the value of the corresponding activation bit.
[0045] During the execution phase of operation following the calculation phase, the precharge transistor P1 is turned off to float the read bit line relative to the node for the common mode voltage VCM. If the activation bit is a binary 1, the closed / open switch state of the L switch and the R switch in the execution phase is opposite to any switch state the L switch and the R switch were in during the calculation phase. In other words, if the L switch or the R switch is closed during the calculation phase, the same switch will be open during the execution phase if the activation bit is a binary 1. Similarly, if the L switch or the R switch is open during the calculation phase, the same switch will be closed during the execution phase if the activation bit is a binary 1. If the activation bit is a binary 0, the closed / open switch state of the L switch and the R switch in the calculation phase remains unchanged during the execution phase.
[0046] Note the difference between the ternary operation of the capacitive CiM bit cell and the XNOR-based operation. In the XNOR-based operation, the accumulation phase always grounds the output node 105. However, in the ternary execution phase, the binary state of the output node 105 can be 1 (charged to the supply voltage VDD) or 0 (discharged to ground). During the execution phase, the output node 105 can therefore be raised from ground to the supply voltage (VDD), remain discharged to ground, remain charged to the supply voltage VDD, or discharge from the supply voltage VDD to ground. Given these four possible outcomes of the output node voltage, it can be understood that the resulting operation is actually ternary, because if the output node voltage is converted from ground to VDD in the execution phase, the read bit line voltage can be raised to above the common mode voltage VCM. In contrast, if the output node remains grounded in both the calculation and execution phases, the read bit line voltage remains unchanged (equal to the common mode voltage). Similarly, if the output node remains charged to the supply voltage VDD in both the calculation and execution phases, the read bit line voltage does not change. Finally, if the output node voltage is converted from the supply voltage VDD in the calculation phase to ground in the execution phase, the read bit line voltage is reduced from the common mode voltage in the execution phase.
[0047] In XNOR-based calculations, the accumulation phase can only discharge the read bit line voltage from the common mode voltage, and there is no increase in the read bit line voltage from the common mode voltage. Therefore, in ternary-based calculations, the output voltage swing of the read bit line is twice the output voltage swing generated according to XNOR-based calculations. This increased output voltage swing of ternary-based calculations is beneficial for reducing analog-to-digital conversion noise in calculations, as will be further explained herein.
[0048] Therefore, ternary computation will operate differently from XNOR-based computation. In XNOR-based computation, there is no sign bit, so both the L switch and the R switch are open during the precharge phase (if present), and then the controller controls these switches based on the activation bit during the computation phase. But in ternary computation, the controller 400 of the L switch and the R switch will look at the sign bit to control the left and right switches during the computation phase, and then look at the activation bit during the execution phase, as shown in FIG. Figure 4As shown in . For example, a signed activation vector can be stored in a buffer 410, which includes a first storage element and a plurality of additional storage elements. A calculation / execution control signal 415 for the controller 400 controls the multiplexer 405 to select from the buffer 410 according to whether the calculation phase or the execution phase is active. In the calculation phase, the control signal 415 controls the multiplexer 405 to select the sign bit from the first storage element in the buffer 410. The sign bit is then used by the logic circuit 425 during the calculation phase to form an activation bit signal (Act) and a complement activation bit signal ActB, which control the open / closed state of the L switch (in the transmission gate implementation). The inverter 420 conceptually represents that the switching state of the R switch is complementary to the switching state of the L switch. As discussed with respect to the bit cell 200, the Act and ActB activation bit signals control the switching state of the R switch in addition to controlling the switching state of the L switch.
[0049] In the execution phase after the calculation phase, the control signal 415 controls the selection of activation bits from multiple additional storage elements in the buffer 410 according to the amplitude (bit significance) of the current calculation. For example, the first execution cycle can start from the LSB activation bit M0. In the continuous execution cycle, the next most effective activation bit is selected. In the buffer 410, the range of the activation bit is from the LSB activation bit M0 to the MSB activation bit M6. However, it should be understood that alternative arrangements of bits can be used, such as selecting from MSB to LSB in other implementations. Therefore, the multiplication of this signed seven-bit wide activation vector involves seven consecutive calculation and execution phases, each execution phase for a corresponding activation bit, and each calculation phase responds to the same sign bit. Depending on the binary value of the selected activation bit, the logic circuit 425 either reverses the open / closed switch state of the L switch and the R switch in the execution phase from its state in the calculation phase, or keeps them unchanged, as discussed with respect to Table 2. After the execution stage, a ternary-based multiplication of the signed activation bit and the stored filter weight bits is sensed from the read bit line voltage, such as by an analog-to-digital converter, as will be discussed further herein.
[0050] about Figure 5The advantageous power consumption reduction of ternary-based computations can be better understood by looking at example switching waveforms for four exemplary computation and execution cycles of FIG. 500. In the first set of waveforms 500, the filter weight bit is a binary 1. In the second set of waveforms 505, the filter weight bit is a binary 0. For consecutive computation and execution phases, the on and off states of pre-charge transistor P1 are common to both waveforms 500 and 505. In these waveforms, the binary state of the bottom plate of capacitor C (charged to supply voltage VDD or ground) is represented as Cbot. The bottom plate is the plate of capacitor C connected to output node 105. In contrast, the plate of capacitor C connected to the read bit line can be represented as the upper plate.
[0051] As previously described, the sign bit is multiplied with the activation bit to form a signed activation bit, resulting in one of four possible values: +1, +0, -0, and -1. Waveform 500 begins in a calculation phase for a -1 value. In this calculation phase, due to the negative sign of the -1 activation, the R switch is turned on and the L switch is turned off. Since the filter weight bit is binary 1, the complement filter weight bit is binary 0. This binary 0 is conducted through the turned-on R switch to ground the Cbot plate. In the execution phase, the binary 1 amplitude of the activation bit forces the switching states of L and R to reverse. Therefore, for the -1 execution phase, the L switch is turned on and the R switch is turned off. The turning on of the L switch allows the binary 1 value of the filter weight bit to charge the Cbot plate of capacitor C to the supply voltage VDD.
[0052] The -1 activation is followed by the +0 activation value of the 500 waveform. Therefore, for the +0 activation, the L switch remains on during the calculation phase, while the R switch remains off. Therefore, during this calculation phase, the Cbot voltage remains charged to the supply voltage VDD. In the subsequent execution phase for the +0 activation, because the amplitude of the +0 activation is 0, the switch state does not change. Therefore, the Cbot voltage remains charged to the supply voltage VDD. Compared with the conventional XNOR-based method where the Cbot voltage is always grounded during the accumulation phase, this means reduced power consumption.
[0053] +1 activation follows +0 activation. Thus, in waveform 500, the L switch is turned on during the calculation phase for +1 activation, while the R switch is turned off. The turning on of the L switch allows the binary 1 value of the filter weight bit to continue charging the Cbot voltage to the supply voltage VDD. In the subsequent execution phase, the switching states of the L switch and the R switch are reversed so that the Cbot voltage is grounded due to the turning on of the R switch.
[0054] A -0 activation is followed by a +1 activation (note that the order of activations depends on the activation vector being processed, waveforms 500 and 505 use a specific activation order in order to show all possible activation values). Since the activations have a negative sign, during the calculation phase for the -0 activation in waveform 500, the R switch is closed and the L switch is open to continue grounding the Cbot voltage. In the subsequent execution phase, due to the binary 0 of the activation magnitude, the R switch remains closed and the L switch remains open. Therefore, the Cbot voltage remains unchanged during the calculation and execution cycle for the -0 activation, which also reduces power consumption compared to the changing Cbot voltage that may occur when performing conventional XNOR-based calculations.
[0055] Waveform 505 will now be discussed. As previously discussed, for waveform 505, the filter weight bits are binary 0. The R and L switch states (on or off) will be as discussed for waveform 500 because these switch states depend only on activation. For +0 activation, waveform 505 represents power savings relative to conventional XNOR-based methods because the Cbot voltage remains grounded. In particular, because the L switch will be turned on, the Cbot voltage is grounded during the calculation phase for +0 activation, which allows the grounded filter weight bits to flow through the L switch to ground the lower plate of capacitor C. This grounded state of the Cbot voltage remains unchanged during the execution phase for +0 activation because the binary 0 amplitude keeps the switch states of the L switch and the R switch unchanged from the calculation phase values. In addition, the -0 activation for the 505 waveform also represents power savings relative to conventional XNOR-based methods. In particular, during the calculation phase for -0 activation, the negative value of the -0 activation turns the R switch on and turns the L switch off, which allows the binary high value of the complement filter weight bit to flow through the closed R switch and charge the Cbot voltage to the supply voltage VDD. Then, due to the binary 0 amplitude of the -0 activation, the switch state remains unchanged during the subsequent execution phase for -0 activation in waveform 505. Therefore, during the execution phase for -0 activation, the Cbot voltage remains charged to the supply voltage. In contrast, the Cbot voltage would be grounded in an XNOR-based calculation.
[0056] Some example CiM bit cell arrays
[0057] CiM bit cells configured for ternary-based computations as disclosed herein may be organized to form multiply and accumulate (MAC) circuits. Figure 6. MAC circuit 600 includes a plurality of CiM bit cells, each CiM bit cell being implemented as discussed with respect to CiM bit cells 100, 200, or 300. In general, the number of bit cells included in MAC circuit 600 will depend on the filter size. For clarity of illustration, MAC circuit 600 is shown as including columns of only seven CiM bit cells, ranging from the zeroth bit cell storing the zeroth filter weight bit W0 to the sixth bit cell storing the sixth filter weight bit W6. The read bit line RBL extends across the columns. During ternary-based computations, each bit cell operates as discussed with respect to bit cells 100, 200, or 300, as discussed with respect to Figure 4 , Figure 5 and discussed in Table 2.
[0058] Multiple MAC circuits may be arranged to form a memory including the memory array 700, such as Figure 7 As shown in . Each column of bit cells 100, 200 or 300 forms a corresponding MAC circuit. For example, the filter size in array 700 is 128, so that each column in array 700 has 128 bit cells 100, 200 or 300. Therefore, activation vector 720 will have 128 input samples. In memory array 700, each input sample is a multi-bit input sample. For any given calculation and execution stage, an activation bit is selected from each multi-bit sample to produce a plurality of activation bits ranging from the first activation bit din1 to the 128th activation bit din128. The sign bit of each activation vector 720 is not illustrated, but will be included as discussed with respect to buffer 410. After the write operation of writing the filter weight bits to memory array 700, each activation vector 720 is sampled sequentially so that each MAC circuit performs a calculation phase in which the corresponding activation bit is multiplied by the corresponding filter weight bit. The calculation phase is followed by an execution phase. Note that in the XNOR-based approach, the execution phase is represented as an accumulation phase because the output node 105 of each bit cell in the MAC circuit is grounded. The charge from each capacitor C in the MAC circuit is therefore accumulated onto the corresponding read bit line. But in ternary-based computation, because each output node 105 can be grounded or remain charged, the charge on the capacitor C is not accumulated in the same manner during the execution phase, as previously discussed. However, the execution phase achieves the same goal because the read bit line voltage will represent the sum (accumulation) of the computation results of all CiM bit cells within the MAC circuit. But unlike XNOR-based computation, note that the activation vector 720 can have both odd and even signed values, so the computation accuracy is increased. In addition, the charging and discharging of the lower plate of capacitor C is reduced, as discussed with respect to Figure 5As indicated. Note that each input sample (such as din1) can be a multi-bit input sample. For example, din1 can be a three-bit wide sample din1. Since each CiM bit cell performs binary multiplication, the individual bits in the multi-bit input sample are processed sequentially by each MAC circuit in the array 700. Therefore, the sequential integrator 705 of each MAC circuit is used to weight the accumulated result according to the weight (bit effective value) of the multi-bit input sample. For example, assume that each sample of the input vector 720 is a three-bit wide sample, ranging from the least significant bit (LSB) sample to the most significant bit (MSB) sample. Therefore, each sequential integrator 705 sums the accumulated result according to the bit effective value of the accumulated result. In addition, the filter weight itself can be a multi-bit filter weight. Since each CiM bit cell stores a binary filter weight, one MAC circuit can be used for one filter weight bit (e.g., LSB weight), and the adjacent MAC circuit can be used for the next most significant filter weight bit, and so on. In this embodiment, three adjacent MAC circuits will be used for an embodiment of a three-bit wide filter weight. The multi-bit weight summing circuit 710 accumulates the corresponding MAC accumulated values (in the case of multi-bit input samples, as required, processed by the corresponding sequential integrator 705) and sums the MAC accumulated values according to the binary weights of the filter weight bits. Finally, the analog-to-digital converter (ADC) 715 digitizes the final accumulated result. However, this digitization is significantly improved due to the doubling of the output voltage swing on the read bit line provided by the ternary activation as described above.
[0059] Now refer to Figure 8The flowchart of discusses a method of controlling a CiM bit cell to multiply a storage bit with a signed input bit. The method includes action 800, which occurs during a first phase of operation and in response to the sign of the signed input bit having a first binary value, and includes: closing a first switch coupled between a node for storing the bit and an output node, and opening a second switch coupled between a node for storing the complement of the bit and the output node. Closing the L switch and opening the R switch in response to activation being positive is an example of action 800. The method also includes action 805, which occurs during the first phase of operation and in response to the sign of the signed bit having a second binary value, and includes opening the first switch and closing the second switch. Opening the L switch and closing the R switch in response to activation being negative is an example of action 805. The method also includes action 810, which occurs during a second phase of operation and in response to the magnitude of the signed input bit having a first binary value, and includes reversing the switch states of the first switch and the second switch established during the first phase of operation. Switching the switch states of the L switch and the R switch in response to the magnitude of the activation being a binary 1 is an example of action 810. Finally, the method includes action 815, which occurs during the second phase of operation and in response to the magnitude of the signed input bit having a second binary value, and includes maintaining the switch states of the first switch and the second switch established during the first phase of operation. During the execution phase, controlling the L switch and the R switch to have the same switch states as established in the calculation phase in response to the magnitude of the activation being a binary 0 is an example of action 815.
[0060] The in-memory computation bit cell with ternary activation as disclosed herein may be advantageously incorporated into any suitable mobile device or electronic system. Fig. 9 As shown in , according to the present disclosure, a cell phone 900, a laptop computer 905, and a tablet PC 910 can all include in-memory computing with in-memory computing bit cells such as for machine learning applications. Other exemplary electronic systems (such as music players, video players, communication devices, and personal computers) can also be configured with in-memory computing constructed according to the present disclosure.
[0061] The present disclosure will now be summarized in the following series of example clauses:
[0062] Clause 1. A memory comprising:
[0063] a bit cell having a switch pair connected to an output node;
[0064] a capacitor coupled to the output node;
[0065] a first storage element and a plurality of additional storage elements; and
[0066] A controller is configured to select a first bit from the first storage element to control the switch pair in response to the first bit during a first phase of operation of the memory, and is configured to select a second bit from the plurality of additional storage elements to control the switch pair in response to the second bit during a second phase of operation of the memory.
[0067] Clause 2. The memory of clause 1, further comprising:
[0068] A bit line is provided wherein the capacitor includes a first terminal coupled to the output node and a second terminal coupled to the bit line.
[0069] Clause 3. A memory as described in any of Clauses 1-2, wherein the first storage element is configured to store a sign bit and the plurality of additional storage elements are configured to store a plurality of magnitude bits.
[0070] Clause 4. The memory of clause 3, wherein the controller comprises a multiplexer configured to select the sign bit during the first phase of operation and to select a second bit from the plurality of amplitude bits during the second phase of operation.
[0071] Item 5. A memory according to item 4, wherein the controller is further configured to: during the first stage of operation, in response to the first bit having a first binary value, close the first switch of the switch pair and open the second switch of the switch pair, and in response to the second bit having a second binary value, open the first switch and close the second switch, wherein the second binary value is the complement of the first binary value.
[0072] Item 6. A memory according to item 5, wherein the controller is further configured to: during the second phase of operation, in response to the second bit having the first binary value, reverse the switching state of the first switch and the second switch, and in response to the second bit having the second binary value, maintain the switching state of the first switch and the second switch.
[0073] Clause 7. A memory according to any one of clauses 2-6, wherein the bit cell includes a first inverter, the first inverter is cross-coupled with a second inverter, and wherein the switch pair includes a first switch coupled between an output node of the first inverter and the output node, and includes a second switch coupled between an output node of the second inverter and the output node.
[0074] Clause 8. The memory of Clause 7, wherein the first switch comprises a first transmission gate, and wherein the second switch comprises a second transmission gate.
[0075] Clause 9. The memory of any of Clauses 7-8, wherein the first switch and the second switch are the only switches coupled to the output node.
[0076] Clause 10. The memory of Clause 7, further comprising a third switch coupled between the output node and ground.
[0077] Item 11. A memory according to item 10, wherein the controller is further configured to: turn on the third switch during a pre-charge phase of operation before the first phase of operation, and turn off the third switch during the first phase of operation and during the second phase of operation.
[0078] Clause 12. A memory as described in any of Clauses 10-11, wherein the first switch is a first p-type metal oxide semiconductor (PMOS) transistor, the second switch is a second PMOS transistor, and the third switch is an n-type metal oxide semiconductor (NMOS) transistor.
[0079] Item 13. The memory of Item 7 further comprising: a third switch coupled between a node for a common mode voltage and the bit line, wherein the controller is further configured to: close the third switch during the first phase of operation, and open the third switch during the second phase of operation.
[0080] Clause 14. The memory of Clause 13, wherein the third switch is a first PMOS transistor.
[0081] Clause 15. A memory as described in any of Clauses 1-14, wherein the memory is incorporated in a cellular telephone.
[0082] Clause 16. A method of controlling a bit cell to multiply a stored bit with a signed input bit, comprising:
[0083] during a first phase of operation and in response to the sign of the signed input bit having a first binary value, closing a first switch coupled between a node for the storage bit and an output node, and opening a second switch coupled between a node for the complement of the storage bit and the output node;
[0084] during said first phase of operation and in response to said sign of said signed bit having a second binary value, opening said first switch and closing said second switch;
[0085] During a second phase of operation and in response to the magnitude of the signed input bit having the first binary value, reversing the switch states of the first switch and the second switch established during the first phase of operation; and
[0086] During the second phase of operation and in response to the magnitude of the signed input bit having the second binary value, the switch states of the first switch and the second switch established during the first phase of operation are maintained unchanged.
[0087] Clause 17. The method of clause 16, wherein the first binary value is a binary 1 value, and wherein the second binary value is a binary 0 value.
[0088] Clause 18. A method according to any one of clauses 16-17, further comprising:
[0089] During the first phase of operation, a bit line is connected to a node for a common mode voltage, wherein the bit line is coupled to the output node through a capacitor.
[0090] Clause 19. The method according to clause 18, further comprising:
[0091] During the second phase of operation, the bit line is disconnected from the node for the common mode voltage.
[0092] Clause 20. A memory comprising:
[0093] a bit cell configured to store a storage bit, the bit cell comprising a first switch coupled between a node for the storage bit and an output node and a second switch coupled between a node for a complement of the storage bit and the output node;
[0094] a capacitor having a first plate connected to the output node; and
[0095] A controller is configured to: in a first stage of operation, in response to a sign bit of an input word having a first binary value, open the second switch and close the first switch, and in response to the sign bit having a second binary value, open the first switch and close the second switch to control the switching states of the first switch and the second switch, wherein the second binary value is a complement of the first binary value.
[0096] Clause 21. The memory of clause 20, further comprising:
[0097] A bit line is coupled to the second plate of the capacitor.
[0098] Clause 22. The memory of any one of Clauses 20-21, further comprising:
[0099] An input buffer for storing the input word, wherein the controller is further configured to: during a second phase of operation, in response to a selected amplitude bit in the input buffer having the first binary value, invert the switch states of the first switch and the second switch.
[0100] Clause 23. A memory according to clause 22, wherein the controller is further configured to: during the second phase of operation, in response to the selected amplitude bit in the input buffer having the second binary value, maintain the switch state of the first switch and the second switch.
[0101] Clause 24. A memory as described in any of Clauses 20-23, wherein the memory is included in a multiplication and accumulation circuit, the multiplication and accumulation circuit including a plurality of additional bit cells, each additional bit cell including a corresponding capacitor.
[0102] Clause 25. The memory of clause 24, further comprising a memory array comprising a plurality of columns, and wherein the multiply and accumulate circuit is configured to form a column of the plurality of columns.
[0103] Clause 26. The memory of clause 25, further comprising:
[0104] A plurality of analog-to-digital converters are provided, the plurality of analog-to-digital converters corresponding one-to-one to the plurality of columns.
[0105] Clause 27. The memory of clause 26, wherein each analog-to-digital converter is a multi-bit analog-to-digital converter.
[0106] Clause 28. A method for performing in-memory computation operations, comprising:
[0107] During a first phase of operation, controlling a pair of switches coupled between a pair of inverters in the bit cell and a plate of a capacitor in response to a sign bit; and
[0108] During a second phase of operation, the switch pair is controlled in response to the amplitude bit.
[0109] Clause 29. The method of clause 28, wherein controlling the switch pair during the first phase of operation comprises: closing a first switch of the switch pair and opening a second switch of the switch pair in response to the sign bit having a first binary value.
[0110] Clause 30. The method of clause 29, wherein controlling the switch pair during the first phase of operation further comprises opening the first switch and closing the second switch in response to the sign bit having a second binary value, the second binary value being the complement of the first binary value.
[0111] It should be understood that many modifications, substitutions and changes may be made to the materials, devices, configurations and methods of use of the equipment disclosed herein without departing from the scope of the present disclosure. In view of this, the scope of the present disclosure should not be limited to the scope of the specific embodiments illustrated and described herein (because they are only examples thereof), but should be fully commensurate with the scope of the claims appended hereto and their functional equivalents.
Claims
1. A memory, comprising: a bit cell having a switch pair connected to an output node; a capacitor coupled to the output node; a first storage element and a plurality of additional storage elements; as well as a controller configured to select a first bit from the first storage element to control the switch pair in response to the first bit during a first phase of operation of the memory, and to select a second bit from the plurality of additional storage elements to control the switch pair in response to the second bit during a second phase of operation of the memory, wherein the controller is further configured to: during the first phase of operation, in response to the first bit having a first binary value, close a first switch of the switch pair and open a second switch of the switch pair, and in response to the first bit having a second binary value, open the first switch and close the second switch, the second binary value being the complement of the first binary value, and The controller is further configured to: during the second phase of operation, in response to the second bit having the first binary value, invert the switch states of the first switch and the second switch, and in response to the second bit having the second binary value, maintain the switch states of the first switch and the second switch.
2. The memory according to claim 1, further comprising: A bit line is provided wherein the capacitor includes a first terminal coupled to the output node and a second terminal coupled to the bit line.
3. The memory of claim 1, wherein the first storage element is configured to store a sign bit and the plurality of additional storage elements are configured to store a plurality of magnitude bits. 4 . The memory of claim 3 , wherein the controller comprises a multiplexer configured to select the first bit during the first phase of operation and to select the second bit during the second phase of operation.
5. The memory of claim 1 , wherein the bit cell comprises a first inverter cross-coupled with a second inverter, and wherein the switch pair comprises a first switch coupled between an output node of the first inverter and the output node, and comprises a second switch coupled between an output node of the second inverter and the output node.
6. The memory of claim 5, wherein the first switch comprises a first transmission gate, and wherein the second switch comprises a second transmission gate.
7. The memory of claim 5, wherein the first switch and the second switch are the only switches coupled to the output node.
8. The memory according to claim 5, further comprising: bit line; as well as A third switch is coupled between a node for a common mode voltage and the bit line, wherein the controller is further configured to close the third switch during the first phase of operation and to open the third switch during the second phase of operation.
9. The memory of claim 8, wherein the third switch is a first PMOS transistor.
10. The memory of claim 1, wherein the memory is incorporated in a cellular telephone.
11. A memory comprising: a bit cell having a switch pair connected to an output node, wherein the bit cell includes a first inverter cross-coupled with a second inverter, and wherein the switch pair includes a first switch coupled between an output node of the first inverter and the output node, and includes a second switch coupled between an output node of the second inverter and the output node; a capacitor coupled to the output node; a first storage element and a plurality of additional storage elements; as well as a controller configured to select a first bit from the first storage element to control the switch pair in response to the first bit during a first phase of operation of the memory, and to select a second bit from the plurality of additional storage elements to control the switch pair in response to the second bit during a second phase of operation of the memory, and A third switch is coupled between the output node and ground.
12. The memory of claim 11, wherein the controller is further configured to: turn on the third switch during a precharge phase of operation prior to the first phase of operation, and turn off the third switch during the first phase of operation and during the second phase of operation.
13. The memory of claim 11, wherein the first switch is a first p-type metal oxide semiconductor (PMOS) transistor, the second switch is a second PMOS transistor, and the third switch is an n-type metal oxide semiconductor (NMOS) transistor.
14. A method of controlling a bit cell to multiply a stored bit with a signed input bit, comprising: during a first phase of operation and in response to the sign of the signed input bit having a first binary value, closing a first switch coupled between a node for the storage bit and an output node, and opening a second switch coupled between a node for the complement of the storage bit and the output node; during said first phase of operation and in response to said sign of said signed input bit having a second binary value, opening said first switch and closing said second switch; during a second phase of operation and in response to the magnitude of the signed input bit having the first binary value, reversing the switch states of the first switch and the second switch established during the first phase of operation; as well as During the second phase of operation and in response to the magnitude of the signed input bit having the second binary value, the switch states of the first switch and the second switch established during the first phase of operation are maintained.
15. The method of claim 14, wherein the first binary value is a binary 1 value, and wherein the second binary value is a binary 0 value.
16. The method according to claim 14, further comprising: During the first phase of operation, a bit line is connected to a node for a common mode voltage, wherein the bit line is coupled to the output node through a capacitor.
17. The method according to claim 16, further comprising: During the second phase of operation, the bit line is disconnected from the node for the common mode voltage.
18. A memory comprising: a bit cell configured to store a storage bit, the bit cell comprising a first switch coupled between a node for the storage bit and an output node and a second switch coupled between a node for a complement of the storage bit and the output node; a capacitor having a first plate connected to the output node; a controller configured to: during a first phase of operation, in response to a sign of an input word having a first binary value, open the second switch and close the first switch, and in response to the sign of the input word having a second binary value, open the first switch and close the second switch to control switch states of the first switch and the second switch, wherein the second binary value is a complement of the first binary value; as well as An input buffer for storing the input word, wherein the controller is further configured to: during a second phase of operation, in response to a selected amplitude bit in the input buffer having the first binary value, invert the switch states of the first switch and the second switch.
19. The memory according to claim 18, further comprising: A bit line is coupled to the second plate of the capacitor.
20. The memory of claim 18, wherein the controller is further configured to: during the second phase of operation, in response to the selected amplitude bit in the input buffer having the second binary value, maintain the switch states of the first switch and the second switch.
21. The memory of claim 20, wherein the memory is included in a multiply and accumulate circuit, the multiply and accumulate circuit comprising a plurality of additional bit cells, each additional bit cell comprising a corresponding capacitor.
22. The memory of claim 21 further comprising a memory array comprising a plurality of columns, and wherein the multiply and accumulate circuit is configured to form a column of the plurality of columns.
23. The memory according to claim 22, further comprising: A plurality of analog-to-digital converters are provided, the plurality of analog-to-digital converters corresponding one-to-one to the plurality of columns.
24. The memory of claim 23, wherein each analog-to-digital converter is a multi-bit analog-to-digital converter.
25. A method for performing in-memory computing operations, comprising: During a first phase of operation, a switch pair coupled between the bit cell and a plate of a capacitor is controlled in response to a sign bit; during a second phase of operation, controlling the switch pair in response to the amplitude bit; wherein controlling the switch pair during the first phase of operation comprises: closing a first switch of the switch pair and opening a second switch of the switch pair in response to the sign bit having a first binary value; and Wherein controlling the switch pair during the first phase of operation further comprises opening the first switch and closing the second switch in response to the sign bit having a second binary value, the second binary value being the complement of the first binary value.
Citation Information
Patent Citations
Realization of neuronal networks with ternary inputs and binary weights in NAND storage arrays
CN110751276A
Voltage accumulation in-memory calculation circuit based on SRAM bit line XNOR
CN111816234A