In-memory calculation circuit and in-memory calculation method based on 10T-SRAM (10 T-Static Random Access Memory)
Through an in-memory computing circuit based on 10T-SRAM, the calculation bit lines are divided into PBL and NBL, and combined with the current mirror structure and pulse width encoding, the full parallel multiplication and accumulation operation of multi-bit input value and weight value is realized, solving the problem of restricted discharge margin and inability to turn on in parallel in existing in-memory computing circuits, and improving the computing efficiency.
Patent Information
- Application Number
- CN202510518680.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-08-08
AI Technical Summary
The existing in-memory computing circuit has the problem of limited discharge margin and the calculation cannot be turned on in parallel, which limits the calculation efficiency and the implementation of parallel computing.
Using an in-memory calculation circuit based on 10T-SRAM, the multiplication and accumulation operation of multi-bit input value and weight value is realized by dividing the calculation bit lines into PBL and NBL, and the consistency of charge and discharge amount is controlled by the current mirror structure, and the charge and discharge time is controlled through pulse width encoding to achieve full parallel calculation.
The voltage fluctuation range on the bit line is improved, the margin of full swing is achieved, the problem of limited discharge margin is solved, and full parallel calculation and efficient multiplication and accumulation operations are realized, and the calculation efficiency is significantly improved.
Smart Images

Figure CN120448338A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an in-memory computing circuit in the technical field of integrated circuits, in particular to an in-memory computing circuit based on 10T-SRAM, and also to an in-memory computing method based on 10T-SRAM. Background Art
[0002] Compute-In-Memory (CIM) is an emerging computing architecture that aims to tightly integrate data storage and computation, thereby reducing data transmission latency and improving computing efficiency. By embedding computational operations directly into the memory unit, CIM technology eliminates the frequent data transfers between memory and the central processing unit (CPU) in traditional von Neumann architectures. This has important applications in large-scale data processing, artificial intelligence, deep learning, and other fields.
[0003] The amount of data used in existing neural network calculations is huge, so fully parallel in-memory calculation modules have emerged. The existing Chinese invention patent with publication number CN117219140A discloses an in-memory calculation circuit based on 8T-SRAM and current mirror, which includes a storage unit, an in-memory calculation unit, a transmission control unit, a current mirror unit, an inverter unit, and a shutdown control unit. On the one hand, this patent stores 1-bit weights in 8T SRAM cells, and on the other hand, divides 5-bit signed numbers into two parts: a 1-bit sign bit and a 4-bit unsigned number, and inputs them into the 8TSRAM cell and the transmission control unit respectively, thereby realizing the multiplication and XOR accumulation of 5-bit signed numbers and 1-bit weights in a near-memory calculation mode. However, this patent only uses one discharge bit line, which limits the discharge margin. The corresponding functional units are arranged near the storage unit, resulting in the inability to start the calculation in parallel. Summary of the Invention
[0004] In order to solve the technical problems of limited discharge margin and inability to start calculation in parallel in existing in-memory calculation circuits, the present invention provides an in-memory calculation circuit and an in-memory calculation method based on 10T-SRAM.
[0005] The present invention is implemented by the following technical solution: an in-memory computing circuit based on 10T-SRAM, which includes an in-memory computing array composed of multiple rows and columns of storage computing units, each unit including:
[0006] The storage part includes NMOS transistors N1 to N4 and PMOS transistors P1 and P2. N1, N2, P1, and P2 are anti-phase cross-coupled to form a pair of storage nodes Q and QB. N3 and N4 are controlled by word lines WL to connect nodes Q and QB to bit lines BL / BLB respectively.
[0007] The calculation part includes NMOS transistors N5 and N6 and PMOS transistors P3 and P4. The gates of P3 and N5 are connected to input nodes VBP and VBN respectively, and the gates of N6 and P4 are connected to nodes Q and QB respectively. N5 and N6, and P3 and P4 form calculation paths connecting bit lines NBL and PBL respectively.
[0008] Among them, the same row stores multi-bit weights. During calculation, the in-memory calculation circuit selects the VBP / VBN input value bit according to the input sign bit, and accumulates the positive / negative results on the bit lines PBL / NBL respectively through column parallel multiplication and accumulation, and finally obtains the common-mode voltage output multiplication and accumulation calculation result through voltage sharing.
[0009] In the present invention, the calculation result of each unit will be reflected on the PBL or NBL. If the input value is a negative number, the NBL will be discharged. If the input value is a positive number, the PBL will be charged. The calculation results of a column will be accumulated on the PBL and NBL. Finally, the voltages on the PBL and NBL are shared to obtain the final common-mode voltage value, realizing the multiplication and accumulation operation of multi-bit input values and weight values. Since the calculation bit line is divided into the PBL and NBL, the voltage fluctuation range on the bit line can be greatly improved, and the full swing margin can be achieved. At the same time, the calculation is performed within the unit, which can realize fully parallel calculation, solving the technical problems of limited discharge margin and inability to start calculation in parallel in existing in-memory calculation circuits.
[0010] As a further improvement to the above scheme, in the Nth row, the gates of N3 and N4 are connected to the word line WL[N], the source of N3 is connected to the node Q[N], the source of N4 is connected to the node QB[N], the gates of P3 and N5 are connected to the input nodes VBP[N] and VBN[N] respectively, the sources of P3 and N5 are connected to a pair of high and low potentials VDD and VSS respectively, the drain of N5 is connected to the source of N6, and the drain of P3 is connected to the source of P4; in each column, the drain of N3 is connected to the bit line BL, the drain of N4 is connected to the bit line BLB, the drain of N6 is connected to the bit line NBL, and the drain of P4 is connected to the bit line PBL; wherein the number of columns of the in-memory calculation array is the same as the number of weight bits, and the number of rows is the same as the number of input value bits.
[0011] Furthermore, the in-memory computing circuit further includes:
[0012] Multiple control switch circuits, each corresponding to a plurality of rows of storage computing units; each control switch circuit includes switches S1 to S4; in the Nth row, one end of S1 is connected to a voltage V bp [N], and the other end is connected to the node VBP[N]; one end of S2 is connected to the high potential VDD, and the other end is connected to the node VBP[N]; one end of S3 is connected to the voltage V bn [N], and the other end is connected to the node VBN[N]; one end of S4 is connected to the low potential VSS, and the other end is connected to the node VBN[N].
[0013] Furthermore, the in-memory computing circuit further includes:
[0014] Current mirror circuit, which is used to provide voltage V to each control switch circuit bp [N] and V bn [N], so that the charge and discharge quantities on the bit lines NBL / PBL are consistent.
[0015] Furthermore, the current mirror circuit includes a main current mirror and multiple slave current mirrors corresponding to multiple control switch circuits respectively; the multiple slave current mirrors in each column share the corresponding main current mirror and output the same current to the storage and computing unit in the corresponding row.
[0016] Furthermore, each slave current mirror includes NMOS transistors N7, N8 and PMOS transistors P5, P6; the drain and gate of N7 are connected, and the source is grounded; the gate of N8 is connected to the gate of N7 and provides a voltage V bn [N], the source is grounded; the source of P5 is connected to the high potential VDD, the gate is connected to the output of the corresponding main current mirror, and the drain is connected to the drain of N7; the source of P6 is connected to the high potential VDD, the gate is connected to the drain of N8 and provides voltage V bp [N].
[0017] Furthermore, each main current mirror includes NMOS transistors N9, N10, PMOS transistor P7 and resistor R B ; The source of N9 is grounded, the gate and drain are connected to the gate of N10, and the drain is connected to R B The source of N10 is grounded, and the drain is connected to the drain of P7; the gate and drain of P7 are connected and serve as the output end of the main current mirror, and the source is connected to the high potential VDD; R B The other end is connected to the input voltage V B .
[0018] Furthermore, during calculation: when the input value is a positive number, S1 and S4 are closed, and S2 and S3 are opened; when the input value is a negative number, S2 and S3 are closed, and S1 and S4 are opened; the closing and opening times of S1 to S4 are determined based on the numerical bits of the multi-bit input value.
[0019] As a further improvement to the above scheme, pulse width coding is used to encode the numerical bits of the multi-bit input value, and the charge and discharge time is controlled by controlling the turn-on time of P3 and N5, and the charge and discharge on PBL and NBL are controlled by controlling the conduction and shutdown of P4 and N6. The same column is turned on for each multiplication and accumulation operation, positive values are accumulated on PBL, and negative values are accumulated on NBL. After the result of a multiplication and accumulation is completed, the voltages of PBL and NBL are shared to obtain a common-mode value; if the common-mode value is higher than the common-mode voltage obtained in the pre-charge stage of PBL and NBL, the result is a positive number, otherwise it is a negative number.
[0020] The present invention also provides an in-memory computing method based on 10T-SRAM, which is applied to any of the above-mentioned in-memory computing circuits based on 10T-SRAM, and comprises the following steps:
[0021] A storage mode is implemented by the storage part: the same row of storage calculation units respectively stores multi-bit weights;
[0022] The calculation mode is implemented by the calculation part: the VBP / VBN input value bit is selected according to the input sign bit, the positive / negative results are accumulated on the bit lines PBL / NBL respectively through column parallel multiplication and accumulation, and finally the common mode voltage output multiplication and accumulation calculation result is obtained through voltage sharing.
[0023] Compared with the existing in-memory computing circuit, the in-memory computing circuit and in-memory computing method based on 10T-SRAM of the present invention have the following beneficial effects:
[0024] 1. In this in-memory calculation circuit based on 10T-SRAM, the calculation result of each unit is reflected on the PBL or NBL. If the input value is a negative number, the NBL is discharged, and if the input value is a positive number, the PBL is charged. The calculation results of a column are accumulated on the PBL and NBL. Finally, the voltages on the PBL and NBL are shared to obtain the final common-mode voltage value, realizing the multiplication and accumulation operation of multi-bit input values and weight values. Since the calculation bit line is divided into the PBL and NBL, the voltage fluctuation range on the bit line can be greatly improved, achieving full swing margin, and solving the problem of limited discharge margin in existing in-memory calculation circuits.
[0025] 2. The in-memory computing circuit based on 10T-SRAM performs calculations within the unit, which can realize fully parallel computing. It also completes multiplication and accumulation calculations while reading data, with high computing efficiency, solving the technical problem that the existing in-memory computing circuit cannot be opened in parallel when performing calculations.
[0026] 3. The in-memory computing circuit based on 10T-SRAM uses a current mirror structure to effectively control the consistency of charging on the PBL and discharging on the NBL, and also controls the consistency of charging and discharging of computing units in the entire array. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 A circuit diagram of an in-memory computing array of an in-memory computing circuit based on 10T-SRAM according to embodiment 1 of the present invention;
[0028] Figure 2 Schematic diagram of the packaging of the in-memory computing circuit based on 10T-SRAM according to Example 1 of the present invention;
[0029] Figure 3 for Figure 2 A circuit diagram of each control switch circuit of the in-memory computing circuit based on 10T-SRAM;
[0030] Figure 4 for Figure 2 The circuit diagram of the current mirror circuit of the in-memory computing circuit based on 10T-SRAM;
[0031] Figure 5 This is a diagram showing the voltage variation of NBL during simulation of the in-memory computing circuit based on 10T-SRAM according to Example 2 of the present invention;
[0032] Figure 6 FIG. 4 is a diagram showing the voltage variation of the PBL during simulation of the in-memory computing circuit based on 10T-SRAM according to Example 2 of the present invention. DETAILED DESCRIPTION
[0033] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0034] Example 1
[0035] See also Figure 1 and Figure 2 This embodiment provides an in-memory computing circuit based on 10T-SRAM. The in-memory computing circuit includes an in-memory computing array, multiple control switch circuits, and a current mirror circuit. In this embodiment, the in-memory computing array primarily performs multiplication and accumulation operations between multi-bit input values and multi-bit weights while storing them. The control switch circuit and current mirror circuit primarily provide a uniform current to the in-memory computing array, ensuring that the currents of each current mirror in the entire array are the same, thereby ensuring that the amount of charge and discharge on the computing bit lines is consistent.
[0036] The in-memory computation array consists of multiple rows and columns of storage and computation cells, which can be defined as a 10T-SRAM. In this embodiment, each storage and computation cell comprises a storage portion and a computation portion. These two portions must work together to implement both storage and computation modes and switch between them as needed. The storage portion includes NMOS transistors N1-N4 and PMOS transistors P1 and P2. N1, N2, P1, and P2 are anti-phase cross-coupled to form a pair of storage nodes Q and QB. N3 and N4, controlled by word line WL, connect nodes Q and QB to bit lines BL and BLB, respectively. The computation portion includes NMOS transistors N5 and N6 and PMOS transistors P3 and P4. The gates of P3 and N5 are connected to input nodes VBP and VBN, respectively, while the gates of N6 and P4 are connected to nodes Q and QB, respectively. N5 and N6, and P3 and P4, respectively, form computation paths connecting bit lines NBL and PBL. In this embodiment, for easy distinction, the NMOS transistors N1 to N6 and the PMOS transistors P1 to P4 in the Nth row are marked as N1[N], N2[N], N3[N], N4[N], N5[N], N6[N] and P1[N], P2[N], P3[N], P4[N], respectively, where N≥1 and is an integer.
[0037] In this embodiment, in row N, the gates of N3[N] and N4[N] are connected to word line WL[N], the source of N3[N] is connected to node Q[N], the source of N4[N] is connected to node QB[N], the gates of P3[N] and N5[N] are connected to input nodes VBP[N] and VBN[N], respectively, the sources of P3[N] and N5[N] are connected to a pair of high and low potentials VDD and VSS, respectively, the drain of N5[N] is connected to the source of N6[N], and the drain of P3[N] is connected to the source of P4[N]. In each column, the drains of all N3s are connected to bit line BL, the drains of all N4s are connected to bit line BLB, the drains of all N6s are connected to bit line NBL, and the drains of all P4s are connected to bit line PBL. The number of columns in the in-memory calculation array is the same as the number of weight bits, and the number of rows is the same as the number of input value bits.
[0038] Among them, the same row stores multi-bit weights. During calculation, the in-memory calculation circuit selects the VBP / VBN input value bit according to the input sign bit, and accumulates the positive / negative results on the bit lines PBL / NBL respectively through column parallel multiplication and accumulation, and finally obtains the common-mode voltage output multiplication and accumulation calculation result through voltage sharing.
[0039] In this embodiment, pulse-width coding is used to encode the numerical bits of a multi-bit input value. The charge and discharge times are controlled by controlling the on-time of P3 and N5, and the charge and discharge of the PBL and NBL are controlled by controlling the on / off state of P4 and N6. Each multiplication-accumulation operation activates the same column, with positive values accumulated on the PBL and negative values accumulated on the NBL. After a multiplication-accumulation operation, the voltages of the PBL and NBL are shared to obtain a common-mode value. If the common-mode value is higher than the common-mode voltage obtained during the precharge phase of the PBL and NBL, the result is positive; otherwise, it is negative. The common-mode voltage of the PBL and NBL is obtained during the precharge phase before calculation. For example, if VDD is 900mV and VSS is 0mV, the common-mode voltage is 450mV. If the calculated common-mode value is greater than 450mV, it is positive; otherwise, it is negative.
[0040] See also Figure 3 There are multiple control switch circuits, each corresponding to a plurality of rows of storage and calculation units. That is, each control switch circuit corresponds to a storage and calculation unit. Each control switch circuit includes switches S1 to S4. In the Nth row, one end of S1 is connected to a voltage V bp [N], and the other end is connected to the node VBP[N]. One end of S2 is connected to the high potential VDD, and the other end is connected to the node VBP[N]. One end of S3 is connected to the voltage V bn [N], and the other end is connected to node VBN[N]. One end of S4 is connected to the low potential VSS, and the other end is connected to node VBN[N]. Each cell has a control switch. The array has only one master current mirror to provide the reference current IB. The SRAM has N rows, and there are N slave current mirrors that copy the reference current IB. If the first column is calculated, the switch in the first column will introduce the Vbp or Vbn signal. At the same time, all switches in the remaining columns open S1 and S3, and close S2 and S4 to connect to VDD and VSS.
[0041] During calculations: When the input value is positive, S1 and S4 are closed, and S2 and S3 are open. When the input value is negative, S2 and S3 are closed, and S1 and S4 are open. The closing and opening times of S1 to S4 are determined based on the number of bits in the multi-bit input value. The times for these switches are encoded according to the number of bits in the input value, for example, "111" is 9t, and "010" is 2t. When no calculation is performed, S1 and S3 are open, and S2 and S3 are connected.
[0042] See also Figure 4 , the current mirror circuit is used to provide voltage V to each control switch circuit bp [N] and V bn[N], ensuring that the charge and discharge on the bit lines NBL / PBL are consistent. The current mirror circuit includes a master current mirror and multiple slave current mirrors corresponding to the multiple control switch circuits. The multiple slave current mirrors in each column share the corresponding master current mirror and output the same current to the storage and computing units in the corresponding row.
[0043] In this embodiment, each slave current mirror includes NMOS transistors N7 and N8 and PMOS transistors P5 and P6. The drain and gate of N7 are connected, and the source is grounded. The gate of N8 is connected to the gate of N7 and provides a voltage V bn [N], the source is grounded. The source of P5 is connected to the high potential VDD, the gate is connected to the output of the corresponding main current mirror, and the drain is connected to the drain of N7. The source of P6 is connected to the high potential VDD, the gate is connected to the drain of N8 and provides voltage V bp Similarly, for easy distinction, the NMOS transistors N7 and N8 and the PMOS transistors P5 and P6 in row N are marked as N7[N], N8[N], P5[N], and P6[N], respectively.
[0044] Each main current mirror includes NMOS transistors N9 and N10, PMOS transistor P7 and resistor R B The source of N9 is grounded, the gate and drain are connected to the gate of N10, and the drain is connected to R B The source of N10 is grounded, and the drain is connected to the drain of P7. The gate and drain of P7 are connected and serve as the output of the main current mirror, and the source is connected to the high potential VDD. B The other end is connected to the input voltage V B .
[0045] In summary, compared with the existing in-memory computing circuit, the in-memory computing circuit based on 10T-SRAM of this embodiment has the following beneficial effects:
[0046] 1. In this in-memory calculation circuit based on 10T-SRAM, the calculation result of each unit is reflected on the PBL or NBL. If the input value is a negative number, the NBL is discharged, and if the input value is a positive number, the PBL is charged. The calculation results of a column are accumulated on the PBL and NBL. Finally, the voltages on the PBL and NBL are shared to obtain the final common-mode voltage value, realizing the multiplication and accumulation operation of multi-bit input values and weight values. Since the calculation bit line is divided into the PBL and NBL, the voltage fluctuation range on the bit line can be greatly improved, achieving full swing margin, and solving the problem of limited discharge margin in existing in-memory calculation circuits.
[0047] 2. The in-memory computing circuit based on 10T-SRAM performs calculations within the unit, which can realize fully parallel computing. It also completes multiplication and accumulation calculations while reading data, with high computing efficiency, solving the technical problem that the existing in-memory computing circuit cannot be opened in parallel when performing calculations.
[0048] 3. The in-memory computing circuit based on 10T-SRAM uses a current mirror structure to effectively control the consistency of charging on the PBL and discharging on the NBL, and also controls the consistency of charging and discharging of computing units in the entire array.
[0049] Example 2
[0050] This embodiment provides an in-memory computing circuit based on 10T-SRAM, which provides an example of the number of rows and columns of the in-memory computing array based on Example 1. In this embodiment, the in-memory computing array includes 4 columns and N rows of storage computing units, N ≥ 1, and every 4 storage computing units in the same row are used to store 1 4-bit weight. N storage computing units in a column share a bit line BL, BLB, PBL, NBL, and the storage computing units in a row share a word line WL. The storage mode of the in-memory computing array is consistent with that of an ordinary 6T-SRAM cell.
[0051] In calculation mode, the data flow is as follows: the input value is 4 bits, with a 1-bit sign bit and 3-bit magnitude bits. The 1-bit sign determines whether the 3-bit input is placed on VBP or VBN. The 3-bit data is pulse-width encoded, for example, "111" is 9t and "010" is 2t. The charge and discharge times are controlled by controlling the on-time of P3 and N5. The weight is 4 bits and stored in the array. Each cell stores 1 bit of the weight. The charge and discharge of the PBL and NBL are controlled by turning on and off transistors P4 and N6. Each multiplication and accumulation operation activates the same column. Positive values are accumulated on the PBL, and negative values are accumulated on the NBL. With a supply voltage of 900mV, the common-mode voltage between the PBL and NBL is 450mV. After a multiplication and accumulation result, the voltages of the PBL and NBL are shared to obtain a common-mode value. Values above 450mV are positive, while values below 450mV are negative.
[0052] To ensure consistent charge and discharge between the PBL and NBL, a row-shared current mirror is used. Each row corresponds to a master current mirror and a slave current mirror. The current mirrors replicate the current, and all slave current mirrors in a row that share a master current mirror have the same current. A row shares a slave current mirror, so the cells in the same row have consistent charge and discharge. By setting the tube size of all slave current mirrors to the same as the single master current mirror, the reference current IB copied from each slave current mirror is identical. This means that when a column of cells is turned on for calculation, the charge and discharge of the cells in that column are consistent, and thus the charge and discharge of the cells in the entire array are consistent.
[0053] The input of the current mirror provides I B This reference current will then generate currents of the same magnitude in N9, N10, and P7, and the slave current mirror will copy the current in the master current mirror and the gate voltage. bp [1]~V bp The value of [N] is the same as the gate voltage of P7, V bn [1]~V bn The value of [N] is the same as the gate voltage of N10. V bp [1]~V bp [N] and V bn [1]~V bn [N] is the output of the current mirror part, which is connected to switches S1 and S3 respectively. When performing calculations, if the input value is positive, S1 and S4 are connected, and S2 and S3 are disconnected. If the input value is negative, S2 and S3 are connected, and S1 and S4 are disconnected. The time for connecting the switches above needs to be encoded according to the 3-bit numerical value of the input value, for example, "111" is 9t, and "010" is 2t. When no calculation is performed, S1 and S3 are disconnected, and S2 and S3 are connected.
[0054] Calculation process: Turn off N3 and N4 transistors to prevent the weight data stored in the unit from being disturbed. Precharge PBL to low potential VSS and NBL to high potential VDD. The current mirror gives the reference current I B , which in turn generates the output V bp and V bn The connection relationship between switches S1, S2, S3, and S4 is selected based on the 1-bit sign bit of the input data, and the switch connection time is determined based on the 3-bit value of the input data. The calculation result of each unit is reflected on the PBL or NBL. If the input value is negative, the NBL is discharged, and if the input value is positive, the PBL is charged. The calculation results of a column are accumulated on the PBL and NBL. Finally, the voltage on the PBL and NBL is shared to obtain the final common-mode voltage value, thereby obtaining the multiplication and accumulation calculation result.
[0055] Example 3
[0056] See also Figure 5 and Figure 6 This embodiment provides an in-memory computation circuit based on 10T-SRAM, which is simulated based on Example 2. The NBL is precharged to 900mV, begins computation at 0ps, and then discharges. When a storage computation unit is enabled and a 4-bit signed number is multiplied by a 1-bit weight, the NBL discharges.
[0057] When the 4-bit signed number is "1001" and the 1-bit weight is "1", the NBL discharges and the voltage drops to 892.97mV.
[0058] When the 4-bit signed number is "1010" and the 1-bit weight is "1", the NBL discharges and the voltage drops to 885.94mV.
[0059] When the 4-bit signed number is "1011" and the 1-bit weight is "1", the NBL discharges and the voltage drops to 878.91mV.
[0060] By turning on one column at a time, you can get the multiplication and accumulation results of the 4-bit input and the 1-bit weight. By turning on four columns in sequence, you can get all the multiplication and accumulation results of the 4-bit input and the 4-bit weight.
[0061] The PBL is precharged to 0mV and starts charging at 0ps. When a storage calculation unit is turned on, the PBL is charged when a 4-bit signed number is multiplied by a 1-bit weight.
[0062] When the 4-bit signed number is "0001" and the 1-bit weight is "1", the PBL is charged and the voltage rises to 7.02mV;
[0063] When the 4-bit signed number is "0010" and the 1-bit weight is "1", the PBL is charged and the voltage rises to 14.03mV;
[0064] When the 4-bit signed number is “0011” and the 1-bit weight is “1”, the PBL is charged and the voltage rises to 21.03mV.
[0065] By turning on one column at a time, you can get the multiplication and accumulation results of the 4-bit input and the 1-bit weight. By turning on four columns in sequence, you can get all the multiplication and accumulation results of the 4-bit input and the 4-bit weight.
[0066] Example 4
[0067] This embodiment provides an in-memory computing method based on 10T-SRAM, which is applied to the in-memory computing circuit based on 10T-SRAM in Embodiments 1 to 3. The in-memory computing method mainly includes the following steps.
[0068] The storage mode is implemented by the storage part: the same row of storage calculation units stores multi-bit weights respectively.
[0069] The calculation mode is implemented by the calculation part: the VBP / VBN input value bit is selected according to the input sign bit, and the positive / negative results are accumulated on the bit lines PBL / NBL respectively through column parallel multiplication and accumulation. Finally, the common-mode voltage output multiplication and accumulation calculation result is obtained through voltage sharing.
[0070] Example 5
[0071] This embodiment provides an in-memory computing chip (CIM chip), which includes the in-memory computing based on 10T-SRAM in Example 1 or Example 2. This circuit can be integrated on a chip. The CIM chip has a storage mode and a computing mode. In the storage mode, the CIM chip is used as a memory. In the computing mode, the CIM chip is used to implement multiplication and accumulation operations on multiple multi-bit input values and multi-bit weights.
[0072] Example 6
[0073] This embodiment provides a static random access memory (SRAM), which uses the in-memory calculation based on 10T-SRAM in Example 1 or Example 2 to implement multiplication and accumulation calculations of multi-bit inputs and multi-bit weights.
[0074] Based on the in-memory calculation based on 10T-SRAM in Example 1 or Example 2, the in-memory calculation of SRAM in this embodiment directly completes the multiplication and accumulation operation in the storage unit, reducing data movement and significantly reducing power consumption. SRAM can process the multiplication and accumulation operations of multiple inputs and weights at the same time, greatly improving computing efficiency, and is particularly suitable for application scenarios requiring high throughput. The read and write speed of SRAM is much higher than that of DRAM and flash memory, and can achieve low-latency multiplication and accumulation calculations, which is suitable for applications with high real-time requirements, such as edge computing and Internet of Things devices. SRAM can be integrated with other computing units (such as CPU, GPU) on the same chip to form an efficient storage and computing integrated architecture.
[0075] The SRAM of this embodiment is suitable for artificial intelligence and machine learning. The reasoning and training process of neural networks involves a large number of multiplication and accumulation operations. SRAM in-memory computing can significantly accelerate these operations and improve overall performance. Multi-bit input and weights enable SRAM to support everything from simple linear models to complex deep neural networks. The in-memory computing of the SRAM of this embodiment reduces the complex interface between memory and processor in traditional computing architectures, simplifying system design. By reducing data movement and simplifying the architecture, SRAM can reduce the overall cost and power consumption of the system.
[0076] Example 7
[0077] This embodiment provides an electronic device comprising a memory and a processor. The memory includes the 10T-SRAM-based in-memory computing described in Example 1 or Example 2. Compared to existing electronic devices, this electronic device can significantly improve computing efficiency, reduce power consumption, and support high-precision computing. It has broad application prospects in fields such as artificial intelligence and edge computing. Although it faces some technical challenges, its advantages make it an important technical direction for in-memory computing.
[0078] Example 8
[0079] This embodiment provides a computer device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor.
[0080] The computer device can take various forms, including embedded chips or modules, or general-purpose data processing devices, such as smart terminals that can execute programs, tablet computers, laptop computers, desktop computers, rack servers, blade servers, tower servers or cabinet servers (including independent servers or server clusters consisting of multiple servers), etc.
[0081] The computer device of this embodiment includes at least, but is not limited to, a memory and a processor that can be interconnected via a system bus. Memory (i.e., readable storage media) includes flash memory, a hard disk, a multimedia card, a card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic storage, a magnetic disk, an optical disk, and the like. In some embodiments, the memory may be an internal storage unit of the computer device, such as the hard disk or memory of the computer device.
[0082] In some embodiments, the processor may be a central processing unit (CPU), a graphics processing unit (GPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor is generally used to control the overall operation of the computer device. In this embodiment, the processor is used to run program code stored in the memory or process data.
[0083] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. An in-memory computing circuit based on 10T-SRAM, characterized in that: It includes an in-memory computing array consisting of multiple rows and columns of storage computing units, each unit including: The storage part includes NMOS transistors N1 to N4 and PMOS transistors P1 and P2. N1, N2, P1, and P2 are anti-phase cross-coupled to form a pair of storage nodes Q and QB. N3 and N4 are controlled by word lines WL to connect nodes Q and QB to bit lines BL / BLB respectively. The calculation part includes NMOS transistors N5 and N6 and PMOS transistors P3 and P4. The gates of P3 and N5 are connected to input nodes VBP and VBN respectively, and the gates of N6 and P4 are connected to nodes Q and QB respectively. N5 and N6, and P3 and P4 form calculation paths connecting bit lines NBL and PBL respectively. Among them, the same row stores multi-bit weights. During calculation, the in-memory calculation circuit selects the VBP / VBN input value bit according to the input sign bit, and accumulates the positive / negative results on the bit lines PBL / NBL respectively through column parallel multiplication and accumulation, and finally obtains the common-mode voltage output multiplication and accumulation calculation result through voltage sharing.
2. The in-memory computing circuit based on 10T-SRAM according to claim 1, characterized in that: In the Nth row, the gates of N3 and N4 are connected to the word line WL[N], the source of N3 is connected to the node Q[N], the source of N4 is connected to the node QB[N], the gates of P3 and N5 are connected to the input nodes VBP[N] and VBN[N] respectively, the sources of P3 and N5 are connected to a pair of high and low potentials VDD and VSS respectively, the drain of N5 is connected to the source of N6, and the drain of P3 is connected to the source of P4; in each column, the drain of N3 is connected to the bit line BL, the drain of N4 is connected to the bit line BLB, the drain of N6 is connected to the bit line NBL, and the drain of P4 is connected to the bit line PBL; wherein the number of columns of the in-memory calculation array is the same as the number of weight bits, and the number of rows is the same as the number of input value bits.
3. The in-memory computing circuit based on 10T-SRAM according to claim 2, characterized in that: The in-memory computing circuit further includes: Multiple control switch circuits, each corresponding to a plurality of rows of storage computing units; each control switch circuit includes switches S1 to S4; in the Nth row, one end of S1 is connected to a voltage V bp [N], and the other end is connected to the node VBP[N]; one end of S2 is connected to the high potential VDD, and the other end is connected to the node VBP[N]; one end of S3 is connected to the voltage V bn [N], and the other end is connected to the node VBN[N]; one end of S4 is connected to the low potential VSS, and the other end is connected to the node VBN[N].
4. The in-memory computing circuit based on 10T-SRAM according to claim 3, characterized in that: The in-memory computing circuit further includes: Current mirror circuit, which is used to provide voltage V to each control switch circuit bp [N] and V bn [N], so that the charge and discharge quantities on the bit lines NBL / PBL are consistent.
5. The in-memory computing circuit based on 10T-SRAM according to claim 4, characterized in that: The current mirror circuit includes a main current mirror and multiple slave current mirrors corresponding to multiple control switch circuits respectively; the multiple slave current mirrors in each column share the corresponding main current mirror and output the same current to the storage and calculation units in the corresponding row.
6. The in-memory computing circuit based on 10T-SRAM according to claim 5, characterized in that: Each slave current mirror includes NMOS transistors N7, N8 and PMOS transistors P5, P6; the drain and gate of N7 are connected, and the source is grounded; the gate of N8 is connected to the gate of N7 and provides a voltage V bn [N], the source is grounded; the source of P5 is connected to the high potential VDD, the gate is connected to the output of the corresponding main current mirror, and the drain is connected to the drain of N7; the source of P6 is connected to the high potential VDD, the gate is connected to the drain of N8 and provides voltage V bp [N].
7. The in-memory computing circuit based on 10T-SRAM according to claim 6, characterized in that: Each main current mirror includes NMOS transistors N9 and N10, PMOS transistor P7 and resistor R B ; The source of N9 is grounded, the gate and drain are connected to the gate of N10, and the drain is connected to R B The source of N10 is grounded, and the drain is connected to the drain of P7; the gate and drain of P7 are connected and serve as the output end of the main current mirror, and the source is connected to the high potential VDD; R B The other end is connected to the input voltage V B .
8. The in-memory computing circuit based on 10T-SRAM according to claim 7, characterized in that: During calculation: when the input value is positive, S1 and S4 are closed, and S2 and S3 are opened; when the input value is negative, S2 and S3 are closed, and S1 and S4 are opened; the closing and opening times of S1 to S4 are determined according to the numerical bits of the multi-bit input value.
9. The in-memory computing circuit based on 10T-SRAM according to claim 1, characterized in that: Pulse width coding is used to encode the numerical bits of the multi-bit input value, and the charge and discharge time is controlled by controlling the turn-on time of P3 and N5. The charge and discharge on PBL and NBL are controlled by controlling the conduction and shutdown of P4 and N6. The same column is turned on for each multiplication and accumulation operation, positive values are accumulated on PBL, and negative values are accumulated on NBL. After the result of a multiplication and accumulation is completed, the voltage of PBL and NBL is shared to obtain a common-mode value; if the common-mode value is higher than the common-mode voltage obtained in the pre-charge stage of PBL and NBL, the result is a positive number, otherwise it is a negative number.
10. An in-memory computing method based on 10T-SRAM, characterized in that: The method is applied to the in-memory computing circuit based on 10T-SRAM as claimed in any one of claims 1 to 9, and comprises the following steps: A storage mode is implemented by the storage part: the same row of storage calculation units respectively stores multi-bit weights; The calculation mode is implemented by the calculation part: the VBP / VBN input value bit is selected according to the input sign bit, the positive / negative results are accumulated on the bit lines PBL / NBL respectively through column parallel multiplication and accumulation, and finally the common mode voltage output multiplication and accumulation calculation result is obtained through voltage sharing.
Citation Information
Patent Citations
In-memory calculation circuit based on 8T-SRAM and current mirror
CN117219140A