A D6T in-memory computing accelerator with always linear discharge and reduced digital steps

Through the always linear discharge convolution mechanism of the D6T in-memory computing accelerator and the bias voltage time converter, the problems of read interference and complex step limitations of existing in-memory computing macro modules at low voltages are solved, and in-memory computing with high energy efficiency and high computing density are achieved.

CN115658009BActive Publication Date: 2025-08-22SHANGHAI TECH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211285251.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-20
Publication Date
2025-08-22
Estimated Expiration
2042-10-20

AI Technical Summary

Technical Problem

Existing in-memory computing macro modules have challenges in improving energy efficiency while maintaining high computation density, including the sensitivity of traditional 6T SRAM cells to read interference at low voltages, the need for pre-charge of bit lines to cause nonlinear currents, and digital implementation steps complexity limiting computation density and energy efficiency.

Method used

Using the D6T in-memory computing accelerator, the bit line voltage is reduced and linear calculation is maintained through an always linear discharge convolution mechanism and a bias voltage time converter, while reducing digital steps, and parallel processing is achieved using a decoupled read transistor and an independent gate signal line.

Benefits of technology

It achieves high energy efficiency and high calculation density, with an average energy efficiency of 8918TOPS/W and a calculation density of 38.6TOPS/mm2, supporting reliable operation and high-precision calculation at low voltage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115658009B_ABST
    Figure CN115658009B_ABST
Patent Text Reader

Abstract

The present invention discloses a D6T in-memory computing accelerator with always linear discharge and reduced digital steps. In the in-memory computing accelerator disclosed by the present invention, three effective technologies are proposed: (1) a decoupled 6T (D6T) bit cell that can operate reliably at 0.4V and standby at 0.26V, supporting parallel processing of decoupled dual ports; (2) an always linear discharge convolution mechanism (ALDCM), which can not only reduce the bit line voltage but also always maintain linear calculations over the entire voltage range of the bit line; (3) the bypass of the bias voltage time converter (BVTC) reduces the digital steps, but still maintains high energy efficiency and computing density at low voltage. The measurement results of the in-memory computing accelerator show that its average energy efficiency is 8918 TOPS / W (8b×8b) and the average computing density is 38.6TOPS / mm in 55nm CMOS process. 2 (8b×8b).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a D6T in-memory computing accelerator. Background Art

[0002] In-memory computing (IMC), as the name suggests, embeds computing units into memory[1-3]. Computers typically run on a von Neumann architecture consisting of two components: memory and computing units. If the memory transfer speed cannot keep up with the CPU's performance, computing power will be limited, resulting in a "memory wall." There are multiple technical approaches to in-memory computing based on different storage media, such as SRAM, RRAM, FeRAM, and other new types of memory[4-7]. In terms of hardware implementation, IMC is a promising computing platform for implementing neural networks for high-throughput artificial intelligence applications. Convolutional neural networks have been rapidly applied to various IMC tasks, including pedestrian detection, face recognition, object segmentation, and object tracking, with great success[8-11]. Recent work on SRAM-based IMC has achieved impressive throughput in both analog and digital domains[12-15]. Furthermore, work on eDRAM-based IMC[16,17] has achieved even higher throughput. In the short term, market opportunities for IMC will primarily arise in edge products, such as wearable devices and smart homes[18-21]. In the long run, as IMC computing power increases by 2-3 orders of magnitude, the application of in-memory computing will expand to more scenarios, such as autonomous driving, cloud computing, etc. [22-25].

[0003] Recent research [12–16] has demonstrated a variety of in-memory computing, mainly including digital [12, 13, 15] and analog [14, 16] implementations.

[0004] Figure 1 This figure shows a 64Kb macroblock with D6T as the basic bit unit. The 64Kb macroblock includes both digital and analog bypasses. We use a common digital bypass architecture as a reference, which includes a Reinforced Luminance (ReLU) sparsification circuit, a 5-bit SAR ADC, on-chip digital processing units (shift, addition, Boolean logic, etc.), and a multi-bit Directed Counting (DTC) circuit. Furthermore, a 64Kb 6T SRAM is used for data transmission and storage in the digital bypass. In contrast, the analog bypass design only has the overhead of a Back-Voltage Counting (BVTC). The 64Kb D6T array serves as the common area for both bypasses, enabling in-memory computation. This array has 256 WWL inputs, organized into 16 groups, to support simultaneous updates of multiple rows of weights. During IMC, 2×256 time-domain input channels (S1[n] and S2[n]) are provided in the row direction. 2×256 bit lines are provided in the column direction to accommodate a sufficient number of filters.

[0005] Existing in-memory computing macromodules [1-5] need to maintain high computing density while improving energy efficiency, which still faces some challenges: (1) Figure 2 As shown in , conventional 6T SRAM cells are sensitive to read disturbance at low voltages and cannot provide more parallelism for in-memory computing; (2) as Figure 3 As shown, for each convolution, the bit line needs to be precharged to a high voltage because the transistors involved in the multiplication and accumulation (MAC) need to be at a sufficient V ds The voltage produces a linear current, but as time goes by, the current still tends to be nonlinear; (3) Figure 4 As shown, in digitally implemented in-memory computing, overly complex steps limit computing density and energy efficiency, and lead to further degradation of its low-voltage computing performance.

[0006] References:

[0007] [1]W.-S.Khwa et al., "A 40-nm,2M-Cell,8b-Precision,Hybrid SLC-MLC PCMComputing-in-Memory Macro with 20.5-65.0TOPS / W for Tiny-Al Edge Devices," 2022IEEE International Solid-State Circuits Conference(ISSCC), 2022, pp.1-3, doi:10.1109 / ISSCC42614.2022.9731670.

[0008] [2]M.Chang et al., "A 40nm 60.64TOPS / W ECC-Capable Compute-in-Memory / Digital 2.25MB / 768KB RRAM / SRAM System with Embedded Cortex M3 Microprocessor for Edge Recommendation Systems," 2022 IEEE International Solid-State CircuitsConference(ISSCC), 2022, pp.1-3, doi:10.1109 / ISSCC42614.2022.9731679.

[0009] [3]D.Wang,C.-T.Lin,G.K.Chen,P.Knag,R.K.Krishnamurthy and M.Seok,"DIMC:2219TOPS / W 2569F2 / b Digital In-Memory Computing Macro in 28nm Based onApproximate Arithmetic Hardware,"2022IEEE International Solid-State CircuitsConference(ISSCC),2022,pp.266-268,doi:10.1109 / ISSCC42614.2022.9731659.

[0010] [4]S.D.Spetalnick et al.,"A 40nm 64kb 26.56TOPS / W 2.37Mb / mm2RRAMBinary / Compute-in-Memory Macro with 4.23x Improvement in Density and>75%Useof Sensing Dynamic Range,"2022IEEE International Solid-State CircuitsConference(ISSCC),2022,pp.1-3,doi:10.1109 / ISSCC42614.2022.9731725.

[0011] [5]M.Chang et al.,"A 40nm 60.64TOPS / W ECC-Capable Compute-in-Memory / Digital 2.25MB / 768KB RRAM / SRAM System with Embedded Cortex M3 Microprocessorfor Edge Recommendation Systems,"2022 IEEE International Solid-State CircuitsConference(ISSCC),2022,pp.1-3,doi:10.1109 / ISSCC42614.2022.9731679.

[0012] [6]Y.-C.Luo,J.Hur,Z.Wang,W.Shim,A.I.Khan and S.Yu,"A Technology Pathfor Scaling Embedded FeRAM to 28 nm and Beyond With 2T1C Structure,"in IEEETransactions on Electron Devices,vol.69,no.1,pp.109-114,Jan.2022,doi:10.1109 / TED.2021.3131108.

[0013] [7]T.Francois et al.,"High-Performance Operation and Solder ReflowCompatibility in BEOL-Integrated 16-kb HfO2:Si-Based 1T-1C FeRAM Arrays,"inIEEE Transactions on Electron Devices,vol.69,no.4,pp.2108-2114,April 2022,doi:10.1109 / TED.2021.3138360.

[0014] [8]Z.Shao,G.Cheng,J.Ma,Z.Wang,J.Wang and D.Li,"Real-Time and AccurateUAV Pedestrian Detection for Social Distancing Monitoring in COVID-19Pandemic,"in IEEE Transactions on Multimedia,vol.24,pp.2069-2083,2022,doi:10.1109 / TMM.2021.3075566.

[0015] [9]C.Fu,X.Wu,Y.Hu,H.Huang and R.He,"DVG-Face:Dual VariationalGeneration for Heterogeneous Face Recognition,"in IEEE Transactions onPattern Analysis and Machine Intelligence,vol.44,no.6,pp.2938-2952,1 June2022,doi:10.1109 / TPAMI.2021.3052549.

[0016]

[10] Y.Chen,L.Li,X.Liu and X.Su,"A Multi-Task Framework for InfraredSmall Target Detection and Segmentation,"in IEEE Transactions on Geoscienceand Remote Sensing,vol.60,pp.1-9,2022,Art no.5003109,doi:10.1109 / TGRS.2022.3195740.

[0017]

[11] B.Yan,E.Paolini,L.Xu and H.Lu,"A Target Detection and TrackingMethod for Multiple Radar Systems,"in IEEE Transactions on Geoscience andRemote Sensing,vol.60,pp.1-21,2022,Art no.5114721,doi:10.1109 / TGRS.2022.3183387.

[0018]

[12] Y.-D.Chih et al.,“An 89TOPS / W and 16.3TOPS / mm2 All-Digital SRAM-Based Full-Precision Compute-In Memory Macro in 22nm for Machine-LearningEdge Applications,”ISSCC,pp.252-253,2021.

[0019]

[13] B.Yan et al.,"A 1.041-Mb / mm2 27.38-TOPS / W Signed-INT8 Dynamic-Logic-Based ADC-less SRAM Compute-in-Memory Macro in 28nm with ReconfigurableBitwise Operation for AI and Embedded Applications,"2022 IEEE InternationalSolid-State Circuits Conference(ISSCC),2022,pp.188-190,doi:10.1109 / ISSCC42614.2022.9731545.

[0020]

[14] Q.Dong et al.,“A 351TOPS / W and 372.4GOPS Compute-in-Memory SRAMMacro in 7nm FinFET CMOS for Machine-Learning Applications,”ISSCC,pp.242-243,2020.

[0021]

[15] H.Fujiwara et al.,"A 5-nm 254-TOPS / W 221-TOPS / mm2 Fully-DigitalComputing-in-Memory Macro Supporting Wide-Range Dynamic-Voltage-FrequencyScaling and Simultaneous MAC and Write Operations,"2022 IEEE InternationalSolid-State Circuits Conference(ISSCC),2022,pp.1-3,doi:10.1109 / ISSCC42614.2022.9731754.

[0022]

[16] S.Xie,C.Ni,A.Sayal,P.Jain,F.Hamzaoglu and J.P.Kulkarni,"16.2eDRAM-CIM:Compute-In-Memory Design with Reconfigurable Embedded-Dynamic-Memory Array Realizing Adaptive Data Converters and Charge-Domain Computing,"2021 IEEE International Solid-State Circuits Conference(ISSCC),2021,pp.248-250,doi:10.1109 / ISSCC42613.2021.9365932.

[0023]

[17] Z.Chen et al.,"15.3 A 65nm 3T Dynamic Analog RAM-Based Computing-in-Memory Macro and CNN Accelerator with Retention Enhancement,AdaptiveAnalog Sparsity and 44TOPS / W System Energy Efficiency,"ISSCC,pp.240-242,2021.

[0024]

[18] N.Momeni,A.A.Valdés,J.Rodrigues,C.Sandi and D.Atienza,"CAFS:Cost-Aware Features Selection Method for Multimodal Stress Monitoring on WearableDevices,"in IEEE Transactions on Biomedical Engineering,vol.69,no.3,pp.1072-1084,March 2022,doi:10.1109 / TBME.2021.3113593.

[0025]

[19] R.Zanetti,A.Arza,A.Aminifar and D.Atienza,"Real-Time EEG-BasedCognitive Workload Monitoring on Wearable Devices,"in IEEE Transactions onBiomedical Engineering,vol.69,no.1,pp.265-277,Jan.2022,doi:10.1109 / TBME.2021.3092206.

[0026]

[20] A.Raj,M.Dubey,L.Gugnani and N.Gupta,"Synergizing Smart Home withSmart Parking using IOT,"2022 Second International Conference on ArtificialIntelligence and Smart Energy(ICAIS),2022,pp.1283-1286,doi:10.1109 / ICAIS53314.2022.9742975.

[0027]

[21] M.Rokonuzzaman,M.I.Akash,M.Khatun Mishu,W.-S.Tan,M.A.Hannan andN.Amin,"IoT-based Distribution and Control System for Smart HomeApplications,"2022IEEE 12th Symposium on Computer Applications&IndustrialElectronics(ISCAIE),2022,pp.95-98,doi:10.1109 / ISCAIE54458.2022.9794497.

[0028]

[22] P.Ghorai,A.Eskandarian,Y.-K.Kim and G.Mehr,"State Estimation andMotion Prediction of Vehicles and Vulnerable Road Users for CooperativeAutonomous Driving:A Survey,"in IEEE Transactions on IntelligentTransportation Systems,doi:10.1109 / TITS.2022.3160932.

[0029]

[23] D.Zhou,X.Song,J.Fang,Y.Dai,H.Li and L.Zhang,"Context-Aware 3DObject Detection From a Single Image in Autonomous Driving,"in IEEETransactions on Intelligent Transportation Systems,doi:10.1109 / TITS.2022.3154022.

[0030]

[24] S.Tuli,S.Ilager,K.Ramamohanarao and R.Buyya,"Dynamic Schedulingfor Stochastic Edge-Cloud Computing Environments Using A3C Learning andResidual Recurrent Neural Networks,"in IEEE Transactions on Mobile Computing,vol.21,no.3,pp.940-954,1March 2022,doi:10.1109 / TMC.2020.3017079.

[0031]

[25] MTIslam, S.Karunasekera and R.Buyya, "Performance and Cost-Efficient Spark Job Scheduling Based on Deep Reinforcement Learning in CloudComputing Environments," in IEEE Transactions on Parallel and DistributedSystems, vol.33, no.7, pp.1695-1710,1July 2022,doi:10.1109 / TPDS.2021.3124670. Summary of the Invention

[0032] The technical problem to be solved by the present invention is: the challenges faced by the existing in-memory computing macro modules pointed out in the background art.

[0033] To solve the above technical problems, the technical solution of the present invention is to provide a D6T in-memory computing accelerator that always performs linear discharge and reduces digital steps. The macro module of the D6T in-memory computing accelerator includes a digital bypass and an analog bypass. The D6T array composed of D6T bits serves as the common area of ​​the digital bypass and the analog bypass, realizing in-memory computing. The characteristics of the present invention are:

[0034] Each D6T bit supports both conventional memory mode and IMC mode, including: a write data transistor N1, used to write logic "0" and "1" to the storage node QB through the word line WWL and the bit line WBL; a transistor P1, used to maintain logic "1" so that the "1" of the storage node QB is directly connected to VDD; an inverter composed of a transistor P2 and a transistor N4, used to provide the gate voltage of the transistor P1 and to increase the parasitic capacitance of the storage node QB; decoupled read transistors N2 and N3, providing dual decoupling ports to the outside, and the read transistors N2 and N3 can also be used independently for convolution calculations; selection signal lines S1 and S2 for providing selection signals for selecting the read transistors N2 and N3; bit lines RBL1 and RBL2 for reading data or obtaining convolution results;

[0035] The bit line voltage is reduced by the convolution mechanism through always linear discharge, but it is kept linear over the entire voltage range of the bit line. Assuming that the longest pulse width of the time domain input is T, the macro module samples the convolution voltage at T / 2+δT, where δT is the ps-level delay: If the convolution voltage is <(VH-VL) / 2, VH represents the bit line precharge voltage (also the voltage value corresponding to the minimum convolution result), and VL represents the minimum voltage of the bit line discharged through the transistor (that is, the voltage value corresponding to the maximum convolution result), the voltage source is used to boost the capacitor top plate to >(VH-VL) / 2, so that the discharged bit lines RBL1 and RBL2 voltage return to a high voltage, maintaining sufficient V ds , so that the read transistors N2 and N3 performing convolution on the bit line maintain a constant current source; otherwise, turn off the voltage source and continue the calculation;

[0036] In the always-on linear discharge convolution mechanism, each bit line RBL1 and RBL2 is connected to a MOM capacitor in a 255C0 capacitor structure for multi-bit symbol calculations. C0 is a unit capacitor, and flexibly configurable switches are distributed between each group of 255C0. These switches can reorganize the capacitance between different bit lines RBL1 and RBL2, redistribute charge, and implement operations between different bits. The convolution result on the bit lines RBL1 and RBL2 enters the 255C0 capacitor structure for charge redistribution, and then the activation result is obtained through the ReLU circuit. This activation result serves as the sampling input of the bias voltage time converter bypass.

[0037] Each row of the D6T bit array has a bias voltage-to-time converter, which is used to collect the activation results of the convolution and ReLU circuits. The activation results are sampled into the Csamp capacitor of the bias voltage-to-time converter to maintain or generate a time signal. This time signal is re-input into the D6T bit array for convolution.

[0038] Preferably, the layout of the D6T bit adopts a symmetrical layout, with the two decoupled read transistors N2 and N3 placed in the upper left corner and the lower left corner respectively, the write data transistor N1 placed in the upper right corner, the bit lines RBL1, RBL2 and WBL extended vertically, and the selection signal lines S1 and S2 extended horizontally.

[0039] Preferably, the strobe signal line S1 and the strobe signal line S2 do not interfere with each other and are completely independent. One set of control logic is used to read data from the bit line RBL1 , and another set of control logic is used to read data from the bit line RBL2 .

[0040] Preferably, in the IMC mode, multi-bit input data is encoded into time signals with different pulse widths and input to the selection signal lines S1 and S2, and multiple transistors on the bit lines RBL1 and RBL2 are discharged simultaneously, and the final result is the convolution result.

[0041] Preferably, the two bit lines RBL1 and RBL2 of each column of D6T bits in the D6T array can simultaneously calculate two different pictures, or can simultaneously calculate two different parts of the same picture.

[0042] Preferably, the always-on linear discharge convolution mechanism provides an N-order linear convolution: the bit line voltage is detected at the time points T / N, 2T / N, 3T / N, ... N-1T / N. If it is lower than VH-(VH-VL) / N, it is boosted multiple times using different voltage sources. The values ​​of these voltage sources can be selected to ensure that the V of the read transistors N2 and N3 after the boost is ds Just enough.

[0043] Preferably, each bias voltage-to-time converter has six Csamp capacitors, and the six Csamp capacitors are sampled through the switches of a CCR, which stores the convolution result just calculated.

[0044] Preferably, different Csamps output time signals with different pulse widths. The larger the voltage value stored in Csamp is, the wider the pulse width of the generated time signal is.

[0045] Preferably, the process of obtaining the time signal is as follows: the transistor P1, the transistor P2, and the data writing transistor N1 in the D6T bit realize linear charging of the capacitor Cx, wherein the transistor P2 is equivalent to a constant current source, and the transistor P1 and the data writing transistor N1 serve as logic switches; the two input terminals of the high-precision high-speed comparator COMPA are respectively connected to the capacitor Cx and the capacitor Csamp, the output terminal of the high-precision high-speed comparator COMPA is connected to one input terminal of a three-input NAND gate, and the other two input terminals of the three-input NAND gate are respectively input with an enable signal EN and the output of an inverter related to shutting down the high-precision high-speed comparator COMPA:

[0046] The enable signal EN is synchronized with the charging start time of the capacitor Cx. When the capacitor Cx begins to charge with a constant current, the enable signal EN switches from "0" to "1". At this time, the three-input NAND gate generates a falling edge of the time signal; when the high-precision, high-speed comparator OMPA generates a transition from "1" to "0", it generates a rising edge of the time signal. The above two edges constitute a complete time signal with a pulse width. At this time, the high-precision, high-speed comparator COMPA can be turned off to save power. At the same time, the inverter always outputs "0" to maintain the logic output of the three-input NAND gate at "1", thereby obtaining a sharp-edge time domain signal through the output of the three-input NAND gate.

[0047] Preferably, the four low-power comparators select a coarse delay of 1 to 4Td to turn on the high-precision and high-speed comparator COMPA.

[0048] The present invention proposes a D6T accelerator for in-memory computing applied to edge artificial intelligence, which has the characteristics of always linear discharge and reduced digital steps. In order to address the challenges faced by existing in-memory computing macro modules, three effective technologies are proposed in the in-memory computing accelerator disclosed in the present invention: (1) a decoupled 6T (D6T) bit cell that can operate reliably at 0.4V and standby at 0.26V, supporting parallel processing of decoupled dual ports; (2) an always linear discharge convolution mechanism (ALDCM), which can not only reduce the bit line voltage but also always maintain linear calculation over the entire voltage range of the bit line; (3) the bypass of the bias voltage time converter (BVTC) reduces the digital steps, but still maintains high energy efficiency and computing density at low voltage. The measurement results of the in-memory computing accelerator show that its average energy efficiency is 8918TOPS / W (8b×8b) and the average computing density is 38.6TOPS / mm in 55nm CMOS process. 2 (8b×8b). BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 The 64Kb macroblock with D6T as the basic bit unit is shown;

[0050] Figures 2 to 4 This paper illustrates the challenges faced by existing in-memory computing macro modules.

[0051] Figure 5 The structure of the D6T bit is shown;

[0052] Figure 6 It shows that when transistor P1 is off, the gate of data write transistor N1 requires a voltage of 0 to 46 mV at room temperature of 25°C;

[0053] Figure 7 The layout of the D6T bit is shown;

[0054] Figure 8 The IMC mode of the D6T bit is shown;

[0055] Figure 9 The two bit lines RBL1 and RBL2 of each column of D6T bits of the present invention are illustrated, and two different pictures can be calculated simultaneously;

[0056] Figure 10 It shows that the computational load of the same graph is divided into two parts and distributed equally to the two bit lines RBL1 and RBL2;

[0057] Figure 11 The measurement results of average standby power consumption are shown;

[0058] Figure 12 It indicates that the standby (leakage) power can be ignored;

[0059] Figure 13 The boost result of N-order linear MAC is shown;

[0060] Figure 14 The 255C0 structure is shown;

[0061] Figure 15 It shows the use of ReLU low-power comparator to output the convolution result;

[0062] Figure 16 It shows that the larger the voltage value stored in Csamp is, the wider the pulse width of the generated time signal is;

[0063] Figure 17 The structure of the bias voltage-to-time converter disclosed in the present invention is illustrated;

[0064] Figure 18 The new process provided by the BVTC of the present invention is illustrated;

[0065] Figure 19 Shows the manual setting of the bit line reference voltage V L The value of

[0066] Figure 20 The measurement results of the 55nm 64Kb D6T chip proposed in this invention are shown;

[0067] Figure 21 A comparison of the present invention with the state-of-the-art work is illustrated;

[0068] Figure 22 Annotated chip photographs and chip summaries are shown. DETAILED DESCRIPTION

[0069] Below in conjunction with specific embodiment, further set forth the present invention.Should be understood that these embodiments are only used to illustrate the present invention and are not used in limiting the scope of the present invention.In addition, should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms fall equally within the scope limited by the appended claims of the application.

[0070] like Figure 5 As shown, the D6T bit of the present invention uses a total of 6 transistors. N1 is a write data transistor, which writes logic "0" and "1" to the storage node QB through the word line WWL and the bit line WBL. P1 is a transistor that maintains logic "1" (weight value W = 0), which can directly connect the "1" of the storage node QB to VDD, thereby enhancing data stability. When the storage node QB is "0" (weight value W = 1), transistor P1 is turned off, and the gate of the write data transistor N1 requires a voltage of 0 to 46mV at room temperature 25°C (such as Figure 6 As shown in the figure, the charge of logic "0" leaks out of the write data transistor N1, preventing "0" from flipping. P2 and N4 form an inverter, providing the gate voltage of transistor P1 and increasing the parasitic capacitance of the storage node QB. N2 and N3 are two decoupled read transistors, or used for independent convolution calculations. S1 and S2 are used to provide a selection signal for selecting N2 or N3. If the selection signal appears "0", it means that the corresponding read transistor is selected; if it appears "1", the corresponding read transistor is turned off. Data can be read on the two vertical bit lines RBL1 and RBL2, or the convolution result can be obtained.

[0071] The D6T bit of the present invention supports a standby mode as low as 0.26V, reliably operates at 0.4V, and has dual decoupled ports. The layout area of ​​this D6T bit is only 1.16 times that of a 6T SRAM under the minimum logic rule. Figure 7As shown, the D6T bit cell adopts a symmetrical layout, with two decoupled read transistors N2 and N3 placed in the upper left and lower left corners, respectively, and the write data transistor N1 placed in the upper right corner. Bit lines RBL1, RBL2, and WBL are extended vertically, while select signal lines S1 and S2 are extended horizontally. Compared to a 1-read / write (1RW) 6T SRAM, this D6T bit cell provides 2 read and 1 write (2R1W) ports to support more complex data transfer and logical calculations. For example, if bit D6T in a row is selected, then bits S1 and S2 in that row are "0." Because bit lines RBL1 and RBL2 are precharged to VDD before each data read, the VDD charge on bit lines RBL1 and RBL2 is discharged to "0" through read transistors N2 and N3, depending on the storage node QB of the row's D6T bit. If storage node QB stores "0," the VDD charge on bit lines RBL1 and RBL2 is discharged to "0" through read transistors N2 and N3. Conversely, if storage node QB stores "1," the VDD charge on bit lines RBL1 and RBL2 remains unchanged, resulting in the bit line being considered a logical "1." Since select signal lines S1 and S2 are independent and independent of each other, two control systems can exist: one reading from bit line RBL1 and the other from bit line RBL2, without any command or timing conflicts. Since writes can only be made through a single path, through write data transistor N1, this achieves a 2R1W function.

[0072] like Figure 8 As shown, in IMC mode (corresponding to logic calculation): the multi-bit input data is encoded into time signals with different pulse widths and input to the selection signal lines S1 and S2. Since the transistor always works in the saturation region (with a sufficient V ds ), so a longer pulse signal means a longer constant current discharge time. Multiple transistors on bit lines RBL1 and RBL2 discharge simultaneously, and the final result is a convolution result.

[0073] Therefore, the D6T bit of the present invention supports both the conventional memory mode and the IMC mode.

[0074] like Figure 9 As shown, the two bit lines RBL1 and RBL2 of each column D6T bit can calculate two different pictures at the same time; or Figure 10As shown, the computational workload for the same image is divided into two parts, evenly distributed across the two bit lines RBL1 and RBL2 of each column of D6T bits. Measurements from ten chip samples show that in standard memory mode, the worst-case read and write times for each row of data at 0.4V are 1.1ns and 6.2ns, respectively. In IMC mode, dual parallel decoupled ports are used, enabling simultaneous processing of two images or a 2x acceleration of a single convolution layer, significantly improving computational density. To reduce convolution bias, the turn-off voltage of MAC transistors N2 and N3 is set to 200mV to suppress current backflow caused by low bit line voltages. Furthermore, the large size of transistors N2 and N3 reduces near-threshold process variations. For data writes, because the threshold voltage of the selected data write transistor N1 (RVT) is significantly lower than that of transistor P1 (HVT), data "0" can be successfully written to the D6T bit even without a charge pump. On the one hand, the measurement results show that the minimum value of WWL at different voltages is 0-46mV at 25°C, and this voltage range can maintain data "0". On the other hand, even if the current is weak, data "1" can still be maintained. As for leakage, the present invention uses a thermostat to evaluate the average standby power consumption at -40°C, 25°C, and 85°C, and obtains 0.17nW / Kb, 2.4nW / Kb, and 32.2nW / Kb at 0.26V, respectively. Figure 11 As shown, especially when the accelerator operates in IMC mode, the dynamic power of IMC is more than 100 times the standby (leakage) power, so this standby (leakage) power can be ignored, as shown in Figure 12 shown.

[0075] The present invention utilizes an always linear discharge convolution mechanism (ALDCM) to reduce the bitline voltage while maintaining linear computation across the entire bitline voltage range. For standard voltage designs, the bitline requires a precharge voltage (VH) of 0.8V to 1.2V. This is to prevent the MAC transistors from entering a nonlinear discharge region, which results in unoptimized bitline energy accounting for 67.3% of the total MAC energy. The present invention develops a novel ALDCM technique, specifically including the following:

[0076] Assuming that the longest pulse width of the time domain input is T, the macro module will sample the convolution voltage at T / 2+δT, where δT is the ps level delay: If the MAC capacitor voltage is <(VH-VL) / 2, the 140mV voltage source will boost the capacitor top plate to >(VH-VL) / 2; otherwise, the voltage source will be turned off and the calculation will continue. By maintaining sufficient V ds (The 140mV voltage source boosts the discharged bit lines RBL1 and RBL2 back to a high voltage, namely the V dsTo further improve computational linearity, ALDCM provides a high-order linear MAC, or N-order linear MAC: the bit line voltage is detected at time points T / N, 2T / N, 3T / N, and so on. If it is lower than VH-(VH-VL) / N, it is boosted multiple times using different voltage sources. The values ​​of these voltage sources can be selected arbitrarily, as long as the boosted V ds It is sufficient (the boosted bit line is higher than VH-(VH-VL) / N again), such as Figure 13 In this way, the V ds The variation range is further reduced, thereby generating a ramp current with an almost constant slope. In addition, the ALDCM disclosed in the present invention has a flexible and reconfigurable metal oxide metal (MOM) capacitor architecture (hereinafter referred to as "MOM capacitor") for multi-bit symbol calculation. The MOM capacitor connected to each bit line is a 255C0 architecture, which can be divided into a variety of 2 n C0, where C0 is the unit capacitance. For this design, the MOM capacitors are fabricated on a high-level metal layer above the D6T array, so there is almost no additional area overhead. When performing multi-bit operations, charge redistribution occurs between the bit lines RBL1[n] of the nth D6T bit, and the same is true for the bit line RBL2[n] of the nth D6T bit. In particular, the capacitor plates for the sign bit and other bits need to be reversed. To reduce the mismatch of capacitance, we use four 4C0 capacitors in series to form the lowest bit, and two 4C0 capacitors in series to form the second lowest bit. For flexible and precise configuration, this capacitor structure can map weights of 2 to 8 bits. Flexible configurable switches are distributed between each group of 255C0. These switches can reorganize the capacitance between different bit lines RBL to implement operations between different bits. For example, for 8-bit operations, the highest bit is the sign bit, corresponding to 2 7 C0, the remaining seven digits are 2 6 C0, 2 5 C0, 2 4 C0, 2 3 C0, 2 2 C0, 2C0, C0. When the calculation is completed, the switches are used to connect the positive poles of the capacitors of the last seven digits to the positive poles and the negative poles to the negative poles to distribute the charge. The capacitor of the sign digit is connected to the negative poles of the last seven digits and to their positive poles to redistribute the charge. Figure 14As shown. Next, the present invention also uses a ReLU low-power comparator to output the activation result. For example: the convolution result range of 8-bit operation is -127VH+128VL / 255 to 128VH-127VL / 255. When the convolution result voltage is lower than VH / 255, 0V is output (strictly speaking, VH / 255 should be output, but VH / 255≈0). Otherwise, the convolution voltage is retained and used as the sampling input of BVTC, as shown in Figure 15 shown.

[0077] The ALDCM proposed in this invention not only performs MAC with virtually no area overhead and at ultra-low bitline voltages, but also provides a consistently linear discharge. Measurements demonstrate that the bitline voltage lower limit (VL) can reach a negative voltage of -60mV. This means that even with an aggressive bitline precharge voltage adjustment of 280mV, a suitable MAC range of 340mV is still achieved.

[0078] The bias voltage time converter (BVTC) bypass disclosed in the present invention significantly reduces the digital steps in the IMC to achieve high energy efficiency while maintaining high computational density at low voltage. The BVTC of the present invention directly converts the result of the bit line RBL into the input of the subsequent convolution.

[0079] like Figure 17 As shown, each row of the D6T bit array has a BVTC, and the input of the BVTC will collect the results after convolution and ReLU. There are 6 Csamps in each BVTC, which means that the results of the most recent 6 convolutions and ReLUs can be collected. Even if the voltage saved by Csamp is not used immediately, the most recent convolution result will be stored in these capacitors in the form of analog voltage for a long time. Therefore, there is no need to convert these analog voltage values ​​into redundant digital codes and store them in external SRAM. Specifically, Csamp[0:5] is sampled through the switch of CCR, and CCR stores the MAC result that has just been calculated, that is, the result calculated by the 255C0 capacitor architecture. During the sampling process, the VH / 2 of the Csamp lower board n The voltage source is turned on. After sampling, the bottom plate is grounded to eliminate the positive common-mode voltage.

[0080] Different Csamps will output time signals with different pulse widths. If the voltage value stored in Csamp is larger, the pulse width of the generated time signal will be wider, such as Figure 16As shown. These time signals serve as the input of the next convolution (MAC). The process of obtaining the time signal is as follows: the transistors P1, P2, and N1 in the D6T bit realize the linear charging of the capacitor Cx, wherein the transistor P2 is equivalent to a constant current source, and the transistors P1 and N1 serve as logic switches. The present invention designs a 0.8V power supply high-precision high-speed comparator COMPA and a three-input NAND gate. The two input ends of the high-precision high-speed comparator COMPA are respectively connected to the capacitor Cx and Csamp[0:5], the output end of the high-precision high-speed comparator COMPA is connected to one input end of the three-input NAND gate, and the other two input ends of the three-input NAND gate are respectively input with the enable signal EN and the output of the inverter related to shutting down the high-precision high-speed comparator COMPA.

[0081] The enable signal EN is synchronized with the charging start time of capacitor Cx. When capacitor Cx begins constant-current charging, the enable signal EN switches from "0" to "1," generating a falling edge (start) of the timing signal on the three-input NAND gate. When the high-precision, high-speed comparator OMPA transitions from "1" to "0," it generates a rising edge (end) of the timing signal. These two edges constitute a complete timing signal with a pulse width. At this point, the high-precision, high-speed comparator COMPA can be shut down to save power. Meanwhile, the inverter continuously outputs "0" to maintain the logic output of the three-input NAND gate at "1," resulting in a sharp-edged time-domain signal at the output of the three-input NAND gate. Five sets of transistors with different sizes and adjustable bias voltages are designed to overcome process variations.

[0082] To reduce BVTC power consumption, we employ the following strategy: To prevent the high-precision, high-speed comparator COMPA from prematurely turning on, particularly when Vsamp (Vsamp refers to the convolution result voltage sampled by capacitor Csamp, which is fed to the input of the high-precision, high-speed comparator COMPA and compared with the voltage on capacitor Cx to generate a variable-width timing signal) is high, which can lead to high power consumption, the four low-power comparators COMPB[0:3] select a coarse 1-4Td delay to turn on the high-precision, high-speed comparator COMPA. When the high-precision, high-speed comparator COMPA flips from "1" to "0," the three-input NAND gate generates a rising edge (the second edge) of the timing signal, indicating that timing signal generation is complete. A shutdown logic module shuts down the high-precision, high-speed comparator COMPA and, through an inverter, outputs "0," maintaining the output logic of the three-input NAND gate at "1."

[0083] Combine Figure 18 , BVTC provides a new process:

[0084] After the convolution (MAC) is completed, the charge redistribution is performed in the 255C0 capacitor structure. Then, an activation result is obtained through a simple ReLU circuit. This activation result is sampled into the Csamp capacitor of the BVTC to maintain or generate a time signal. This time signal is re-input into the D6T bit array for convolution (MAC).

[0085] like Figure 19 As shown, in order to manually set the reference voltage V L To further reduce the BVTC deviation, we check the output of any column: Csamp in each row BVTC samples and saves the voltage of (128VH-127VL) / 255, and converts the voltage into a time signal. The voltage of the storage node QB of all memory cells in the tested column is 0V (weight value W=1). We can use the bit line voltage and the reference voltage V L Comparison is used to determine whether the lower limit of the bit line is the reference voltage V L If the comparator flips, it means that the lower limit is indeed the reference voltage V L .

[0086] Combine Figure 20 We measured digital and analog bypassing in ten chips, using the former as a hardware reference for comparison. When the D6T bits were powered at 0.4V, the analog bypassing achieved an energy efficiency of 8918 TOPS / W (8b×8b), which is 17.4 times and 132.7 times higher than the 0.4V and 1V digital bypassing, respectively. Simultaneously, this approach maintained a high compute density of 38.6 TOPS / mm² (8b×8b) in a 55nm CMOS process, which is 32.2 times and 2.9 times higher than the 0.4V and 1V digital bypassing, respectively. For the ResNet-20 model, digital bypassing is more susceptible to accuracy loss below 0.5V. We analyze that this may be due to the excessive number of steps leading to error accumulation. However, the analog bypassing still achieves 91.85% accuracy on CIFAR-10 and 67.94% accuracy on CIFAR-100 at 0.4V. Regarding energy composition, digital bypass has a high digital correlation of 64.77%, while analog bypass has a BVTC of 20.57%. In particular, analog bypass for ResNet-20 consumes only 471.5nJ / frame. We also measured temperature characteristics: excessively high temperatures (>70°C) can cause a sharp increase in on-chip capacitor leakage, leading to product failure; excessively low temperatures (<-20°C) can cause nodes storing data "0" to lose their correctness.

[0087] Combine Figure 21, this paper implements a 64Kb D6T acceleration chip in 55nm CMOS, improving energy efficiency while maintaining high computing density. Both digital bypass and analog bypass are implemented and compared in ten chips. This D6T array achieves an energy efficiency of 8918TOPS / W (8b×8b), which is 141 times higher than the most advanced research [5]. Compared with more advanced CMOS processes (<10nm), this 55nm work can still achieve a computing density of 38.6TOPS / mm2 (8b×8b).

Claims

1. A D6T in-memory computing accelerator with always linear discharge and reduced digital steps. The macromodule of the D6T in-memory computing accelerator includes a digital bypass and an analog bypass. A D6T array composed of D6T bits serves as the common area of ​​the digital bypass and the analog bypass, realizing in-memory computing. The characteristics are: Each D6T bit supports both conventional memory mode and IMC mode, including: a write data transistor N1, used to write logic "0" and "1" to the storage node QB through the word line WWL and the bit line WBL; a transistor P1, used to maintain logic "1" so that the "1" of the storage node QB is directly connected to VDD; an inverter composed of a transistor P2 and a transistor N4, used to provide the gate voltage of the transistor P1 and to increase the parasitic capacitance of the storage node QB; decoupled read transistors N2 and N3, providing dual decoupling ports to the outside, and the read transistors N2 and N3 can also be used independently for convolution calculations; selection signal lines S1 and S2 for providing selection signals for selecting the read transistors N2 and N3; and bit lines RBL1 and RBL2 for reading data or obtaining convolution results. The convolution mechanism reduces the bit line voltage by always discharging linearly, but maintains linearity over the entire voltage range of the bit line. Calculation: Assuming the longest pulse width of the time domain input is T, the macromodule samples the convolution voltage at T / 2+δT, where δT is a ps-level delay: If the convolution voltage is <(VH-VL) / 2, VH represents the bit line precharge voltage, and VL represents the minimum voltage of the bit line discharged through the transistor, the capacitor top plate is boosted to >(VH-VL) / 2 through the voltage source, so that the discharged bit lines RBL1 and RBL2 voltage return to a high voltage to maintain sufficient V ds , so that the read transistors N2 and N3 performing convolution on the bit line maintain a constant current source; otherwise, turn off the voltage source and continue the calculation; In the always-on linear discharge convolution mechanism, each bit line RBL1 and RBL2 is connected to a MOM capacitor in a 255C0 capacitor structure for multi-bit symbol calculations. C0 is a unit capacitor, and flexibly configurable switches are distributed between each group of 255C0. These switches can reorganize the capacitance between different bit lines RBL1 and RBL2, redistribute charge, and implement operations between different bits. The convolution result on the bit lines RBL1 and RBL2 enters the 255C0 capacitor structure for charge redistribution, and then the activation result is obtained through the ReLU circuit. This activation result serves as the sampling input of the bias voltage time converter bypass. Each row of the D6T bit array has a bias voltage-to-time converter, which is used to collect the activation results of the convolution and ReLU circuits. The activation results are sampled into the Csamp capacitor of the bias voltage-to-time converter to maintain or generate a time signal. This time signal is re-input into the D6T bit array for convolution.

2. The D6T in-memory computing accelerator with always linear discharge and reduced digital steps as claimed in claim 1, characterized in that: The D6T bit cell adopts a symmetrical layout, with the two decoupled read transistors N2 and N3 placed in the upper left corner and lower left corner respectively, the write data transistor N1 placed in the upper right corner, the bit lines RBL1, RBL2 and WBL extended vertically, and the selection signal lines S1 and S2 extended horizontally.

3. The D6T in-memory computing accelerator with always linear discharge and reduced digital steps as claimed in claim 1, characterized in that: The strobe signal line S1 and the strobe signal line S2 do not interfere with each other and are completely independent. One set of control logic is used to read data from the bit line RBL1 , and another set of control logic is used to read data from the bit line RBL2 .

4. The D6T in-memory computing accelerator with always linear discharge and reduced digital steps as claimed in claim 1, characterized in that: In the IMC mode, multi-bit input data is encoded into time signals with different pulse widths and input to the selection signal lines S1 and S2. Multiple transistors on the bit lines RBL1 and RBL2 are discharged simultaneously, and the final result is the convolution result.

5. The D6T in-memory computing accelerator with always linear discharge and reduced digital steps as claimed in claim 1, characterized in that: The two bit lines RBL1 and RBL2 of each column of D6T bits in the D6T array can simultaneously calculate two different pictures, or can simultaneously calculate two different parts of the same picture.

6. The D6T in-memory computing accelerator with always linear discharge and reduced digital steps as claimed in claim 1, characterized in that: The always-on linear discharge convolution mechanism provides N-order linear convolution: the bit line voltage is detected at the time points T / N, 2T / N, 3T / N, ... N-1T / N. If it is lower than VH-(VH-VL) / N, it is boosted multiple times using different voltage sources. The values ​​of these voltage sources can be selected to ensure that the V of the read transistors N2 and N3 after the boost is within the range of VH-(VH-VL) / N. ds Just enough.

7. The D6T in-memory computing accelerator with always linear discharge and reduced digital steps as claimed in claim 1, characterized in that: Each bias voltage-to-time converter has six Csamp capacitors, and the six Csamp capacitors are sampled through the switches of the CCR, which stores the convolution result just calculated.

8. The D6T in-memory computing accelerator with always linear discharge and reduced digital steps as claimed in claim 1, characterized in that: Different Csamps output time signals with different pulse widths. The larger the voltage value stored in Csamp, the wider the pulse width of the generated time signal.

9. The D6T in-memory computing accelerator with always linear discharge and reduced digital steps as claimed in claim 1, characterized in that: The process of obtaining the time signal is as follows: the transistors P1, P2, and data writing transistor N1 in the D6T bit realize linear charging of the capacitor Cx, wherein the transistor P2 acts as a constant current source, and the transistors P1 and data writing transistor N1 act as logic switches; the two input terminals of the high-precision high-speed comparator COMPA are respectively connected to the capacitor Cx and the capacitor Csamp, and the output terminal of the high-precision high-speed comparator COMPA is connected to one input terminal of the three-input NAND gate, and the other two input terminals of the three-input NAND gate are respectively input with the enable signal EN and the output of the inverter related to shutting down the high-precision high-speed comparator COMPA: The enable signal EN is synchronized with the charging start time of the capacitor Cx. When the capacitor Cx begins constant-current charging, the enable signal EN switches from "0" to "1." At this point, the three-input NAND gate generates a falling edge of the timing signal. When the high-precision, high-speed comparator OMPA transitions from "1" to "0," it generates a rising edge of the timing signal. These two edges constitute a complete timing signal with a pulse width. At this point, the high-precision, high-speed comparator COMPA can be shut down to save power. Meanwhile, the inverter continuously outputs "0" to maintain the logic output of the three-input NAND gate at "1." This results in a sharp-edged time-domain signal at the output of the three-input NAND gate.

10. The D6T in-memory computing accelerator with always linear discharge and reduced digital steps as claimed in claim 9, characterized in that: The four low-power comparators select a coarse 1-4Td delay to turn on the high-precision high-speed comparator COMPA.

Citation Information

Patent Citations

  • In-memory computing device suitable for binary convolutional neural network computing

    CN111126579A

  • In-memory computing device, chip and electronic equipment

    CN114298297A