Artificial intelligence acceleration system based on 8T-SRAM calculation and storage integrated macro cell
By integrating computing transistors in the 6T-SRAM cell, an integrated macro unit of 8T-SRAM computing and storage is formed, the problem of insufficient multiplication and accumulation computing efficiency and accuracy in the prior art is solved, and high-efficiency and low-power AI computing is realized.
Patent Information
- Application Number
- CN202510214695.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-02-26
AI Technical Summary
The existing integrated computing and storage design based on 6T-SRAM is difficult to achieve efficient multiplication and accumulation operations, and the calculation accuracy is insufficient, which cannot meet the efficiency and accuracy requirements of artificial intelligence computing.
An artificial intelligence acceleration system based on 8T-SRAM is designed to integrate computing and storage macro cells. By integrating two computing transistors in each 6T-SRAM cell, a digital-analog hybrid computing array is formed to achieve efficient on-chip multiplication and accumulation operations.
It realizes efficient multiplication and accumulation operations, improves AI computing efficiency and calculation accuracy, is simple to control, low power consumption and high accuracy, and is suitable for fields such as artificial intelligence accelerators and edge computing.
Smart Images

Figure CN120144530A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of integrated circuit technology, and particularly to an artificial intelligence acceleration system based on an 8T-SRAM computing-in-memory macrocell. Background Art
[0002] In recent years, with the rapid development of artificial intelligence technology, deep neural networks have achieved remarkable breakthroughs in fields such as image recognition and speech processing. The calculations in deep learning mainly rely on a large number of multiply-accumulate operations (MAC), which pose extremely high requirements on the computing power and memory bandwidth of the processor. In the traditional von Neumann architecture, the processing unit and the memory unit are physically separated, resulting in frequent data movement between the two, causing huge power consumption overhead and performance bottlenecks.
[0003] In the prior art, a computing-in-memory design based on 6T-SRAM is designed, and by integrating simple analog computing circuits in the memory unit, the in-situ computing function is realized.
[0004] However, the above technical solution still has obvious deficiencies. The design based on 6T-SRAM is limited by the traditional SRAM cell structure, making it difficult to achieve efficient multiply-accumulate operations and having insufficient computing accuracy. Summary of the Invention
[0005] An artificial intelligence acceleration system based on an 8T-SRAM computing-in-memory macrocell provided by the present invention can achieve efficient multiply-accumulate operations and improve the AI computing efficiency and computing accuracy.
[0006] To achieve the above object, the key of an artificial intelligence acceleration system based on an 8T-SRAM computing-in-memory macrocell provided by the present invention is: a digital-analog hybrid computing array is provided, the digital-analog hybrid computing array is provided with m×n 8T-SRAM computing-in-memory units, and each 8T-SRAM computing-in-memory unit is provided with a 6T-SRAM unit, a first computing transistor T7, and a second computing transistor T8;
[0007] The 6T-SRAM unit is provided with a memory unit, a first access transistor T5, and a second access transistor T6. The gate of the second access transistor T6 is connected to the word line WL, the source of the second access transistor T6 is connected to the input terminal Q of the memory unit, and the drain of the second access transistor T6 is connected to the first column of computing lines BL; the gate of the first access transistor T5 is connected to the word line WL, the source of the first access transistor T5 is connected to the output terminal Qb of the memory unit, and the drain of the first access transistor T5 is connected to the second column of computing lines BLb;
[0008] The gate of the first computing transistor T7 is connected to the output terminal Qb of the storage unit, the source of the first computing transistor T7 is grounded, and the drain of the first computing transistor T7 is connected to the source of the second computing transistor T8; the gate terminal of the second computing transistor T8 is connected to the read word line RWL, and the drain of the second computing transistor T8 is connected to the read bit line RBL;
[0009] One end of the first column computing line BL and the second column computing line BLb is connected to the column address decoder, and the other end is connected to the read-write computing control unit RWCCU; the word line WL and the read word line RWL are connected to the row address decoder; one end of the read bit line RBL is connected to the power supply, and the other end is connected to the column analog accumulation unit CAAU, and the column analog accumulation unit CAAU outputs the calculation result through the analog-to-digital converter ADC.
[0010] Through the above design, the 6T-SRAM cell includes: a basic storage cell composed of two cross-coupled inverters, and two access transistors for data read and write operations; the present invention innovatively integrates two computing transistors on the basis of the traditional 6T-SRAM, wherein the first computing transistor T7 is connected to the Qb output terminal of the 6T-SRAM core to represent the weight, and the second computing transistor T8 is controlled by the read word line RWL to represent the input data, realizing efficient on-chip multiply-accumulate operation.
[0011] The read-write computing control unit RWCCU is responsible for controlling the precharging process of the read bit line and activating the precise weighted capacitor array. The column analog accumulation unit CAAU and the analog-to-digital converter ADC constitute a data processing path. The CAAU performs analog domain accumulation operation, and the ADC converts the result into a digital signal to complete the final digital accumulation.
[0012] Preferably: the storage unit is provided with a first transistor T1, a second transistor T2, a third transistor T3 and a fourth transistor T4;
[0013] The gate of the first transistor T1 is connected to the input terminal Q of the storage unit, the drain of the first transistor T1 is connected to the output terminal Qb of the storage unit, and the source of the first transistor T1 is grounded;
[0014] The gate of the second transistor T2 is connected to the input terminal Q of the storage unit, the drain of the second transistor T2 is connected to the output terminal Qb of the storage unit, and the source of the second transistor T2 is connected to the power supply;
[0015] The gate of the third transistor T3 is connected to the output terminal Qb of the storage unit, the drain of the third transistor T3 is connected to the input terminal Q of the storage unit, and the source of the third transistor T3 is grounded;
[0016] The gate of the fourth transistor T4 is connected to the output terminal Qb of the storage unit, the drain of the fourth transistor T4 is connected to the input terminal Q of the storage unit, and the source of the fourth transistor T4 is connected to the power supply.
[0017] Preferably, the read-write calculation control unit RWCCU is provided with a first AND gate AND1, a second AND gate AND2, a third AND gate AND3, a fourth AND gate AND4, a first read-write isolation transistor A1, a second read-write isolation transistor A2, a third read-write isolation transistor A3, and a transistor A4;
[0018] The first input terminals of the first AND gate AND1, the second AND gate AND2, and the third AND gate AND3 are connected to a first control signal CS', the second input terminals of the first AND gate AND1 and the second AND gate AND2 are connected to a second control signal R / W', and the second input terminal of the third AND gate AND3 is connected to a third control signal CE;
[0019] The output terminal of the first AND gate AND1 is connected to the gates of the second read-write isolation transistor A2 and the third read-write isolation transistor A3. The sources of the second read-write isolation transistor A2 and the third read-write isolation transistor A3 are connected to the I / O port. The drain of the second read-write isolation transistor A2 is connected to the second column calculation line BLb, and the drain of the third read-write isolation transistor A3 is connected to the first column calculation line BL;
[0020] The output terminal of the second AND gate AND2 is connected to the column analog accumulation unit CAAU and the first input terminal of the fourth AND gate AND4. The second input terminal of the fourth AND gate AND4 is connected to the output terminal of the third AND gate AND3. The output terminal of the fourth AND gate AND4 is connected to the gate of the first read-write isolation transistor A1. The source of the first read-write isolation transistor A1 is connected to the second column calculation line BLb, and the drain of the first read-write isolation transistor A1 is connected to the I / O port;
[0021] The output terminal of the third AND gate AND3 is connected to the column analog accumulation unit CAAU, the read word line RWL, and the gate of the transistor A4. The source of the transistor A4 is connected to the output terminal of the first AND gate AND1, and the drain of the transistor A4 is grounded.
[0022] The read-write computing control unit RWCCU adopts an efficient on-chip logic control structure, including synchronously integrated digital logic units (AND2, AND3, and AND5), a high-precision voltage comparator CMP1 with hysteresis, and a complementary MOS transistor network optimized with a low threshold. Among them, the output of AND3, VAND3 = CE·CS', controls the read-write isolation transistor group (A1 - A3) through a negative feedback loop; the output of AND5, VAND5 = VAND3·CMP1', drives the charging transistors (T12 - T14) through a buffer stage; the output of AND2, VAND2 = R / W'·CS', realizes the differential control of the charge equalization transistors (T15 - T16); CMP1 adopts a high common-mode rejection ratio structure and realizes zero-delay automatic phase switching from pre-charge to discharge by monitoring the voltage of the compensation capacitor; at the same time, a reset control transistor (A5 - A7) with sub-threshold optimization is integrated to realize the fast discharge function of the capacitor.
[0023] Preferably, the column analog accumulation unit CAAU is provided with a transistor A5, a transistor A6, a main computing capacitor, a secondary computing capacitor, a transistor T12, a transistor T14, a transistor T15, and a transistor T16;
[0024] The gates of the transistor T15 and the transistor T16 are connected to the output terminal of the second AND gate AND2. The drain of the transistor T15 is connected to the read bit line RBL, the source of the transistor T15 is connected to the drain of the transistor T16, and the source of the transistor T16 is connected to the input terminal of the analog-to-digital converter ADC through a charge equalization path;
[0025] The gates of the transistor A5 and the transistor A6 are connected between the output terminal of the third AND gate AND3 and the read word line RWL. The source of the transistor A5 is connected to the drain of the transistor T16, and the source of the transistor A5 is also grounded through a series main computing capacitor. The drain of the transistor A5 is grounded; the source of the transistor A6 is connected to the source of the transistor T14, and the source of the transistor A6 is also grounded through a series secondary computing capacitor. The drain of the transistor A6 is grounded; the drain of the transistor T14 is connected to the drain of the transistor T15, the gate of the transistor T14 is connected to the gate of the transistor T12, the source of the transistor T12 is connected to the power supply, and the drain of the transistor T12 is connected to the read bit line RBL.
[0026] Preferably, a stage automatic conversion circuit is further provided. The stage automatic conversion circuit is provided with a transistor A7, a voltage comparator CMP1, a fifth AND gate AND5, a transistor T13, and a compensation capacitor;
[0027] The first input terminal of the fifth AND gate AND5 is connected between the output terminal of the third AND gate AND3 and the read word line RWL. The output terminal of the fifth AND gate AND5 is connected to the gate of the transistor T12. The second input terminal of the fifth AND gate AND5 is connected to the output terminal of the voltage comparator CMP1. The first input terminal of the voltage comparator CMP1 is connected to the power supply. The second input terminal of the voltage comparator CMP1 is grounded after being connected in series with a compensation capacitor. The second input terminal of the voltage comparator CMP1 is also connected to the drain of the transistor T13. The source of the transistor T13 is connected to the power supply. The gate of the transistor T13 is connected to the gate of the transistor T12. The second input terminal of the voltage comparator CMP1 is also connected to the source of the transistor A7. The drain of the transistor A7 is grounded. The gate of the transistor A7 is connected between the output terminal of the third AND gate AND3 and the read word line RWL.
[0028] The phase automatic conversion circuit realizes automatic phase switching by using a compensation capacitor: the upper plate of the compensation capacitor is connected to the second input terminal of CMP1. The first input terminal of CMP1 is connected to the power supply VDD. When the compensation capacitor is fully charged, it indicates that both the main calculation capacitor and the secondary calculation capacitor are fully charged. The output of CMP1 flips from 0 to 1, and the inverted input through AND5 causes its output to become 0. At the same time, the three transistors T12, T13, and T14 are turned off. The turn-off of T14 ensures the isolation between the secondary calculation capacitor and the read bit line RBL, and the turn-off of T13 keeps the compensation capacitor in a charged state.
[0029] The voltage comparator CMP1 adopts a fully differential self-biased comparison structure. The differential inputs are respectively connected to the upper plate of the compensation capacitor and the VDD reference voltage generated by the bandgap reference. By real-time monitoring of the differential voltage ΔV = VCC - VDD, the state flip is triggered at the zero crossing point, realizing the switching from the pre-charge stage to the discharge stage within sub-nanoseconds, significantly optimizing the critical path delay. The voltage comparator CMP1 realizes automatic phase switching by comparing the voltage VCC of the compensation capacitor with the reference voltage VDD.
[0030] Preferably, it includes four-stage operation processes, which are the RBL pre-charge stage, the column multiplication and accumulation discharge stage, the charge equalization stage, and the analog-to-digital conversion stage in sequence;
[0031] First, in the read bit line pre-charge stage, the control signals CS' = 0, CE = 1, and R / W' = 0 are set, so that the output control signal VAND3 of AND3 = CE·CS' = 1, controlling the read-write isolation transistors A1 - A3 to turn off to achieve read-write disable. At the same time, the output control signal VAND5 of AND5 = VAND3·CMP1' = 1, controlling the transistors T12 and T13 to conduct, and charging the main calculation capacitor, the secondary calculation capacitor, and the compensation capacitor;
[0032] Subsequently, it enters the column multiplication and accumulation discharge stage. When the compensation capacitor is fully charged, VCC = VDD, and the output of the voltage comparator CMP1 flips from 0 to 1, making VAND5 = 0, turning off transistors T12 - T14, and waiting for the input signal of the read word line RWL. When transistors T7 and T8 are turned on simultaneously, the main computing capacitor completes multiplication and accumulation through the discharge path, and its discharge amount ΔQBL is proportional to the product of the weight W and the input signal X;
[0033] Then, in the charge equalization stage, set R / W' = 1. Through the AND2 output control signal VAND2 = R / W'·CS' = 1, control transistor T15 to turn off and T16 to turn on, establish the connection between the main computing capacitor and the charge equalization circuit, and complete the charge redistribution process;
[0034] Finally, in the analog - to - digital conversion stage, the master control MOS transistor is turned on, and the equalized charge is introduced into the analog - to - digital converter ADC. By measuring the charge loss, the partial multiplication and accumulation calculation results pMACVn of each column are obtained. According to the 4 - bit binary weight, weighted summation is performed, and finally the digital result DOUT = Σ(pMACVn×2^n), n = 0, 1, 2, 3;
[0035] During reset, set CE = 0 to turn on transistors A5, A6, and A7 simultaneously, and connect the main computing capacitor, secondary computing capacitor, and compensation capacitor to ground respectively to complete the discharge reset operation.
[0036] Through the above design, the automatic switching from pre - charging to discharging is realized through the cooperation of the compensation capacitor and the comparator, and the precise timing control of each stage is ensured through the optimized transistor control network, simplifying the control logic and improving the computing efficiency.
[0037] The present invention realizes a pipelined four - phase operation process: (1) RBL pre - charging stage: Through the synchronous control signal configuration (CS' = 0, CE = 1, R / W' = 0), establish a parallel charging path to achieve zero - deviation charging of the capacitor array; (2) RBL discharge stage (column accumulation): Trigger the state flip of the Schmidt trigger of CMP1 when the compensation capacitor is fully charged, and complete the high - linearity current - mode multiplication operation through the series conduction path of T7 - T8; (3) Charge equalization stage: Set R / W' to 1 to activate the complementary switch pair T15 - T16 to achieve lossless inter - column charge redistribution; (4) Analog - to - digital conversion stage: Adopt a high - speed SAR - ADC architecture to measure the charge loss amount, and output the final digital result DOUT = Σ(pMACVn×2^n), n = 0, 1, 2, 3.
[0038] Preferably: The main computing capacitor includes four specifications of 8Cu, 4Cu, 2Cu, and 1Cu, and the size of the compensation capacitor is 9Cu.
[0039] The present invention adopts a multi-stage capacitor network configuration with optimized area: the main computing capacitor uses a complementary MIM structure to carry out high-bit weight calculation, the secondary computing capacitor uses a high-linearity MOM structure to perform low-bit weight operations, and the compensation capacitor uses a low-leakage metal interconnection structure to achieve automatic phase switching control. Through precise dynamic timing control and parasitic effect compensation, this capacitor network ensures high linearity and temperature stability of analog computing.
[0040] The compensation capacitor is used to ensure uniform charging time and compensate for parasitic effects.
[0041] Preferably, a transistor T10 is provided on the first column computing line BL. The drain of the transistor T10 is connected to the first column computing line BL, the source of the transistor T10 is connected to the drain of the third read / write isolation transistor A3, and the gate of the transistor T10 is connected to the column address line strobe control interface.
[0042] A transistor T9 is provided on the second column computing line BLb. The drain of the transistor T9 is connected to the second column computing line BLb, the source of the transistor T9 is connected to the source of the first read / write isolation transistor A1 and the drain of the second read / write isolation transistor A2, and the gate of the transistor T9 is connected to the column address line strobe control interface.
[0043] A transistor T11 is provided on the read word line RWL. The source and drain of the transistor T11 are connected to the read word line RWL, and the gate of the transistor T11 is connected to the read / write computing control unit RWCCU and the column analog accumulation unit CAAU.
[0044] The transistors T10 and T9 are used to realize the strobe control of the column computing line.
[0045] Preferably, the digital-analog hybrid computing array supports 16×4-bit MAC operations.
[0046] The design of the present invention supports 16×4-bit MAC operations, can be flexibly extended to larger-scale arrays, and is suitable for various neural network applications. The system adopts a standard CMOS process, having good process compatibility and cost advantages.
[0047] The beneficial effects of the present invention: The structure is simple and clear, realizing efficient on-chip multiply-accumulate operations, having the advantages of simple control, low power consumption, and high precision; the innovative automatic phase conversion mechanism and precise weighted capacitor design ensure the reliability and accuracy of the calculation; this solution is fully compatible with the standard CMOS process, easy for large-scale integration, and has important application value in fields such as AI accelerators and edge computing, being an ideal new computing paradigm. Description of the Drawings
[0048] Figure 1Circuit diagram of the artificial intelligence acceleration system based on the 8T-SRAM computing-in-memory macro cell in the embodiment;
[0049] Figure 2 Block diagram of the artificial intelligence acceleration system based on the 8T-SRAM computing-in-memory macro cell in the embodiment;
[0050] Figure 3 Schematic diagram of the system circuit in the RBL pre-charge stage in the embodiment;
[0051] Figure 4 Schematic diagram of the system circuit in the RBL discharge stage in the embodiment;
[0052] Figure 5 Schematic diagram of the system circuit in the charge equalization stage in the embodiment;
[0053] Figure 6 Schematic diagram of the system circuit in the analog-to-digital conversion stage in the embodiment;
[0054] Figure 7 Waveform schematic diagram of the timing control signal in the embodiment. Detailed implementation manners
[0055] The present invention will be further described in detail below with reference to the accompanying drawings and specific examples. The following examples or drawings are used to illustrate the present invention, but not to limit the scope of the present invention.
[0056] As Figure 1 、 Figure 2 shown: An artificial intelligence acceleration system based on an 8T-SRAM computing-in-memory macro cell is provided with a digital-analog hybrid computing array. The digital-analog hybrid computing array is provided with m×n 8T-SRAM computing-in-memory units. Each 8T-SRAM computing-in-memory unit is provided with a 6T-SRAM unit, a first computing transistor T7 and a second computing transistor T8; the digital-analog hybrid computing array supports 16×4-bit MAC operations.
[0057] The 6T-SRAM unit is provided with a storage unit, a first access transistor T5 and a second access transistor T6. The gate of the second access transistor T6 is connected to the word line WL. The source of the second access transistor T6 is connected to the input terminal Q of the storage unit. The drain of the second access transistor T6 is connected to the first column computing line BL; the gate of the first access transistor T5 is connected to the word line WL. The source of the first access transistor T5 is connected to the output terminal Qb of the storage unit. The drain of the first access transistor T5 is connected to the second column computing line BLb;
[0058] The gate of the first computing transistor T7 is connected to the output terminal Qb of the storage cell, the source of the first computing transistor T7 is grounded, and the drain of the first computing transistor T7 is connected to the source of the second computing transistor T8; the gate terminal of the second computing transistor T8 is connected to the read word line RWL, and the drain of the second computing transistor T8 is connected to the read bit line RBL;
[0059] One end of the first column computing line BL and the second column computing line BLb in each column is connected to the column address decoder, and the other end is connected to the read / write computing control unit RWCCU; the word line WL and the read word line RWL are connected to the row address decoder; one end of the read bit line RBL is connected to the power supply, and the other end is connected to the column analog accumulation unit CAAU, and the column analog accumulation unit CAAU outputs a computing result through the analog-to-digital converter ADC.
[0060] As Figure 1 shown: The storage cell is provided with a first transistor T1, a second transistor T2, a third transistor T3 and a fourth transistor T4;
[0061] The gate of the first transistor T1 is connected to the input terminal Q of the storage cell, the drain of the first transistor T1 is connected to the output terminal Qb of the storage cell, and the source of the first transistor T1 is grounded;
[0062] The gate of the second transistor T2 is connected to the input terminal Q of the storage cell, the drain of the second transistor T2 is connected to the output terminal Qb of the storage cell, and the source of the second transistor T2 is connected to the power supply;
[0063] The gate of the third transistor T3 is connected to the output terminal Qb of the storage cell, the drain of the third transistor T3 is connected to the input terminal Q of the storage cell, and the source of the third transistor T3 is grounded;
[0064] The gate of the fourth transistor T4 is connected to the output terminal Qb of the storage cell, the drain of the fourth transistor T4 is connected to the input terminal Q of the storage cell, and the source of the fourth transistor T4 is connected to the power supply.
[0065] The read / write computing control unit RWCCU is provided with a first AND gate AND1, a second AND gate AND2, a third AND gate AND3, a fourth AND gate AND4, a first read / write isolation transistor A1, a second read / write isolation transistor A2, a third read / write isolation transistor A3 and a transistor A4;
[0066] The first input terminals of the first AND gate AND1, the second AND gate AND2 and the third AND gate AND3 are connected to the first control signal CS’, the second input terminals of the first AND gate AND1 and the second AND gate AND2 are connected to the second control signal R / W’, and the second input terminal of the third AND gate AND3 is connected to the third control signal CE;
[0067] The output terminal of the first AND gate AND1 is connected to the gates of the second read / write isolation transistor A2 and the third read / write isolation transistor A3. The sources of the second read / write isolation transistor A2 and the third read / write isolation transistor A3 are connected to the I / O port. The drain of the second read / write isolation transistor A2 is connected to the second column of calculation lines BLb. The drain of the third read / write isolation transistor A3 is connected to the first column of calculation lines BL;
[0068] The output terminal of the second AND gate AND2 is connected to the first input terminal of the column analog accumulation unit CAAU and the fourth AND gate AND4. The second input terminal of the fourth AND gate AND4 is connected to the output terminal of the third AND gate AND3. The output terminal of the fourth AND gate AND4 is connected to the gate of the first read / write isolation transistor A1. The source of the first read / write isolation transistor A1 is connected to the second column of calculation lines BLb. The drain of the first read / write isolation transistor A1 is connected to the I / O port;
[0069] The output terminal of the third AND gate AND3 is connected to the column analog accumulation unit CAAU, the read word line RWL, and the gate of the transistor A4. The source of the transistor A4 is connected to the output terminal of the first AND gate AND1. The drain of the transistor A4 is grounded.
[0070] The read / write calculation control unit RWCCU adopts a simplified logic control structure to achieve efficient operation mode switching and capacitor charge and discharge control. In the bit line pre-charging stage and the column multiplication and accumulation discharge stage, the time-domain control signal is automatically converted to CS'(t)=0, CE(t)=1, R / W'(t)=0; in the charge equalization stage and the analog-to-digital conversion stage, the time-domain control signal is automatically converted to CS'(t)=0, CE(t)=1, R / W'(t)=1. After the calculation is completed, the time-domain control signal is automatically converted to CS'(t)=0, CE(t)=0, R / W'(t)=1.
[0071] The column analog accumulation unit CAAU is provided with a transistor A5, a transistor A6, a main calculation capacitor, a secondary calculation capacitor, a transistor T12, a transistor T14, a transistor T15, and a transistor T16;
[0072] The gates of the transistor T15 and the transistor T16 are connected to the output terminal of the second AND gate AND2. The drain of the transistor T15 is connected to the read bit line RBL. The source of the transistor T15 is connected to the drain of the transistor T16. The source of the transistor T16 is connected to the input terminal of the analog-to-digital converter ADC through a charge equalization path;
[0073] The gates of the transistors A5 and A6 are connected between the output terminal of the third AND gate AND3 and the read word line RWL. The source of the transistor A5 is connected to the drain of the transistor T16. The source of the transistor A5 is also connected to the ground after being serially connected with the main computing capacitor. The drain of the transistor A5 is grounded. The source of the transistor A6 is connected to the source of the transistor T14. The source of the transistor A6 is also connected to the ground after being serially connected with the secondary computing capacitor. The drain of the transistor A6 is grounded. The drain of the transistor T14 is connected to the drain of the transistor T15. The gate of the transistor T14 is connected to the gate of the transistor T12. The source of the transistor T12 is connected to the power supply. The drain of the transistor T12 is connected to the read bit line RBL.
[0074] A stage automatic conversion circuit is further provided. The stage automatic conversion circuit is provided with a transistor A7, a voltage comparator CMP1, a fifth AND gate AND5, a transistor T13, and a compensation capacitor.
[0075] The first input terminal of the fifth AND gate AND5 is connected between the output terminal of the third AND gate AND3 and the read word line RWL. The output terminal of the fifth AND gate AND5 is connected to the gate of the transistor T12. The second input terminal of the fifth AND gate AND5 is connected to the output terminal of the voltage comparator CMP1. The first input terminal of the voltage comparator CMP1 is connected to the power supply. The second input terminal of the voltage comparator CMP1 is grounded after being serially connected with the compensation capacitor. The second input terminal of the voltage comparator CMP1 is also connected to the drain of the transistor T13. The source of the transistor T13 is connected to the power supply. The gate of the transistor T13 is connected to the gate of the transistor T12. The second input terminal of the voltage comparator CMP1 is also connected to the source of the transistor A7. The drain of the transistor A7 is grounded. The gate of the transistor A7 is connected between the output terminal of the third AND gate AND3 and the read word line RWL.
[0076] A transistor T10 is provided on the first column computing line BL. The drain of the transistor T10 is connected to the first column computing line BL. The source of the transistor T10 is connected to the drain of the third read / write isolation transistor A3. The gate of the transistor T10 is connected to the column address line strobe control interface.
[0077] A transistor T9 is provided on the second column computing line BLb. The drain of the transistor T9 is connected to the second column computing line BLb. The source of the transistor T9 is connected to the source of the first read / write isolation transistor A1 and the drain of the second read / write isolation transistor A2. The gate of the transistor T9 is connected to the column address line strobe control interface.
[0078] A transistor T11 is provided on the read word line RWL. The source and drain of the transistor T11 are connected to the read word line RWL. The gate of the transistor T11 is connected to the read / write computing control unit RWCCU and the column analog accumulation unit CAAU.
[0079] The specific operation process of the present invention is as follows:
[0080] During the RBL pre-charge phase, the basic control signals are configured as CS' = 0, CE = 1, R / W' = 0. The main control logic includes two AND gate circuits, AND3 and AND5: AND3 receives the inverted inputs of the CE signal and the CS' signal, and its output is high. On the one hand, it controls the read / write isolation transistor banks A1 - A3 to turn off to disable the read / write operation; on the other hand, it is jointly input to AND5 together with the inverted output (initially 0) of the voltage comparator CMP1. At the same time, R / W' = 0 ensures that the signal inverted with CS' generates a low-level output through AND2, turning on the PMOS transistor T15 and turning off the NMOS transistor T16, disconnecting the 8Cu capacitor from the charge sharing circuit. When the output of AND5 is high, it controls both T12 and T13 to turn on. Among them, T12 is connected to the VDD power supply and is responsible for charging the main computing capacitor, that is, the main computing capacitor and the secondary computing capacitor; T13 is responsible for charging the compensation capacitor, as Figure 1 、 Figure 3 shown.
[0081] During the RBL discharge phase, wait for the input signal of the RWL modulation pulse. When the weight of the storage unit is 1 (T7 is on) and the input data is high level (T8 is on), the main computing capacitor completes the multiplication and accumulation operation through the T7 - T8 discharge path, as Figure 1 、 Figure 4 shown.
[0082] During the charge sharing process, the control signal R / W' is set to 1, and the signal inverted with CS' generates a control signal through AND2. When the output of AND2 is high, it controls the PMOS transistor T15 to turn off and the NMOS transistor T16 to turn on, establishing a path between the main computing capacitor and the charge sharing circuit to achieve charge redistribution, as Figure 1 、 Figure 5 shown.
[0083] Finally, the evenly shared charge is introduced into the analog-to-digital converter ADC through a master control MOS transistor. The ADC completes the digital conversion of the multiplication and accumulation result by measuring the charge loss, as Figure 1 、 Figure 6 shown.
[0084] During reset, set CE = 0 to turn on transistors A5, A6, and A7 simultaneously, connecting the main computing capacitor, the secondary computing capacitor, and the compensation capacitor to ground respectively to complete the discharge reset operation.
[0085] The main computing capacitor includes four specifications: 8Cu, 4Cu, 2Cu, and 1Cu, and the size of the compensation capacitor is 9Cu.
[0086] The present invention adds two computing transistors on the basis of a traditional 6T-SRAM to achieve efficient on-chip multiply-accumulate operations while maintaining full compatibility with the standard CMOS process; designs an accurate weighted capacitor array to achieve high-precision analog computing, and innovatively introduces compensation capacitors to eliminate the unevenness of charge and discharge; adopts an automatic phase conversion mechanism to simplify the control logic, significantly improving the operation efficiency. Specifically, it includes the following advantages:
[0087] 1. The innovative four-multiply-accumulate operation process combined with the automatic phase conversion mechanism realizes efficient on-chip computing, with simple control logic, low power consumption, and high reliability; the time-domain pipelining control ensures the precise coordination of operations in each stage, improving the computing efficiency.
[0088] 2. The accurate weighted capacitor array adopts a binary weight design and combines compensation capacitor technology to effectively solve the problems of computing accuracy and reliability; the innovative voltage comparison structure realizes automatic phase conversion, reducing the dependence on external control.
[0089] 3. The proposed digital-analog hybrid computing architecture fully combines the advantages of digital storage and analog computing, supports flexible data mapping strategies, and is easy to expand to large-scale arrays.
[0090] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, various changes and modifications can be made to the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. An artificial intelligence acceleration system based on 8T-SRAM computing and storage integrated macro unit, characterized by: A digital-analog hybrid computing array is provided, wherein the digital-analog hybrid computing array is provided with m×n 8T-SRAM storage-computation integrated units, and each 8T-SRAM storage-computation integrated unit is provided with a 6T-SRAM unit, a first computing transistor T7, and a second computing transistor T8; The 6T-SRAM cell is provided with a storage cell, a first access transistor T5 and a second access transistor T6, wherein the gate of the second access transistor T6 is connected to the word line WL, the source of the second access transistor T6 is connected to the input terminal Q of the storage cell, and the drain of the second access transistor T6 is connected to the first column calculation line BL; the gate of the first access transistor T5 is connected to the word line WL, the source of the first access transistor T5 is connected to the output terminal Qb of the storage cell, and the drain of the first access transistor T5 is connected to the second column calculation line BLb; The gate of the first calculation transistor T7 is connected to the output terminal Qb of the storage unit, the source of the first calculation transistor T7 is grounded, and the drain of the first calculation transistor T7 is connected to the source of the second calculation transistor T8; the gate of the second calculation transistor T8 is connected to the read word line RWL, and the drain of the second calculation transistor T8 is connected to the read bit line RBL; One end of the first column calculation line BL and the second column calculation line BLb are connected to the column address decoder, and the other end is connected to the read-write calculation control unit RWCCU; the word line WL and the read word line RWL are connected to the row address decoder; one end of the read bit line RBL is connected to the power supply, and the other end is connected to the column analog accumulation unit CAAU, and the column analog accumulation unit CAAU outputs the calculation result through the analog-to-digital converter ADC.
2. The artificial intelligence acceleration system based on the 8T-SRAM computing and storage integrated macro unit according to claim 1 is characterized in that: The memory cell is provided with a first transistor T1, a second transistor T2, a third transistor T3 and a fourth transistor T4; The gate of the first transistor T1 is connected to the input terminal Q of the storage unit, the drain of the first transistor T1 is connected to the output terminal Qb of the storage unit, and the source of the first transistor T1 is grounded; The gate of the second transistor T2 is connected to the input terminal Q of the storage unit, the drain of the second transistor T2 is connected to the output terminal Qb of the storage unit, and the source of the second transistor T2 is connected to the power supply; The gate of the third transistor T3 is connected to the output terminal Qb of the storage unit, the drain of the third transistor T3 is connected to the input terminal Q of the storage unit, and the source of the third transistor T3 is grounded; The gate of the fourth transistor T4 is connected to the output terminal Qb of the storage unit, the drain of the fourth transistor T4 is connected to the input terminal Q of the storage unit, and the source of the fourth transistor T4 is connected to the power supply.
3. The artificial intelligence acceleration system based on the 8T-SRAM computing and storage integrated macro unit according to claim 1 is characterized in that: The read-write calculation control unit RWCCU is provided with a first AND gate AND1, a second AND gate AND2, a third AND gate AND3, a fourth AND gate AND4, a first read-write isolation transistor A1, a second read-write isolation transistor A2, a third read-write isolation transistor A3 and a transistor A4; The first input terminals of the first AND gate AND1, the second AND gate AND2 and the third AND gate AND3 are connected to the first control signal CS', the second input terminals of the first AND gate AND1 and the second AND gate AND2 are connected to the second control signal R / W', and the second input terminal of the third AND gate AND3 is connected to the third control signal CE; The output end of the first AND gate AND1 is connected to the gates of the second read-write isolation transistor A2 and the third read-write isolation transistor A3, the sources of the second read-write isolation transistor A2 and the third read-write isolation transistor A3 are connected to the I / O port, the drain of the second read-write isolation transistor A2 is connected to the second column calculation line BLb, and the drain of the third read-write isolation transistor A3 is connected to the first column calculation line BL; The output end of the second AND gate AND2 is connected to the column analog accumulator unit CAAU and the first input end of the fourth AND gate AND4, the second input end of the fourth AND gate AND4 is connected to the output end of the third AND gate AND3, the output end of the fourth AND gate AND4 is connected to the gate of the first read-write isolation transistor A1, the source of the first read-write isolation transistor A1 is connected to the second column calculation line BLb, and the drain of the first read-write isolation transistor A1 is connected to the I / O port; The output terminal of the third AND gate AND3 is connected to the column analog accumulating unit CAAU, the read word line RWL and the gate of the transistor A4, the source of the transistor A4 is connected to the output terminal of the first AND gate AND1, and the drain of the transistor A4 is grounded.
4. The artificial intelligence acceleration system based on the 8T-SRAM computing and storage integrated macro unit according to claim 3 is characterized in that: The column analog accumulating unit CAAU is provided with a transistor A5, a transistor A6, a primary calculation capacitor, a secondary calculation capacitor, a transistor T12, a transistor T14, a transistor T15 and a transistor T16; The gates of the transistors T15 and T16 are connected to the output end of the second AND gate AND2, the drain of the transistor T15 is connected to the read bit line RBL, the source of the transistor T15 is connected to the drain of the transistor T16, and the source of the transistor T16 is connected to the input end of the analog-to-digital converter ADC via a charge sharing path; The gates of the transistor A5 and the transistor A6 are connected between the output end of the third AND gate AND3 and the read word line RWL, the source of the transistor A5 is connected to the drain of the transistor T16, the source of the transistor A5 is also connected in series with the main calculation capacitor and then grounded, and the drain of the transistor A5 is grounded; the source of the transistor A6 is connected to the source of the transistor T14, the source of the transistor A6 is also connected in series with the secondary calculation capacitor and then grounded, and the drain of the transistor A6 is grounded; the drain of the transistor T14 is connected to the drain of the transistor T15, the gate of the transistor T14 is connected to the gate of the transistor T12, the source of the transistor T12 is connected to the power supply, and the drain of the transistor T12 is connected to the read bit line RBL.
5. The artificial intelligence acceleration system based on the 8T-SRAM computing and storage integrated macro unit according to claim 4 is characterized in that: A phase automatic conversion circuit is also provided, and the phase automatic conversion circuit is provided with a transistor A7, a voltage comparator CMP1, a fifth AND gate AND5, a transistor T13 and a compensation capacitor; The first input terminal of the fifth AND gate AND5 is connected between the output terminal of the third AND gate AND3 and the read word line RWL, the output terminal of the fifth AND gate AND5 is connected to the gate of the transistor T12, the second input terminal of the fifth AND gate AND5 is connected to the output terminal of the voltage comparator CMP1, the first input terminal of the voltage comparator CMP1 is connected to the power supply, the second input terminal of the voltage comparator CMP1 is connected in series with a compensation capacitor and then grounded, the second input terminal of the voltage comparator CMP1 is also connected to the drain of the transistor T13, the source of the transistor T13 is connected to the power supply, and the gate of the transistor T13 is connected to the gate of the transistor T12; the second input terminal of the voltage comparator CMP1 is also connected to the source of the transistor A7, the drain of the transistor A7 is grounded, and the gate of the transistor A7 is connected between the output terminal of the third AND gate AND3 and the read word line RWL.
6. The artificial intelligence acceleration system based on 8T-SRAM computing and storage integrated macro unit according to claim 5 is characterized in that: It includes four phases of operation, namely, RBL pre-charging phase, column multiplication and accumulation discharge phase, charge averaging phase and analog-to-digital conversion phase. First, in the read bit line precharge stage, the control signals CS'=0, CE=1, and R / W'=0 are set, so that AND3 outputs the control signal VAND3=CE·CS'=1, controls the read-write isolation transistors A1-A3 to be turned off to achieve read-write disable, and at the same time, AND5 outputs the control signal VAND5=VAND3·CMP1'=1, controls the transistors T12 and T13 to be turned on, and charges the main calculation capacitor, the secondary calculation capacitor, and the compensation capacitor; Then it enters the column multiplication and accumulation discharge stage. When the compensation capacitor is full, VCC=VDD, the output of the voltage comparator CMP1 flips from 0 to 1, making VAND5=0, turning off transistors T12-T14, waiting for the input signal of the read word line RWL. When transistors T7 and T8 are turned on at the same time, the main calculation capacitor completes the multiplication and accumulation through the discharge path, and its discharge amount ΔQBL is proportional to the product of the weight W and the input signal X; Then, in the charge balancing stage, R / W'=1 is set, and the control signal VAND2=R / W'·CS'=1 is output through AND2, which controls the transistor T15 to be turned off and T16 to be turned on, and the connection between the main calculation capacitor and the charge balancing circuit is established, completing the charge redistribution process; Finally, in the analog-to-digital conversion stage, the master control MOS tube is turned on, and the amortized charge is introduced into the analog-to-digital converter ADC. The partial multiplication and accumulation calculation results pMACVn of each column are obtained by measuring the charge loss, and weighted summation is performed according to the 4-bit binary weight, and the final output digital result DOUT = Σ(pMACVn×2^n), n = 0, 1, 2, 3; During reset, CE=0 is set to turn on transistors A5, A6, and A7 at the same time, and the primary calculation capacitor, the secondary calculation capacitor, and the compensation capacitor are connected to the ground respectively to complete the discharge reset operation.
7. The artificial intelligence acceleration system based on 8T-SRAM computing and storage integrated macro unit according to claim 6 is characterized in that: The main calculation capacitor includes four specifications: 8Cu, 4Cu, 2Cu, and 1Cu, and the size of the compensation capacitor is 9Cu.
8. The artificial intelligence acceleration system based on 8T-SRAM computing and storage integrated macro unit according to claim 1 is characterized in that: A transistor T10 is provided on the first column calculation line BL, the drain of the transistor T10 is connected to the first column calculation line BL, the source of the transistor T10 is connected to the drain of the third read-write isolation transistor A3, and the gate of the transistor T10 is connected to the column address line selection control interface; A transistor T9 is provided on the second column calculation line BLb, the drain of the transistor T9 is connected to the second column calculation line BLb, the source of the transistor T9 is connected to the source of the first read-write isolation transistor A1 and the drain of the second read-write isolation transistor A2, and the gate of the transistor T9 is connected to the column address line selection control interface; A transistor T11 is provided on the read word line RWL, a source and a drain of the transistor T11 are connected to the read word line RWL, and a gate of the transistor T11 is connected to a read-write calculation control unit RWCCU and a column analog accumulation unit CAAU.
9. The artificial intelligence acceleration system based on 8T-SRAM computing and storage integrated macro unit according to claim 1 is characterized in that: The digital-analog hybrid computing array supports 16×4-bit MAC operations.
Citation Information
Patent Citations
Voltage margin enhanced capacitance coupling storage and calculation integrated unit, subarray and device
CN113255904A
SRAM storage and calculation integrated chip based on capacitive coupling
CN115048075A
SRAM storage and computing integrated chip based on capacitive coupling
WO2023207441A1
Cited By
10T-SRAM unit, read-write damage resistant dual-port SRAM circuit and chip
CN121641112A