An in-memory computing eDRAM accelerator for convolutional neural networks

By using the 5T1C ping-pong eDRAM bit cell array and noise compensation circuit in the CIM CNN accelerator, high throughput and high parallelism convolutional neural network calculation is realized, solving the problem of improving throughput and weight accuracy in the existing technology, and achieving efficient computing density and low power consumption effect.

CN113946310BActive Publication Date: 2025-08-12SHANGHAI TECH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202111169936.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-08
Publication Date
2025-08-12
Estimated Expiration
2041-10-08

AI Technical Summary

Technical Problem

The existing CIM CNN accelerators based on multi-bit eDRAM have challenges in improving throughput, and the throughput of SAR ADC has not been effectively improved, and it is difficult to improve weight accuracy and parallelism without increasing area.

Method used

The 5T1C ping-pong eDRAM bit cell array is used for parallel in-memory convolution operations, combined with the dual 2T readout port and noise compensation circuit, supports parallel in-image and parallel convolution between images, and realizes non-2 and 2's complement calculation through the cross connection of SAR ADC units, and the input sampling capacitance of the accumulated bit lines is allocated without increasing the area.

Benefits of technology

The peak calculation density is improved by about 30 times, reaching 59.1TOPS/mm2, significantly improving throughput and weight accuracy, parallelism, while reducing power and energy consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113946310B_ABST
    Figure CN113946310B_ABST
Patent Text Reader

Abstract

The present invention provides an in-memory computing eDRAM accelerator for convolutional neural networks. The accelerator is characterized by comprising four P2ARAM blocks, each of which includes a 5T1C ping-pong eDRAM bit cell array consisting of 64x16 5T1C ping-pong eDRAM bit cells. Within each P2ARAM block, 64x2 digital-to-time converters convert 4-bit activation values in the row direction into different pulse widths, which are then input into the 5T1C ping-pong eDRAM bit cell array for computation. A total of 16x2 convolution results are output in the column direction of the 5T1C ping-pong eDRAM bit cell array. The convolutional neural accelerator proposed in the present invention utilizes parallel multi-bit storage and convolution in the 5T1C ping-pong eDRAM bit cells. Without adding additional area overhead, the input sampling capacitors of the accumulation bit lines are distributed to the sign-value SAR ADC units of the CDAC array, thus proposing an S2M-ADC solution. In this way, the eDRAM-based in-memory computing neural network accelerator disclosed in the present invention achieves a peak computational density of 59.1TOPS / mm. 2 , which is about 30 times higher than previous work.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a CIM CNN accelerator based on eDRAM. Background Art

[0002] Among various deep neural network structures, the convolutional neural network (CNN) is the most widely used one. It was proposed by LeCun in 1989 [1]. Early CNNs were successfully applied to handwritten character image recognition [1-3]. In 2012, the deepening of the AlexNet network [4] was successful. Since then, CNN has flourished and has been widely used in various fields, achieving state-of-the-art results on many problems. After the emergence of AlexNet, convolutional neural networks were quickly applied to various tasks in machine vision, including pedestrian detection, face recognition, object segmentation, object tracking, etc., and have achieved success [5-8]. In terms of hardware implementation, compute-in-memory (CIM) is a promising computing platform for implementing neural networks for high-throughput constrained artificial intelligence applications. Although recent SRAM-based CIM works have achieved considerable throughput in both analog and digital domains [9-11], some eDRAM-based CIM works [12,13] have achieved higher throughput.

[0003] However, as shown in Figure 1, current multi-bit eDRAM-based CIM CNN accelerators still face some challenges in further improving throughput.

[0004] References:

[0005] [1]Y.LeCun,B.Boser,JSDenker,D.Henderson,REHoward,W.Hubbard,andL.D.Jackel.Backpropagation applied to handwritten zip code recognition.NeuralComputation,1989.

[0006] [2]Y.LeCun,B.Boser,J.S.Denker,D.Henderson,R.E.Howard,W.Hubbard,andL.D.Jackel.Handwritten digit recognition with a back-propagation network.InDavid Touretzky,editor,Advances in Neural Information Processing Systems 2(NIPS*89),Denver,CO,1990,Morgan Kaufman.

[0007] [3]Y.LeCun,L.Bottou,Y.Bengio,and P.Haffner.Gradient-based learningapplied to document recognition.Proceedings of the IEEE,november 1998.

[0008] [4]Alex Krizhevsky,Ilya Sutskever,Geoffrey E.Hinton.ImageNetClassification with Deep Convolutional Neural Networks.

[0009] [5]P.Yang,G.Zhang,L.Wang,L.Xu,Q.Deng and M.-H.Yang,"A Part-AwareMulti-Scale Fully Convolutional Network for Pedestrian Detection,"in IEEETransactions on Intelligent Transportation Systems,vol.22,no.2,pp.1125-1137,Feb.2021.

[0010] [6]Y.Huang and H.Hu,"A Parallel Architecture of Age AdversarialConvolutional Neural Network for Cross-Age Face Recognition,"in IEEETransactions on Circuits and Systems for Video Technology,vol.31,no.1,pp.148-159,Jan.2021.

[0011] [7]L.Ma,Y.Li,J.Li,W.Tan,Y.Yu and M.A.Chapman,"Multi-Scale Point-WiseConvolutional Neural Networks for 3D Object Segmentation From LiDAR PointClouds in Large-Scale Environments,"in IEEE Transactions on IntelligentTransportation Systems,vol.22,no.2,pp.821-836,Feb.2021.

[0012] [8]J.Fang and G.Liu,"Visual Object Tracking Based on Mutual LearningBetween Cohort Multiscale Feature-Fusion Networks With Weighted Loss,"in IEEETransactions on Circuits and Systems for Video Technology,vol.31,no.3,pp.1055-1065,March 2021.

[0013] [9]J.Yue et al.,"14.3 A 65nm Computing-in-Memory-Based CNN Processorwith 2.9-to-35.8TOPS / W System Energy Efficiency Using Dynamic-SparsityPerformance-Scaling Architecture and Energy-Efficient Inter / Intra-Macro DataReuse,"ISSCC,pp.234-236,2020.

[0014]

[10] Q.Dong et al.,“A 351TOPS / W and 372.4GOPS Compute-in-Memory SRAMMacro in 7nm FinFET CMOS for Machine-Learning Applications”,ISSCC,pp.242-243,2020.

[0015]

[11] Y.-D.Chih et al.,"16.4 An 89TOPS / W and 16.3TOPS / mm2 All-DigitalSRAM-Based FullPrecision Compute-In Memory Macro in 22nm for Machine-LearningEdge Applications,"ISSCC,pp.252-254,2021.

[0016]

[12] Z.Chen et al.,"15.3 A 65nm 3T Dynamic Analog RAM-Based Computing-in-Memory Macro and CNN Accelerator with Retention Enhancement,AdaptiveAnalog Sparsity and 44TOPS / W System Energy Efficiency,"ISSCC,pp.240-242,2021.

[0017]

[13] S.Xie et al., "16.2 eDRAM-CIM: Compute-In-Memory Design withReconfigurable EmbeddedDynamic-Memory Array Realizing Adaptive DataConverters and Charge-Domain Computing," ISSCC, pp.248-250, 2021. Summary of the Invention

[0018] The purpose of the present invention is to provide a CIM CNN accelerator to achieve: (1) improve macro throughput with fewer transistors, improve weight accuracy and parallelism; and (2) improve the throughput of SAR ADC without increasing additional area.

[0019] To achieve the above objectives, the technical solution of the present invention is to provide an in-memory computing eDRAM accelerator for convolutional neural networks, characterized in that it includes four P2ARAM blocks, each P2ARAM block includes a 5T1C ping-pong eDRAM bit cell array consisting of 64x16 5T1C ping-pong eDRAM bit cells, each 5T1C ping-pong eDRAM bit cell adopts a 5T1C circuit structure and has dual 2T readout ports, the two readout ports are respectively connected to accumulation bit line 1 and accumulation bit line 2, and the two readout ports respectively correspond to two activation value input terminals;

[0020] The dual 2T readout ports of the 5T1C Ping-Pong eDRAM bit cell support parallel in-memory convolution operations at the bit cell level. Within a single cycle, the two readout ports perform convolution and bit line reset in parallel. The two parallel readout ports operate in a ping-pong manner. The readout port performing bit line reset completes convolution in the next cycle, while the readout port completing convolution completes bit line reset in the next cycle. The readout port performing convolution calculations hides the bit line pre-discharge overhead.

[0021] The eDRAM cell storage node of each 5T1C ping-pong eDRAM bit cell is used to store an analog weight value and a voltage value with reverse turn-off noise. The reverse turn-off noise is generated by a noise compensation circuit. When the write transistor of each eDRAM cell storage node is turned off, the forward turn-off noise and the reverse turn-off noise stored in the current eDRAM cell storage node cancel each other out, thereby reducing the impact of noise on the analog weight value stored on the eDRAM cell storage node.

[0022] In each P2ARAM block, 64x2 digital-to-time converters convert the 4-bit activation values into different pulse widths along the row direction, which are then fed into the 5T1C ping-pong eDRAM bit cell array for calculation. A total of 16x2 convolution results are output along the column direction of the 5T1C ping-pong eDRAM bit cell array. The convolution is achieved by accumulating multiple 5T1C ping-pong eDRAM bit cells on the bit line to simultaneously charge the input sampling capacitor of the SAR ADC unit. The SAR ADC then reads the voltage on the input sampling capacitor.

[0023] The input sampling capacitors on the accumulation bit lines are merged into the SAR ADC unit connected to the current accumulation bit line, and the area of the input sampling capacitors on the accumulation bit lines is shared for the C-DAC capacitors of the SAR ADC unit; the 16 columns of 5T1C ping-pong eDRAM bit cells in the 5T1C ping-pong eDRAM bit cell array are grouped into two columns, and in one group, one column of 5T1C ping-pong eDRAM bit cells is a sign column, and the other column of 5T1C ping-pong eDRAM bit cells is a value column. The accumulation bit line 1 and the accumulation bit line 2 of the sign column are respectively connected to three SAR ADC units, and the SAR ADC unit is redefined as an RS ADC unit; the accumulation bit line 1 and the accumulation bit line 2 of the value column are respectively connected to three SAR ADC units, and the SAR ADC unit is redefined as an RM ADC unit; the 12 related SAR ADC units corresponding to a group of 5T1C ping-pong eDRAM bit cell columns are divided and interleaved, wherein the three RS ADC units connected to the accumulation bit line 1 of the sign column and the three RM ADC units connected to the accumulation bit line 1 of the value column are respectively connected to the accumulation bit line 1 of the value column. The ADC units are interleaved. The three RS ADC units connected to the accumulation bit line 2 of the sign column are interleaved with the three RM ADC units connected to the accumulation bit line 2 of the value column. The modes of the two interleaved SAR ADC units are configured to support non-2's and 2's complement calculations:

[0024] When performing a two's complement calculation: the interleaved RM ADC units and RS ADC units are merged in pairs, and the merged RM ADC units and RS ADC units are used as one ADC for conversion. At this time, the sign bit column is used to store a 1-bit sign value, and the value bit column is used to store a 5-bit other bit value. The input sampling capacitors of the RS ADC units obtain the result of the sign bit multiplication, and the input sampling capacitors of the RM ADC units obtain the result of the value bit multiplication. The input sampling capacitors of the RS ADC units and the input sampling capacitors of the RM ADC units directly read out a 6-bit two's complement value through the RS ADC units.

[0025] When performing non-two's complement calculations: the RM ADC unit and the RS ADC unit perform the conversion independently. In this case, the sign bit column and the value bit column are calculated independently, and the sign bit column and the value bit column respectively store 5-bit non-two's complement values. The RM ADC unit and the RS ADC unit simultaneously read the 5-bit non-two's complement values from their respective input sampling capacitors.

[0026] The SAR ADC unit operation and jump control logic are tightly coupled in a bit-serial manner, supporting cross-layer calculation and early termination of convolutional layers, activation function layers, and maximum pooling layers.

[0027] Preferably, the 5T1C ping-pong eDRAM bit cell uses an NMOS transistor as a write transistor, and the dual 2T readout ports are provided by a PMOS transistor.

[0028] Preferably, the noise compensation circuit includes an operational amplifier and a write noise compensation unit. The target current permutations and combinations are superimposed to obtain a unit current of 0 to 32 times. After setting the target current multiple, the operational amplifier calculates the analog voltage required by the eDRAM unit storage node. The analog voltage is written into 20 write noise compensation units respectively through a write transistor of the write noise compensation unit. Subsequently, the write transistor of each noise compensation unit is turned off, and the read transistor of each noise compensation unit is turned on. At this time, the analog voltage stored in the 20 write noise compensation units obtains reverse turn-off noise. This analog voltage with reverse turn-off noise drives each write bit line WBL through a subsequent voltage follower and is written into each eDRAM unit storage node of the 5T1C ping-pong eDRAM bit cell array in rows.

[0029] Preferably, the 5T1C ping-pong eDRAM bit cell supports two parallel convolution modes: intra-image parallelism and inter-image parallelism;

[0030] In intra-image parallel convolution mode, the 5T1C Ping-Pong eDRAM bit cells perform segmented convolution on the same image. The pixels or activation values in the upper half of the image are convolved using the accumulation bit line 1 corresponding to one activation input, while the pixels or activation values in the lower half of the image are convolved using the accumulation bit line 2 corresponding to the other activation input.

[0031] In the inter-image parallel convolution mode, the first image obtains the convolution operation result from the accumulation bit line 1 corresponding to one activation value input terminal, and the second image obtains the convolution operation result from the accumulation bit line 2 corresponding to the other activation value input terminal.

[0032] Preferably, the working phases of the three SAR ADC units connected to the same accumulation bit line differ by exactly 2 cycles, so that the convolution results on the corresponding accumulation bit lines are cyclically sampled.

[0033] The convolutional neural accelerator proposed in this invention uses: 5T1C ping-pong eDRAM bit cells for parallel multi-bit storage and convolution; without adding additional area overhead, the input sampling capacitors of the accumulation bit lines are distributed to the sign-value SAR ADC units of the CDAC array, proposing an S2M-ADC solution. In this way, the eDRAM-based in-memory computing neural network accelerator disclosed in this invention achieves a peak computing density of 59.1TOPS / mm 2 , which is about 30 times higher than previous work [9]. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1A 、 Figure 1B 、 Figure 1C and Figure 2 The challenges and comparisons between the current state-of-the-art design and the design disclosed in this invention are illustrated, wherein: Figure 1A 、 Figure 1B and Figure 1C This shows that the present invention uses fewer transistors than the most advanced designs to achieve higher weight accuracy and parallelism, thereby improving throughput. Figure 1A The difference in specific structure between the present invention and the most advanced design is shown. Figure 1B and Figure 1C The difference between the present invention and the state-of-the-art design in terms of transistor count and weight precision × parallelism is shown.

[0035] Figure 2 It is shown that the present invention improves the throughput of SAR ADC without additional area overhead;

[0036] Figure 3 The overall structure of the present invention;

[0037] Figure 4 This is the schematic diagram of the 5T1C ping-pong eDRAM bit cell. In the figure, 1Cal. represents the DTC1 convolution operation, 1PD represents the ABL1 pre-discharge operation, 2Cal. represents the DTC2 convolution operation, and 2PD represents the ABL1 pre-discharge operation.

[0038] Figure 5 Schematic diagram of ping-pong convolution method;

[0039] Figure 6 Supplementary circuit diagram for writing noise;

[0040] Figure 7 It is a schematic diagram of dual working mode;

[0041] Figure 8 Design a schematic diagram for the S2M-ADC combination;

[0042] Figure 9 Schematic diagram of the pipeline operation S2M solution based on SAR ADC proposed in this invention

[0043] Figure 10 This is a schematic diagram of the principle of the S2M-ADC solution;

[0044] Figure 11 Schematic diagram for implementing 2's complement / non-2's complement;

[0045] Figure 12 Schematic diagram of ReLU+Max-Pooling early termination

[0046] Figure 13 For measurement results;

[0047] Figure 14 For comparison results. DETAILED DESCRIPTION

[0048] Below in conjunction with specific embodiment, further set forth the present invention.Should be understood that these embodiments are only used to illustrate the present invention and are not used in limiting the scope of the present invention.In addition, should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms fall equally within the scope limited by the appended claims of the application.

[0049] like Figure 3 As shown, the present invention provides an in-memory computing eDRAM accelerator for convolutional neural networks, including four P2ARAM blocks, each P2ARAM block including a 5T1C ping-pong eDRAM bit cell array consisting of 64x16 5T1C ping-pong eDRAM bit cells, and the 5T1C ping-pong eDRAM bit cells are used for multi-bit storage and parallel convolution.

[0050] Combine Figure 4 , each 5T1C ping pong eDRAM bit cell uses an NMOS transistor as a write transistor for write operations, and the dual 2T read ports use PMOS transistors. The two PMOS read ports of the 5T1C ping pong eDRAM bit cell are respectively connected to the two accumulation bit lines ABL1, ABL2 or ABL3, ABL4, and the two PMOS read ports correspond to the two PMOS activation value input terminals DTC1, DTC2. The main purpose of choosing PMOS transistor readout is to reduce the overall size of the sampling capacitor on the accumulation bit line, because compared with NMOS transistors, PMOS transistors can provide lower calculation (charging) current under large DTC pulse width. In this way, all accumulation bit lines need to be pre-discharged to GND before the calculation operation. In the present invention, the SN node capacitor connected to VDD has stronger noise resistance than the SN node capacitor connected to GND.

[0051] The dual 2T readout ports used in the 5T1C Ping-Pong eDRAM bit cell of this invention support parallel in-memory convolution operations at the bit cell level. For a conventional single port, the following operations are required: convolution -> bit line reset -> convolution -> bit line reset -> convolution -> bit line reset -> ..., so a convolution operation requires two cycles to complete. For the dual 2T port of the present invention, assuming that the PMOS read port connected to the accumulation bit line ABL1 or ABL3 is read port No. 1, and the PMOS read port connected to the accumulation bit line ABL2 or ABL4 is read port No. 2, the present invention can implement the following operations: convolution (read port No. 1) + bit line reset (read port No. 2) -> bit line reset (read port No. 1) + convolution (read port No. 2) -> ..., that is, within the same cycle, the two PMOS read ports complete convolution and bit line reset in parallel, the PMOS read port in the bit line reset state (that is, the PMOS read port in the pre-discharge state) completes convolution in the next cycle, and the PMOS read port that completes convolution completes the bit line reset in the next cycle. Therefore, in the present invention, the convolution operation can be completed in each cycle.

[0052] like Figure 5 As shown in Figure 1, two parallel PMOS read ports operate in a ping-pong manner. The PMOS read port in the convolution calculation stage hides the bit line pre-discharge overhead. Therefore, the 5T1C ping-pong eDRAM bit cell proposed in this invention increases the throughput by 2 times. At the same time, the ping-pong operation can significantly eliminate the noise coupling from the source (s) and drain (d) to the SN node (g), which is caused by the parasitic capacitance C gd and C gs produce.

[0053] The eDRAM cell storage node (hereinafter referred to as "SN node") of each 5T1C ping-pong eDRAM bit cell is used to store analog weight values and voltage values with reverse turn-off noise. The reverse turn-off noise is generated by a noise compensation circuit composed of an operational amplifier and a write noise compensation unit WNCC. Figure 6The target current permutations and combinations can be combined to produce a unit current multiplier of 0 to 32. After setting the target current multiplier, the operational amplifier calculates the required analog voltage at the SN node. This analog voltage is then written into each of the 20 write noise compensation cells WNCC via a write transistor in the write noise compensation cell WNCC. Subsequently, the control signal NCC_bar for each noise compensation cell WNCC switches from 1 to 0, turning off the write transistor in the noise compensation cell WNCC. The control signal NCC_bar for each noise compensation cell WNCC switches from 0 to 1, turning on the read transistor in each noise compensation cell WNCC. At this point, the analog voltages stored in the 20 write noise compensation cells WNCC acquire reverse turn-off noise. This analog voltage with reverse turn-off noise drives each write bit line WBL through a subsequent voltage follower, writing it row-by-row to each SN node in the 5T1C ping-pong eDRAM bit cell array. When the write transistor in each SN node is turned off, the forward turn-off noise and the reverse turn-off noise stored in the SN node cancel each other out. The present invention effectively suppresses write operation coupling noise through the write noise compensation circuit unit, generating a reverse noise amplitude during the write phase to compensate for the write noise. This minimizes the impact of noise on the analog weight values stored on the SN node, which is crucial for the accuracy of inference.

[0054] Before convolution, the pre-trained digital weight values are converted into analog weight values through the noise compensation circuit. Similar to the above, the analog weight values are stored in each SN node of the 5T1C ping-pong eDRAM bit cell array by row through the control signal on the word line WL.

[0055] In addition, the 5T1C ping-pong eDRAM bit cell proposed in the present invention supports two parallel convolution modes: intra-image parallel and inter-image parallel. Figure 7 As shown in the figure. In intra-image parallel convolution mode, the 5T1C ping-pong eDRAM bit cells perform segmented convolution on the same image. The pixels or activation values in the upper half of the image are convolved using the accumulation bit line ABL1 corresponding to the PMOS activation value input terminal DTC1, while the pixels or activation values in the lower half of the image are convolved using the accumulation bit line ABL2 corresponding to the PMOS activation value input terminal DTC2. Inter-image parallel convolution mode is used in dual-camera scenarios. The first image is convolved using the accumulation bit line ABL1 corresponding to the PMOS activation value input terminal DTC1, while the second image is convolved using the accumulation bit line ABL2 corresponding to the PMOS activation value input terminal DTC2.

[0056] In each P2ARAM block, 64x2 digital-to-time converters (DTCs) convert the 4-bit activation value in the row direction into pulses of varying widths, which are then fed into the 5T1C ping-pong eDRAM bit cell array for computation. A total of 16x2 convolution (CONV) outputs are generated in the column direction of the 5T1C ping-pong eDRAM bit cell array. Convolution is achieved by accumulating the simultaneous charging of the input sampling capacitor of the SAR ADC unit by multiple 5T1C ping-pong eDRAM bit cells on the bit line. The constant current charge value of each 5T1C ping-pong eDRAM bit cell is determined by the voltage stored on the SN node: a smaller stored voltage value results in a larger constant current value, while a larger stored voltage value results in a smaller constant current value. The constant current discharge time of each 5T1C ping-pong eDRAM bit cell is determined by the DTC pulse width: a wider pulse length results in a longer charging time. This results in a mixed charge on the input sampling capacitor, which serves as the final convolution result. Finally, the voltage value of the input sampling capacitor is read out by the SAR ADC.

[0057] In this invention, the input sampling capacitor on the accumulation bit line is merged into the SAR ADC unit connected to the current accumulation bit line, and an S2M-ADC scheme is proposed. In this invention, the connection method between the SAR ADC unit and the 5T1C ping-pong eDRAM bit cell array is referenced. Figure 7 The 16 columns of 5T1C ping-pong eDRAM bit cells in the 5T1C ping-pong eDRAM bit cell array are grouped into two columns. Within each group, one column of 5T1C ping-pong eDRAM bit cells is the sign column, and the other column of 5T1C ping-pong eDRAM bit cells is the magnitude column. Accumulation bit lines ABL1 and ABL2 for the sign column are each connected to three SAR ADC cells, which are redefined as RS ADC cells, where RS stands for result of sign. Accumulation bit lines ABL3 and ABL4 for the magnitude column are each connected to three SAR ADC cells, which are redefined as RM ADC cells, where RM stands for result of magnitude. These 12 related SAR ADC cells are divided and interleaved. The three RS ADC cells connected to accumulation bit line ABL1 intersect with the three RM ADC cells connected to accumulation bit line ABL3, and the three RS ADC cells connected to accumulation bit line ABL2 intersect with the three RM ADC cells connected to accumulation bit line ABL4. Non-2's and 2's complement calculations are supported by configuring the interleaved mode of the two SAR ADC units.

[0058] The operating phases of the three SAR ADC units connected to the same accumulation bit line differ by exactly two cycles. That is, after the first SAR ADC unit starts working, the second SAR ADC unit starts working in the third cycle, and the third SAR ADC unit starts working in the fifth cycle, thereby cyclically sampling the convolution results on the corresponding accumulation bit line.

[0059] When performing 2's complement calculation: merge the RM ADC unit and the RS ADC unit that implement the crossover (switch φ = 0, ), the combined RM ADC unit and the RS ADC unit perform conversion as a single ADC. In this case, the sign bit column stores the 1-bit sign value, and the magnitude bit column stores the 5-bit other bit value. The RS ADC unit's input sampling capacitors receive the result of the sign bit multiplication, while the RM ADC unit's input sampling capacitors receive the result of the magnitude bit multiplication. The RS ADC unit directly reads a 6-bit 2's complement value from both the RS and RM ADC unit's input sampling capacitors.

[0060] When performing non-2's complement calculations: the RM ADC unit and the RS ADC unit perform conversions independently (switch φ = 1, At this time, the sign and magnitude columns are calculated independently, so each column stores a 5-bit non-2's complement value. The RM ADC unit and the RS ADC unit simultaneously read the 5-bit non-2's complement value from their respective input sampling capacitors.

[0061] The SAR ADC unit operation and jump control logic are tightly coupled in a bit-serial manner, supporting simultaneous cross-layer calculation and early termination of convolutional layers, activation function layers, and maximum pooling layers (CONV, ReLU, Max-Pooling), saving energy without losing accuracy, thereby achieving full on-chip calculation while reconfiguring the kernel size of the maximum pooling layer.

[0062] Take VGG16 as an example: (1) If the convolution layer is followed by only the ReLU layer. If the first bit is "1" (indicating that a negative number has been sampled), the ADC readout is terminated in advance, because no matter how large the negative value is, the result obtained by the ReLU function must be 0. If the first bit read out is "0" (indicating that a positive number has been sampled), then the subsequent bit conversion is performed. (2) If the convolution layer is followed by both the ReLU layer and the maximum pooling layer. If the first bit is "1" (indicating that a negative number has been sampled), the ADC readout is terminated in advance, because no matter how large the negative value is, the result obtained by the ReLU function must be 0. If the first bit read out is "0" (indicating that a positive number has been sampled), then the subsequent bit conversion is performed. If the kernel of the maximum pooling layer is 2X2, then the largest value needs to be selected in 2X2, and the reading of other values needs to be terminated in advance.

[0063] In the subsequent comparison process, the SAR ADC unit that first outputs a result stores it in a digital register. The remaining three numbers are then compared bit by bit with the value in the register. If a bit in a value is found to be greater than the value in the register, it is replaced by that value (possibly the maximum value). If a bit in a value is found to be less than the register value, the reading of that value is terminated prematurely (it is definitely not the maximum value). If a bit in a value is found to be equal to the register value, the next bit is read for comparison.

[0064] Compared with the most advanced design, the area of the MOM capacitor for cumulative bit line sampling is shared with the C-DAC capacitor of the SAR ADC, which allows three SAR ADCs to be implemented per bit line without excess area overhead. The three SAR ADC units on the same bit line are operated in a pipeline manner. Under the delay determined by the SAR ADC unit, all SAR ADC units work in parallel under non-2's complement operation to improve the overall throughput. Through the mode conversion switch, two adjacent SAR ADC units are combined to implement 2's complement operation. The sampling switch is implemented using a local NMOS device (Zero-VT), and the sampling switch is turned off using a -200mV power rail. Figure 12 The skip control principles for ReLU and Max-Pooling are presented, enabling seamless implementation of different computational layers, such as CONV, ReLU, and Max-Pooling. Because the SAR ADC unit outputs sequentially from the most significant bit to the least significant bit, the skip controller can perform sign bit detection and bit-serial comparison, leading to early termination. The termination signal is also used to shut down the SAR ADC unit conversion, ensuring maximum energy conservation during this phase. When implemented on VGG-16 and tested on the CIFAR-10 dataset, the proposed S2M-ADC solution achieved approximately 1.82x power reduction.

[0065] Figure 13 The measurement results of a 4Kb CIM-based P2ARAM accelerator test chip made in 55nm CMOS process are shown. The chip can operate reliably with a power supply of 0.8~1.2V and a power consumption of 10.4~84.7mW at a frequency of 91.2~354.6MHz. Figure 13 As shown in the upper left, since the memory weight needs to be refreshed regularly after losing 1LSB, the refresh power consumption is 2.1 to 16mW, from 0.8V to 1.2V, accounting for about 21% of the total system power consumption.

[0066] Figure 13 The upper right panel shows the test accuracy of the CIFAR-10 and CIFAR-100 datasets under different configurations of activation and weight precision. It can be seen that the inference accuracy of CIFAR-10 shows better tolerance than that of CIFAR-100 in both the 2's (4b×6b) and non-2's (4b×5b) modes, while the inference accuracy of CIFAR-100 is more susceptible to degradation in activation and weight precision. In addition, a 1LSB voltage drift at the SN node results in significant accuracy loss. In the 2's mode, the overall accuracy of CIFAR-10 reaches 90.68% with a bit precision of 4b×6b. Compared to the baseline accuracy, the accuracy of CIFAR-100 is an acceptable 66.82% with a bit precision of 8b×6b. Furthermore, the ReLU and Max-Pooling skip scheme achieves an average 1.82× power reduction, with a peak system energy efficiency of 305.4 TOPS / W in non-2's complement mode at 0.8 V, while this energy efficiency is reduced by half (152.7 TOPS / W) in 2's complement mode at 0.8 V. In non-2's complement mode at 1.2 V, the peak computational density reaches 59.1 TOPS / mm2, which is approximately 30 times higher than previous work.

[0067] Figure 14 This paper summarizes the comparison between the present invention and previous studies. The present invention supports 2's and non-2's complement CONV operations, can calculate 4b / 8b type activations and 3 / 4 / 5 / 6b type weights, and is commonly used in multi-bit quantized CNN models. The peak operating frequency is 355MHz at 1.2V, which is due to the sampling capacitor fusion of the SAR ADC solution. Compared with the state-of-the-art work, the CIM-based P2ARAM macromodule achieves a computational density of 59.1TOPS / mm2, which is about 30X higher than the current best work, and the highest energy efficiency is 305.4TOPS / W.

[0068] The 64X64 CIM-based P2ARAM accelerator provided by the present invention is manufactured using a 55nm CMOS process. The peak classification accuracy of the accelerator on the CIFAR-10 and CIFAR-100 datasets is 90.68% and 66.92% respectively.

Claims

1. An in-memory computing eDRAM accelerator for convolutional neural networks, characterized in that: The P2ARAM includes four P2ARAM blocks, each of which includes a 5T1C ping-pong eDRAM bit cell array consisting of 64x16 5T1C ping-pong eDRAM bit cells. Each 5T1C ping-pong eDRAM bit cell adopts a 5T1C circuit structure and has dual 2T readout ports. The two readout ports are respectively connected to accumulation bit line 1 and accumulation bit line 2, and the two readout ports correspond to two activation value input terminals. The dual 2T readout ports of the 5T1C Ping-Pong eDRAM bit cell support parallel in-memory convolution operations at the bit cell level. Within a single cycle, the two readout ports perform convolution and bit line reset in parallel. The two parallel readout ports operate in a ping-pong manner. The readout port performing bit line reset completes convolution in the next cycle, while the readout port completing convolution completes bit line reset in the next cycle. The readout port performing convolution calculations hides the bit line pre-discharge overhead. The eDRAM cell storage node of each 5T1C ping-pong eDRAM bit cell is used to store an analog weight value and a voltage value with reverse turn-off noise. The reverse turn-off noise is generated by a noise compensation circuit. When the write transistor of each eDRAM cell storage node is turned off, the forward turn-off noise and the reverse turn-off noise stored in the current eDRAM cell storage node cancel each other out, thereby reducing the impact of noise on the analog weight value stored on the eDRAM cell storage node. In each P2ARAM block, 64x2 digital-to-time converters convert the 4-bit activation values into different pulse widths along the row direction, which are then fed into the 5T1C ping-pong eDRAM bit cell array for computation. A total of 16x2 convolution outputs are generated along the column direction of the 5T1C ping-pong eDRAM bit cell array. The convolution is accomplished by accumulating multiple 5T1C ping-pong eDRAM bit cells on the bit line to simultaneously charge the input sampling capacitors of the SAR ADC unit. The SAR ADC then reads the voltage on the input sampling capacitors. Merging the input sampling capacitor on the accumulation bit line into the SAR ADC unit connected to the current accumulation bit line, so that the area of the input sampling capacitor on the accumulation bit line is shared by the C-DAC capacitor of the SAR ADC unit; The 16 columns of 5T1C ping-pong eDRAM bit cells in the 5T1C ping-pong eDRAM bit cell array are grouped into two columns. In one group, one column of 5T1C ping-pong eDRAM bit cells is a sign column, and the other column of 5T1C ping-pong eDRAM bit cells is a value column. The accumulation bit line 1 and the accumulation bit line 2 of the sign column are respectively connected to three SAR ADC units, and the SAR ADC unit is redefined as an RS ADC unit; the accumulation bit line 1 and the accumulation bit line 2 of the value column are respectively connected to three SAR ADC units, and the SAR ADC unit is redefined as an RM ADC unit. The 12 related SAR ADC units corresponding to a group of 5T1C ping-pong eDRAM bit cell columns are divided and interleaved, wherein the three RS ADC units connected to the accumulation bit line 1 of the sign column are interleaved with the three RM ADC units connected to the accumulation bit line 1 of the value column, and the three RS ADC units connected to the accumulation bit line 2 of the sign column are interleaved with the three RM ADC units connected to the accumulation bit line 2 of the value column. The ADC units are interleaved to support non-2's and 2's complement calculations by configuring the modes of the two interleaved SAR ADC units: When performing a two's complement calculation: the interleaved RM ADC units and RS ADC units are merged in pairs, and the merged RMADC units and RS ADC units are used as one ADC for conversion. At this time, the sign bit column is used to store a 1-bit sign value, and the value bit column is used to store a 5-bit other bit value. The input sampling capacitors of the RS ADC units obtain the result of the sign bit multiplication, and the input sampling capacitors of the RMADC units obtain the result of the value bit multiplication. The input sampling capacitors of the RS ADC units and the input sampling capacitors of the RM ADC units directly read out a 6-bit two's complement value through the RS ADC units. When performing non-two's complement calculations: the RM ADC unit and the RS ADC unit perform conversions independently. In this case, the sign bit column and the value bit column are calculated independently, and the sign bit column and the value bit column respectively store 5-bit non-two's complement values. The RM ADC unit and the RS ADC unit simultaneously read the 5-bit non-two's complement values from their respective input sampling capacitors. The SAR ADC unit operation and jump control logic are tightly coupled in a bit-serial manner, supporting cross-layer calculation and early termination of convolutional layers, activation function layers, and maximum pooling layers.

2. The in-memory computing eDRAM accelerator for convolutional neural networks according to claim 1, wherein: The 5T1C ping-pong eDRAM bit cell uses an NMOS transistor as a write transistor, and the dual 2T readout ports are provided by PMOS transistors.

3. The in-memory computing eDRAM accelerator for convolutional neural networks according to claim 1, wherein: The noise compensation circuit includes an operational amplifier and a write noise compensation unit. Target current permutations and combinations are superimposed to obtain a unit current of 0 to 32 times. After the target current multiple is set, the operational amplifier calculates the analog voltage required by the eDRAM cell storage node. This analog voltage is written into each of the 20 write noise compensation units through a write transistor of the write noise compensation unit. Subsequently, the write transistor of each noise compensation unit is turned off, and the read transistor of each noise compensation unit is turned on. At this time, the analog voltage stored in the 20 write noise compensation units obtains reverse turn-off noise. This analog voltage with reverse turn-off noise drives each write bit line WBL through a subsequent voltage follower and is written into each eDRAM cell storage node of the 5T1C ping-pong eDRAM bit cell array row by row.

4. The in-memory computing eDRAM accelerator for convolutional neural networks according to claim 1, wherein: The 5T1C ping-pong eDRAM bit cell supports two parallel convolution modes: intra-image parallelism and inter-image parallelism; In intra-image parallel convolution mode, the 5T1C Ping-Pong eDRAM bit cells perform segmented convolution on the same image. The pixels or activation values in the upper half of the image are convolved using the accumulation bit line 1 corresponding to one activation input, while the pixels or activation values in the lower half of the image are convolved using the accumulation bit line 2 corresponding to the other activation input. In the inter-image parallel convolution mode, the first image obtains the convolution operation result from the accumulation bit line 1 corresponding to one activation value input terminal, and the second image obtains the convolution operation result from the accumulation bit line 2 corresponding to the other activation value input terminal.

5. The in-memory computing eDRAM accelerator for convolutional neural networks according to claim 1, wherein: The working phases of the three SAR ADC units connected to the same accumulation bit line differ by exactly two cycles, thereby performing cyclic sampling on the convolution results on the corresponding accumulation bit line.