In-memory computing architecture with full-simulation-domain multi-bit signed number computing
Through the in-memory computing architecture of multi-bit signed number calculation in full simulation domain, the problems of slow computing speed and high power consumption in edge devices are solved, efficient and reliable multi-bit computing is achieved, circuit design is simplified, cost reduction is reduced, and a variety of application scenarios are adapted.
Patent Information
- Application Number
- CN202510414459.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-08
AI Technical Summary
The prior art has problems such as slow computing speed, high power consumption, waste of memory cell area and multiple dynamic buffer power consumption in artificial intelligence neural network calculations in edge devices. Especially when computing multi-bit signed numbers, frequent data transfer and high-latency ADC conversion are required.
The in-memory computing architecture adopts a full-analog domain multi-bit signed number calculation, and through high-bit and low-bit in-memory computing units, a single-row in-memory computing array and data weight configuration circuit, the multi-bit signed number calculation in the analog domain is realized, which simplifies circuit design, reduces data transfer, reduces power consumption, and improves calculation speed.
Simplify circuit design, improve computing speed, reduce power consumption, enhance system integration, adapt to a variety of application scenarios, provide efficient and reliable multi-bit computing solutions, improve computing accuracy and system flexibility, and reduce costs.
Smart Images

Figure CN120278098A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of in-memory computing architectures, and more particularly to an in-memory computing architecture with all-analog-domain multi-bit signed number calculation. Background Art
[0002] In the era of big data, neural networks deployed by artificial intelligence are widely deployed in edge devices and applied to fields such as image processing and perception recognition. The amount of computation generated is also increasing continuously. In the traditional von Neumann architecture, data needs to be frequently transferred between the storage unit and the computing unit, which limits the speed of data processing. At the same time, the process of data transfer also causes a large amount of power consumption.
[0003] In the prior art (charge-domain SRAM memory macro for edge inference in 22nm FinFET process based on C-2C trapezoidal 8-bit MAC unit), a DAC is used to convert the digital input activation (IA) into analog IA, the MAC operation is completed in the analog domain, and finally an ADC conversion is performed. The MAC operation of multiple-bit IA and multiple-bit weights can be completed at one time in the analog domain, with high parallelism; however, its DAC requires precise and high-drive external voltage supply, and the cost is relatively high. In the prior art (4-bit mixed-signal MAC macro with single ADC conversion), nine partial products are mapped to five wires according to their relative weights, the voltages of the five wires are dynamically buffered and sampled on SAR ADC capacitors of appropriate size to achieve one conversion. This structure has no DAC structure design and only requires one AD conversion, reducing the latency and power consumption of AD conversion. However, there is a problem of wasted area of the memory-computation unit, and the power consumption of multiple dynamic buffers cannot be ignored.
[0004] In the prior art (an in-memory computing structure with channel-parallel full-effective storage), the highest bit of the C2C array is processed through switch switching, and weighted and signed number calculations can be performed without special processing of the storage unit. The highest bit of the C2C array is processed through switch switching, and weighted and signed number calculations can be performed without special processing of the storage unit. However, it takes one time to complete 8w1a calculations, and shift addition needs to be performed in the digital domain, and the power consumption and latency of ADC conversion are large. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to overcome the above technical defects and provide an in-memory computing architecture with all-analog-domain multi-bit signed number calculation.
[0006] To solve the above problems, the technical solution of the present invention is: an in-memory computing architecture with all-analog-domain multi-bit signed number calculation, including: High-bit-in-memory computing unit: The high-bit-in-memory computing unit is composed of a computing unit and a storage unit. The computing unit of a high-bit-in-memory computing unit is composed of 3 NMOS transistors, 1 MOM capacitor C0, four complementary CMOS switches S1, S0, S_gnd, and S_vdd, and the storage unit is a 6T-SRAM; Low-bit-in-memory computing unit: The low-bit-in-memory computing unit is composed of a computing unit and a storage unit. The computing unit of a low-bit-in-memory computing unit is composed of 3 NMOS transistors, 1 MOM capacitor C0, and 2 complementary CMOS switches S1 and S0, and the storage unit is a 6T-SRAM; Single-row in-memory computing array: The single-row in-memory computing array contains 16 in-memory computing units, 4 complementary CMOS switches S2, 1 complementary CMOS switch S3, and 4 three-state gates; each single-row in-memory computing array can perform DAC conversion on a 4-bit activation value of an input in the analog domain to obtain the analog voltage value corresponding to the 4-bit activation value, and its voltage value is obtained through charge sharing on the upper plates of 16 capacitors C0 in the single-row in-memory computing unit. These voltages can be used to perform multiplication and accumulation operations with the weights of the storage unit in the charge domain by relying on switch timing control in the in-memory computing unit; the in-memory computing array is set with 144 rows of memory computing arrays, and each row of in-memory computing arrays converts the digital activation value into an analog signal and inputs it into the memory array for calculation; between each single-row in-memory computing array, according to charge redistribution, the shared charge of the capacitors between columns is obtained, and multi-bit weighted calculation is performed. After passing through multiple rows of in-memory computing arrays, the weighted activation value and the MAC result with unweighted weights are obtained, and then passed to the next-level data weight configuration circuit to weight the weights of the MAC value to obtain the MAC result with both the activation value and the weights weighted; Data weight configuration circuit: Utilizing the unique weighting characteristic of the C2C circuit, the MAC result with only the activation value weighted passed down from the previous-level in-memory computing array is weighted with weights to obtain the complete result of multiplying and accumulating multiple 4-bit weights and 4-bit activation values, and finally passed to the ADC to convert the analog signal into a digital signal. The analog-domain weighting structure is composed of 11 complementary CMOS switches, 5 1fF, and 3 2fF capacitors; The steps implemented by the in-memory computing architecture with full analog-domain multi-bit signed number calculation are as follows: S1. Each 4-bit digital activation value is input into the single-row in-memory computing array for DAC conversion to weight the activation value; S2. The DAC conversion result performs MAC calculation in the memory computing unit, and the MAC results of multiple rows of in-memory computing arrays are accumulated by column in the charge domain; S3. The MAC result access data weight configuration circuit weights the weights of the MAC results to obtain multiple multiply-accumulation results of 4W4A; S4. Finally, after passing through the ADC, the analog signal is finally transmitted to the ADC to be converted into a digital signal.
[0007] Furthermore, the weight data read from the 6T-SRAM is subjected to a multiply-accumulation operation with the activation value in the memory-computation unit, and writing the data into the 6T-SRAM realizes the storage of the weights.
[0008] Furthermore, in the high-bit in-memory computing unit, S1, S0, S_gnd, and S_vdd represent the signals for controlling the switches, S1_b, S0_b, S_gnd_b, and S_vdd_b represent the inversions of the signals for controlling the switches, W_sign represents the weight sign bit, W_sign_b represents the inversion of the weight sign bit, W[3:0] represents the weight bits, and A[3:0] represents the activation value; The calculation principle of the high-bit in-memory computing unit is as follows: When S1 is closed, S0 is open, S_gnd is closed, S_vdd is open, and RL is pulled down, the activation value is input, and the upper plate of the capacitor C0 is charged to vdd; When S1 is closed, S0 is open, S_gnd is closed, S_vdd is open, and RL is pulled up, the switch connected to the SRAM decides whether to release the charge in the capacitor C0 according to each bit of the weight W[i] (Q) (retain "1", release "0"); When S1 is open, S0 is closed, RL is pulled down, S_gnd is open, and S_vdd is closed, the lower plate of the capacitor is connected to vdd, and the upper plate of the capacitor is coupled to send the charge to the output line OL.
[0009] Furthermore, in the low-bit in-memory computing unit, S1, S0, S_gnd, and S_vdd represent the signals for controlling the switches, S1_b, S0_b, S_gnd_b, and S_vdd_b represent the inversions of the signals for controlling the switches, W_sign represents the weight sign bit, W_sign_b represents the inversion of the weight sign bit, W[3:0] represents the weight bits, and A[3:0] represents the activation value; The calculation principle of the low-bit in-memory computing unit: When S1 is closed, S0 is open, and RL is pulled down, the activation value is input (the input of the activation value will be explained below). The upper plate of the capacitor C0 is charged to vdd; When S1 is closed, S0 is open, and RL is pulled up, the switch connected to the SRAM decides whether to release the charge in the capacitor C0 according to each bit of the weight W[i] (Q) (retain "1", release "0"); When S1 is open, S0 is closed, and RL is pulled down, the capacitor C0 discharges to send the charge to the output line OL.
[0010] Furthermore, the complementary CMOS switch consists of 1 NMOS transistor and 1 PMOS transistor.
[0011] Furthermore, in the single-row in-memory computing array, the activation value A[3:0] is an unsigned number, and the weight value W[3:0] is a signed number. Among them, the 4-bit weight value is stored in 4 in-memory computing units in descending order of bit positions. Every 3 low-bit in-memory computing units and 1 high-bit in-memory computing unit form a group of memory-computation units. The single-row in-memory computing array has a total of four groups of memory-computation units. The 4-bit in-memory computing unit group can represent weight values in the data range of -7 to +7 (1111 to 0111); Each binary number in the 4-bit activation value A[3:0] controls the enable switch of the tri-state gate. The S2 switch groups the 16 in-memory computing units in a 1:2:4:8:1 ratio. The purpose of grouping is to achieve separate charging of the in-memory computing units for the input voltage vdd.
[0012] Furthermore, the computing principle of the single-row in-memory computing array is as follows: 1) Activation value input: All in-memory computing units in the same row share one input line. The input lines between adjacent groups are controlled and connected through the switch S2. Each group of input lines is connected to the output of the tri-state gate. The input of the tri-state gate is vdd, and the EN terminal is controlled by one activation value. Each row inputs a 4-bit activation value A[3:0], and each row stores 4 different 4-bit weight values W0[3:0] - W3[3:0]; When the activation value A[3:0] is input, the switch S2 is disconnected, EN is enabled by the activation value, and according to the input binary activation value A[3:0], the capacitors C0 inside different numbers of in-memory computing capacitors in the same row are charged. Then, the tri-state gate is closed. When the redundant activation value is input, S3 is closed and connected to gnd; 2) Charge sharing: When S3 is disconnected and S2 is closed, at this time, all the capacitors C0 in the same row share the un-released charge on the upper plate and obtain the same voltage ; In the i-th row, after the activation value input is completed, the lower plate of the C0 capacitor is grounded, and the voltage of the shared upper plate is :
[0013] IN[i] is the decimal representation of the activation value A [i] (for example, the activation value input in the 0th row is A [0] = 1011, and its decimal representation is 11, then IN [0] = 11, N is the input bit width equal to 4, then ).
[0014] The advantages of the present invention compared with the existing technologies are as follows: 1. Simplify the circuit design, improve energy efficiency, enhance the computing speed, increase the system integration level, reduce costs, improve reliability, and adapt to various application scenarios, providing an efficient and reliable solution for multi-bit operations.
[0015] 2. Have multiple advantages such as simplifying the system architecture, improving energy efficiency, enhancing performance, reducing costs, enhancing reliability, optimizing resource utilization, strong adaptability, and easy expansion. This not only improves the overall efficiency of the in-memory computing system but also provides a more flexible and reliable solution for future high-performance computing applications.
[0016] 3. Have multiple advantages such as enhancing the computing accuracy, wide application scenarios, reducing costs, supporting high-parallel computing, and enhancing system flexibility. This not only provides more powerful computing capabilities for in-memory computing technologies but also provides efficient solutions for the high-performance computing requirements in fields such as artificial intelligence, signal processing, and scientific computing.
[0017] 4. Significantly reduce the computing latency. The shift-and-add operation is one of the potential fault points in the system, especially in high-load or complex scheduling scenarios. The circuit structure provided by the present invention avoids this problem and improves the system reliability. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 It is the structure diagram of a 6T-SRAM memory cell.
[0019] Figure 2 It is the structure diagram of a high-bit in-memory computing unit.
[0020] Figure 3 It is the structure diagram of a low-bit in-memory computing unit.
[0021] Figure 4 It is the structure diagram of a complementary CMOS switch.
[0022] Figure 5 It is the structure diagram of a single-row in-memory computing array.
[0023] Figure 6 It is the structure diagram of an analog-domain weighting.
[0024] Figure 7 It is OL [0] and OL [3] The schematic diagram of the charge redistribution on the OL line.
[0025] Figure 8 It is the overall circuit diagram of a multi-bit MAC circuit.
[0026] Figure 9 It is the correspondence diagram between the MAC digital value and the analog voltage value. Detailed implementation manners
[0027] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0028] As Figures 1 to 9 shown, a memory-in-computation architecture with all-analog-domain multi-bit signed number calculation includes: I. 6T-SRAM The weight data read from the 6T-SRAM is subjected to multiply-accumulate operation with the activation value in the memory-computation unit, and the data written into the 6T-SRAM realizes the storage of the weight.
[0029] 6T-SRAM read and write principle 1. Read operation (BL and BLB = 1) (init Q = 1, QB = 0) (1) Pre-charge BL and BLB to 1; WL is pulled high (2) M6 is turned on, BL = Q; M5 is turned on, BLB = QB (3) Initially Q = 1, M1 is conducting; initially QB = 0, M4 is conducting; (4) BL = Q = 1, BLB -> QB = 0 (since Q remains unchanged, so QB will not change, but BLB will change instead) (5) A voltage difference is generated between BL and BLB, and SA amplifies it (the voltage of BL remains unchanged, the voltage of BLB drops, and the readout is 1; otherwise, the readout is 0) 2. Write operation (when writing 1, BL = 1, BLB = 0) (init Q = 1, QB = 0) (1) Pre-charge BL = 0, BLB = 1 (write 0); WL is pulled high (2) M6 is turned on, Q -> BL = 0; M5 is turned on, QB -> BLB = 1 (3) Q and QB change in the opposite direction at the same time, and the latch stores the numerical change (4) Q = 0, QB = 1, and the write is successful 3. Hold operation (1) WL is pulled low, and at this time Q and QB are interlocked and unchanged Since the weight W stored in the 6T-SRAM is used, the weight W used in the following calculations is the Q value stored in the 6T-SRAM, and W_b is Q_b.
[0030] II. High-bit memory-in-computation unit In specific applications, the high-bit in-memory computing unit is composed of a computing unit and a storage unit 6T-SRAM. The integration of storage and computing avoids the frequent transfer of data between the processor and the memory, reducing latency and energy consumption. To support large-scale parallel computing and improve data accuracy, a high-bit in-memory computing unit is set up so that the circuit architecture can complete signed multiply-accumulate calculations.
[0031] The computing unit of 1 high-bit in-memory computing unit is composed of 3 NMOS transistors, 1 MOM capacitor C0, and four complementary CMOS switches S1, S0, S_gnd, and S_vdd. The storage unit is composed of 1 6T-SRAM. The capacitor uses a MOM computing capacitor that is more robust to PVT variations.
[0032] S1, S0, S_gnd, and S_vdd represent the signals controlling the switches. S1_b, S0_b, S_gnd_b, and S_vdd_b represent the inverted signals of the switch control signals. W_sign represents the weight sign bit, and W_sign_b represents the inverted weight sign bit. W[3:0] represents the weight bits, and A[3:0] represents the activation values.
[0033] Computing principle of the high-bit in-memory computing unit: When S1 is closed, S0 is open, S_gnd is closed, S_vdd is open, and RL is pulled down, the activation value is input (the activation value will be explained below). The upper plate of capacitor C0 is charged to vdd.
[0034] When S1 is closed, S0 is open, S_gnd is closed, S_vdd is open, and RL is pulled up, the switch connecting to the SRAM decides whether to release the charge in capacitor C0 according to each bit of the weight W[i] (Q) (retain "1", release "0").
[0035] When S1 is open, S0 is closed, RL is pulled down, S_gnd is open, and S_vdd is closed, the lower plate of the capacitor is connected to vdd, and the upper plate of the capacitor is coupled to send the charge to the output line OL.
[0036] III. Low-bit in-memory computing unit 1 low-bit in-memory computing unit consists of 3 NMOS transistors, 1 MOM capacitor C0, 1 6T-SRAM, and 2 complementary CMOS switches S1 and S0. The capacitor uses a MOM computing capacitor that is more robust to PVT variations.
[0037] S1, S0, S_gnd, S_vdd represent the signals for controlling switches, S1_b, S0_b, S_gnd_b, S_vdd_b represent the inverted signals for controlling switches, W_sign represents the weight sign bit, W_sign_b represents the inverted weight sign bit, W[3:0] represents the weight bits, and A[3:0] represents the activation value.
[0038] Calculation principle of the low-bit in-memory computing unit: When S1 is closed, S0 is open, and RL is pulled down, the activation value is input (the activation value input will be explained below). The upper plate of capacitor C0 is charged to vdd.
[0039] When S1 is closed, S0 is open, and RL is pulled up, the switch connected to the SRAM decides whether to release the charge in capacitor C0 according to each bit weight W[i] (Q) (retain "1", release "0").
[0040] When S1 is open, S0 is closed, and RL is pulled down, capacitor C0 discharges and sends the charge to the output line OL.
[0041] IV. Structure of the complementary CMOS switch The complementary CMOS switch consists of 1 NMOS transistor and 1 PMOS transistor. In the T40 process, the minimum dimensions are W = 40nm and L = 120nm, which can prevent leakage, has good robustness, a large noise margin, an extremely high input impedance, and basically no static power consumption.
[0042] V. Single-row in-memory computing array Each single-row in-memory computing array can perform DAC conversion on a 4-bit activation value of an input in the analog domain to obtain the analog voltage value corresponding to the 4-bit activation value. The voltage value is obtained through charge sharing on the upper plates of the 16 capacitors C0 in the single-row in-memory computing unit. These voltages can perform multiplication and accumulation operations with the weight of the storage unit in the charge domain by relying on switch timing control in the in-memory computing unit, avoiding the overhead of transferring data from the memory to the processor and significantly improving the computing efficiency. The in-memory computing array sets 144 rows of in-memory computing arrays as shown in Figure 6 and its quantity is set according to a specific convolutional network. Each row of the in-memory computing array converts the digital activation value into an analog signal and inputs it into the memory array for calculation, realizing the efficient docking of digital signals and analog computing units.
[0043] Between each single-row in-memory computing array, charge redistribution can be carried out, and then the shared charge of the capacitors between columns can be obtained, and multi-bit weighted calculations can be performed. After passing through multiple rows of in-memory computing arrays, the weighted activation value and the MAC result with unweighted weights are obtained, and then passed to the next-level data weight configuration circuit to weight the MAC value, and the MAC result with both the activation value and the weight weighted can be obtained.
[0044] The single - row in - memory computing array includes 16 in - memory computing units, 4 complementary CMOS switches S2, 1 complementary CMOS switch S3, and 4 tri - state gates.
[0045] The activation value A[3:0] is an unsigned number, and the weight value W[3:0] is a signed number. Among them, the 4 - Bit weight values are stored in 4 in - memory computing units in descending order of bit position. Every 3 low - bit in - memory computing units and 1 high - bit in - memory computing unit form a group of memory - computing units. The single - row in - memory computing array has a total of four groups of memory - computing units. The 4 - Bit memory - computing unit group can represent weight values in the data range of - 7 to +7 (1111 to 0111).
[0046] Each binary number in the 4 - Bit activation value A[3:0] controls the enable switch of the tri - state gate. The S2 switch groups the 16 in - memory computing units in a 1:2:4:8:1 ratio. The purpose of grouping is to achieve the charging of in - memory computing units in groups by the input voltage vdd respectively.
[0047] The computing principle of the single - row in - memory computing array: Activation value input: All in - memory computing units in the same row share one input line. The input lines between adjacent groups are controlled and connected by the switch S2. Each group of input lines is connected to the output of the tri - state gate. The input of the tri - state gate is vdd, and the EN terminal is controlled by one bit of the activation value. Each row inputs a 4 - bit activation value A[3:0]. Each row stores 4 different 4 - bit weight values W0[3:0] - - W3[3:0].
[0048] When the activation value A[3:0] is input, the switch S2 is disconnected, and EN is enabled by the activation value. According to the input binary activation value A[3:0], the capacitors C0 inside different numbers of in - memory computing capacitors in the same row are charged, and then the tri - state gate is closed. When the redundant - bit activation value is input, S3 is closed and connected to gnd.
[0049] Charge sharing When S3 is disconnected and S2 is closed. At this time, all the capacitors C0 in the same row share the unreleased charge on the upper plate and obtain the same voltage .
[0050] In the i - th row, after the activation value input is completed, the lower plate of the C0 capacitor is grounded, and the voltage of the shared upper plate is :
[0051] IN[i] is the decimal representation of the activation value A [i] (For example, the activation value input in the 0 - th row is A [0]= 1011, whose decimal representation is 11, then IN [0] = 11, N is the input bit width equal to 4, then ).
[0052] VI. Data Weight Configuration Circuit The present invention proposes an analog-domain weighting structure, which can complete the MAC multiplication and accumulation operation of 4-bit activation values and 4-bit weight values through only one analog-domain weighting. The main advantages are low power consumption, high speed, and small delay.
[0053] The present invention is an improvement on the C2C array. By utilizing the unique weighting characteristic of the C2C circuit, the MAC result that only weights the activation value passed down from the upper-level in-memory computing array is weighted by the weight value, and then the complete result of multiplying and accumulating multiple 4-bit weight values and 4-bit activation values is obtained, and finally passed to the ADC to convert the analog signal into a digital signal.
[0054] The analog-domain weighting structure consists of 11 complementary CMOS switches, 5 1fF, and 3 2fF capacitors.
[0055] Low-order bits: Taking OL [0] as an example: When the switch C2C_Lowbits_sample is closed and C2C_Lowbits_couple is open, the sampling capacitor C is connected to the output line OL [0] ; the switch C2C_bottom is closed, so that the lower plate of the sampling capacitor C is connected to gnd, and the upper plate of the sampling capacitor C in the data weight configuration circuit is charged.
[0056] After that, the switch C2C_bottom is opened and C2C_lowbits_couple is closed, so that the upper plate of the capacitor C is connected to V DD , and the lower plate couples out the voltage V couple_LSB[0] .
[0057] Since V sample_LSB[0] = V IN (V IN is the voltage of the upper plate of C0 when shared, V sample_LSB[0] is the voltage of the upper plate of the capacitor C connected to the OL [0] line in the data weight configuration circuit) According to the law of conservation of charge: (V sample_LSB[0] - 0)C = (0 – V couple_LSB[0] )C, we get: V couple_LSB[0] = -V sample_LSB[0] After passing through the C2C array, the voltage value generated at the V 4w4a terminal is -1 / 16 VIN 。
[0058] And so on: The voltage generated by the low bit on the V4w4a terminal is: -1 / 4 V IN -1 / 8 V IN -1 / 16 V IN High bit: Taking OL [3] as an example: When the switch C2C_MSB_sample is closed and C2C_MSB_couple is open, the sampling capacitor C is connected to the output line OL [3] , and the upper plate of the sampling capacitor C in the data weight configuration circuit is charged.
[0059] After that, the switch C2C_MSB_sample is opened. Since the lower plate of the sampling capacitor C is connected to gnd and the upper plate couples out the voltage V couple_MSB[3] Since V sample_MSB[3] = V DD + V IN (V IN is the voltage of the upper plate of the capacitor C0 in the shared time-memory-computation unit, and V sample_MSB[3] is the voltage of the upper plate of the capacitor C connected to the OL [3] line in the data weight configuration circuit).
[0060] According to the law of conservation of charge: (V sample_MSB[3] - V DD )C = (V couple_MSB[3] - V DD )C, we get: V couple_MSB[3] = V DD + V IN After passing through the C2C array, a voltage of 1 / 2 V3 = 1 / 2(V 4w4a + V DD + V IN ) is generated at the V Total voltage generated at the V4w4a port: V 4w4a = 1 / 2(V DD + V IN ) - 1 / 4 V IN - 1 / 8 V IN - 1 / 16 V IN After considering charge redistribution, as shown in Figure 7 Figure.
[0061] Analog domain addition circuit sampling: High-bit sampling principle: According to the law of conservation of charge:
[0062]
[0063]
[0064] Low-bit sampling principle: According to the law of conservation of charge:
[0065] Solve to get:
[0066] Solve to get:
[0067]
[0068] Analog domain addition circuit output expression: Taking the high bit as OL [3] as an example: According to the law of conservation of charge: Get: Solve to get:
[0069] According to the weighting characteristic of the C2C circuit, the high-bit voltage output result at the 4W4A end is:
[0070] Taking the low bit as OL [2-0] as an example: According to the law of conservation of charge: Get: Solve to get:
[0071] According to the weighting characteristic of the C2C circuit, the high-bit voltage output result at the 4W4A end is:
[0072] Analog domain addition circuit 4W4A port output voltage expression:
[0073]
[0074]
[0075]
[0076] Figure 8 It is the overall circuit diagram of the multi-bit MAC circuit.
[0077] The charges in the same column of the IMCC are subjected to charge redistribution, and the charges are evenly distributed to the capacitors of each column through the output line OL, and the voltage is transferred to the upper plate of the sampling capacitor C in the data weight configuration circuit. After the voltage weighting of the data weight configuration circuit, the calculation result of 4W4A is obtained. Finally, the in-memory calculation of 4-bit signed numbers and 4-bit unsigned numbers in the all-analog domain is realized, and the final result is transferred to the ADC.
[0078] Figure 9 It is the correspondence between the MAC digital value and the analog voltage value.
[0079] Among them, LSB (Least Significant Bit) represents that in digital signal processing, when the digital value changes by a minimum unit, the corresponding change in the analog signal voltage. It can be seen from the figure that when the MAC value is positive, as the MAC value increases linearly, the voltage decreases linearly; when the MAC value is negative, as the MAC value decreases linearly, the voltage increases linearly, showing an overall linear correspondence.
[0080] The steps for implementing the in-memory computing architecture with all-analog-domain multi-bit signed number calculation are as follows: S1. Each 4-bit digital activation value is input into the single-row in-memory computing array for DAC conversion to weight the activation value; S2. The DAC conversion result is subjected to MAC calculation in the memory computing unit, and the MAC results of the multi-row in-memory computing array are accumulated column by column in the charge domain; S3. The MAC result is connected to the data weight configuration circuit to weight the weight value of the MAC result to obtain multiple multiply-accumulate results of 4W4A; S4. Finally, through the ADC, and finally transferred to the ADC to convert the analog signal into a digital signal; The parts not disclosed in the present invention are all prior arts, and their specific structures and working principles will not be elaborated.
[0081] It should be noted that, in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device.
[0082] Although embodiments of the present invention have been shown and described, those of ordinary skill in the art will understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
[0083] The above describes the present invention and its implementation manners. Such description is not restrictive. What is shown in the drawings is only one of the implementation manners of the present invention, and the actual structure is not limited thereto. In general, if those of ordinary skill in the art are inspired by it and, without departing from the gist of the present invention, design similar structural manners and embodiments to this technical solution without creative efforts, they shall fall within the protection scope of the present invention.
Claims
1. An in-memory computing architecture with all-analog-domain multi-bit signed number calculation, characterized in that , including: High-bit-in-memory computing unit: The high-bit-in-memory computing unit is composed of a computing unit and a storage unit. The computing unit of a high-bit-in-memory computing unit is composed of 3 NMOS transistors, 1 MOM capacitor C0, four complementary CMOS switches S1, S0, S_gnd, and S_vdd. The storage unit is a 6T-SRAM; Low-bit-in-memory computing unit: The low-bit-in-memory computing unit is composed of a computing unit and a storage unit. The computing unit of a low-bit-in-memory computing unit is composed of 3 NMOS transistors, 1 MOM capacitor C0, and 2 complementary CMOS switches S1 and S0. The storage unit is a 6T-SRAM; Single-row in-memory computing array: The single-row in-memory computing array contains 16 in-memory computing units, 4 complementary CMOS switches S2, 1 complementary CMOS switch S3, and 4 tri-state gates; Each single-row in-memory computing array can perform DAC conversion on a 4-bit activation value of an input in the analog domain to obtain the analog voltage value corresponding to the 4-bit activation value. The voltage value is obtained through charge sharing on the upper plates of the 16 capacitors C0 in the single-row in-memory computing unit. These voltages can be used to perform multiply-accumulation operations with the weights of the storage unit in the charge domain by relying on switch timing control in the in-memory computing unit; The in-memory computing array is set with 144 rows of memory-computing arrays. Each row of the in-memory computing array converts the digital activation value into an analog signal and inputs it into the memory array for calculation; Between each single-row in-memory computing array, according to charge redistribution, the shared charge of the capacitors between columns is obtained, and multi-bit weighted calculations are performed. After multiple rows of in-memory computing arrays, the weighted activation value and the MAC result with unweighted weights are obtained, and then passed to the next-level data weight configuration circuit to weight the weights of the MAC value to obtain the MAC result with both the activation value and the weights weighted; Data weight configuration circuit: Utilizing the unique weighting characteristic of the C2C circuit, it weights the weights of the MAC result with only the activation value weighted passed down from the previous-level in-memory computing array to obtain the complete result of multiplying and accumulating multiple 4-bit weights and 4-bit activation values, and finally passes it to the ADC to convert the analog signal into a digital signal. The analog-domain weighting structure is composed of 11 complementary CMOS switches, 5 1fF, and 3 2fF capacitors; The steps for implementing the in-memory computing architecture with full analog-domain multi-bit signed number calculation are as follows: S1. Each 4-bit digital activation value is input into the single-row in-memory computing array for DAC conversion to weight the activation value; S2. The DAC conversion result performs MAC calculation in the memory-computing unit, and the MAC results of multiple rows of in-memory computing arrays are accumulated by column in the charge domain; S3. The MAC result is connected to the data weight configuration circuit to weight the weights of the MAC result to obtain the multiply-accumulation result of multiple 4W4A; S4. Finally, through the ADC, it is finally passed to the ADC to convert the analog signal into a digital signal.
2. The in-memory computing architecture with full analog-domain multi-bit signed number calculation according to claim 1, wherein: The weight data read from the 6T-SRAM is multiplied and accumulated with the activation value in the memory and computing unit, and the data written into the 6T-SRAM realizes the storage of the weight.
3. The in-memory computing architecture with full analog domain multi-bit signed number calculation according to claim 1, characterized in that: In the high-bit in-memory computing unit, S1, S0, S_gnd, and S_vdd represent the signals for controlling the switches, S1_b, S0_b, S_gnd_b, and S_vdd_b represent the inverted signals of the switch control signals, W_sign represents the weight sign bit, W_sign_b represents the inverted weight sign bit, W[3:0] represents the weight bits, and A[3:0] represents the activation value; The calculation principle of the high-bit in-memory computing unit is as follows: When S1 is closed, S0 is open, S_gnd is closed, S_vdd is open, and RL is pulled down, the activation value is input, and the upper plate of capacitor C0 is charged to vdd. When S1 is closed, S0 is open, S_gnd is closed, S_vdd is open, and RL is pulled up, the switch connected to the SRAM decides whether to release the charge in capacitor C0 according to each bit of the weight W[i] (Q) (retain "1", release "0"); When S1 is open, S0 is closed, RL is pulled down, S_gnd is open, and S_vdd is closed, the lower plate of the capacitor is connected to vdd, and the upper plate of the capacitor is coupled to send the charge to the output line OL.
4. A in-memory computing architecture with full analog-domain multi-bit signed number calculation according to claim 1, characterized in that: In the low-bit in-memory computing unit, S1, S0, S_gnd, and S_vdd represent the signals for controlling the switches, S1_b, S0_b, S_gnd_b, and S_vdd_b represent the inverted signals of the switch control signals, W_sign represents the weight sign bit, W_sign_b represents the inverted weight sign bit, W[3:0] represents the weight bits, and A[3:0] represents the activation value; The calculation principle of the low-bit in-memory computing unit: When S1 is closed, S0 is open, and RL is pulled down, the activation value is input (the input of the activation value will be explained below). The upper plate of capacitor C0 is charged to vdd; When S1 is closed, S0 is open, and RL is pulled up, the switch connected to the SRAM decides whether to release the charge in capacitor C0 according to each bit of the weight W[i] (Q) (retain "1", release "0"); When S1 is open, S0 is closed, and RL is pulled down, capacitor C0 discharges to send the charge to the output line OL.
5. The in-memory computing architecture with full analog-domain multi-bit signed number calculation according to claim 1, characterized in that: The complementary CMOS switch consists of 1 NMOS transistor and 1 PMOS transistor.
6. The in-memory computing architecture with full analog domain multi-bit signed number calculation according to claim 1, characterized in that: In the single-row in-memory computing array, the activation value A[3:0] is an unsigned number, and the weight W[3:0] is a signed number. Among them, the 4-bit weight is stored in 4 in-memory computing units in order from high to low. Every 3 low-bit in-memory computing units and 1 high-bit in-memory computing unit form a group of memory and computing units. There are four groups of memory and computing units in the single-row in-memory computing array. The 4-bit in-memory computing unit group can represent weights in the data range of -7 to +7 (1111 to 0111); Each binary number in the 4-bit activation value A[3:0] controls the enable switch of the tri-state gate. The S2 switch groups the 16 in-memory computing units in a ratio of 1:2:4:8:
1. The purpose of grouping is to realize the charging of the in-memory computing units in groups by the input voltage vdd respectively.
7. The in-memory computing architecture with full analog-domain multi-bit signed number calculation according to claim 6, wherein The calculation principle of the single - row in - memory computing array is as follows: 1) Activation value input: All in - memory computing units in the same row share one input line. The input lines between adjacent groups are controlled and connected through switches S2. Each group of input lines is connected to the output of a tri - state gate. The input of the tri - state gate is vdd, and the EN terminal is controlled by a single activation value. Each row inputs a 4 - bit activation value A[3:0], and each row stores 4 different 4 - bit weight values W0[3:0] - - W3[3:0]. When the activation value A[3:0] is input, switch S2 is disconnected, and EN is enabled by the activation value. According to the input binary activation value A[3:0], the capacitors C0 inside different numbers of in - memory computing capacitors in the same row are charged. Then, the tri - state gate is closed. When the redundant - bit activation value is input, S3 is closed and connected to gnd; 2) Charge sharing: When S3 is disconnected and S2 is closed, at this time, all the capacitors C0 in the same row share the un-released charge on the upper plate and obtain the same voltage ; In the i-th row, the activation value input is completed, the lower plate of the C0 capacitor is grounded, and the voltage of the upper plate after sharing is : IN[i] is the activation value A [i] in decimal representation. (For example, the activation value input in the 0th row is A [0] = 1011, and its decimal representation is 11, then IN [0] = 11. If N, the input bit width, is equal to 4, then )