Mixed-Signal In-Memory MAC Array With Shared Sub-Cells
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional MAC arrays in mixed-signal in-memory computing suffer from high transistor count in each computing cell, inefficient energy usage due to unconditional transistor driving, and high computation errors, limiting their application scope and energy efficiency.
Innovation Solution
A bit-width reconfigurable mixed-signal in-memory computing module with a sub-cell design that uses a 6T SRAM cell, a complementary transmission gate, and shared NMOS and computing capacitors, reducing the number of transistors and minimizing computation errors by avoiding unconditional driving of the transmission gate and sharing transistors and capacitors among sub-cells.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Area of moving object
If conventional 10-transistor computing cells are used for 1-bit MAC operations, then multiplication and accumulation functions are achieved, but area efficiency is reduced and power consumption increases
Solution Approach 1:
The patent merges multiple computing functions into a single integrated structure where 6T SRAM cells perform both storage and computation. The computing cell combines the SRAM cell with a transmission gate and capacitor to achieve 1-bit MAC operations, eliminating the need for separate 10-transistor computing circuits and reducing overall transistor count while maintaining functionality.
Solution Approach 2:
The 6T SRAM cell is designed to serve multiple functions: it stores weight bits, performs multiplication through conditional charge transfer controlled by the transmission gate, and accumulates results in the capacitor. This multi-functional design eliminates dedicated computing transistors and reduces the transistor count from 10 to 6 per computing unit.
2Use of energy by moving object
If transmission gates are unconditionally driven for every accumulation operation, then complete charge transfer is ensured, but energy efficiency deteriorates due to unnecessary switching
Solution Approach 1:
The transmission gate is designed with dynamic control where the control signal is generated based on the actual data values being processed. The gate is activated only when multiplication is needed (when both weight and input are 1), and remains inactive otherwise, adapting its operation to the computational requirements rather than operating unconditionally.
Solution Approach 2:
The computing operation is divided into distinct phases: multiplication phase where transmission gates are activated based on data values, and accumulation phase where charge is transferred to the capacitor. This periodic activation pattern ensures energy efficiency by avoiding unnecessary gate switching while maintaining computation accuracy through proper phase sequencing.
3Area of moving object
If individual capacitors are allocated to each 1-bit multiplication cell, then charge storage for accumulation is ensured, but area efficiency is reduced
Solution Approach 1:
Multiple computing units share a common capacitor structure for charge accumulation. Instead of allocating individual capacitors to each 1-bit multiplication cell, the patent implements a shared capacitor array where capacitors serve multiple computing units, significantly reducing the total capacitor area while maintaining the ability to store and accumulate charge from multiple multiplication operations.
Solution Approach 2:
The capacitor array is designed to serve multiple functions: storing charge from multiplication operations, accumulating results from multiple computing units, and providing the basis for subsequent analog-to-digital conversion. This universal capacitor structure eliminates the need for dedicated capacitors per computing cell while maintaining full functionality.
4Reliability
If top plates of capacitors are connected for accumulation operations, then charge sharing is achieved, but computation errors increase due to non-ideal effects
Solution Approach 1:
Instead of connecting the top plates of capacitors for accumulation as in conventional designs, the patent inverts the approach by connecting the bottom plates of the capacitors. This inversion avoids the non-ideal effects that plague top-plate connections, such as charge injection and clock feedthrough, while still achieving the desired charge sharing and accumulation functionality through the bottom plate connection network.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This design achieves reduced area usage, improved energy efficiency, and enhanced computation accuracy by minimizing transistor count and errors, enabling faster computation speeds and higher throughput while supporting a broader range of neural network applications.
Implementation Method 1
a computing capacitor, wherein a multiplication result of the input signal and the filter parameter is stored as a voltage on the bottom plate of the computing capacitor
Implementation Method 2
a complementary transmission gate, wherein during computation, a gate of an NMOS transistor in the complementary transmission gate is connected to an input signal and a gate of a PMOS transistor in the complementary transmission gate is connected to a complementary input signal
Implementation Method 3
the conventional 6T SRAM cell consisting of MOS transistors M1, M2, M3, M4, M5, M6, in which a complementary metal-oxide-semiconductor (CMOS) inverter consisting of the MOS transistors M1, M2 is cross-coupled to a CMOS inverter consisting of the MOS transistors M3, M4
Data Source
AI summary
A mixed-signal in-memory computing sub-cell only requires 9 transistors for 1-bit multiplication. A computing cell is constructed from a plurality of such sub-cells that share a common computing capacitor and a common transistor. A MAC array for performing MAC operations, includes a plurality of the computing cells each activating the sub-cells therein in a time-multiplexed manner. A differential version of the MAC array provides improved computation error tolerance and an in-memory mixed-signal computing module for digitalizing parallel analog outputs of the MAC array and for performing other tasks in the digital domain. An ADC block in the computing module makes full use of capacitors in the MAC array, allowing the computing module to have a reduced area and suffer from fewer computational errors.


