Mixed-Signal MAC Array With Shared Capacitors and Differential Readout
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional MAC arrays in mixed-signal in-memory computing suffer from high energy consumption, large area requirements, and high computation errors due to unnecessary transistor driving, inefficient capacitor usage, and vulnerability to parasitic capacitance and leakage, limiting their application to large-scale neural networks.
Innovation Solution
A bit-width reconfigurable mixed-signal in-memory computing module with a sub-cell design featuring shared CMOS inverters and transistors, reducing the number of control elements and capacitors, and implementing a differential architecture to minimize computation errors and enhance energy efficiency by avoiding unconditional transistor driving and optimizing capacitor usage.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Area of stationary object
If conventional 1-bit MAC operation is implemented with separate capacitors for each computing cell, then multiplication results can be stored, but the area efficiency is reduced due to MOM capacitors being located outside the SRAM array
Solution Approach 1:
The patent merges the computing capacitor with the SRAM cell structure by using the storage node capacitance of the SRAM cell itself as the computing capacitor. This integration eliminates the need for separate MOM capacitors outside the array, improving area efficiency while maintaining computation accuracy through the inherent stability of SRAM storage nodes.
Solution Approach 2:
The SRAM cell is designed to serve dual purposes: storing weight data and performing multiplication computation. The storage nodes of the SRAM cell function both as memory elements for weight storage and as computing capacitors for holding multiplication results, enabling the same hardware structure to perform both memory and computation functions.
2Loss of energy
If transmission gates are unconditionally driven for every accumulation operation, then charge can be shared among capacitors, but power consumption increases and sparsity of input data cannot be utilized
Solution Approach 1:
The patent implements dynamic control of transmission gates based on the actual values of input data and stored weights. Transmission gates are activated only when necessary (when input is 1 and weight is 1), allowing the circuit to adapt its operation to the data sparsity pattern. This dynamic activation reduces power consumption significantly while maintaining computing speed through selective operation.
Solution Approach 2:
The control signals for transmission gates are changed based on the logical values of input data and stored weights. The gates switch between active and inactive states dynamically, changing their electrical parameters (conductance) according to the computation requirements, thereby reducing unnecessary power dissipation while preserving computation accuracy.
3Reliability
If accumulation operation connects top plates of capacitors, then charge sharing can be performed, but computation errors increase due to non-ideal effects like charge injection, clock feedthrough, and leakage
Solution Approach 1:
The patent extracts the accumulation function from the traditional capacitor top-plate connection approach and implements it through a differential read circuit that senses voltage differences without direct charge sharing. This eliminates the harmful non-ideal effects associated with charge injection and clock feedthrough while maintaining the accumulation functionality through differential sensing.
Solution Approach 2:
The patent introduces a differential read circuit as an intermediary between the storage nodes and the output. Instead of directly connecting capacitor top plates for charge sharing, the differential read circuit mediates the read operation by sensing voltage differences through high-impedance paths, thereby isolating the storage nodes from harmful effects while still enabling accurate accumulation.
4Area of stationary object
If mismatch between computing capacitors and DAC capacitors occurs due to physical layout, then area can be reduced, but computation errors increase
Solution Approach 1:
The patent uses the inherent symmetry and matching of differential SRAM cell structures to ensure that the effective computing capacitors are well-matched. The differential read circuit copies the same sensing operation for both differential sides, and any layout-induced mismatches are common-mode rejected, thereby maintaining high computation accuracy without requiring additional matching capacitors.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
The solution achieves reduced power consumption, increased storage density, and improved computation accuracy by minimizing transistor count and shared capacitor usage, enabling efficient operation in larger neural networks with reduced area and error tolerance.
Implementation Method 1
the memory element comprising two cross-coupled CMOS inverters and a complementary transmission gate... one of the CMOS inverters comprising an output connected to an input of the complementary transmission gate
Implementation Method 2
a complementary transmission gate comprising an NMOS transistor comprising a gate connected to an input signal, the complementary transmission gate further comprising a PMOS transistor comprising a gate connected to a complementary input signal... an output of the complementary transmission gate connected to both a bottom plate of the computing capacitor and the control element
Data Source
AI summary
A mixed-signal in-memory computing sub-cell requires only 9 transistors for 1-bit multiplication. In one aspect, there is a computing cell is constructed from a plurality of such sub-cells that share a common computing capacitor and common transistors. As a result, the average number of transistors in each sub-cell is close to 6. Also proposed is a MAC array for performing MAC operations, which includes a plurality of the computing cells each activating the sub-cells therein in a time-multiplexed manner. Also proposed is a differential version of the MAC array with improved computation error tolerance. Also proposed is an in-memory mixed-signal computing module for digitalizing parallel analog outputs of the MAC array and for performing other tasks in the digital domain. An ADC block in the computing module makes full use of capacitors in the MAC array, thus allowing the computing module to have a reduced area and suffer from less computation errors.


