Chopper-Stabilized MAC Circuit With Binary Charge Transfer ADC
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing multiplier-accumulator architectures for machine learning applications face challenges in power consumption due to synchronous operation and increased gate complexity, particularly in forming dot products for large matrices, which results in high power dissipation and inefficiency.
Innovation Solution
A scalable asynchronous multiplier-accumulator architecture with a common charge transfer bus for MAC, Bias, and ADC unit elements, utilizing NAND-groups and binary weighted charge transfer capacitors to minimize displacement currents and power consumption, and incorporating a Successive Approximation Register (SAR) controller for efficient charge conversion.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Power
If synchronous clocked stages are used for multiplier operation, then operation timing is controlled, but power dissipation increases
Solution Approach 1:
The patent employs periodic chopping action at frequency fc to modulate the multiplier output, converting DC offset and low-frequency 1/f noise into AC components at the chopping frequency and its harmonics. This periodic modulation enables subsequent filtering to remove noise components while preserving the signal, thereby reducing power dissipation without sacrificing operational control.
Solution Approach 2:
The patent replaces the traditional synchronous clocked control mechanism with an asynchronous chopping-based control system. Instead of using clocked stages that generate displacement currents, the invention uses periodic chopping signals to control the multiplier operation, substituting the mechanical clocking system with a signal-processing approach that reduces power consumption while maintaining operational precision.
2Productivity
If large n×n multiplier is implemented for machine learning, then computational capability increases, but gate complexity increases as n2
Solution Approach 1:
The patent segments the large n×n multiplier into multiple smaller multiplier units, each handling a portion of the computational task. By dividing the overall computation into manageable segments that can be processed in parallel or sequentially through chopping cycles, the gate complexity of individual units is reduced while maintaining the total computational capability through coordinated operation of multiple units.
Solution Approach 2:
The patent uses periodic chopping action to enable time-multiplexed operation of multiplier units. Instead of requiring all n² gates to operate simultaneously, the chopping mechanism allows different segments of the computation to be processed in alternating cycles, reducing the peak gate complexity requirement while maintaining overall computational throughput.
3Measurement precision
If more adders are added for multiply-accumulate operations, then computational accuracy improves, but power consumption increases
Solution Approach 1:
The patent introduces chopping as an intermediary process between multiplication and accumulation operations. The chopping mechanism modulates the multiplier output, allowing the accumulation to occur in the frequency domain rather than requiring multiple simultaneous adders. This intermediary transformation enables high-precision multiply-accumulate operations with reduced power consumption by sequential processing through the chopping cycle.
Data Source
AI summary
An architecture for a chopper stabilized multiplier-accumulator (MAC) uses a chop clock and common Unit Element (UE), the MAC formed as a plurality of MAC UEs receiving X and W values and a sign bit exclusive ORed with the chop clock, a plurality of Bias UEs receiving E value and a sign bit exclusive ORed with the chop clock, and a plurality of Analog to Digital Conversion (ADC) UEs which collectively perform a scalable MAC operation and generate a binary result. Each MAC UE, BIAS UE and ADC UE comprises groups of NAND gates with complementary outputs arranged in NAND-groups, each NAND gate coupled to a differential charge transfer bus through a binary weighted charge transfer capacitor. The analog charge transfer bus is coupled to groups of ADC UEs with an ADC controller which enables and disables the ADC UEs using successive approximation to determine the accumulated multiplication result.


