Shared Charge-Bus Multiplier-Accumulator for Low-Power Scaling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing multiplier-accumulator architectures for machine learning applications face challenges in power consumption due to synchronous clocked operations and increased gate complexity, particularly in scalable hardware designs for multiply-accumulate operations.
Innovation Solution
An asynchronous multiplier-accumulator architecture utilizing AND gates and charge transfer capacitors to minimize internal state changes, with a charge summing unit and analog-to-digital converter to generate digital outputs, allowing for cascaded operations and reduced power consumption.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If synchronous clocked stages are used for multiplication operations, then operational reliability is improved, but power consumption increases
Solution Approach 1:
The patent inverts the conventional synchronous clocked approach by implementing an asynchronous multiplication architecture. Instead of using clock signals to coordinate operations, the system uses hand-shaking signals and ready/valid flags to indicate when operations are complete, thereby eliminating continuous clocking and reducing power consumption while maintaining operational reliability
Solution Approach 2:
The patent employs periodic action through the use of reset signals that periodically clear the accumulation register and through the periodic hand-shaking protocol between multiplier and accumulator stages. This periodic resetting and signal exchange ensures reliable operation without requiring continuous clocking, thus reducing power consumption
2Productivity
If n×n multiplier size is increased for machine learning applications, then computational capability is improved, but gate complexity increases as n2
Solution Approach 1:
The patent segments large n×n multiplication operations into smaller m×m unit element blocks. Each block performs local multiplication and accumulation, and the results are combined to produce the final output. This segmentation reduces the gate complexity of individual units while maintaining the overall computational capability through parallel block operations
Solution Approach 2:
The patent implements a nested structure where multiple m×m unit element blocks are nested within a larger n×n multiplier-accumulator array. The unit elements are arranged in a grid pattern where blocks are interconnected through shared buses and accumulation paths, allowing hierarchical computation that scales efficiently with increasing n
3Adaptability or versatility
If multiple adders are added for multiply-accumulate operations, then computational functionality is improved, but device complexity increases
Solution Approach 1:
The patent merges the multiplication and accumulation functions into a single integrated unit element block. The multiply-accumulate operation is performed by combining the AND gate multiplication stage with an accumulation register and addition logic within the same block, eliminating the need for separate adder circuits and reducing overall device complexity while maintaining full computational functionality
Solution Approach 2:
The unit element block is designed as a universal module that can perform multiple functions: multiplication via AND gates, accumulation via the accumulation register, and output via the ready/valid flag system. This multi-functional design eliminates the need for dedicated separate circuits for each operation, reducing device complexity while maintaining adaptability
4Use of energy by moving object
If internal state changes are minimized for static weighting matrices, then power consumption is reduced, but operational flexibility is constrained
Solution Approach 1:
The patent implements a dynamic architecture where the multiplication and accumulation operations are triggered by hand-shaking signals and ready/valid flags rather than continuous clocking. The system dynamically activates computation only when input data is ready and the accumulator is prepared, minimizing unnecessary state changes and power consumption while maintaining operational flexibility through signal-based control
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
The solution enables low-power, scalable, and efficient multiply-accumulate operations by minimizing internal state changes and power dissipation, suitable for machine learning applications with static weighting matrices, while maintaining flexibility and accuracy in digital output generation.
Implementation Method 1
each AND gate output coupled through a capacitor of value Cu to a particular analog charge line of an analog charge bus according to bit order
Implementation Method 2
the charge summing capacitors having a second terminal which are coupled together and also coupled to the input of an analog to digital converter
Data Source
AI summary
A plurality of unit elements share a charge transfer bus, each unit element accepts A and B digital inputs and generates a product P as an analog charge transferred to the charge transfer bus, each unit element comprised of groups of AND gates coupled to charge transfer lines through a capacitor Cu. Each unit element receives one bit of the B input applied to all of the AND gates of the unit element, and each unit element having the bits of A applied to each associated AND gate input of each unit element. The AND gates of each unit element are coupled to charge transfer lines through a capacitor Cu, and the charge transfer lines couple to binary weighted charge summing capacitors which sum and scale the charges contributed by all unit elements to the charge transfer lines according to a bit weight and converted to a digital value output.


