Concatenated MAC Circuit Architecture for Flexible Pipelining
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing multiplier-accumulator circuitry and integrated circuits face challenges in efficiently performing multiply and accumulate operations due to limitations in pipelining and concatenation architectures, which affect system performance and adaptability to temporal-based requirements.
Innovation Solution
The implementation of a multiplier-accumulator circuitry with a plurality of separate circuits and shadow registers, organized into rows and connected via a switch interconnect network, allowing for flexible pipelining and concatenation of operations, enabling adjustment of the number of interconnected circuits during operation to meet system requirements.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If multiple multiplier-accumulator circuits are interconnected in concatenation architecture, then the speed of multiply and accumulate operations is improved, but the device complexity increases
Solution Approach 1:
The system is divided into multiple independent multiplier-accumulator circuits (MAC0, MAC1, MAC2, etc.) that can be individually configured and interconnected. Each MAC circuit contains its own registers and control logic, allowing the system to be segmented into functional units that can be independently optimized and reconfigured based on performance requirements.
Solution Approach 2:
The concatenation architecture allows dynamic reconfiguration of the number and arrangement of MAC circuits during operation. The system can adjust the extent of concatenation (i.e., how many MAC circuits are interconnected) in situ to meet changing temporal-based requirements, making the device complexity adaptive rather than fixed.
2Productivity
If the extent of concatenation is increased to meet temporal-based requirements, then the productivity is improved, but the device complexity increases
Solution Approach 1:
The switch interconnect network provides universal connectivity that can accommodate different numbers and configurations of MAC circuits. The same basic architecture can be used for various productivity levels by simply changing the interconnection configuration, rather than requiring different hardware designs for different performance levels.
Solution Approach 2:
The system allows changing the parameter of concatenation extent (number of interconnected MAC circuits) to adjust productivity. By modifying the interconnection configuration through the switch network, the system can scale its computational throughput to match temporal-based requirements without fundamental architectural changes.
3Speed
If separate multiplier-accumulator circuits with shadow registers are used for pipelining, then the speed of operations is improved, but the device complexity increases
Solution Approach 1:
Shadow registers are used to pre-load and hold input data before it is processed by the MAC circuits. This preliminary action allows the system to pipeline operations efficiently, as data is already prepared and staged in the shadow registers while previous operations are completing, thereby increasing operational speed without proportionally increasing complexity.
4Adaptability or versatility
If the number of interconnected multiplier-accumulator circuits is adjusted in situ, then the adaptability to temporal-based requirements is improved, but the device complexity increases
Solution Approach 1:
The switch interconnect network enables dynamic reconfiguration of the MAC circuit array during operation. Control logic can adjust which MAC circuits are active and how they are interconnected based on real-time temporal-based requirements, providing adaptability while using a standardized hardware template that limits the growth of device complexity.
Solution Approach 2:
The system changes the operational parameters (number of active MAC circuits, interconnection topology) in situ to adapt to different temporal-based requirements. This parameter-based adaptability allows the same physical hardware to serve multiple performance levels without requiring additional complex control mechanisms for each configuration.
Data Source
AI summary
An integrated circuit comprising a plurality of multiply-accumulator circuitry interconnected in a concatenation architecture. Each multiply-accumulator circuitry includes first and second MAC circuits and a load-store register. The first MAC circuit includes a multiplier to multiply first data by a first multiplier weight data and generate a first product data, and an accumulator to add first input data and the first product data to generate first sum data. The second MAC circuit includes a multiplier to multiply second data by a second multiplier weight data and generate a second product data, and an accumulator, coupled to the multiplier of the second MAC circuit and the accumulator of the first MAC circuit, to add the first sum data and the second product data to generate second sum data. The load-store register is coupled to the accumulator of the second MAC circuit to temporarily store the second sum data.


