Reconfigurable MAC Pipeline Architecture for Higher Throughput
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing multiplier-accumulator circuitry and integrated circuits face challenges in efficiently performing multiply and accumulate operations due to limitations in pipelining and concatenation architectures, which affect system performance and adaptability to temporal-based requirements.
Innovation Solution
The implementation of a multiplier-accumulator circuitry with a plurality of separate circuits and shadow registers, organized into rows and connected via a switch interconnect network, allowing for flexible pipelining and concatenation of operations, enabling adjustment of the number of interconnected circuits during operation to meet system requirements.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If multiple multiplier-accumulator circuits are interconnected in concatenation architecture, then the speed of multiply and accumulate operations is improved, but the device complexity increases
Solution Approach 1:
The system is divided into multiple separate multiplier-accumulator circuits (first MAC circuit, second MAC circuit, etc.) that can be independently configured and interconnected. Each MAC circuit contains its own multipliers and accumulators, allowing the system to segment complex operations into smaller units that can be pipelined and concatenated to achieve higher speeds while managing complexity through modular design
Solution Approach 2:
The interconnection between MAC circuits is made dynamic and reconfigurable through switch interconnect networks and control logic. The system can adjust the number and arrangement of interconnected circuits during operation to optimize performance for different computational requirements, allowing the complexity to be adapted rather than fixed
2Productivity
If pipelining is implemented with multiple registers and shadow registers, then the productivity of operations is improved, but the device complexity increases
Solution Approach 1:
Shadow registers are used to pre-load and stage input data before it reaches the main computation pipelines. This preliminary action allows data to be prepared and organized in advance, enabling continuous operation of the MAC circuits without waiting for data readiness, thereby improving throughput while organizing register complexity in a structured manner
Solution Approach 2:
The pipelining architecture with multiple registers ensures that while one MAC circuit is performing computation, other circuits are simultaneously loading data, performing intermediate calculations, or preparing outputs. This continuous overlapping of operations across multiple stages maintains high productivity by eliminating idle time in the computational pipeline
3Speed
If the number of interconnected multiplier-accumulator circuits is increased, then the speed of operations is improved, but the adaptability to changing system requirements decreases
Solution Approach 1:
The system employs reconfigurable switch interconnect networks that allow the number and arrangement of MAC circuits to be dynamically adjusted during operation. Control logic enables the system to reconfigure the pipeline depth and circuit interconnections based on changing computational requirements, maintaining both high speed performance and adaptability to different system demands
Solution Approach 2:
Each MAC circuit is designed with universal functionality to handle various types of multiply and accumulate operations. The circuits can be configured to perform different computational tasks through programming and control signals, allowing the same hardware structure to adapt to different algorithms and processing requirements while maintaining high-speed operation
Data Source
AI summary
An integrated circuit comprising a plurality of multiply-accumulator circuitry, configurable in a concatenation architecture, to perform a plurality of multiply and accumulate operations, wherein the plurality of multiply-accumulator circuitry is organized into a plurality of groups, including a first group of multiply-accumulator circuitry and a second group of multiply-accumulator circuitry, wherein each group includes: a plurality of MAC circuits, each including a multiplier to multiply data by a multiplier weight data and generate a product data, and an accumulator to add input data and the product data to generate sum data, and wherein the plurality of MAC circuits of each group is organized in at least one row and connected in series to perform a plurality of concatenated multiply and accumulate operations. The integrated circuit also includes configurable interface circuitry to connect and/or disconnect the plurality of MAC circuits of the first and second groups of multiply-accumulator circuitry.


