Reconfigurable MAC Pipeline and Memory Switching in Logic Tiles
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current multiplier-accumulator circuitry and integrated circuits face challenges in efficiently performing multiply and accumulate operations due to limitations in pipelining and concatenation architectures, which affect system performance and adaptability to temporal-based requirements.
Innovation Solution
The implementation of a plurality of separate multiplier-accumulator circuits with shadow registers and a switch interconnect network that facilitates pipelining and concatenation of operations, allowing for flexible configuration and adjustment of the number of interconnected circuits during operation to meet system requirements.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If multiplier-accumulator circuits are interconnected in concatenation architecture to perform operations faster, then productivity is improved, but device complexity increases
Solution Approach 1:
The system divides the multiplier-accumulator functionality into multiple separate circuits (e.g., 2-NMAX clusters) that can be independently configured and interconnected. Each cluster contains multiple MAC circuits that can operate in parallel or sequence, allowing the system to achieve high productivity through concatenation while managing complexity through modular segmentation.
Solution Approach 2:
The interconnection architecture is designed to be dynamically reconfigurable, allowing the number and arrangement of interconnected MAC circuits to be adjusted during operation. This dynamic capability enables the system to optimize productivity for different computational tasks while adapting the complexity level to match actual performance requirements.
2Productivity
If the number of interconnected multiplier-accumulator circuits is increased to meet temporal-based system requirements, then productivity is improved, but device complexity increases
Solution Approach 1:
The system employs dynamic reconfiguration capabilities that allow the interconnection of MAC circuits to be adjusted in real-time based on temporal performance requirements. Configuration registers and control logic enable the system to activate only the necessary number of MAC circuits and interconnections needed to meet specific timing constraints, avoiding the need for permanently high complexity.
Solution Approach 2:
The system allows changing operational parameters such as the number of active MAC circuits, pipeline depth, and interconnection patterns to match temporal performance requirements. By varying these parameters dynamically, the system achieves high productivity when needed while maintaining lower complexity for less demanding tasks.
3Productivity
If separate multiplier-accumulator circuits with shadow registers are used to facilitate pipelining, then productivity is improved, but device complexity increases
Solution Approach 1:
The pipelining structure is segmented into multiple stages with shadow registers strategically placed at pipeline boundaries. Each 2-NMAX cluster contains multiple MAC circuits with associated shadow registers that enable independent pipeline stages to operate concurrently. This segmentation allows high productivity through efficient pipelining while managing register complexity through systematic placement.
Solution Approach 2:
The shadow registers enable continuous operation by allowing data to be transferred between pipeline stages without interruption. While one MAC circuit performs multiplication, shadow registers hold intermediate results for the next stage, ensuring continuous useful action across all pipeline stages and maximizing productivity.
Data Source
AI summary
An integrated circuit including configurable multiplier-accumulator circuitry, wherein, during processing operations, a plurality of the multiplier-accumulator circuits are serially connected into pipelines to perform concatenated multiply and accumulate operations. The integrated circuit includes a first memory and a second memory, and a switch interconnect network, including configurable multiplexers arranged in a plurality of switch matrices. The first and second memories are configurable as either a dedicated read memory or a dedicated write memory and connected to a given pipeline, via the switch interconnect network, during a processing operation performed thereby; wherein, during a first processing operations, the first memory is dedicated to write data to a first pipeline and the second memory is dedicated to read data therefrom and, during a second processing operation, the first memory is dedicated to read data from a second pipeline and the second memory is dedicated to write data thereto.


