Reconfigurable Crossbar Switch for CNN Accelerator Adaptability
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current convolutional accelerators face challenges in efficiently managing the varied memory and compute requirements across different layers of convolutional neural networks (CNNs), particularly in resource-constrained hardware, which hampers real-time performance and throughput for applications like image recognition.
Innovation Solution
The proposed solution involves a hardware accelerator with a reconfigurable crossbar switch and stream engines that enable flexible interconnection of processing elements, allowing for dynamic mode switching between compute and memory operations, and a hierarchical data storage hierarchy to optimize data locality and reuse, thereby accommodating the diverse needs of CNN layers.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If a fixed architecture is used in convolutional accelerators, then the hardware design is simple, but it cannot adapt to varied memory and compute requirements across different CNN layers
Solution Approach 1:
The patent implements a dynamic architecture where processing elements can switch between compute mode and memory mode based on operational requirements. The reconfigurable crossbar switch dynamically connects processing elements to different functional units (compute units or memory units) depending on whether a given layer requires computation or data storage, enabling the hardware to adapt to varied memory and compute requirements across different CNN layers without requiring separate fixed architectures for each function
2Productivity
If resource-constrained hardware is used, then the device is more efficient and cheaper, but real-time performance and throughput are hampered
Solution Approach 1:
The patent creates a universal processing element that can function as either a compute unit or a memory unit depending on configuration. This multi-functionality allows a single hardware resource to serve multiple purposes, increasing the effective computational capacity and data storage capacity within the same physical footprint. By enabling processing elements to be dynamically assigned to compute or memory roles, the system achieves higher throughput and real-time performance without proportionally increasing hardware resource constraints
Solution Approach 2:
The dynamic reconfiguration capability allows the system to optimize resource allocation in real-time based on the specific computational tasks. When compute-intensive operations are needed, more processing elements can be configured as compute units; when data storage is critical, more elements can be configured as memory units. This dynamic adaptation enables the system to maintain high productivity and real-time performance while operating within constrained hardware resources
Data Source
AI summary
A hardware accelerator includes a plurality of functional circuits, a stream switch, and a plurality of stream engines. The stream engines are coupled to the functional circuits via the stream switch, and in operation, generate data streaming requests to stream data to and from the functional circuits. The functional circuits include at least one convolutional cluster, which includes a plurality of processing elements coupled together via a reconfigurable crossbar switch. The reconfigurable crossbar switch is coupled to the stream switch, and in operation, streams data to, from, and between processing elements of the processing cluster.


