Reconfigurable Accelerator Architecture for Versatile Processing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Application-specific accelerators require costly redesign and verification for new processing algorithms and have limited versatility due to their narrow functionality, making them inefficient for a wide range of applications.
Innovation Solution
A reconfigurable accelerator architecture combining a microcontroller, stream processor, and dataflow processor that can automatically handle simple memory access patterns and computational intensity, allowing for rapid reconfiguration and implementation of various special-purpose accelerator functions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If application-specific accelerators are used for dedicated processing tasks, then processing speed and energy efficiency are improved, but versatility and adaptability to new algorithms deteriorate
Solution Approach 1:
The patent implements a reconfigurable accelerator architecture that can be dynamically programmed to perform different processing algorithms. The system uses a configuration interface that allows loading of different configuration data to reconfigure the processing elements, enabling a single hardware device to function as multiple specialized accelerators for different applications such as image processing, video encoding, and machine learning workloads.
Solution Approach 2:
The accelerator employs dynamic reconfiguration capability where the processing elements can change their functionality during operation or between tasks. The configuration data stored in memory can be loaded at different times to adapt the hardware architecture to match the requirements of different algorithms, making the system flexible rather than fixed.
2Productivity
If application-specific accelerators are designed for narrow functionality, then processing efficiency for specific tasks is improved, but implementation of new algorithms requires costly redesign and verification
Solution Approach 1:
The reconfigurable accelerator uses a universal set of processing elements that can be programmed through software to implement different algorithms. Instead of requiring hardware redesign for new algorithms, the system loads new configuration data that reprograms the existing processing elements, dramatically reducing the cost and time required to implement new processing functions.
Solution Approach 2:
The system changes its operational parameters by loading different configuration data that modifies the behavior of processing elements. This allows the same physical hardware to exhibit different functional characteristics depending on the loaded configuration, enabling efficient implementation of new algorithms without physical redesign.
3Adaptability or versatility
If reconfigurable accelerator architecture is implemented, then versatility is improved, but circuit area and power consumption increase
Solution Approach 1:
The accelerator is divided into multiple independent processing elements that can be selectively activated. Each processing element is a self-contained unit with its own arithmetic logic unit and configuration capabilities. This segmentation allows the system to activate only the processing elements needed for a given task, reducing the effective circuit area in use and improving power efficiency.
Solution Approach 2:
The system implements partial reconfiguration where only the necessary processing elements are activated for each specific task. Rather than requiring all processing elements to be permanently active and interconnected, the system activates subsets of processing elements as needed, reducing the overall circuit area and power consumption while maintaining versatility.
Data Source
AI summary
A reconfigurable hardware accelerator for computers combines a high-speed dataflow processor, having programmable functional units rapidly reconfigured in a network of programmable switches, with a stream processor that may autonomously access memory in predefined access patterns after receiving simple stream instructions. The result is a compact, high-speed processor that may exploit parallelism associated with many application-specific programs susceptible to acceleration.


