Reconfigurable AI Accelerator Processing Elements
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional AI accelerators are limited by their fixed dataflow architecture, which does not optimally support diverse neural network workloads, leading to inefficient performance, energy consumption, and increased area overhead.
Innovation Solution
The implementation of reconfigurable processing elements (PEs) within AI accelerators, equipped with multiplexors and controlled by various signals, allows for support of multiple dataflows, such as output stationary, input stationary, and weight stationary workflows, enabling efficient energy use and reduced area consumption by using lower precision adders and registers without compromising accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If fixed dataflow architecture is used in AI accelerators, then device complexity is reduced and ease of manufacture is improved, but adaptability to diverse neural network workloads deteriorates and productivity is limited
Solution Approach 1:
The patent implements reconfigurable processing elements that can dynamically change their dataflow configuration between different neural network workloads. The processing elements include multiplexors and configurable connections that allow the architecture to adapt its dataflow pattern (e.g., row-stationary, column-stationary, output-stationary) based on the specific computational requirements, thereby resolving the contradiction between adaptability and complexity by making the system dynamically adjustable rather than statically fixed
Solution Approach 2:
The patent designs processing elements with universal functionality that can perform multiple dataflow operations through reconfiguration. The same physical hardware structure can support different neural network layer types and computational patterns by changing its operational mode, eliminating the need for separate dedicated hardware for each workload type while maintaining ease of manufacture through standardized reusable components
2Manufacturing precision
If higher precision adders and registers are used, then manufacturing precision and reliability are improved, but area overhead and energy consumption increase
Solution Approach 1:
The patent applies local quality by using lower precision adders and registers in specific locations where full precision is not required, while maintaining higher precision only where necessary for accuracy-critical operations. The reconfigurable architecture allows precision to be locally optimized based on the specific computational workload, reducing overall area overhead while maintaining sufficient manufacturing precision for the application requirements
Solution Approach 2:
The patent changes the precision parameter of adders and registers dynamically based on the operational mode and workload requirements. The processing elements can adjust their precision levels through reconfiguration, using lower precision when adequate for the current task to reduce area overhead, and switching to higher precision when computational accuracy is critical, thereby resolving the contradiction between precision and area
3Productivity
If reconfigurable processing elements with multiplexors are implemented, then adaptability and productivity are improved, but device complexity and energy consumption increase
Solution Approach 1:
The patent segments the control logic and reconfiguration functions into modular components within each processing element. The multiplexors and configuration mechanisms are divided into manageable segments that can be independently controlled and optimized, reducing the overall device complexity while maintaining high adaptability and productivity through coordinated operation of these segmented components
Data Source
AI summary
A reconfigurable processing circuit of an AI accelerator and a method of operating the same are disclosed. In one aspect, the reconfigurable processing circuit includes a first memory configured to store an input activation state, a second memory configured to store a weight, a multiplier configured to multiply the weight and the input activation state and output a product, a first multiplexer (mux) configured to, based on a first selector, output a previous sum from a previous reconfigurable processing element, a third memory configured to store a first sum, a second mux configured to, based on a second selector, output the previous sum or the first sum, an adder configured to add the product and the previous sum or the first sum to output a second sum, and a third mux configured to, based on a third selector, output the second sum or the previous sum.


