Reconfigurable AI Accelerator Processing Elements

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional AI accelerators are limited by their fixed dataflow architecture, which does not optimally support diverse neural network workloads, leading to inefficient performance, energy consumption, and increased area overhead.

Innovation Solution

The implementation of reconfigurable processing elements (PEs) within AI accelerators, equipped with multiplexors and controlled by various signals, allows for support of multiple dataflows, such as output stationary, input stationary, and weight stationary workflows, enabling efficient energy use and reduced area consumption by using lower precision adders and registers without compromising accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If fixed dataflow architecture is used in AI accelerators, then device complexity is reduced and ease of manufacture is improved, but adaptability to diverse neural network workloads deteriorates and productivity is limited

Engineering Contradiction:
Improveadaptability to diverse neural network workloadsVSAvoiddataflow architecture complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements reconfigurable processing elements that can dynamically change their dataflow configuration between different neural network workloads. The processing elements include multiplexors and configurable connections that allow the architecture to adapt its dataflow pattern (e.g., row-stationary, column-stationary, output-stationary) based on the specific computational requirements, thereby resolving the contradiction between adaptability and complexity by making the system dynamically adjustable rather than statically fixed

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent designs processing elements with universal functionality that can perform multiple dataflow operations through reconfiguration. The same physical hardware structure can support different neural network layer types and computational patterns by changing its operational mode, eliminating the need for separate dedicated hardware for each workload type while maintaining ease of manufacture through standardized reusable components

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Manufacturing precision

If higher precision adders and registers are used, then manufacturing precision and reliability are improved, but area overhead and energy consumption increase

Engineering Contradiction:
Improvecomputational accuracyVSAvoidarea overhead
Core Design Contradiction:
Manufacturing precisionVSArea of stationary object

Solution Approach 1:

The patent applies local quality by using lower precision adders and registers in specific locations where full precision is not required, while maintaining higher precision only where necessary for accuracy-critical operations. The reconfigurable architecture allows precision to be locally optimized based on the specific computational workload, reducing overall area overhead while maintaining sufficient manufacturing precision for the application requirements

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent changes the precision parameter of adders and registers dynamically based on the operational mode and workload requirements. The processing elements can adjust their precision levels through reconfiguration, using lower precision when adequate for the current task to reduce area overhead, and switching to higher precision when computational accuracy is critical, thereby resolving the contradiction between precision and area

Inventive Principle:
Principle #35Parameter changes

3Productivity

If reconfigurable processing elements with multiplexors are implemented, then adaptability and productivity are improved, but device complexity and energy consumption increase

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidprocessing element complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments the control logic and reconfiguration functions into modular components within each processing element. The multiplexors and configuration mechanisms are divided into manageable segments that can be independently controlled and optimized, reducing the overall device complexity while maintaining high adaptability and productivity through coordinated operation of these segmented components

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20240028869A1Reconfigurable processing elements for artificial intelligence accelerators and methods for operating the same
Publication Date: 2024.01.25 TAIWAN SEMICONDUCTOR MANUFACTURING CO LTD
  • US20240028869A1 patent drawing
  • US20240028869A1 patent drawing
  • US20240028869A1 patent drawing

AI summary

A reconfigurable processing circuit of an AI accelerator and a method of operating the same are disclosed. In one aspect, the reconfigurable processing circuit includes a first memory configured to store an input activation state, a second memory configured to store a weight, a multiplier configured to multiply the weight and the input activation state and output a product, a first multiplexer (mux) configured to, based on a first selector, output a previous sum from a previous reconfigurable processing element, a third memory configured to store a first sum, a second mux configured to, based on a second selector, output the previous sum or the first sum, an adder configured to add the product and the previous sum or the first sum to output a second sum, and a third mux configured to, based on a third selector, output the second sum or the previous sum.