AI Accelerator Mini-Buffer Layout for Lower Latency
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Larger monolithic on-chip buffers in ASIC-based AI accelerator devices suffer from increased latency, greater access energy, and reduced performance due to increased wire routing complexity and limited memory bandwidth, leading to underutilization of processing elements.
Innovation Solution
Partitioning on-chip buffers into mini buffers associated with subsets of rows and columns of the processing element array, using a distributor circuit to direct data to the appropriate mini buffers, reducing wire routing complexity and increasing bandwidth.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If monolithic on-chip buffers are used to store neural network data, then data storage capacity is improved, but wire routing complexity increases and memory bandwidth is limited
Solution Approach 1:
The patent divides the monolithic on-chip buffer into multiple smaller sub-buffers distributed across the processing element array. Each sub-buffer is associated with specific processing elements, reducing the wiring complexity required for any single buffer while maintaining total storage capacity. The distributor circuit segments data routing to appropriate sub-buffers based on processing requirements.
2Quantity of substance
If monolithic on-chip buffers are used, then data storage capacity is improved, but access energy increases due to limited memory bandwidth
Solution Approach 1:
By segmenting the monolithic buffer into distributed sub-buffers, the patent enables parallel access paths for multiple processing elements. This increases effective memory bandwidth by allowing simultaneous data retrieval for different processing operations, thereby reducing the energy required per access operation.
Solution Approach 2:
The patent transitions from a centralized monolithic buffer architecture to a distributed two-dimensional buffer array that mirrors the processing element layout. This spatial reorganization enables shorter data paths and parallel access, reducing both access time and energy consumption while maintaining total storage capacity.
3Quantity of substance
If monolithic on-chip buffers are used, then data storage capacity is improved, but processing element utilization decreases due to underutilization
Solution Approach 1:
The patent associates specific sub-buffers with specific processing elements or groups of processing elements, enabling each processing element to efficiently access its dedicated data storage. This reduces access conflicts and waiting time, thereby improving processing element utilization and overall system productivity while maintaining adequate data storage capacity.
4Quantity of substance
If monolithic on-chip buffers are used, then data storage capacity is improved, but latency increases due to increased wire routing complexity
Solution Approach 1:
By dividing the monolithic buffer into distributed sub-buffers located near relevant processing elements, the patent reduces the physical distance data must travel. This segmentation creates shorter data paths and reduces propagation latency while maintaining adequate storage capacity through the collective capacity of all sub-buffers.
Solution Approach 2:
The patent introduces distributor circuits as intermediary components that efficiently route data from external memory or between buffers and processing elements. These distributors optimize data flow paths, reducing routing complexity and minimizing latency by directing data to the appropriate sub-buffer or processing element through optimized pathways.
Data Source
AI summary
An artificial intelligence (AI) accelerator device may include a plurality of on-chip mini buffers that are associated with a processing element (PE) array. Each mini buffer is associated with a subset of rows or a subset of columns of the PE array. Partitioning an on-chip buffer of the AI accelerator device into the mini buffers described herein may reduce the size and complexity of the on-chip buffer. The reduced size of the on-chip buffer may reduce the wire routing complexity of the on-chip buffer, which may reduce latency and may reduce access energy for the AI accelerator device. This may increase the operating efficiency and/or may increase the performance of the AI accelerator device. Moreover, the mini buffers may increase the overall bandwidth that is available for the mini buffers to transfer data to and from the PE array.


