Tiled Streaming Processor Architecture for Deep Learning Efficiency
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional chip multiprocessor (CMP) architectures face challenges in scalability, performance, and energy efficiency due to increasing workload complexity and size, particularly in vector and matrix calculations required for deep learning and neural network applications.
Innovation Solution
The Tensor Streaming Processor (TSP) architecture employs a tiled design with computational tiles and data storage/switching tiles interconnected in a Superlane structure, enabling data transfers and leveraging dataflow locality to enhance performance and reduce energy consumption.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If traditional chip multiprocessor (CMP) architectures are used to handle increasing workload complexity and size, then processing capacity is maintained, but scalability and energy efficiency deteriorate
Solution Approach 1:
The processor is divided into multiple independent streaming processors organized in a grid layout, where each streaming processor can independently execute streaming instructions. This segmentation allows the system to scale processing capacity by activating only the necessary number of streaming processors for a given workload, improving energy efficiency by avoiding the energy consumption of idle cores while maintaining scalability.
Solution Approach 2:
The patent transitions from traditional vertical core arrangements to a two-dimensional grid layout of streaming processors with associated row and column interconnect networks. This dimensional change enables more efficient data routing and reduces the number of memory access hops required, thereby improving energy efficiency while maintaining high processing capacity through parallel operations across the grid.
2Productivity
If traditional CMP architectures are used, then processing power is maintained, but scalability deteriorates due to increasing workload complexity
Solution Approach 1:
The system dynamically activates or deactivates streaming processors based on workload requirements. The control logic can enable or disable specific streaming processors in the grid according to the computational demands of different applications, allowing the processing power to scale dynamically with workload complexity while maintaining adaptability to various computational patterns.
Solution Approach 2:
Each streaming processor is designed with a unified architecture that can handle multiple types of operations including integer arithmetic, floating-point operations, and memory access through a single instruction stream. This universal design allows the same hardware structure to scale effectively across different workload types without requiring specialized architectures for each application domain.
3Productivity
If more processing cores are added to increase performance, then computational requirements are met, but device complexity increases
Solution Approach 1:
Multiple streaming processors share common resources including the interconnect network infrastructure, memory controllers, and control logic. By merging these resources across the grid of processing units, the system achieves high performance through parallel processing while avoiding the linear increase in device complexity that would result from providing dedicated resources to each core.
Solution Approach 2:
The system uses replicated streaming processor units with identical or similar architectures throughout the grid. This copying approach simplifies the overall design by using a standardized building block that can be replicated to achieve the desired performance level, rather than designing and managing increasingly complex heterogeneous core architectures as performance requirements grow.
Data Source
AI summary
Improved placement of memory and functional modules, ‘tiles’, within a tiled processor architecture are disclosed for linear algebra calculations involving vectors and matrices comprising large amounts of data. The improved placement places the data in close proximity to the functional modules performing calculations using the data. These modules enable these calculations to be performed more quickly while using less energy. These modules, in particular, improve the efficiency of the training and application of deep learning and artificial neural network systems. This Abstract and the independent Claims are concise signifiers of embodiments of the claimed inventions. The Abstract does not limit the scope of the claimed inventions.


