Tiled Streaming Processor Architecture for Deep Learning Efficiency

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional chip multiprocessor (CMP) architectures face challenges in scalability, performance, and energy efficiency due to increasing workload complexity and size, particularly in vector and matrix calculations required for deep learning and neural network applications.

Innovation Solution

The Tensor Streaming Processor (TSP) architecture employs a tiled design with computational tiles and data storage/switching tiles interconnected in a Superlane structure, enabling data transfers and leveraging dataflow locality to enhance performance and reduce energy consumption.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If traditional chip multiprocessor (CMP) architectures are used to handle increasing workload complexity and size, then processing capacity is maintained, but scalability and energy efficiency deteriorate

Engineering Contradiction:
Improveprocessing capacityVSAvoidenergy efficiency
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The processor is divided into multiple independent streaming processors organized in a grid layout, where each streaming processor can independently execute streaming instructions. This segmentation allows the system to scale processing capacity by activating only the necessary number of streaming processors for a given workload, improving energy efficiency by avoiding the energy consumption of idle cores while maintaining scalability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transitions from traditional vertical core arrangements to a two-dimensional grid layout of streaming processors with associated row and column interconnect networks. This dimensional change enables more efficient data routing and reduces the number of memory access hops required, thereby improving energy efficiency while maintaining high processing capacity through parallel operations across the grid.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Productivity

If traditional CMP architectures are used, then processing power is maintained, but scalability deteriorates due to increasing workload complexity

Engineering Contradiction:
Improveprocessing powerVSAvoidscalability
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The system dynamically activates or deactivates streaming processors based on workload requirements. The control logic can enable or disable specific streaming processors in the grid according to the computational demands of different applications, allowing the processing power to scale dynamically with workload complexity while maintaining adaptability to various computational patterns.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

Each streaming processor is designed with a unified architecture that can handle multiple types of operations including integer arithmetic, floating-point operations, and memory access through a single instruction stream. This universal design allows the same hardware structure to scale effectively across different workload types without requiring specialized architectures for each application domain.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Productivity

If more processing cores are added to increase performance, then computational requirements are met, but device complexity increases

Engineering Contradiction:
ImproveperformanceVSAvoiddevice complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

Multiple streaming processors share common resources including the interconnect network infrastructure, memory controllers, and control logic. By merging these resources across the grid of processing units, the system achieves high performance through parallel processing while avoiding the linear increase in device complexity that would result from providing dedicated resources to each core.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system uses replicated streaming processor units with identical or similar architectures throughout the grid. This copying approach simplifies the overall design by using a standardized building block that can be replicated to achieve the desired performance level, rather than designing and managing increasingly complex heterogeneous core architectures as performance requirements grow.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS12340300B1Streaming processor architecture
Publication Date: 2025.06.24 GROQ INC
  • US12340300B1 patent drawing
  • US12340300B1 patent drawing
  • US12340300B1 patent drawing

AI summary

Improved placement of memory and functional modules, ‘tiles’, within a tiled processor architecture are disclosed for linear algebra calculations involving vectors and matrices comprising large amounts of data. The improved placement places the data in close proximity to the functional modules performing calculations using the data. These modules enable these calculations to be performed more quickly while using less energy. These modules, in particular, improve the efficiency of the training and application of deep learning and artificial neural network systems. This Abstract and the independent Claims are concise signifiers of embodiments of the claimed inventions. The Abstract does not limit the scope of the claimed inventions.