Multi-DPU Scaling Architecture for Neural Network Clusters
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing DRAM-based processing units (DPUs) fall short of replicating the neural network capabilities of a human brain, requiring hundreds to thousands of DPUs to implement a human brain-like neural network, and face communication overhead challenges compared to CPU/GPU scaling-out.
Innovation Solution
A multi-DPU scaling-out architecture is employed, where each memory unit is configurable to operate as memory, a computation unit, or a hybrid memory-computation unit, allowing for job partitioning, data distribution, and collection across multiple DPUs in a scalable cluster architecture.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Power
If DPU capacity is increased to match human brain neural network capabilities, then computational capability is improved, but device complexity increases due to requiring hundreds to thousands of DPUs
Solution Approach 1:
The system divides the computational task into distributed DPUs, where each DPU handles a portion of the neural network processing. This segmentation allows the system to achieve human brain-like computational capability by aggregating the processing power of multiple DPUs while maintaining manageable complexity at each individual unit level.
Solution Approach 2:
The patent transitions from a single-DPU architecture to a multi-DPU distributed architecture, adding the dimension of spatial distribution. This dimensional change enables the system to scale computational capability by adding more DPUs to the cluster rather than increasing the complexity of a single DPU.
2Power
If DPU cluster size is increased to provide human brain-like neural network, then computational capability is improved, but communication overhead increases
Solution Approach 1:
The system architecture is designed to minimize communication overhead by creating an efficient distributed computing environment where DPUs can operate with reduced communication dependencies. The host controller strategically partitions jobs and distributes data to optimize the computational-communication tradeoff across the DPU cluster.
Data Source
AI summary
A processor includes a plurality of memory units, each of the memory units including a plurality of memory cells, wherein each of the memory units is configurable to operate as memory, as a computation unit, or as a hybrid memory-computation unit.


