Multi-DPU Scaling Architecture for Neural Network Clusters

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing DRAM-based processing units (DPUs) fall short of replicating the neural network capabilities of a human brain, requiring hundreds to thousands of DPUs to implement a human brain-like neural network, and face communication overhead challenges compared to CPU/GPU scaling-out.

Innovation Solution

A multi-DPU scaling-out architecture is employed, where each memory unit is configurable to operate as memory, a computation unit, or a hybrid memory-computation unit, allowing for job partitioning, data distribution, and collection across multiple DPUs in a scalable cluster architecture.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Power

If DPU capacity is increased to match human brain neural network capabilities, then computational capability is improved, but device complexity increases due to requiring hundreds to thousands of DPUs

Engineering Contradiction:
Improvecomputational capabilityVSAvoidsystem complexity
Core Design Contradiction:
PowerVSDevice complexity

Solution Approach 1:

The system divides the computational task into distributed DPUs, where each DPU handles a portion of the neural network processing. This segmentation allows the system to achieve human brain-like computational capability by aggregating the processing power of multiple DPUs while maintaining manageable complexity at each individual unit level.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transitions from a single-DPU architecture to a multi-DPU distributed architecture, adding the dimension of spatial distribution. This dimensional change enables the system to scale computational capability by adding more DPUs to the cluster rather than increasing the complexity of a single DPU.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Power

If DPU cluster size is increased to provide human brain-like neural network, then computational capability is improved, but communication overhead increases

Engineering Contradiction:
Improvecomputational capabilityVSAvoidcommunication overhead
Core Design Contradiction:
PowerVSLoss of energy

Solution Approach 1:

The system architecture is designed to minimize communication overhead by creating an efficient distributed computing environment where DPUs can operate with reduced communication dependencies. The host controller strategically partitions jobs and distributes data to optimize the computational-communication tradeoff across the DPU cluster.

Inventive Principle:
Principle #12Equipotentiality

Data Source

PatentUS12340101B2Scaling out architecture for dram-based processing unit (DPU)
Publication Date: 2025.06.24 SAMSUNG ELECTRONICS CO LTD
  • US12340101B2 patent drawing
  • US12340101B2 patent drawing
  • US12340101B2 patent drawing

AI summary

A processor includes a plurality of memory units, each of the memory units including a plurality of memory cells, wherein each of the memory units is configurable to operate as memory, as a computation unit, or as a hybrid memory-computation unit.