Distributed GPU Architecture for Near-Storage Processing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current GPU architectures are limited in their memory storage capabilities and processing efficiency, particularly in handling large datasets and performing near-storage processing.

Innovation Solution

A distributed graphics processor unit (GPU) architecture that consists of an array of processing nodes, each equipped with a GPU node, fast memory, and storage unit, allowing for near-storage processing and improved memory storage capabilities.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If GPUs use bandwidth processors with small capacity to interface with host memory, then processing speed is improved, but memory storage capability deteriorates

Engineering Contradiction:
Improveprocessing speedVSAvoidmemory storage capability
Core Design Contradiction:
SpeedVSQuantity of substance

Solution Approach 1:

The system is divided into multiple processing nodes, each with its own GPU, fast memory, and storage unit. This segmentation allows each node to independently process data locally while collectively providing large storage capacity and high processing throughput across the distributed system.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transitions from a single-GPU architecture to a distributed array of processing nodes, adding a dimensional aspect to the system. This distributed architecture enables the system to achieve both high processing speed at the node level and large storage capacity at the system level by operating across multiple spatial dimensions.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Device complexity

If data is processed remotely from storage location, then system simplicity is maintained, but data traffic and processing efficiency deteriorate

Engineering Contradiction:
Improvesystem simplicityVSAvoiddata traffic and processing efficiency
Core Design Contradiction:
Device complexityVSProductivity

Solution Approach 1:

Each processing node is equipped with local fast memory and storage units that are directly coupled to its GPU. This local quality enables data to be processed close to where it is stored, reducing the need for data movement across the system and improving processing efficiency without significantly increasing overall system complexity.

Inventive Principle:
Principle #3Local quality

3Device complexity

If single chip GPU architecture is used, then device simplicity is maintained, but aggregate bandwidth and throughput deteriorate

Engineering Contradiction:
Improvedevice simplicityVSAvoidaggregate bandwidth and throughput
Core Design Contradiction:
Device complexityVSPower

Solution Approach 1:

The patent merges multiple processing nodes into a unified distributed system where each node contains a GPU, fast memory, and storage unit. By combining the resources of multiple nodes, the system achieves high aggregate bandwidth and multi-teraflop throughput while maintaining relative simplicity at the individual node level.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS20250028676A1Distributed graphics processor unit architecture
Publication Date: 2025.01.23 MICRON TECHNOLOGY INC
  • US20250028676A1 patent drawing
  • US20250028676A1 patent drawing
  • US20250028676A1 patent drawing

AI summary

The present disclosure is directed to a distributed graphics processor unit (GPU) architecture that includes an array of processing nodes. Each processing node may include a GPU node that is coupled to its own fast memory unit and its own storage unit. The fast memory unit and storage unit may be integrated into a single unit or may be separately coupled to the GPU node. The processing node may have its fast memory unit coupled to both the GPU node and the storage node. The various architectures provide a GPU-based system that may be treated as a storage unit, such as solid state drive (SSD) that performs onboard processing to perform memory-oriented operations. In this respect, the system may be viewed as a “smart drive” for big-data near-storage processing.