Heterogeneous ML Accelerator Clusters for Flexible Memory-Compute Balance

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

State-of-the-art machine learning models require hundreds to tens of thousands of accelerators with fixed computing and memory resources, leading to suboptimal resource allocation as different models need varying balances of computing and memory, resulting in stranded resources.

Innovation Solution

A heterogeneous machine learning accelerator system with non-uniformly distributed memory and compute nodes connected by high-speed chip-to-chip interconnects, utilizing prefetch, intelligent compression, and memory swapping to dynamically balance resources without remote processing units.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If fixed computing and memory resources are allocated to each accelerator, then system structure is simplified, but resource utilization efficiency deteriorates due to stranded resources for different models

Engineering Contradiction:
Improvesystem structureVSAvoidresource utilization efficiency
Core Design Contradiction:
Device complexityVSProductivity

Solution Approach 1:

The system segments memory resources from compute resources by introducing separate memory nodes that can be independently allocated. Memory is divided into multiple memory nodes with different capacities and speeds, allowing flexible assignment to different compute nodes based on model requirements, thus eliminating stranded resources while maintaining structural clarity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system implements dynamic resource allocation where memory nodes can be dynamically assigned to different compute nodes based on the specific machine learning model being executed. This dynamic configuration allows the system to adapt resource balances to match different model requirements, improving resource utilization efficiency without requiring a complex reconfiguration of the overall system architecture.

Inventive Principle:
Principle #15Dynamics

2Quantity of substance

If memory capacity is expanded using remote processing units, then memory capacity increases, but system complexity and cost increase

Engineering Contradiction:
Improvememory capacityVSAvoidsystem complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The system introduces universal memory nodes that can serve multiple compute nodes simultaneously. These memory nodes are designed with standardized interfaces and protocols, allowing them to be shared across different compute nodes without requiring dedicated remote processing units for each memory expansion, thereby reducing system complexity while expanding memory capacity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system introduces a high-speed interconnect as an intermediary between memory nodes and compute nodes. This interconnect acts as a mediator that enables direct, high-bandwidth communication between memory and compute resources without requiring remote processing units, simplifying the system architecture while achieving memory expansion.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Speed

If high speed interconnects are used to connect thousands of accelerators, then computation speed improves, but system cost and complexity increase

Engineering Contradiction:
Improvecomputation speedVSAvoidsystem complexity
Core Design Contradiction:
SpeedVSDevice complexity

Solution Approach 1:

The system segments the interconnect architecture into hierarchical levels, with high-speed interconnects used only where critical for performance (between memory nodes and compute nodes), while lower-speed connections are used for less time-sensitive communications. This segmentation maintains computation speed where needed while reducing overall system complexity and cost.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system applies high-speed interconnects locally at critical paths between memory nodes and compute nodes, rather than uniformly across all system connections. This localized application of high-speed connectivity optimizes computation speed for memory-access-intensive operations while avoiding the cost and complexity of high-speed interconnects throughout the entire system.

Inventive Principle:
Principle #3Local quality

4Device complexity

If uniform memory distribution is used across compute nodes, then system balance is simplified, but performance deteriorates for models with varying memory requirements

Engineering Contradiction:
Improvesystem balanceVSAvoidmodel performance
Core Design Contradiction:
Device complexityVSProductivity

Solution Approach 1:

The system implements non-uniform memory distribution where different compute nodes can access memory nodes with different capacities and speeds based on their specific requirements. This allows each compute node to have optimized memory characteristics tailored to the models it executes, improving model performance while maintaining manageable system balance through standardized memory node interfaces.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system enables dynamic memory allocation where the distribution and assignment of memory nodes to compute nodes can be adjusted based on the specific model being executed. This dynamic configuration allows the system to optimize memory distribution for different workloads, improving model performance without requiring a permanently complex static configuration.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS12417047B2Heterogeneous ML accelerator cluster with flexible system resource balance
Publication Date: 2025.09.16 GOOGLE LLC
  • US12417047B2 patent drawing
  • US12417047B2 patent drawing
  • US12417047B2 patent drawing

AI summary

Aspects of the disclosure are directed to a heterogeneous machine learning accelerator system with compute and memory nodes connected by high speed chip-to-chip interconnects. While existing remote/disaggregated memory may require memory expansion via remote processing units, aspects of the disclosure add memory nodes into machine learning accelerator clusters via the chip-to-chip interconnects without needing assistance from remote processing units to achieve higher performance, simpler software stack, and/or lower cost. The memory nodes may support prefetch and intelligent compression to enable the use of low cost memory without performance degradation.