Heterogeneous Accelerator with Stacked HBM for Learning Systems

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Emerging applications like deep neural networks require massive computational and memory resources for efficient training and learning, while also demanding energy efficiency and low latency, which existing technologies struggle to meet due to power consumption and latency issues in data-intensive tasks.

Innovation Solution

A heterogeneous computing environment is established, comprising a processing unit, a reprogrammable processing unit, and a stack of high-bandwidth memory dies, where a task scheduler coordinates computational tasks between these units to leverage processing-in-memory functionality, optimizing energy efficiency and latency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If processing-in-memory functionality is implemented using conventional memory architectures, then computational ability is improved, but power consumption increases

Engineering Contradiction:
Improvecomputational abilityVSAvoidpower consumption
Core Design Contradiction:
ProductivityVSUse of energy by stationary object

Solution Approach 1:

The patent merges memory and processing functions by stacking processing-in-memory dies directly on top of high-bandwidth memory dies, creating an integrated compute-memory system that reduces data movement and associated power consumption while maintaining high computational ability

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent transitions from conventional 2D memory-processor architecture to a 3D stacked architecture, placing processing dies vertically above memory dies to minimize data access distance and reduce power consumption while enhancing computational throughput

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Speed

If data is frequently transferred between memory and processing units, then computational speed is improved, but latency increases

Engineering Contradiction:
Improvecomputational speedVSAvoidlatency
Core Design Contradiction:
SpeedVSLoss of time

Solution Approach 1:

By integrating processing units directly on the memory stack, the patent eliminates traditional memory-bus interfaces and reduces data transfer latency, allowing high-speed computational access without the overhead of conventional memory access protocols

Inventive Principle:
Principle #5Merging (Combining)

3Speed

If high-bandwidth memory is used to increase data access speed, then memory bandwidth is improved, but device complexity increases

Engineering Contradiction:
Improvememory bandwidthVSAvoiddevice complexity
Core Design Contradiction:
SpeedVSDevice complexity

Solution Approach 1:

The patent implements a heterogeneous computing environment where a single HBM stack can serve multiple processing units with different functionalities (fixed-function and reprogrammable), allowing high memory bandwidth to be shared across diverse computational workloads without proportionally increasing device complexity

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20240193111A1Heterogeneous accelerator for highly efficient learning systems
Publication Date: 2024.06.13 SAMSUNG ELECTRONICS CO LTD
  • US20240193111A1 patent drawing
  • US20240193111A1 patent drawing
  • US20240193111A1 patent drawing

AI summary

An apparatus may include a heterogeneous computing environment that may be controlled, at least in part, by a task scheduler in which the heterogeneous computing environment may include a processing unit having fixed logical circuits configured to execute instructions; a reprogrammable processing unit having reprogrammable logical circuits configured to execute instructions that include instructions to control processing-in-memory functionality; and a stack of high-bandwidth memory dies in which each may be configured to store data and to provide processing-in-memory functionality controllable by the reprogrammable processing unit such that the reprogrammable processing unit is at least partially stacked with the high-bandwidth memory dies. The task scheduler may be configured to schedule computational tasks between the processing unit, and the reprogrammable processing unit.