3D Hybrid Bonding Memory-Logic Layout for AI Weight Bandwidth

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional memory devices suffer from insufficient bandwidth and high power consumption during numerous reads/writes, particularly in AI inferencing applications, leading to poor performance and increased costs.

Innovation Solution

A 3D hybrid bonding architecture is implemented using Copper to Copper (Cu-to-Cu) hybrid bond vias to directly connect processing units and memory cells, forming a high read endurance cell/array design with reduced read/write iterations, thereby enhancing memory bandwidth and reducing power consumption.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If conventional memory devices are used for AI inferencing applications, then storage capacity is provided, but bandwidth is insufficient and power consumption is high

Engineering Contradiction:
Improvememory bandwidthVSAvoidpower consumption
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The patent merges memory and processing functions by implementing processing elements directly on the memory die, creating a unified memory-compute architecture that eliminates separate memory modules and interconnects, thereby increasing bandwidth and reducing power consumption

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent transitions from conventional 2D memory architecture to 3D stacked architecture with multiple memory planes stacked vertically, enabling higher bandwidth through parallel access to multiple planes simultaneously and reducing the physical distance for data transfer

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Loss of time

If conventional memory architecture is used, then data storage is provided, but read/write operations consume excessive power and time

Engineering Contradiction:
Improveread/write timeVSAvoidpower consumption
Core Design Contradiction:
Loss of timeVSUse of energy by moving object

Solution Approach 1:

The patent implements cache memory structures that pre-load and store frequently accessed data, reducing the need for repeated read/write operations to main memory and thereby decreasing both time and power consumption for data access

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent enables continuous data flow between memory and processing elements through dedicated data paths and buffers, eliminating idle cycles and ensuring that data transfer operations occur continuously without interruption, improving efficiency and reducing overall operation time

Inventive Principle:
Principle #20Continuity of useful action

3Productivity

If conventional memory devices are used, then basic storage functionality is provided, but performance is poor for AI applications requiring high bandwidth

Engineering Contradiction:
ImprovebandwidthVSAvoidmemory architecture complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent divides the memory system into multiple independent planes and processing elements that can operate in parallel, with each plane having dedicated data paths to processing elements, thereby increasing total bandwidth while maintaining manageable complexity through modular design

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent designs processing elements that can handle multiple types of operations (integer arithmetic, floating-point operations, neural network computations) and memory structures that can serve multiple functions (storage, caching, buffering), reducing the need for specialized components and simplifying the overall architecture

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12585931B23D hybrid bonding 3D memory devices with NPU/CPU for AI inference application
Publication Date: 2026.03.24 MACRONIX INTERNATIONAL CO LTD
  • US12585931B2 patent drawing
  • US12585931B2 patent drawing
  • US12585931B2 patent drawing

AI summary

An AI inference platform comprises a logic die including an array of AI processing elements. Each AI processing element including an activation memory storing activation data for use in neural network computations. The platform includes a memory die that includes an array of 3D memory cells and a page buffer that facilitates storage and retrieval of neural network weights for use in neural network computations. A plurality of vertical connections can directly connect AI processing elements in the logic die and page buffers of corresponding ones of the memory cells in the memory die, enabling storage or retrieval of a neural network weight to and from a particular page buffer of a corresponding 3D memory cell for use in neural network computations conducted by a corresponding AI processing element in the logic die.