3D FeRAM Memory-Compute Stack for Low-Latency AI Processing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing AI processor systems face challenges in reducing latency and power consumption during the training and inference processes, which are hardware-intensive activities.
Innovation Solution
The proposed solution involves a packaging technology that stacks a compute die on top of a memory die, utilizing ferroelectric random access memory (FeRAM) for improved performance. This configuration enhances data locality, reduces memory access latency, and lowers power consumption.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of time
If traditional memory and compute architectures are used, then device complexity is reduced, but memory access latency increases and power consumption increases
Solution Approach 1:
The patent merges memory and compute functions into a single integrated device, where memory cells directly perform computational operations. This integration eliminates the need for separate memory access operations, reducing latency while consolidating device complexity into a unified architecture that operates more efficiently.
Solution Approach 2:
The patent transitions from traditional two-dimensional planar memory architectures to three-dimensional vertical stacking, where memory layers are stacked above compute layers. This dimensional change enables simultaneous memory access and computation operations, reducing latency without proportionally increasing device complexity.
2Use of energy by stationary object
If traditional memory access methods are used, then device complexity is maintained, but power consumption increases
Solution Approach 1:
By merging memory storage and compute operations into a single integrated structure, the patent eliminates redundant data transfer operations between separate memory and compute units. This reduces power consumption significantly while the integrated architecture manages complexity through unified control mechanisms.
3Productivity
If memory and compute are integrated, then productivity increases, but device complexity increases
Solution Approach 1:
The patent uses three-dimensional vertical stacking to integrate memory and compute layers, enabling high-density integration that increases AI processing throughput. The vertical architecture allows multiple operations to occur simultaneously in different layers, improving productivity while the modular layering approach manages integration complexity.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
The use of FeRAM in the memory die significantly accelerates matrix multiplication processes by 15 to 20 times compared to traditional methods, while also reducing power consumption and interconnect energy, thus enhancing the overall performance and efficiency of AI processing systems.
Implementation Method 1
The first die includes a ferroelectric random access memory (FeRAM) having bit-cells, wherein each bit-cell includes an access transistor and a capacitor including ferroelectric material
Data Source
AI summary
Described is a packaging technology to improve performance of an AI processing system. An IC package is provided which comprises: a substrate; a first die on the substrate, and a second die stacked over the first die. The first die includes memory and the second die includes computational logic. The first die comprises a ferroelectric RAM (FeRAM) having bit-cells. Each bit-cell comprises an access transistor and a capacitor including ferroelectric material. The access transistor is coupled to the ferroelectric material. The FeRAM can be FeDRAM or FeSRAM. The memory of the first die may store input data and weight factors. The computational logic of the second die is coupled to the memory of the first die. The second die is an inference die that applies fixed weights for a trained model to an input data to generate an output. In one example, the second die is a training die that enables learning of the weights.


