Function-Based Activation of Memory Hierarchies

The use of a 3D compute-in-memory architecture addresses the inefficiencies in large deep learning models by performing computations directly in-memory, reducing the need for DRAM and enhancing inference speed and energy efficiency.

JP2025533740APending Publication Date: 2025-10-09INTERNATIONAL BUSINESS MACHINE CORPORATION
0 Cites 0 Cited by

Patent Information

Application Number
JP2025514236
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-09-13
Filing Date
2023-08-08
Publication Date
2025-10-09

AI Technical Summary

Technical Problem

Current deep learning architectures with large numbers of parameters face challenges in training and inference due to high computational costs and the need for large amounts of dynamic random access memory (DRAM), preventing fast and energy-efficient operations.

Method used

Implementing a neural network model system with a hierarchy of compute-in-memory structures using a 3D memory architecture, such as resistive memory or 3D NAND flash, to perform computations in-memory, thereby eliminating the need for constant weight passing between memory and CPU/GPU and reducing DRAM requirements.

Benefits of technology

Enables fast and energy-efficient inference for models with billions of parameters by leveraging 3D compute-in-memory systems, which store weights within the memory architecture and perform computations directly in-memory, thus enhancing computational speed and throughput.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

A 3D compute in-memory accelerator system and method for efficient inference of Mixture of Experts (MoE) neural network models are provided. The system includes multiple compute in-memory cores, each including multiple hierarchies of in-memory computational cells. One or more hierarchies of in-memory computational cells correspond to expert submodels of the MoE model. One or more expert submodels are selected for activation propagation based on a function-based routing, and the corresponding hierarchies of experts are activated based on the function. In one embodiment, the function is a hash-based hierarchical selection function used for dynamic routing of input and output activations. In one embodiment, the function is applied to select a single expert or multiple experts using an input database or layer activation-based MoE for activation of a single hierarchical layer. Furthermore, the system can be configured as a multi-model system with single expert model selection or a multi-model system with multiple expert selection.
Need to check novelty before this filing date? Find Prior Art