In-Memory Compute Chiplet for Transformer Workloads

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Transformer-based neural network models, particularly in natural language processing, face significant challenges due to high computational intensity and memory requirements, making it inefficient to serve large-scale NLP models effectively.

Innovation Solution

The implementation of AI accelerator apparatuses using chiplet devices with digital in-memory compute (DIMC) functionality, which integrates computational functions and memory fabric, and includes SIMD devices to accelerate attention functions and softmax computations, along with die-to-die interconnects and PCIe buses for efficient communication, enabling scalable and efficient processing of transformer workloads.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If traditional computing architectures are used to process transformer workloads, then general-purpose computing capability is maintained, but computational performance is insufficient and power consumption is high

Engineering Contradiction:
Improvecomputational performanceVSAvoidpower consumption
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The system is divided into multiple chiplet devices, each containing specialized DIMC engines and memory fabric. These chiplets can be independently configured and scaled to match specific transformer workload requirements, enabling efficient parallel processing while maintaining modular power management

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent replaces traditional von Neumann architecture with an in-memory compute architecture where computation is performed directly within the memory fabric. This eliminates the need for continuous data movement between separate CPU and memory units, dramatically reducing energy consumption while increasing computational throughput for transformer operations

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If model size and compute requirements are increased to improve NLP performance, then inference accuracy is improved, but training time and computational resource requirements increase significantly

Engineering Contradiction:
Improveinference accuracyVSAvoidtraining time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary configuration of chiplet devices to match specific transformer model architectures. The modular chiplet design allows pre-optimization of compute and memory resources for particular model sizes and types, enabling rapid deployment of large-scale models without requiring linear increases in training time

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

Multiple chiplet devices are combined and interconnected through high-speed interfaces to form a unified computing system. This parallel architecture enables large transformer models to be processed across multiple devices simultaneously, reducing training time while maintaining the ability to handle trillion-parameter models

Inventive Principle:
Principle #5Merging (Combining)

3Productivity

If computational resources are increased to serve large-scale NLP models, then model serving capability is improved, but device complexity and scalability challenges increase

Engineering Contradiction:
Improvemodel serving capabilityVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system uses segmented chiplet devices that can be independently configured for specific transformer workload types. Each chiplet contains self-contained DIMC engines and memory fabric, allowing the system to scale by simply adding identical or heterogeneous chiplets rather than redesigning the entire system architecture

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The chiplet devices are designed with universal interfaces and configurable DIMC engines that can handle various transformer operations including attention mechanisms, feedforward networks, and normalization layers. This multi-functionality allows a single chiplet design to serve multiple model types and sizes, reducing overall system complexity

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12260223B2Generative AI accelerator apparatus using in-memory compute chiplet devices for transformer workloads
Publication Date: 2025.03.25 D-MATRIX CORP
  • US12260223B2 patent drawing
  • US12260223B2 patent drawing
  • US12260223B2 patent drawing

AI summary

An AI accelerator apparatus using in-memory compute chiplet devices. The apparatus includes one or more chiplets, each of which includes a plurality of tiles. Each tile includes a plurality of slices, a central processing unit (CPU), and a hardware dispatch device. Each slice can include a digital in-memory compute (DIMC) device configured to perform high throughput computations. In particular, the DIMC device can be configured to accelerate the computations of attention functions for transformer-based models (a.k.a. transformers) applied to machine learning applications, including generative AI. A single input multiple data (SIMD) device configured to further process the DIMC output and compute softmax functions for the attention functions. The chiplet can also include die-to-die (D2D) interconnects, a peripheral component interconnect express (PCIe) bus, a dynamic random access memory (DRAM) interface, and a global CPU interface to facilitate communication between the chiplets, memory and a server or host system.