DIMC Chiplet Architecture for Low-Latency AI Inference

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional Network Interface Cards (NICs) face challenges in scaling across distributed AI accelerators for applications like NLP and computer vision, with RDMA over Converged Ethernet and InfiniBand solutions being complex and limited by PCIe fabric topologies, hindering low-latency performance in Generative AI inferences.

Innovation Solution

A digital in-memory compute (DIMC) accelerator system using a chiplet architecture with modular chiplets, each comprising tiles and slices, optimized for high throughput computations and memory bandwidth, dynamically switching precision levels, and integrating computational functions and memory fabric to accelerate transformer models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If RDMA over Converged Ethernet combined with InfiniBand is used to facilitate multi-node GPU accelerator communication, then communication capability is improved, but system complexity increases due to complex shared address space requirements

Engineering Contradiction:
Improvecommunication capabilityVSAvoidsystem complexity
Core Design Contradiction:
SpeedVSDevice complexity

Solution Approach 1:

The system is divided into discrete chiplet devices that can be independently configured and deployed. Each chiplet represents a modular unit that can be scaled and organized in different topologies, simplifying the overall system architecture while maintaining communication capabilities.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The chiplet architecture provides a universal interface that can work across different network configurations and topologies. The standardized chiplet design allows the same building blocks to serve multiple functions and be deployed in various scenarios without requiring complex specialized configurations.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Loss of time

If RDMA over Converged Ethernet or InfiniBand fabric solutions are used to meet low-latency demands, then communication speed is improved, but deployment complexity increases

Engineering Contradiction:
ImprovelatencyVSAvoiddeployment complexity
Core Design Contradiction:
Loss of timeVSEase of operation

Solution Approach 1:

By segmenting the system into standardized chiplet devices, deployment becomes more manageable. Each chiplet can be independently tested, configured, and deployed, reducing the overall deployment complexity while maintaining low-latency performance through optimized inter-chiplet communication.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system allows dynamic adjustment of communication parameters and precision levels to optimize performance for different workloads. This flexibility enables the system to adapt to varying latency requirements without requiring complex reconfiguration or redeployment.

Inventive Principle:
Principle #35Parameter changes

3Adaptability or versatility

If PCIe fabric topologies are used for intra-switch communication, then connectivity is improved, but scalability is limited by available PCIe lanes

Engineering Contradiction:
ImproveconnectivityVSAvoidscalability
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The system transitions from a centralized PCIe fabric to a distributed chiplet-based architecture. This segmentation allows each chiplet to have direct communication paths, eliminating the bottleneck of shared PCIe lanes and enabling linear scaling as more chiplets are added to the system.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The architecture moves from a two-dimensional PCIe switch fabric to a multi-dimensional chiplet network topology. This allows communication paths to extend in multiple directions and layers, providing scalability beyond the constraints of traditional PCIe lane limitations.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

4Measurement precision

If precision is increased to maintain computational accuracy, then computation quality is improved, but power consumption increases

Engineering Contradiction:
Improvecomputational accuracyVSAvoidpower consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The system dynamically adjusts precision levels based on the specific computational workload and accuracy requirements. This allows the system to use higher precision only when necessary for maintaining computational accuracy, while using lower precision for operations where it is sufficient, thereby optimizing power consumption.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The chiplet architecture supports configurable precision parameters that can be changed based on workload characteristics. This enables the system to adapt precision levels to match computational needs, reducing power consumption by avoiding unnecessary high-precision computations while maintaining required accuracy.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20250209017A1Accelerator system using digital in-memory compute chiplet devices for computational workloads
Publication Date: 2025.06.26 D-MATRIX CORP
  • US20250209017A1 patent drawing
  • US20250209017A1 patent drawing
  • US20250209017A1 patent drawing

AI summary

A digital in-memory compute (DIMC) accelerator system using a chiplet architecture. The system includes a host device is configured to compile computational workload data for a target application obtained from data gathering devices into an instruction set architecture (ISA) graph to be executed by a plurality of accelerator apparatuses. Each such accelerator includes a plurality of chiplets, each of which includes a plurality of tiles, and each such tile includes a plurality of slices, a central processing unit (CPU), and a DIMC device configured to perform high throughput computations using the ISA graph to process the computational workload. The target application can include natural language processing (NLP), autonomous reasoning/decision-making, video/image processing, cybersecurity/fraud detection, manufacturing/industrial processes, agentic artificial intelligence (AI), or smart cities/Internet of Things (IoT). And the data gathering devices can include a web-scrapers, a dataset readers, a crowdsourcing devices, a sensors, a simulation devices, an IoT network, and others.